AI Content Detection
Quick facts
- What it affects
- Spam and trust filtering at step 3 (grounding) and step 4 (synthesis) of the answer loop.
- Is there a literal 'AI detector' in production?
- No major AI engine or search engine has confirmed that it uses one. OpenAI retired its own classifier in July 2023 because its accuracy was too low.
- What is actually penalized
- Patterns associated with mass production, including manufactured statistics, fabricated bylines, over-chunking, and unsupported schema claims, regardless of who or what produced them.
- AI use versus production patterns
- AI use itself is not penalized. Human review, original framing, and evidence of first-hand experience help AI-assisted work remain suitable for grounding.
- Industry-standard term?
- The general pattern is well established, but engines call it 'scaled content abuse' or 'low-quality content,' not 'AI detection.'
1. What “AI content detection” actually means
The phrase covers two different practices. Confusion about whether AI-assisted content risks penalties arises when these practices are treated as one.
Sense A: external classifiers. Third-party tools such as GPTZero, Originality.ai, Copyleaks, Pangram Labs, Turnitin, and Hive analyze features in a passage and predict whether a model wrote it. Publishers, schools, recruiters, and compliance teams buy these products. The tools do not operate inside AI engines, and their scores do not determine whether ChatGPT, Perplexity, Google AI Overviews, or Bing Copilot will display or cite a page.
Sense B: quality systems in search and AI engines. These are the spam and trust filters applied during retrieval, grounding, and synthesis. They are not designed to classify text as either human or AI. Instead, they penalize patterns associated with low effort or scaled abuse, regardless of whether the text was produced by a model, a content farm, or an agency. Google’s 2023 guidance on AI-generated content and its March 2024 Scaled Content Abuse policy state that the policy applies to “automation, human efforts, or some combination.”
The distinction can be summarized as follows.
| Sense A: external classifiers | Sense B: engine quality systems | |
|---|---|---|
| What it is | A vendor tool that scores features in a passage and predicts whether a model wrote it. | Spam and trust filters that engines apply during retrieval, grounding, and synthesis. |
| Examples | GPTZero, Originality.ai, Copyleaks, Pangram Labs, and Turnitin. | Google, Bing, OpenAI, and Perplexity use these systems inside their engines. |
| What it affects | The score has no direct effect inside an engine. People decide how to act on it. | The filters act on production patterns such as scaled abuse, manufactured statistics, fabricated bylines, and unsupported schema claims. |
| Why it matters for GEO | The classifier cannot make a page appear in or disappear from an AI answer. | The filters can exclude a page during grounding, before the Citability gate and the E-E-A-T trust filter. |
The distinction is: engines do not detect “AI.” They detect patterns associated with AI production at scale, and they penalize the same patterns when people produce them.
2. Where the anti-signal appears in the answer loop
These quality checks affect two stages of the four-step answer loop, along with a separate check at index time.
- Step 3: grounding. Over-optimized structures, including excessive chunking, FAQ stuffing, and template spam, can be detected here. A page may enter the candidate set but never be selected. Citability §6 explains why good structure is necessary but not sufficient.
- Step 4: trust filtering and synthesis. Low-effort mass content, manufactured statistics, and fabricated bylines can fail the source-worthiness check. E-E-A-T §7 explains why trust must be earned rather than declared.
- Before retrieval: index time. Schema or markup that asserts properties unsupported by the page can trigger anti-abuse checks before retrieval begins. This check is separate from the trust filter used during synthesis (see Schema for AI).
These checks do not form a separate “AI detector” stage. They are trust and spam systems operating at greater scale across a larger volume of AI-generated material. The same mechanism can reject over-optimized structure during grounding, fabricated authority during synthesis, or unsupported schema at index time.
3. Why classifier-based detection is unreliable
No peer-reviewed evidence shows that a commercial AI-text classifier can sustain its advertised accuracy on adversarial, real-world inputs. Vendor claims and published research tell very different stories.
| Evidence | What it shows |
|---|---|
| OpenAI shut down its own classifier, July 20, 2023 | OpenAI wrote that “the AI classifier is no longer available due to its low rate of accuracy.” At launch, the classifier identified AI text correctly 26% of the time and incorrectly labeled human text as AI-generated 9% of the time. OpenAI considered those results inadequate even before adversarial use (see OpenAI). |
| Liang et al., Patterns 2023 | ”More than half of the non-native-authored TOEFL essays are incorrectly classified as ‘AI-generated,’ while detectors exhibit near-perfect accuracy for US 8th-grade essays” (see arXiv:2304.02819). The headline finding is bias against non-native English writers. More broadly, the detectors confuse stylistic features such as lower perplexity and a more restricted vocabulary with model output. |
| Sadasivan et al., arXiv 2023 | ”Paraphrasing attacks can break a range of detectors, including those using watermarking schemes and neural network-based detectors” (see arXiv:2303.11156). The paper also establishes a theoretical limit: as language models more closely reproduce human text, even the best possible detector approaches random classification. |
The major vendors make ambitious claims. GPTZero (founded 2023-01) advertises “99% accuracy.” Originality.ai advertises “99% accuracy” and combines detection with plagiarism and fact-checking. Copyleaks advertises “99%+ accuracy” with a “0.2% false positive rate,” Pangram Labs advertises “99.98% accuracy,” and Turnitin advertises an “under 1% false positive rate” for documents containing at least 20% AI-generated text. Each figure comes from the vendor’s own benchmark. None has published a peer-reviewed evaluation that reproduces its advertised performance under the adversarial conditions studied in the research above.
For GEO work, an external classifier score is therefore too unreliable to use as audit evidence. Engine quality systems are the relevant concern.
4. Which patterns AI engines penalize
Google’s policy is based on purpose and pattern, not on the use of AI. It prohibits using any form of automation, including AI, to produce content primarily to manipulate rankings. The March 2024 core update expanded the spam policies to name expired domain abuse, scaled content abuse, and site reputation abuse. The current Spam Policies page defines scaled content abuse as “when many pages are generated for the primary purpose of manipulating search rankings and not helping users… using generative AI tools or other similar tools to generate many pages without adding value for users.” Bing’s Webmaster Guidelines and AI Performance preview use the same quality-based framing. OpenAI, Anthropic, and Perplexity do not publish a separate rule for AI tool use, and their answer systems rely on signals about source authority.
Google expanded the definition on 2026-05-15. The opening line of its spam policies previously referred to manipulating Search into “ranking content highly.” It now covers manipulating Search into “featuring content prominently … or attempting to manipulate generative AI responses in Google Search.” The catalog of prohibited patterns did not change, nor did the distinction between tools and production patterns. The change extended the policy’s reach to AI Overviews and AI Mode. A page that fails the trust filter during synthesis now violates a written rule rather than an inferred one. Manipulating an answer is now explicitly addressed in policy, not just in research (see GEO spam and manipulation).
Engine quality systems evaluate the following patterns.
| Pattern | What it looks like | Why it’s penalized |
|---|---|---|
| Mass-generated content | Many superficially complete pages produced at low marginal cost across unrelated topics. | Engines can detect low-effort patterns at scale, and Google’s scaled-content-abuse policy names the practice explicitly. E-E-A-T §7 explains the corresponding trust concerns. |
| Manufactured statistics | Unsourced numbers, suspiciously round figures, or citations to studies that do not exist. | Unsourced numbers fail trust filtering. Citability §6 and E-E-A-T §7 describe the same problem. |
| Fabricated bylines or credentials | Author profiles without sameAs corroboration or Knowledge Graph presence, paired with bios written to sound authoritative. | The claimed identity cannot be resolved, so the source can fail the trust filter described in E-E-A-T §7. |
| Over-chunking or FAQ stuffing | Many short, question-shaped fragments that do not match real queries. | The page resembles citable content, but its fragments lack context and its questions do not reflect real demand. The pattern can be detected as boilerplate during grounding. |
| Template or boilerplate spam | The same structure repeated across many topics or domains. | This mass-production pattern is named in both Google’s scaled-content-abuse policy and Bing’s quality guidance. |
| Unsupported schema or markup claims | Structured data asserts details about authors, ratings, or an organization’s sameAs identities that the page does not support. | Unsupported claims can trigger anti-abuse checks just as fabricated authority can trigger trust filters (see Schema for AI). |
| Expired-domain abuse | A previously trusted domain is purchased and repurposed for unrelated content. | The March 2024 spam-policy expansion names the practice directly. |
| Citation stuffing without substance | A page contains many citations, but they do not support the claims beside them. | Engines recognize and down-weight mismatches between citations and claims. E-E-A-T §7 describes the corresponding trust problem. |
Remove the word “AI,” and these patterns still describe scaled production. Human-written pages with the same patterns can therefore fail the same quality checks.
5. What GEO research shows about keyword stuffing
Aggarwal et al. tested nine content rewrites against GEO-bench. Rewrites that added substantive elements, including sources, statistics, and quotations, measurably improved answer visibility. Keyword Stuffing, the familiar SEO tactic, did not improve visibility and could make it worse. See Aggarwal et al., KDD ‘24 and the paper entry.
The limits of that finding matter. Using its own metric and an internal test system, the paper reported a lift of “up to 40%” for a single rewrite. On the live Perplexity.ai engine, the reported lift was around 22%. The paper entry’s critique attributes part of that gap to live-engine trust filtering: rewrites that “manufacture statistics” can perform well in the test system but lose ground on a live engine running Sense B quality checks. Puerto et al.’s C-SEO Bench (NeurIPS ‘25 D&B) extends this result to competitive settings. Many of the rewrites became ineffective or counterproductive when more than one author applied them.
The evidence supports this conclusion: engines actively penalize SEO-spam-style patterns, and the same mechanism limits the gains from substantive rewrites. Pattern detection and the ceiling on rewrite performance are two effects of the same quality filter.
6. What watermarking can and cannot do
As of 2026-05, text watermarking remains a research frontier rather than a useful source of audit evidence.
- Scott Aaronson’s 2022 proposal. In a Microsoft Research talk, Aaronson described a cryptographic approach that biases token sampling to create a pattern detectable only by the holder of a key.
- Google DeepMind’s SynthID-Text. This is the most credible attempt deployed in production. Google released it through Hugging Face Transformers in late 2024 and deployed it in Gemini. The system changes token-probability scores in a way that is “imperceptible to humans but visible to a trained model” (see SynthID). A Nature paper by Dathathri et al. reports a live A/B test involving about 20 million Gemini users with no decline in quality.
- Limits on detection. Paraphrasing through another model, light human editing, translation by a model that does not apply watermarks, and mixing watermarked and unwatermarked text all reduce detectability. The Nature paper likewise reports lower confidence for short or heavily edited outputs.
- No shared enforcement. OpenAI, Anthropic, Meta, and Google use different schemes or no scheme at all. Nothing requires their approaches to work together. A page processed by two models will almost certainly not retain a usable watermark.
No production AI engine grounds answers on watermark signals. A watermark is not an indexable sign of trust, an audit input, or a citation signal.
7. Can I use AI to write content?
Using AI is not penalized. Patterns associated with AI production at scale are. A draft produced with AI and then improved through human editing, original framing, and verifiable expertise does not fit the failure pattern. Mass AI content published without human review does, but engines can detect it through the same quality systems that have identified content farms for a decade. They evaluate the result rather than whether a model or a person wrote it.
Include details that are difficult to produce credibly at scale, such as evidence of first-hand contact with the subject. Those details support the Experience component of E-E-A-T. Specific observations, original data, named places, dated events, and verifiable claims are often absent from mass AI content. The limitation is not that a model cannot write such details. It is that producing them credibly at scale requires genuine experience.
You can stop worrying about:
- Whether GPTZero or Originality.ai will flag the page. As Section 3 explains, their scores do not enter an engine’s decision.
- Whether a ChatGPT-assisted draft is automatically penalized. Google’s policy states that it is not.
- Whether a machine-translated draft will trigger detection. As the FAQ explains, the risk comes from publishing without human review, not from machine translation itself.
You should focus instead on:
- Whether the content contains evidence of experience that could not have been produced without real contact with the subject.
- Whether every statistic is sourced and verifiable rather than merely plausible.
- Whether the byline identifies a real person whose
sameAsreferences and Knowledge Graph presence support the claim of authorship. - Whether the structure is substantive rather than merely appearing well structured. Citability §6 explains why structure is necessary but not sufficient.
A better question is not, “Did a human write this?” It is, “Is a human accountable for the claims?“
8. How to respond to the anti-signal
Quality signals and abuse patterns influence the same decision in opposite ways. Evidence and substance, as explained in E-E-A-T §9 and Citability §8, can improve selection, while patterns associated with scaled abuse can prevent it.
| Your intent | First stop |
|---|---|
| Review content for signs of over-optimization | Citability §6 |
| Check authorship, credentials, and sourcing | E-E-A-T |
| Confirm that schema does not make unsupported claims | Schema for AI |
| Identify where the anti-signal appears in the process | Answer Loop |
| Understand the broader method | Generative Engine Optimization |
| Review the supporting experiment | Aggarwal et al. (KDD ‘24) |
References
Official platform documentation (as of 2026-05):
- Google Search Central: Google Search’s guidance about AI-generated content (2023-02-08) · Using AI-generated content · What web creators should know about our March 2024 core update and new spam policies (2024-03-05) · Spam Policies for Google Web Search (spam definition revised 2026-05-15 to cover generative AI responses; see Search Engine Land) · An update to our site reputation abuse policy (2024-11-19)
- OpenAI: New AI classifier for indicating AI-written text (2023-01-31; discontinued 2023-07-20)
- Google DeepMind: SynthID — text watermarking
- Microsoft Bing: Webmaster Guidelines · Introducing AI Performance in Bing Webmaster Tools (Public Preview) (2026-02-09)
Academic:
- Dathathri, S., et al. (2024). Scalable watermarking for identifying large language model outputs. Nature 634, 818–823. doi:10.1038/s41586-024-08025-4
- Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns 4(7), 100779. arXiv:2304.02819
- Sadasivan, V. S., Kumar, A., Balasubramanian, S., Wang, W., & Feizi, S. (2023). Can AI-Generated Text be Reliably Detected? arXiv:2303.11156
- Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., & Deshpande, A. (2024). GEO: Generative Engine Optimization. KDD ‘24. arXiv:2311.09735 · ACM DL · paper summary
- Puerto, H., Gubri, M., Green, C., Oh, S. J., & Yun, S. (2025). C-SEO Bench: Does Conversational SEO Work? NeurIPS ‘25 Datasets & Benchmarks. arXiv:2506.11097
Vendor pages (listed for reference only; see Section 3 for the evidence on reliability):
Frequently asked questions
Will Google detect that I used ChatGPT to write this article?
Are GPTZero, Originality.ai, Copyleaks, Pangram, and Turnitin reliable?
Does watermarking solve this?
Can I use AI for first drafts of articles?
What about AI-translated content?
See also
Sources
Primary
- Google Search's guidance about AI-generated content · Google Search Central · 2023-02-08
- Using AI-generated content · Google Search Central
- What web creators should know about our March 2024 core update and new spam policies · Google Search Central · 2024-03-05
- Spam Policies for Google Web Search · Google Search Central
- An update to our site reputation abuse policy · Google Search Central · 2024-11-19
- New AI classifier for indicating AI-written text · OpenAI · 2023-01-31
- SynthID — text watermarking · Google DeepMind
- Bing Webmaster Guidelines · Microsoft Bing
- Introducing AI Performance in Bing Webmaster Tools (Public Preview) · Microsoft Bing · 2026-02-09
- GEO: Generative Engine Optimization (Aggarwal et al., KDD '24) · arXiv · 2024-06-28
- GEO: Generative Engine Optimization (KDD '24 Proceedings) · ACM SIGKDD · 2024-08-25
Secondary
- Google updates Search spam policies to clarify it applies to generative AI responses · Search Engine Land
- Scalable watermarking for identifying large language model outputs (Dathathri et al., Nature 2024) · Nature
- GPT detectors are biased against non-native English writers (Liang et al., Patterns 2023) · arXiv / Patterns (Cell Press)
- Can AI-Generated Text be Reliably Detected? (Sadasivan et al. 2023) · arXiv
- C-SEO Bench: Does Conversational SEO Work? (Puerto et al., NeurIPS '25 D&B) · arXiv / NeurIPS '25 D&B