Skip to content

AI Content Detection

Quick facts

What it affects
Spam and trust filtering at step 3 (grounding) and step 4 (synthesis) of the answer loop.
Is there a literal 'AI detector' in production?
No major AI engine or search engine has confirmed that it uses one. OpenAI retired its own classifier in July 2023 because its accuracy was too low.
What is actually penalized
Patterns associated with mass production, including manufactured statistics, fabricated bylines, over-chunking, and unsupported schema claims, regardless of who or what produced them.
AI use versus production patterns
AI use itself is not penalized. Human review, original framing, and evidence of first-hand experience help AI-assisted work remain suitable for grounding.
Industry-standard term?
The general pattern is well established, but engines call it 'scaled content abuse' or 'low-quality content,' not 'AI detection.'

1. What “AI content detection” actually means

The phrase covers two different practices. Confusion about whether AI-assisted content risks penalties arises when these practices are treated as one.

Sense A: external classifiers. Third-party tools such as GPTZero, Originality.ai, Copyleaks, Pangram Labs, Turnitin, and Hive analyze features in a passage and predict whether a model wrote it. Publishers, schools, recruiters, and compliance teams buy these products. The tools do not operate inside AI engines, and their scores do not determine whether ChatGPT, Perplexity, Google AI Overviews, or Bing Copilot will display or cite a page.

Sense B: quality systems in search and AI engines. These are the spam and trust filters applied during retrieval, grounding, and synthesis. They are not designed to classify text as either human or AI. Instead, they penalize patterns associated with low effort or scaled abuse, regardless of whether the text was produced by a model, a content farm, or an agency. Google’s 2023 guidance on AI-generated content and its March 2024 Scaled Content Abuse policy state that the policy applies to “automation, human efforts, or some combination.”

The distinction can be summarized as follows.

Sense A: external classifiersSense B: engine quality systems
What it isA vendor tool that scores features in a passage and predicts whether a model wrote it.Spam and trust filters that engines apply during retrieval, grounding, and synthesis.
ExamplesGPTZero, Originality.ai, Copyleaks, Pangram Labs, and Turnitin.Google, Bing, OpenAI, and Perplexity use these systems inside their engines.
What it affectsThe score has no direct effect inside an engine. People decide how to act on it.The filters act on production patterns such as scaled abuse, manufactured statistics, fabricated bylines, and unsupported schema claims.
Why it matters for GEOThe classifier cannot make a page appear in or disappear from an AI answer.The filters can exclude a page during grounding, before the Citability gate and the E-E-A-T trust filter.

The distinction is: engines do not detect “AI.” They detect patterns associated with AI production at scale, and they penalize the same patterns when people produce them.

2. Where the anti-signal appears in the answer loop

These quality checks affect two stages of the four-step answer loop, along with a separate check at index time.

  • Step 3: grounding. Over-optimized structures, including excessive chunking, FAQ stuffing, and template spam, can be detected here. A page may enter the candidate set but never be selected. Citability §6 explains why good structure is necessary but not sufficient.
  • Step 4: trust filtering and synthesis. Low-effort mass content, manufactured statistics, and fabricated bylines can fail the source-worthiness check. E-E-A-T §7 explains why trust must be earned rather than declared.
  • Before retrieval: index time. Schema or markup that asserts properties unsupported by the page can trigger anti-abuse checks before retrieval begins. This check is separate from the trust filter used during synthesis (see Schema for AI).

These checks do not form a separate “AI detector” stage. They are trust and spam systems operating at greater scale across a larger volume of AI-generated material. The same mechanism can reject over-optimized structure during grounding, fabricated authority during synthesis, or unsupported schema at index time.

3. Why classifier-based detection is unreliable

No peer-reviewed evidence shows that a commercial AI-text classifier can sustain its advertised accuracy on adversarial, real-world inputs. Vendor claims and published research tell very different stories.

EvidenceWhat it shows
OpenAI shut down its own classifier, July 20, 2023OpenAI wrote that “the AI classifier is no longer available due to its low rate of accuracy.” At launch, the classifier identified AI text correctly 26% of the time and incorrectly labeled human text as AI-generated 9% of the time. OpenAI considered those results inadequate even before adversarial use (see OpenAI).
Liang et al., Patterns 2023”More than half of the non-native-authored TOEFL essays are incorrectly classified as ‘AI-generated,’ while detectors exhibit near-perfect accuracy for US 8th-grade essays” (see arXiv:2304.02819). The headline finding is bias against non-native English writers. More broadly, the detectors confuse stylistic features such as lower perplexity and a more restricted vocabulary with model output.
Sadasivan et al., arXiv 2023”Paraphrasing attacks can break a range of detectors, including those using watermarking schemes and neural network-based detectors” (see arXiv:2303.11156). The paper also establishes a theoretical limit: as language models more closely reproduce human text, even the best possible detector approaches random classification.

The major vendors make ambitious claims. GPTZero (founded 2023-01) advertises “99% accuracy.” Originality.ai advertises “99% accuracy” and combines detection with plagiarism and fact-checking. Copyleaks advertises “99%+ accuracy” with a “0.2% false positive rate,” Pangram Labs advertises “99.98% accuracy,” and Turnitin advertises an “under 1% false positive rate” for documents containing at least 20% AI-generated text. Each figure comes from the vendor’s own benchmark. None has published a peer-reviewed evaluation that reproduces its advertised performance under the adversarial conditions studied in the research above.

For GEO work, an external classifier score is therefore too unreliable to use as audit evidence. Engine quality systems are the relevant concern.

4. Which patterns AI engines penalize

Google’s policy is based on purpose and pattern, not on the use of AI. It prohibits using any form of automation, including AI, to produce content primarily to manipulate rankings. The March 2024 core update expanded the spam policies to name expired domain abuse, scaled content abuse, and site reputation abuse. The current Spam Policies page defines scaled content abuse as “when many pages are generated for the primary purpose of manipulating search rankings and not helping users… using generative AI tools or other similar tools to generate many pages without adding value for users.” Bing’s Webmaster Guidelines and AI Performance preview use the same quality-based framing. OpenAI, Anthropic, and Perplexity do not publish a separate rule for AI tool use, and their answer systems rely on signals about source authority.

Google expanded the definition on 2026-05-15. The opening line of its spam policies previously referred to manipulating Search into “ranking content highly.” It now covers manipulating Search into “featuring content prominently … or attempting to manipulate generative AI responses in Google Search.” The catalog of prohibited patterns did not change, nor did the distinction between tools and production patterns. The change extended the policy’s reach to AI Overviews and AI Mode. A page that fails the trust filter during synthesis now violates a written rule rather than an inferred one. Manipulating an answer is now explicitly addressed in policy, not just in research (see GEO spam and manipulation).

Engine quality systems evaluate the following patterns.

PatternWhat it looks likeWhy it’s penalized
Mass-generated contentMany superficially complete pages produced at low marginal cost across unrelated topics.Engines can detect low-effort patterns at scale, and Google’s scaled-content-abuse policy names the practice explicitly. E-E-A-T §7 explains the corresponding trust concerns.
Manufactured statisticsUnsourced numbers, suspiciously round figures, or citations to studies that do not exist.Unsourced numbers fail trust filtering. Citability §6 and E-E-A-T §7 describe the same problem.
Fabricated bylines or credentialsAuthor profiles without sameAs corroboration or Knowledge Graph presence, paired with bios written to sound authoritative.The claimed identity cannot be resolved, so the source can fail the trust filter described in E-E-A-T §7.
Over-chunking or FAQ stuffingMany short, question-shaped fragments that do not match real queries.The page resembles citable content, but its fragments lack context and its questions do not reflect real demand. The pattern can be detected as boilerplate during grounding.
Template or boilerplate spamThe same structure repeated across many topics or domains.This mass-production pattern is named in both Google’s scaled-content-abuse policy and Bing’s quality guidance.
Unsupported schema or markup claimsStructured data asserts details about authors, ratings, or an organization’s sameAs identities that the page does not support.Unsupported claims can trigger anti-abuse checks just as fabricated authority can trigger trust filters (see Schema for AI).
Expired-domain abuseA previously trusted domain is purchased and repurposed for unrelated content.The March 2024 spam-policy expansion names the practice directly.
Citation stuffing without substanceA page contains many citations, but they do not support the claims beside them.Engines recognize and down-weight mismatches between citations and claims. E-E-A-T §7 describes the corresponding trust problem.

Remove the word “AI,” and these patterns still describe scaled production. Human-written pages with the same patterns can therefore fail the same quality checks.

5. What GEO research shows about keyword stuffing

Aggarwal et al. tested nine content rewrites against GEO-bench. Rewrites that added substantive elements, including sources, statistics, and quotations, measurably improved answer visibility. Keyword Stuffing, the familiar SEO tactic, did not improve visibility and could make it worse. See Aggarwal et al., KDD ‘24 and the paper entry.

The limits of that finding matter. Using its own metric and an internal test system, the paper reported a lift of “up to 40%” for a single rewrite. On the live Perplexity.ai engine, the reported lift was around 22%. The paper entry’s critique attributes part of that gap to live-engine trust filtering: rewrites that “manufacture statistics” can perform well in the test system but lose ground on a live engine running Sense B quality checks. Puerto et al.’s C-SEO Bench (NeurIPS ‘25 D&B) extends this result to competitive settings. Many of the rewrites became ineffective or counterproductive when more than one author applied them.

The evidence supports this conclusion: engines actively penalize SEO-spam-style patterns, and the same mechanism limits the gains from substantive rewrites. Pattern detection and the ceiling on rewrite performance are two effects of the same quality filter.

6. What watermarking can and cannot do

As of 2026-05, text watermarking remains a research frontier rather than a useful source of audit evidence.

  • Scott Aaronson’s 2022 proposal. In a Microsoft Research talk, Aaronson described a cryptographic approach that biases token sampling to create a pattern detectable only by the holder of a key.
  • Google DeepMind’s SynthID-Text. This is the most credible attempt deployed in production. Google released it through Hugging Face Transformers in late 2024 and deployed it in Gemini. The system changes token-probability scores in a way that is “imperceptible to humans but visible to a trained model” (see SynthID). A Nature paper by Dathathri et al. reports a live A/B test involving about 20 million Gemini users with no decline in quality.
  • Limits on detection. Paraphrasing through another model, light human editing, translation by a model that does not apply watermarks, and mixing watermarked and unwatermarked text all reduce detectability. The Nature paper likewise reports lower confidence for short or heavily edited outputs.
  • No shared enforcement. OpenAI, Anthropic, Meta, and Google use different schemes or no scheme at all. Nothing requires their approaches to work together. A page processed by two models will almost certainly not retain a usable watermark.

No production AI engine grounds answers on watermark signals. A watermark is not an indexable sign of trust, an audit input, or a citation signal.

7. Can I use AI to write content?

Using AI is not penalized. Patterns associated with AI production at scale are. A draft produced with AI and then improved through human editing, original framing, and verifiable expertise does not fit the failure pattern. Mass AI content published without human review does, but engines can detect it through the same quality systems that have identified content farms for a decade. They evaluate the result rather than whether a model or a person wrote it.

Include details that are difficult to produce credibly at scale, such as evidence of first-hand contact with the subject. Those details support the Experience component of E-E-A-T. Specific observations, original data, named places, dated events, and verifiable claims are often absent from mass AI content. The limitation is not that a model cannot write such details. It is that producing them credibly at scale requires genuine experience.

You can stop worrying about:

  • Whether GPTZero or Originality.ai will flag the page. As Section 3 explains, their scores do not enter an engine’s decision.
  • Whether a ChatGPT-assisted draft is automatically penalized. Google’s policy states that it is not.
  • Whether a machine-translated draft will trigger detection. As the FAQ explains, the risk comes from publishing without human review, not from machine translation itself.

You should focus instead on:

  • Whether the content contains evidence of experience that could not have been produced without real contact with the subject.
  • Whether every statistic is sourced and verifiable rather than merely plausible.
  • Whether the byline identifies a real person whose sameAs references and Knowledge Graph presence support the claim of authorship.
  • Whether the structure is substantive rather than merely appearing well structured. Citability §6 explains why structure is necessary but not sufficient.

A better question is not, “Did a human write this?” It is, “Is a human accountable for the claims?“

8. How to respond to the anti-signal

Quality signals and abuse patterns influence the same decision in opposite ways. Evidence and substance, as explained in E-E-A-T §9 and Citability §8, can improve selection, while patterns associated with scaled abuse can prevent it.

Your intentFirst stop
Review content for signs of over-optimizationCitability §6
Check authorship, credentials, and sourcingE-E-A-T
Confirm that schema does not make unsupported claimsSchema for AI
Identify where the anti-signal appears in the processAnswer Loop
Understand the broader methodGenerative Engine Optimization
Review the supporting experimentAggarwal et al. (KDD ‘24)

References

Official platform documentation (as of 2026-05):

Academic:

  • Dathathri, S., et al. (2024). Scalable watermarking for identifying large language model outputs. Nature 634, 818–823. doi:10.1038/s41586-024-08025-4
  • Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns 4(7), 100779. arXiv:2304.02819
  • Sadasivan, V. S., Kumar, A., Balasubramanian, S., Wang, W., & Feizi, S. (2023). Can AI-Generated Text be Reliably Detected? arXiv:2303.11156
  • Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., & Deshpande, A. (2024). GEO: Generative Engine Optimization. KDD ‘24. arXiv:2311.09735 · ACM DL · paper summary
  • Puerto, H., Gubri, M., Green, C., Oh, S. J., & Yun, S. (2025). C-SEO Bench: Does Conversational SEO Work? NeurIPS ‘25 Datasets & Benchmarks. arXiv:2506.11097

Vendor pages (listed for reference only; see Section 3 for the evidence on reliability):

Frequently asked questions

Will Google detect that I used ChatGPT to write this article?
Google does not use a classifier to sort pages into 'human' and 'AI' categories. Its published position in 2023 and 2024 was that appropriate AI use is acceptable, while using any form of automation to generate content primarily to manipulate rankings violates its spam policies. The March 2024 core update made this explicit through the Scaled Content Abuse policy, which applies whether content was produced 'through automation, human efforts, or some combination.' Google evaluates production patterns and intent, not the choice of tool.
Are GPTZero, Originality.ai, Copyleaks, Pangram, and Turnitin reliable?
These commercial products report accuracy of 99% or higher, but independent peer-reviewed evaluations have not reproduced those results. Liang et al. (Patterns, 2023) found that commonly used detectors misclassified more than half of TOEFL essays written by non-native English speakers, while achieving near-perfect accuracy on essays by native speakers. Sadasivan et al. (2023) found that simple paraphrasing reduced detector accuracy to nearly random. OpenAI retired its own classifier in July 2023 'due to its low rate of accuracy.' Treat any classifier score, including a score of 99%, as low-confidence input rather than evidence.
Does watermarking solve this?
Not yet, and not through any single vendor. Google DeepMind's SynthID-Text (Nature, 2024) is the most credible production scheme. It is used in Gemini and is open source, but its watermark can be detected statistically only in raw model output. Paraphrasing, light human editing, translation by a model that does not apply watermarks, or mixing watermarked and unwatermarked text all reduce detectability. There is also no shared enforcement across vendors. OpenAI, Anthropic, Meta, and Google use different schemes or none at all, and nothing requires those schemes to work together. As of 2026-05, no production AI engine grounds answers on watermark signals.
Can I use AI for first drafts of articles?
Yes, with two conditions. First, engines penalize scaled abuse, manufactured statistics, fabricated bylines, FAQ stuffing, and unsupported schema claims, not the use of an AI tool. An AI-assisted draft that receives human editing, original framing, and verifiable expertise does not fit those patterns. Second, the content needs evidence of experience, such as first-hand product use, specific named places, dated events, or original data. Those details cannot be produced credibly at scale without real knowledge. The practical question is not whether a human wrote every word, but whether a human is accountable for the claims.
What about AI-translated content?
Machine translation alone is not the problem. Risk arises when no one reviews the translation, the source material is thin, or the translated page claims expertise that it cannot support. The same signals apply to original and translated content: corroborated authorship, accurate claims, source citations, and freshness. A carefully edited translation of a strong source can be used for grounding. An unreviewed machine translation of a thin source can resemble scaled abuse, just as the original would.

See also

Sources

Primary

  1. Google Search's guidance about AI-generated content · Google Search Central · 2023-02-08
  2. Using AI-generated content · Google Search Central
  3. What web creators should know about our March 2024 core update and new spam policies · Google Search Central · 2024-03-05
  4. Spam Policies for Google Web Search · Google Search Central
  5. An update to our site reputation abuse policy · Google Search Central · 2024-11-19
  6. New AI classifier for indicating AI-written text · OpenAI · 2023-01-31
  7. SynthID — text watermarking · Google DeepMind
  8. Bing Webmaster Guidelines · Microsoft Bing
  9. Introducing AI Performance in Bing Webmaster Tools (Public Preview) · Microsoft Bing · 2026-02-09
  10. GEO: Generative Engine Optimization (Aggarwal et al., KDD '24) · arXiv · 2024-06-28
  11. GEO: Generative Engine Optimization (KDD '24 Proceedings) · ACM SIGKDD · 2024-08-25

Secondary

  1. Google updates Search spam policies to clarify it applies to generative AI responses · Search Engine Land
  2. Scalable watermarking for identifying large language model outputs (Dathathri et al., Nature 2024) · Nature
  3. GPT detectors are biased against non-native English writers (Liang et al., Patterns 2023) · arXiv / Patterns (Cell Press)
  4. Can AI-Generated Text be Reliably Detected? (Sadasivan et al. 2023) · arXiv
  5. C-SEO Bench: Does Conversational SEO Work? (Puerto et al., NeurIPS '25 D&B) · arXiv / NeurIPS '25 D&B
First published: 2026-05-21 Last updated: 2026-07-26 Authors: Ray Yang Topic: Signals