How AI Detectors Work: A Practical 2026 Guide

Learn how AI text detectors and provider watermarks work, why ChatGPT, Claude and Gemini differ, and how to evaluate a result responsibly.

Follow in Google Search

AI text detectors estimate whether a passage resembles writing produced by a language model. They do not search for a hidden ChatGPT label, recover the prompt or prove who typed the words. A modern detector usually combines a machine learning classifier with patterns learned from labeled human and AI text, then turns those signals into a document score and, in some products, sentence-level highlights.

The useful question is therefore not whether a detector can read authorship like a fingerprint. It is whether the tool was trained for the language, model family, document type and editing conditions in front of it, and whether the result is strong enough to justify a closer review. Longer samples, current training data and a transparent review process make the evidence more useful. A score on its own should not decide a grade, disciplinary case or publication decision.

Question

Short answer

Practical consequence

What does a detector measure?

Patterns associated with human and AI training examples

The result is probabilistic, not proof of authorship

Does every detector use perplexity?

No

Some expose simple statistical features; others use proprietary neural classifiers

Does more text help?

Usually

Short passages contain less stable evidence and should be treated cautiously

Can edited AI text be detected?

Sometimes

Heavy editing can weaken the original machine pattern

Can human text be flagged?

Yes

Formal, repetitive or highly constrained prose can create false positives

What an AI detector is actually classifying

A text detector is a supervised classification system. Developers assemble examples labeled as human-written or AI-generated, divide them into training and evaluation sets, and teach a model to separate the two classes. The model does not memorize one universal AI style. It learns combinations of features that helped distinguish the examples in its dataset.

Those examples determine the boundary the detector learns. A system trained mainly on English essays from one set of generators may perform differently on technical documentation, translated prose, creative fiction, newer models or text revised by a human. Strong vendors continually refresh evaluation sets as generators change. Strong users also test the detector on their own document types instead of relying on one vendor accuracy percentage.

The five stages of AI text detection

1. The document is cleaned and divided into units

The detector first normalizes the submission. It may remove formatting, identify prose, split the text into sentences or overlapping windows, and exclude structures the model was not designed to score. References, code, tables, bullet lists and very short fragments may be ignored by one detector and included by another. This preprocessing step changes the denominator behind the final percentage.

2. The system creates a numerical representation

Language must be converted into numbers before a model can classify it. Some systems calculate explicit features such as word predictability, sentence-length variation, repeated constructions or vocabulary distribution. Neural systems may instead create embeddings that capture relationships among tokens, sentences and larger passages. Many commercial products use a mixture of engineered and learned features without publishing every component.

3. A classifier looks for learned combinations

The classifier weighs many signals together. One predictable sentence is not enough, and a varied sentence is not automatically human. The model asks whether the pattern across the sample is closer to the human or AI examples it learned from. Transformer-based classifiers can capture contextual relationships that a simple perplexity threshold would miss.

4. Sentence or segment scores are combined

A document can contain human writing, generated text and later edits. Detectors often score smaller windows and aggregate them into a document result. Overlapping windows give surrounding context, but they can also make boundaries fuzzy. A human sentence beside an AI passage may inherit some of the surrounding signal, while a short document may produce an unstable all-or-nothing result.

5. The probability is converted into a report

The interface may show an AI probability, a predicted share of AI-like prose, a human score or a category such as likely AI. These labels are not interchangeable. Users should read the product definition before interpreting the number. Sentence highlights are useful for locating evidence, but they still represent model predictions rather than a record of the writing process.

Signal or component

What it captures

Why it cannot decide alone

Perplexity

How surprising a sequence is to a language model

Formal human prose can also be predictable

Burstiness

Variation in sentence length and structure

Writers and genres differ naturally

Stylometric patterns

Vocabulary, syntax, punctuation and recurring habits

Editing and assignment constraints change style

Embeddings

Contextual and semantic relationships in the text

The representation depends on training data

Neural classifier

Combinations of learned features across a passage

A confident output can still be wrong outside its training distribution

Perplexity and burstiness are useful explanations, not a universal recipe

Perplexity is often used to explain detection because generators choose words from probability distributions. A smooth, highly predictable passage can receive lower perplexity than writing with unusual phrasing. Burstiness describes variation across sentences or passages. Human drafts often shift rhythm, while unedited generated text can appear more uniform.

These concepts make the problem understandable, but they should not be mistaken for the complete implementation of every product. A detector can use a transformer classifier without calculating a user-facing perplexity score. Another can combine dozens of explicit features with embeddings. Turnitin, for example, describes a proprietary transformer-based deep-learning architecture, while Winston AI describes a layered analysis that includes statistical and machine learning signals. The output must be judged by measured performance, not by how simple the marketing explanation sounds.

Watermarking is a different kind of AI detection

The recent move toward text watermarking changes the detection landscape, but it does not make ordinary AI detectors obsolete. A classifier studies patterns that tend to appear in AI writing without cooperation from the model provider. A watermark is deliberately placed during generation by a provider that controls the model. Detecting it requires knowledge of the watermarking method and, for keyed systems, access to the correct verification key.

Text watermarking does not usually mean adding an invisible Unicode character, a hidden account number or a marker that appears when text is pasted. Modern statistical watermarks make tiny changes to token selection while the model writes. When several next words would all be acceptable, the generator uses a secret or defined pattern to influence the choice. Across a long passage, those choices form evidence that a compatible verifier can test.

Provider or system

Current text position in September 2026

What users should understand

Google Gemini

SynthID watermarks and identifies text generated in the Gemini app and web experience

The signal is created through token-probability adjustments and is not visible to readers

Anthropic Claude

Newer Claude models use a SynthID-Text-derived keyed watermark, with older-model coverage being rolled forward

Verification estimates Claude involvement and does not identify a user, organization or chat

OpenAI ChatGPT

Official OpenAI provenance checks currently document supported signals for images and audio, not ordinary text

Do not assume copied ChatGPT text contains a deployed, publicly verifiable text watermark

Independent AI detector

Classifies writing patterns without a provider watermark key

A detector score and a watermark result are different forms of evidence

How token-level text watermarks work

A language model repeatedly chooses the next token from a probability distribution. In many positions, several choices preserve the meaning and quality of the response. A watermarking scheme can use a key and the preceding context to divide or rank acceptable token choices, then gently favor choices that encode a statistical pattern. The finished paragraph reads normally, but a verifier with the right method can test whether the sequence is unusually consistent with that pattern.

The watermark needs enough eligible choices to accumulate evidence. A long explanatory passage offers many opportunities. A short answer, a fixed fact, a mathematical result or exact code offers fewer. This is why provider documentation emphasizes confidence rather than certainty and why watermark detection becomes stronger as the sample grows.

Google Gemini and SynthID Text

Google DeepMind says SynthID Text is used to watermark and identify text generated by the Gemini app and web experience. The system changes token probability scores during generation without producing a visible label or intentionally changing output quality. Google has also open-sourced SynthID Text components for developers who want to add compatible watermarking to their own models.

Scope matters. A statement about the Gemini app and web experience should not automatically be extended to every historical Gemini model, API route, third-party host or transformed copy. Verification should match the originating product and date. A negative result means the expected signal was not detected; it does not prove that a person wrote the text.

Claude watermarking

Anthropic announced a keyed text watermark for newer Claude models based on the SynthID-Text approach. It changes the source of randomness used when Claude chooses among similarly suitable words. Anthropic says it adds no hidden characters, carries no user or chat identifier and has no practical effect on price or output quality. Detection asks whether Claude was likely involved in producing part of the text, not who used it or who legally authored it.

Anthropic also documents important limits. Small samples contain too few choices. Factual passages, code and narrow proofreading tasks offer less room for a mark. Light edits may preserve enough of the pattern, while a complete rewrite can remove it. Its dedicated detection API is in private preview for eligible organizations, so a normal third-party AI detector cannot simply claim to read Claude's secret key.

ChatGPT and OpenAI provenance

OpenAI's current official Content Provenance API checks images for C2PA Content Credentials and SynthID signals, and checks supported audio for SynthID. The documented endpoint accepts image and audio files, not plain text. That means the safe current statement is narrow: OpenAI provides provenance verification for supported media, while its official developer documentation does not establish a comparable public watermark for ordinary ChatGPT text.

This also separates three ideas that are often mixed together. C2PA is signed file metadata and can be removed when a file is converted or stripped. An embedded media watermark lives in image or audio content and may survive some transformations. A statistical text watermark is a pattern distributed across token choices. None of these is the same as a general classifier deciding that prose sounds AI-generated.

Why watermarking does not settle authorship

A detected watermark can provide stronger source evidence than a generic style classifier because the provider deliberately created the signal. It still does not show whether using AI was allowed, whether a person substantially revised the work, who operated the account or who owns the output. A missing watermark is even less conclusive because the text may be too short, may come from an unsupported model, may have been edited or may have been generated before deployment.

Schools and publishers should therefore keep watermark checks and AI detector scores in separate fields. Record which provider signal was tested, which verifier and key were used, the text length and whether editing occurred. Then examine the writing process and policy context. Watermarking improves provenance for supported output; it does not turn a final document into a complete history of how the work was created.

Technical background: Google DeepMind's SynthID-Text research paper describes the statistical method, while the C2PA standard covers signed media provenance.

Why detector accuracy changes

Condition

Likely effect

Reason

Very short text

Lower confidence

Too little evidence for stable document-level patterns

New generator or model update

Possible false negatives

The output may differ from the detector training set

Heavy human editing

Possible false negatives

Revision can replace machine-associated patterns

Formal or formulaic human writing

Possible false positives

Predictability and repeated structure can resemble training examples

Translated or second-language writing

Variable performance

Language distributions may differ from the training population

Mixed human and AI passages

Uncertain boundaries

Segment aggregation can blur where one source ends

Accuracy is not one permanent property of a detector. It is a result measured on a particular dataset, at a particular threshold, with particular versions of the detector and generators. A vendor can report excellent performance on balanced, clean examples while a school sees different results on authentic student work. False-positive rate, false-negative rate, precision and recall all matter, especially when the real proportion of AI-assisted work is unknown.

Compare the current product landscape in Best AI Detectors in 2026. For Winston AI specifically, its score interpretation guide explains how to read sentence and document evidence.

A repeatable way to test an AI detector

A fair evaluation starts with known provenance. Collect human writing created before widespread generative AI use or under a documented process, and generate matched AI samples for the same prompts. Add hybrid and edited versions because those are the cases users will actually encounter. Keep the evaluator blind to the source while recording the exact tool version, date, threshold and result.

Test set

Minimum design

What it reveals

Human control

Multiple writers, genres and proficiency levels

False-positive risk and demographic variation

Unedited AI

Several current models using matched prompts

Baseline recall across generators

Hybrid document

Known inserted passages at several proportions

Sensitivity to mixed authorship and boundaries

Human-edited AI

Documented revisions of different intensity

Robustness after realistic editing

Short and long versions

Matched samples at several lengths

How evidence changes with document size

Out-of-domain text

Code, lists, technical and creative samples

Where the detector should decline or warn

Do not average every output into a single accuracy figure and stop. Report false positives separately from missed AI text. Show results by length, language, generator and editing condition. If a product returns a continuous score, choose the decision threshold before looking at final test labels. Retesting after major product updates is essential because both sides of the classification problem change.

How to interpret a detector result responsibly

Result pattern

Reasonable next step

What not to conclude

Low score on a long supported sample

Treat as one weak signal and continue normal review

That no AI assistance occurred

High score with concentrated highlights

Inspect the passages, assignment and drafts

That misconduct is proven

Conflicting tools

Check tool scope, versions and document length

That majority vote creates certainty

Score on a short fragment

Request more context or decline to judge

That the fragment identifies its author

Human writer disputes the result

Review notes, sources, version history and explanation

That the tool outranks process evidence

In education, a detector should open an inquiry rather than close one. Compare the highlighted language with the student's drafts, notes and earlier work. Ask the student to explain the argument, sources and revision choices. Apply the same process to every student, document the evidence considered and provide an appeal route. In publishing or hiring, disclose the review policy and avoid treating a proprietary score as a hidden automatic rejection rule.

AI detection and plagiarism detection answer different questions

Plagiarism checking searches for matching language in websites, publications or stored submissions. AI detection classifies patterns inside the submitted text. Generated prose can be original in the matching sense while still violating an assignment rule. Human prose can contain copied passages while looking entirely human to an AI classifier. A complete integrity review keeps the two reports separate.

See the distinction in practice in Best Plagiarism Checkers and the guide to AI tools for teachers. Winston AI also provides both an AI content detector and plagiarism review within one workflow.

The central limitation is missing process evidence

Text classification examines the final artifact. It cannot see brainstorming, drafting, source use, classroom discussion or the sequence of edits that produced it. That is why version history, oral explanation and clear assignment rules are stronger companions than a second detector score. Newer writing-transparency systems attempt to capture process evidence directly, but they raise their own privacy and implementation questions.

Independent studies continue to find that performance varies sharply across tools and conditions. Research has also documented risks for non-native English writers and for text transformed after generation. The practical response is neither blind trust nor blanket dismissal. Use detectors within their documented scope, validate them locally and reserve consequential decisions for a human review with multiple kinds of evidence.

Frequently asked questions

Can AI detectors prove that ChatGPT wrote something?

No. They estimate whether textual patterns resemble known AI-generated examples. They do not recover the prompt, account or writing history that created the document.

What is a good AI detection score?

There is no universal cutoff. The meaning depends on the product, document length, supported language, threshold and purpose. Read the tool definition and examine passage-level evidence.

Why do different AI detectors disagree?

They use different training data, model architectures, thresholds, preprocessing rules and update schedules. A disagreement is a reason to inspect scope and evidence, not to take a majority vote.

Can a humanizer fool an AI detector?

Editing and paraphrasing can reduce detectable patterns, but results vary by detector and method. This is one reason evaluations should include human-edited and transformed samples.

Do AI detectors work on code or bullet points?

Some products are designed only for long-form prose. Code, lists and fragments may be unsupported or excluded, so check the product requirements before interpreting a result.

Should schools use AI detectors?

They can be one part of an academic-integrity process if the school validates the tool, trains reviewers, protects student data and prohibits automatic penalties based on a score alone.

Sources and methodology notes

This guide was checked on September 16, 2026 against current vendor documentation and recent research. Product mechanisms are described only at the level their publishers document. Vendor accuracy claims are not treated as directly comparable because datasets, thresholds and versions differ.

Independent research: Evaluating the accuracy and reliability of AI content detectors in academic contexts.

Bias research: GPT detectors are biased against non-native English writers.

Winston AI: How AI detectors work.

Turnitin sources reviewed without outbound product links: AI writing detection capabilities FAQ, Using the AI Writing Report, File requirements for an AI Writing Report, and AI writing detection model release notes.

Author

Dr. Elena Vasquez

PhD in Computer Science, Stanford University (2018); MS in Machine Learning, Carnegie Mellon University. Research on scaling laws, evaluation methodologies, and robustness in large neural models.