AI text detectors estimate whether a passage resembles writing produced by a language model. They do not search for a hidden ChatGPT label, recover the prompt or prove who typed the words. A modern detector usually combines a machine learning classifier with patterns learned from labeled human and AI text, then turns those signals into a document score and, in some products, sentence-level highlights.
The useful question is therefore not whether a detector can read authorship like a fingerprint. It is whether the tool was trained for the language, model family, document type and editing conditions in front of it, and whether the result is strong enough to justify a closer review. Longer samples, current training data and a transparent review process make the evidence more useful. A score on its own should not decide a grade, disciplinary case or publication decision.
Question | Short answer | Practical consequence |
|---|---|---|
What does a detector measure? | Patterns associated with human and AI training examples | The result is probabilistic, not proof of authorship |
Does every detector use perplexity? | No | Some expose simple statistical features; others use proprietary neural classifiers |
Does more text help? | Usually | Short passages contain less stable evidence and should be treated cautiously |
Can edited AI text be detected? | Sometimes | Heavy editing can weaken the original machine pattern |
Can human text be flagged? | Yes | Formal, repetitive or highly constrained prose can create false positives |
What an AI detector is actually classifying
A text detector is a supervised classification system. Developers assemble examples labeled as human-written or AI-generated, divide them into training and evaluation sets, and teach a model to separate the two classes. The model does not memorize one universal AI style. It learns combinations of features that helped distinguish the examples in its dataset.
Those examples determine the boundary the detector learns. A system trained mainly on English essays from one set of generators may perform differently on technical documentation, translated prose, creative fiction, newer models or text revised by a human. Strong vendors continually refresh evaluation sets as generators change. Strong users also test the detector on their own document types instead of relying on one vendor accuracy percentage.
The five stages of AI text detection
1. The document is cleaned and divided into units
The detector first normalizes the submission. It may remove formatting, identify prose, split the text into sentences or overlapping windows, and exclude structures the model was not designed to score. References, code, tables, bullet lists and very short fragments may be ignored by one detector and included by another. This preprocessing step changes the denominator behind the final percentage.
2. The system creates a numerical representation
Language must be converted into numbers before a model can classify it. Some systems calculate explicit features such as word predictability, sentence-length variation, repeated constructions or vocabulary distribution. Neural systems may instead create embeddings that capture relationships among tokens, sentences and larger passages. Many commercial products use a mixture of engineered and learned features without publishing every component.
3. A classifier looks for learned combinations
The classifier weighs many signals together. One predictable sentence is not enough, and a varied sentence is not automatically human. The model asks whether the pattern across the sample is closer to the human or AI examples it learned from. Transformer-based classifiers can capture contextual relationships that a simple perplexity threshold would miss.
4. Sentence or segment scores are combined
A document can contain human writing, generated text and later edits. Detectors often score smaller windows and aggregate them into a document result. Overlapping windows give surrounding context, but they can also make boundaries fuzzy. A human sentence beside an AI passage may inherit some of the surrounding signal, while a short document may produce an unstable all-or-nothing result.
5. The probability is converted into a report
The interface may show an AI probability, a predicted share of AI-like prose, a human score or a category such as likely AI. These labels are not interchangeable. Users should read the product definition before interpreting the number. Sentence highlights are useful for locating evidence, but they still represent model predictions rather than a record of the writing process.
Signal or component | What it captures | Why it cannot decide alone |
|---|---|---|
Perplexity | How surprising a sequence is to a language model | Formal human prose can also be predictable |
Burstiness | Variation in sentence length and structure | Writers and genres differ naturally |
Stylometric patterns | Vocabulary, syntax, punctuation and recurring habits | Editing and assignment constraints change style |
Embeddings | Contextual and semantic relationships in the text | The representation depends on training data |
Neural classifier | Combinations of learned features across a passage | A confident output can still be wrong outside its training distribution |
Perplexity and burstiness are useful explanations, not a universal recipe
Perplexity is often used to explain detection because generators choose words from probability distributions. A smooth, highly predictable passage can receive lower perplexity than writing with unusual phrasing. Burstiness describes variation across sentences or passages. Human drafts often shift rhythm, while unedited generated text can appear more uniform.
These concepts make the problem understandable, but they should not be mistaken for the complete implementation of every product. A detector can use a transformer classifier without calculating a user-facing perplexity score. Another can combine dozens of explicit features with embeddings. Turnitin, for example, describes a proprietary transformer-based deep-learning architecture, while Winston AI describes a layered analysis that includes statistical and machine learning signals. The output must be judged by measured performance, not by how simple the marketing explanation sounds.
Watermarking is a different kind of AI detection
The recent move toward text watermarking changes the detection landscape, but it does not make ordinary AI detectors obsolete. A classifier studies patterns that tend to appear in AI writing without cooperation from the model provider. A watermark is deliberately placed during generation by a provider that controls the model. Detecting it requires knowledge of the watermarking method and, for keyed systems, access to the correct verification key.
Text watermarking does not usually mean adding an invisible Unicode character, a hidden account number or a marker that appears when text is pasted. Modern statistical watermarks make tiny changes to token selection while the model writes. When several next words would all be acceptable, the generator uses a secret or defined pattern to influence the choice. Across a long passage, those choices form evidence that a compatible verifier can test.
Provider or system | Current text position in September 2026 | What users should understand |
|---|---|---|
Google Gemini | SynthID watermarks and identifies text generated in the Gemini app and web experience | The signal is created through token-probability adjustments and is not visible to readers |
Anthropic Claude | Newer Claude models use a SynthID-Text-derived keyed watermark, with older-model coverage being rolled forward | Verification estimates Claude involvement and does not identify a user, organization or chat |
OpenAI ChatGPT | Official OpenAI provenance checks currently document supported signals for images and audio, not ordinary text | Do not assume copied ChatGPT text contains a deployed, publicly verifiable text watermark |
Independent AI detector | Classifies writing patterns without a provider watermark key | A detector score and a watermark result are different forms of evidence |
How token-level text watermarks work
A language model repeatedly chooses the next token from a probability distribution. In many positions, several choices preserve the meaning and quality of the response. A watermarking scheme can use a key and the preceding context to divide or rank acceptable token choices, then gently favor choices that encode a statistical pattern. The finished paragraph reads normally, but a verifier with the right method can test whether the sequence is unusually consistent with that pattern.
The watermark needs enough eligible choices to accumulate evidence. A long explanatory passage offers many opportunities. A short answer, a fixed fact, a mathematical result or exact code offers fewer. This is why provider documentation emphasizes confidence rather than certainty and why watermark detection becomes stronger as the sample grows.
Google Gemini and SynthID Text
Google DeepMind says SynthID Text is used to watermark and identify text generated by the Gemini app and web experience. The system changes token probability scores during generation without producing a visible label or intentionally changing output quality. Google has also open-sourced SynthID Text components for developers who want to add compatible watermarking to their own models.
Scope matters. A statement about the Gemini app and web experience should not automatically be extended to every historical Gemini model, API route, third-party host or transformed copy. Verification should match the originating product and date. A negative result means the expected signal was not detected; it does not prove that a person wrote the text.
Claude watermarking
Anthropic announced a keyed text watermark for newer Claude models based on the SynthID-Text approach. It changes the source of randomness used when Claude chooses among similarly suitable words. Anthropic says it adds no hidden characters, carries no user or chat identifier and has no practical effect on price or output quality. Detection asks whether Claude was likely involved in producing part of the text, not who used it or who legally authored it.
Anthropic also documents important limits. Small samples contain too few choices. Factual passages, code and narrow proofreading tasks offer less room for a mark. Light edits may preserve enough of the pattern, while a complete rewrite can remove it. Its dedicated detection API is in private preview for eligible organizations, so a normal third-party AI detector cannot simply claim to read Claude's secret key.
ChatGPT and OpenAI provenance
OpenAI's current official Content Provenance API checks images for C2PA Content Credentials and SynthID signals, and checks supported audio for SynthID. The documented endpoint accepts image and audio files, not plain text. That means the safe current statement is narrow: OpenAI provides provenance verification for supported media, while its official developer documentation does not establish a comparable public watermark for ordinary ChatGPT text.
This also separates three ideas that are often mixed together. C2PA is signed file metadata and can be removed when a file is converted or stripped. An embedded media watermark lives in image or audio content and may survive some transformations. A statistical text watermark is a pattern distributed across token choices. None of these is the same as a general classifier deciding that prose sounds AI-generated.
Why watermarking does not settle authorship
A detected watermark can provide stronger source evidence than a generic style classifier because the provider deliberately created the signal. It still does not show whether using AI was allowed, whether a person substantially revised the work, who operated the account or who owns the output. A missing watermark is even less conclusive because the text may be too short, may come from an unsupported model, may have been edited or may have been generated before deployment.
Schools and publishers should therefore keep watermark checks and AI detector scores in separate fields. Record which provider signal was tested, which verifier and key were used, the text length and whether editing occurred. Then examine the writing process and policy context. Watermarking improves provenance for supported output; it does not turn a final document into a complete history of how the work was created.
Technical background: Google DeepMind's SynthID-Text research paper describes the statistical method, while the C2PA standard covers signed media provenance.
Why detector accuracy changes
Condition | Likely effect | Reason |
|---|---|---|
Very short text | Lower confidence | Too little evidence for stable document-level patterns |
New generator or model update | Possible false negatives | The output may differ from the detector training set |
Heavy human editing | Possible false negatives | Revision can replace machine-associated patterns |
Formal or formulaic human writing | Possible false positives | Predictability and repeated structure can resemble training examples |
Translated or second-language writing | Variable performance | Language distributions may differ from the training population |
Mixed human and AI passages | Uncertain boundaries | Segment aggregation can blur where one source ends |
Accuracy is not one permanent property of a detector. It is a result measured on a particular dataset, at a particular threshold, with particular versions of the detector and generators. A vendor can report excellent performance on balanced, clean examples while a school sees different results on authentic student work. False-positive rate, false-negative rate, precision and recall all matter, especially when the real proportion of AI-assisted work is unknown.
Compare the current product landscape in Best AI Detectors in 2026. For Winston AI specifically, its score interpretation guide explains how to read sentence and document evidence.
A repeatable way to test an AI detector
A fair evaluation starts with known provenance. Collect human writing created before widespread generative AI use or under a documented process, and generate matched AI samples for the same prompts. Add hybrid and edited versions because those are the cases users will actually encounter. Keep the evaluator blind to the source while recording the exact tool version, date, threshold and result.
Test set | Minimum design | What it reveals |
|---|---|---|
Human control | Multiple writers, genres and proficiency levels | False-positive risk and demographic variation |
Unedited AI | Several current models using matched prompts | Baseline recall across generators |
Hybrid document | Known inserted passages at several proportions | Sensitivity to mixed authorship and boundaries |
Human-edited AI | Documented revisions of different intensity | Robustness after realistic editing |
Short and long versions | Matched samples at several lengths | How evidence changes with document size |
Out-of-domain text | Code, lists, technical and creative samples | Where the detector should decline or warn |
Do not average every output into a single accuracy figure and stop. Report false positives separately from missed AI text. Show results by length, language, generator and editing condition. If a product returns a continuous score, choose the decision threshold before looking at final test labels. Retesting after major product updates is essential because both sides of the classification problem change.
How to interpret a detector result responsibly
Result pattern | Reasonable next step | What not to conclude |
|---|---|---|
Low score on a long supported sample | Treat as one weak signal and continue normal review | That no AI assistance occurred |
High score with concentrated highlights | Inspect the passages, assignment and drafts | That misconduct is proven |
Conflicting tools | Check tool scope, versions and document length | That majority vote creates certainty |
Score on a short fragment | Request more context or decline to judge | That the fragment identifies its author |
Human writer disputes the result | Review notes, sources, version history and explanation | That the tool outranks process evidence |
In education, a detector should open an inquiry rather than close one. Compare the highlighted language with the student's drafts, notes and earlier work. Ask the student to explain the argument, sources and revision choices. Apply the same process to every student, document the evidence considered and provide an appeal route. In publishing or hiring, disclose the review policy and avoid treating a proprietary score as a hidden automatic rejection rule.
AI detection and plagiarism detection answer different questions
Plagiarism checking searches for matching language in websites, publications or stored submissions. AI detection classifies patterns inside the submitted text. Generated prose can be original in the matching sense while still violating an assignment rule. Human prose can contain copied passages while looking entirely human to an AI classifier. A complete integrity review keeps the two reports separate.
See the distinction in practice in Best Plagiarism Checkers and the guide to AI tools for teachers. Winston AI also provides both an AI content detector and plagiarism review within one workflow.
The central limitation is missing process evidence
Text classification examines the final artifact. It cannot see brainstorming, drafting, source use, classroom discussion or the sequence of edits that produced it. That is why version history, oral explanation and clear assignment rules are stronger companions than a second detector score. Newer writing-transparency systems attempt to capture process evidence directly, but they raise their own privacy and implementation questions.
Independent studies continue to find that performance varies sharply across tools and conditions. Research has also documented risks for non-native English writers and for text transformed after generation. The practical response is neither blind trust nor blanket dismissal. Use detectors within their documented scope, validate them locally and reserve consequential decisions for a human review with multiple kinds of evidence.
Frequently asked questions
Can AI detectors prove that ChatGPT wrote something?
No. They estimate whether textual patterns resemble known AI-generated examples. They do not recover the prompt, account or writing history that created the document.
What is a good AI detection score?
There is no universal cutoff. The meaning depends on the product, document length, supported language, threshold and purpose. Read the tool definition and examine passage-level evidence.
Why do different AI detectors disagree?
They use different training data, model architectures, thresholds, preprocessing rules and update schedules. A disagreement is a reason to inspect scope and evidence, not to take a majority vote.
Can a humanizer fool an AI detector?
Editing and paraphrasing can reduce detectable patterns, but results vary by detector and method. This is one reason evaluations should include human-edited and transformed samples.
Do AI detectors work on code or bullet points?
Some products are designed only for long-form prose. Code, lists and fragments may be unsupported or excluded, so check the product requirements before interpreting a result.
Should schools use AI detectors?
They can be one part of an academic-integrity process if the school validates the tool, trains reviewers, protects student data and prohibits automatic penalties based on a score alone.
Sources and methodology notes
This guide was checked on September 16, 2026 against current vendor documentation and recent research. Product mechanisms are described only at the level their publishers document. Vendor accuracy claims are not treated as directly comparable because datasets, thresholds and versions differ.
Independent research: Evaluating the accuracy and reliability of AI content detectors in academic contexts.
Bias research: GPT detectors are biased against non-native English writers.
Winston AI: How AI detectors work.
Turnitin sources reviewed without outbound product links: AI writing detection capabilities FAQ, Using the AI Writing Report, File requirements for an AI Writing Report, and AI writing detection model release notes.