Can AI Detectors Detect Claude, ChatGPT and Gemini?

Alex Halpin
7/27/2026

There is a belief that circulates in student forums and Reddit threads with real confidence: Claude does not get caught. Write it in ChatGPT and Turnitin flags you; write the same thing in Claude and you sail through.
There is something behind this. It is smaller, less reliable, and less useful than the confident version suggests.
The short answer
Yes, AI detectors detect Claude. They detect ChatGPT and Gemini too. None of the three is invisible, and no detector can tell you which model produced a piece of text.
What varies is the rate, modestly, and inconsistently between studies.
Why detectors do not care which model you used
Detectors do not maintain a database of model fingerprints. They measure two statistical properties of the text in front of them:
- Perplexity — how predictable each word is given the words before it. Language models select likely next tokens by design, so their output tends to sit in a narrow, low-surprise band.
- Burstiness — how much sentence length and rhythm vary. Human writing lurches between short sentences and long ones. Model output tends to flatten out.
Both properties are consequences of how all transformer language models generate text. Claude, GPT and Gemini differ in training data, alignment and decoding defaults, but they share the fundamental mechanism. That is why none of them escapes detection categorically.
It also means the detector's output is a probability that text is machine-like, not an identification of authorship or model. Any tool advertising "this was written by Claude specifically" is selling you something it cannot deliver.
Where the Claude reputation comes from
The belief is not baseless. Two things feed it.
Claude's default register is more varied. Anecdotally and in several informal comparisons, Claude produces more variation in sentence length and structure than GPT models at default settings. Higher burstiness means a lower detector score. If that holds for your prompt, your text scores better — but it is a tendency of one model's default style, not a guarantee, and it changes with every model release.
Selection bias in the reports. People post when they get away with something. The student whose Claude essay was flagged does not start a Reddit thread about it. The visible evidence is filtered toward success.
There is also a measurement point worth making: some third-party testing reports lower detection accuracy on Claude output than on ChatGPT output. Those tests are almost all published by companies selling detectors or humanizers, with no stated methodology, sample size, or detector version. Treat the direction as weakly informative and the specific numbers as marketing.
What actually moves detection rates
If you rank the factors that change whether text gets flagged, model choice is far from the top.
| Factor | Effect on detection | Notes |
|---|---|---|
| How much it was edited | Very large | Reported accuracy drops from ~94% to ~78% on lightly edited text, and as low as ~40% once personal examples and restructuring are added |
| Length of the sample | Large | Documents under 500 words produce materially more errors in both directions |
| Whether it is mixed human/AI | Large | Mixed authorship confuses segment-level scoring badly |
| Which detector is used | Large | Detectors disagree with each other constantly on the same text |
| Which model wrote it | Small and unstable | Real but inconsistent between studies and detector versions |
The gap between the first row and the last is the practical lesson. Choosing Claude over ChatGPT is a rounding error compared to whether the text was genuinely revised.
The three models, specifically
With the caveat that all of this shifts on every model release, here is what the differences actually amount to.
Claude
Claude's default output tends toward more structural variation than its peers — longer subordinate clauses, more mid-sentence qualification, occasional short declaratives for emphasis. Higher burstiness, lower score.
It also has recognisable verbal habits. A tendency to open with a restatement of the question. A fondness for "It's worth noting that" and "That said." Tricolons — three parallel items where two would do. These are stylistic tells a careful human reader may catch even where a detector does not, and graders increasingly read for them.
ChatGPT
The most-detected model, for two reasons that have little to do with the model itself.
First, it is by far the most used, so detector training corpora contain more of its output than anything else. Detectors are better at what they have seen most.
Second, its default register is highly consistent: uniform paragraph lengths, heavy signposting, the closing summary paragraph nobody asked for. That regularity is precisely the low-burstiness signature detectors key on.
Prompted well, GPT models produce far more varied output than their defaults suggest. The detection gap between models is partly a gap between default outputs, which is not the same thing.
Gemini
Sits between the two in most informal comparisons, with a tendency toward list-heavy, information-dense output. Bulleted content is interesting from a detection standpoint because list items are short and structurally parallel, which reads as low burstiness — but many detectors handle non-prose segments inconsistently, so results on list-heavy text are especially noisy.
What this is worth
Not much, operationally. Every one of these observations is a tendency of a default configuration that changes with the next release, measured against detectors that also change. Building a strategy on "Claude is safer" is building on a number that was never stable and is not verifiable.
What your instructor actually sees
A practical point that gets lost in the model-comparison discussion.
If you are a student, the detector that matters is the one your institution runs, and you do not see its output. Turnitin's AI indicator appears in the instructor's view, not yours. Whatever GPTZero or a free checker tells you is a different model, with a different threshold, on a different day.
This has a few consequences worth being clear-eyed about:
- A clean score on a public detector is not evidence of anything. Different model, different threshold. It tells you your text is not obviously machine-like, and nothing more.
- You cannot iterate against the detector that counts. The optimisation loop people imagine — check, tweak, recheck — does not exist for the one detector with consequences.
- The instructor sees a number without context. They rarely see methodology, error rates, or the research on false positives. That asymmetry is why the evidence you keep matters more than the score you achieve.
The detector disagreement problem
Before optimising for any of this, understand what you are optimising against.
Liang et al., "GPT detectors are biased against non-native English writers" in Patterns, tested seven detectors on 91 TOEFL essays written by humans. The average false positive rate was 61.22%. Rewriting those same human essays with more elaborate vocabulary dropped it to 11.6%.
The instruments are noisy enough that they flag most non-native human writing. A tool with that error profile cannot support fine-grained conclusions about which model is 8% harder to catch.
If you have been flagged for work you wrote, why AI detectors flag human essays covers what to do about it.
What about Turnitin specifically?
Turnitin is the one that matters for students, because it is the only detector your institution actually sees.
It does not publish per-model detection rates, so anyone quoting you a "Turnitin catches Claude X% of the time" figure is inventing it. What Turnitin does claim is a false positive rate under 1% on documents with more than 20% AI content — a number that independent research has repeatedly failed to reproduce.
Its scoring works at the sentence level and aggregates upward, which is why mixed human-and-AI documents produce such erratic results. How Turnitin AI detection works breaks down the mechanism.
The "Claude humanizer" confusion
Worth clearing up, because search results for this term have shifted underneath it.
If you search "Claude humanizer" today, most of what you get is about a Claude Skill — a downloadable extension that makes Claude itself write less like a model, installed into Claude Code or the Claude app. That is a different thing from a web tool that takes finished text and rewrites it.
Both approaches exist and they solve the problem at different points:
- A Claude Skill shapes the output as it is generated. Useful if you are drafting inside Claude anyway.
- A humanizer tool processes text after the fact, whatever produced it. Useful if your text came from somewhere else, or if you are working from a mixed draft.
Neither is a detector-defeating device. Both are ways of getting less formulaic prose.
The mixed-authorship problem
The scenario almost nobody plans for, and the one that produces the worst outcomes.
Most real AI-assisted work is not wholly generated. It is a draft you wrote, with a paragraph the model tightened. Or a model outline you filled in yourself. Or your own prose with a model-written literature summary in the middle.
Detectors handle this badly, because they score segments and aggregate. A document that is 80% yours and 20% model-written does not reliably return "20% AI." It can return almost anything:
- The model-written section scores high and drags the document average up
- Your own careful writing in the same document also scores high, because it is well-structured, and the total climbs further
- Or the aggregate washes out entirely and the document reads clean despite containing generated text
None of these is a correct answer, and you cannot predict which you will get. This is also why "what percentage of AI is acceptable" is not a well-formed question — the percentage is not measuring what the policy assumes it measures.
If you are working this way, the version history argument becomes more important, not less. A document that visibly grew across sessions, with your revisions visible, is the only artefact that survives this ambiguity.
The honest guidance
If you are using AI models in work that will be checked:
- Do not pick a model based on detectability. The difference is small, unstable, and will change with the next release. Pick the model that produces the best draft.
- Edit substantively. This is the factor with by far the largest effect, and it is the only one that also improves the work.
- Add what a model cannot. Specific sources you actually read, numbers from your own data, examples from your own experience. This is both the strongest human signal and the thing that makes writing good.
- Keep your version history. If you are wrongly accused, a revision timeline is worth more than any score. Detector output is contestable; a document that visibly grew over three days is not.
- Check your institution's policy. Detection is a technical question. Whether AI assistance is permitted is a different one, and getting the second wrong is the more serious problem.
Summary
Claude is detectable. So are ChatGPT and Gemini. The differences between them are real but small, unstable across detector updates, and swamped by the effect of actual editing.
Anyone telling you a specific model is safe is selling something or repeating a forum rumour. The durable answer has not changed: do the work, revise it properly, and keep the drafts that prove you did.
If you are working from an AI-assisted draft and want it to carry your own voice rather than a model's default cadence, RewriteAI's humanizer restores sentence-level variation without the synonym-swapping that mangles technical writing. There are also model-specific guides for Claude, ChatGPT and Gemini.
Keep reading
- Does Canvas Have an AI Detector in 2026?
Canvas doesn't have a built-in AI detector — but Turnitin, Copyleaks, and GPTZero integrate directly into Canvas LMS. Here's what actually detects AI in Canvas submissions.
- iThenticate Free: What You Actually Get
iThenticate is not free, and the workarounds circulating on Reddit mostly are not either. Here is what access actually exists, what your institution probably already pays for, and how to check a manuscript properly before you submit it.
- AI Detector Comparison: Turnitin vs GPTZero vs Grammarly (2026)
We tested Turnitin, GPTZero, and Grammarly's AI detectors head-to-head to find which is most accurate, which generates false positives, and how to bypass each one.
Humanize AI Text and Improve Your Writing Right Now
Rewrite for clarity, flow, and readability while keeping your original meaning and writing style.


