Can AI Detectors Detect Claude, ChatGPT and Gemini?

Alex Halpin

Alex Halpin

7/27/2026

#AI Detection#Claude#ChatGPT#Gemini#turnitin
Overhead view of a person working at a desk with a laptop, folders and documents

There is a belief that circulates in student forums and Reddit threads with real confidence: Claude does not get caught. Write it in ChatGPT and Turnitin flags you; write the same thing in Claude and you sail through.

There is something behind this. It is smaller, less reliable, and less useful than the confident version suggests.

The short answer

Yes, AI detectors detect Claude. They detect ChatGPT and Gemini too. None of the three is invisible, and no detector can tell you which model produced a piece of text.

What varies is the rate, modestly, and inconsistently between studies.

Why detectors do not care which model you used

Detectors do not maintain a database of model fingerprints. They measure two statistical properties of the text in front of them:

  • Perplexity — how predictable each word is given the words before it. Language models select likely next tokens by design, so their output tends to sit in a narrow, low-surprise band.
  • Burstiness — how much sentence length and rhythm vary. Human writing lurches between short sentences and long ones. Model output tends to flatten out.

Both properties are consequences of how all transformer language models generate text. Claude, GPT and Gemini differ in training data, alignment and decoding defaults, but they share the fundamental mechanism. That is why none of them escapes detection categorically.

It also means the detector's output is a probability that text is machine-like, not an identification of authorship or model. Any tool advertising "this was written by Claude specifically" is selling you something it cannot deliver.

Where the Claude reputation comes from

The belief is not baseless. Two things feed it.

Claude's default register is more varied. Anecdotally and in several informal comparisons, Claude produces more variation in sentence length and structure than GPT models at default settings. Higher burstiness means a lower detector score. If that holds for your prompt, your text scores better — but it is a tendency of one model's default style, not a guarantee, and it changes with every model release.

Selection bias in the reports. People post when they get away with something. The student whose Claude essay was flagged does not start a Reddit thread about it. The visible evidence is filtered toward success.

There is also a measurement point worth making: some third-party testing reports lower detection accuracy on Claude output than on ChatGPT output. Those tests are almost all published by companies selling detectors or humanizers, with no stated methodology, sample size, or detector version. Treat the direction as weakly informative and the specific numbers as marketing.

What actually moves detection rates

If you rank the factors that change whether text gets flagged, model choice is far from the top.

FactorEffect on detectionNotes
How much it was editedVery largeReported accuracy drops from ~94% to ~78% on lightly edited text, and as low as ~40% once personal examples and restructuring are added
Length of the sampleLargeDocuments under 500 words produce materially more errors in both directions
Whether it is mixed human/AILargeMixed authorship confuses segment-level scoring badly
Which detector is usedLargeDetectors disagree with each other constantly on the same text
Which model wrote itSmall and unstableReal but inconsistent between studies and detector versions

The gap between the first row and the last is the practical lesson. Choosing Claude over ChatGPT is a rounding error compared to whether the text was genuinely revised.

The three models, specifically

With the caveat that all of this shifts on every model release, here is what the differences actually amount to.

Claude

Claude's default output tends toward more structural variation than its peers — longer subordinate clauses, more mid-sentence qualification, occasional short declaratives for emphasis. Higher burstiness, lower score.

It also has recognisable verbal habits. A tendency to open with a restatement of the question. A fondness for "It's worth noting that" and "That said." Tricolons — three parallel items where two would do. These are stylistic tells a careful human reader may catch even where a detector does not, and graders increasingly read for them.

ChatGPT

The most-detected model, for two reasons that have little to do with the model itself.

First, it is by far the most used, so detector training corpora contain more of its output than anything else. Detectors are better at what they have seen most.

Second, its default register is highly consistent: uniform paragraph lengths, heavy signposting, the closing summary paragraph nobody asked for. That regularity is precisely the low-burstiness signature detectors key on.

Prompted well, GPT models produce far more varied output than their defaults suggest. The detection gap between models is partly a gap between default outputs, which is not the same thing.

Gemini

Sits between the two in most informal comparisons, with a tendency toward list-heavy, information-dense output. Bulleted content is interesting from a detection standpoint because list items are short and structurally parallel, which reads as low burstiness — but many detectors handle non-prose segments inconsistently, so results on list-heavy text are especially noisy.

What this is worth

Not much, operationally. Every one of these observations is a tendency of a default configuration that changes with the next release, measured against detectors that also change. Building a strategy on "Claude is safer" is building on a number that was never stable and is not verifiable.

What your instructor actually sees

A practical point that gets lost in the model-comparison discussion.

If you are a student, the detector that matters is the one your institution runs, and you do not see its output. Turnitin's AI indicator appears in the instructor's view, not yours. Whatever GPTZero or a free checker tells you is a different model, with a different threshold, on a different day.

This has a few consequences worth being clear-eyed about:

  • A clean score on a public detector is not evidence of anything. Different model, different threshold. It tells you your text is not obviously machine-like, and nothing more.
  • You cannot iterate against the detector that counts. The optimisation loop people imagine — check, tweak, recheck — does not exist for the one detector with consequences.
  • The instructor sees a number without context. They rarely see methodology, error rates, or the research on false positives. That asymmetry is why the evidence you keep matters more than the score you achieve.

The detector disagreement problem

Before optimising for any of this, understand what you are optimising against.

Liang et al., "GPT detectors are biased against non-native English writers" in Patterns, tested seven detectors on 91 TOEFL essays written by humans. The average false positive rate was 61.22%. Rewriting those same human essays with more elaborate vocabulary dropped it to 11.6%.

The instruments are noisy enough that they flag most non-native human writing. A tool with that error profile cannot support fine-grained conclusions about which model is 8% harder to catch.

If you have been flagged for work you wrote, why AI detectors flag human essays covers what to do about it.

What about Turnitin specifically?

Turnitin is the one that matters for students, because it is the only detector your institution actually sees.

It does not publish per-model detection rates, so anyone quoting you a "Turnitin catches Claude X% of the time" figure is inventing it. What Turnitin does claim is a false positive rate under 1% on documents with more than 20% AI content — a number that independent research has repeatedly failed to reproduce.

Its scoring works at the sentence level and aggregates upward, which is why mixed human-and-AI documents produce such erratic results. How Turnitin AI detection works breaks down the mechanism.

The "Claude humanizer" confusion

Worth clearing up, because search results for this term have shifted underneath it.

If you search "Claude humanizer" today, most of what you get is about a Claude Skill — a downloadable extension that makes Claude itself write less like a model, installed into Claude Code or the Claude app. That is a different thing from a web tool that takes finished text and rewrites it.

Both approaches exist and they solve the problem at different points:

  • A Claude Skill shapes the output as it is generated. Useful if you are drafting inside Claude anyway.
  • A humanizer tool processes text after the fact, whatever produced it. Useful if your text came from somewhere else, or if you are working from a mixed draft.

Neither is a detector-defeating device. Both are ways of getting less formulaic prose.

The mixed-authorship problem

The scenario almost nobody plans for, and the one that produces the worst outcomes.

Most real AI-assisted work is not wholly generated. It is a draft you wrote, with a paragraph the model tightened. Or a model outline you filled in yourself. Or your own prose with a model-written literature summary in the middle.

Detectors handle this badly, because they score segments and aggregate. A document that is 80% yours and 20% model-written does not reliably return "20% AI." It can return almost anything:

  • The model-written section scores high and drags the document average up
  • Your own careful writing in the same document also scores high, because it is well-structured, and the total climbs further
  • Or the aggregate washes out entirely and the document reads clean despite containing generated text

None of these is a correct answer, and you cannot predict which you will get. This is also why "what percentage of AI is acceptable" is not a well-formed question — the percentage is not measuring what the policy assumes it measures.

If you are working this way, the version history argument becomes more important, not less. A document that visibly grew across sessions, with your revisions visible, is the only artefact that survives this ambiguity.

The honest guidance

If you are using AI models in work that will be checked:

  1. Do not pick a model based on detectability. The difference is small, unstable, and will change with the next release. Pick the model that produces the best draft.
  2. Edit substantively. This is the factor with by far the largest effect, and it is the only one that also improves the work.
  3. Add what a model cannot. Specific sources you actually read, numbers from your own data, examples from your own experience. This is both the strongest human signal and the thing that makes writing good.
  4. Keep your version history. If you are wrongly accused, a revision timeline is worth more than any score. Detector output is contestable; a document that visibly grew over three days is not.
  5. Check your institution's policy. Detection is a technical question. Whether AI assistance is permitted is a different one, and getting the second wrong is the more serious problem.

Summary

Claude is detectable. So are ChatGPT and Gemini. The differences between them are real but small, unstable across detector updates, and swamped by the effect of actual editing.

Anyone telling you a specific model is safe is selling something or repeating a forum rumour. The durable answer has not changed: do the work, revise it properly, and keep the drafts that prove you did.

If you are working from an AI-assisted draft and want it to carry your own voice rather than a model's default cadence, RewriteAI's humanizer restores sentence-level variation without the synonym-swapping that mangles technical writing. There are also model-specific guides for Claude, ChatGPT and Gemini.

Humanize AI Text and Improve Your Writing Right Now

Rewrite for clarity, flow, and readability while keeping your original meaning and writing style.