Are AI Detectors Accurate? Every Published Number, in One Table

Alex Halpin
9/15/2026

In June 2026, three researchers at the Vrije Universiteit Brussel took 40 research papers written end to end by GPT-4o and ran them through four commercial AI detectors. Turnitin flagged none of them. GPTZero flagged none. Copyleaks flagged none. Only Pangram called them what they were.
All four of these vendors have published the accuracy of their products on their own websites. Three of the four report an accuracy of 99% or better.
Both things are true and the gap between them is the whole story. The question of whether AI detectors are accurate doesn't have a single answer. Accuracy is not one number. It depends on who wrote the text, how long it is, whether it was edited and what the test set is made of. A detector that is 99% accurate on one pile of documents can score zero on another.
Here is every published number I could find for AI detectors with the date of publication and who published it. These numbers are the vendor's claims first, followed by what happened when others tested them.
What the vendors claim about themselves
These are the figures on each company's own pages, read on 15 September 2026.
| Tool | Claimed accuracy | Claimed false positive rate | Source and date |
|---|---|---|---|
| Copyleaks | "over 99%"; English 99.97% on human text, 99.20% on AI | 0.03% | Product page, read Sep 2026 |
| Pangram | 99.9%+ overall; 99.7% across a 26-model benchmark | 0.01% (1 in 10,000) | Product page, read Sep 2026 |
| Originality.ai Turbo 3.0.2 | 99%+ | 1.5% | Model notes, Sep 2025 |
| Originality.ai Lite 1.0.2 | 99% | 0.5% | Model notes, Sep 2025 |
| Originality.ai Academic 0.0.5 | 99%+ | under 1% | Model notes, Sep 2025 |
| GPTZero | 95.7% of AI texts on the RAID benchmark; "over 99%" once discontinued models like GPT-3.5 are filtered out | 1% | Accuracy page, read Sep 2026 |
| Turnitin | no overall accuracy figure published | under 1% at document level, for documents scored 20% AI or higher; about 4% at sentence level | Company blog, 14 June 2023 |
Notice what varies. Copyleaks and Pangram quote a false positive rate two orders of magnitude apart from Originality.ai's and all three are describing the same task. Turnitin declines to publish a headline accuracy number at all. That looks like the most defensible choice on the list.
None of these figures is a lie. Each one describes a real measurement on a real test set. The test set is the variable nobody markets.
What happened when other people ran the tests
The most cited independent evaluation of these tools is by Weber-Wulff et al., published in the International Journal for Educational Integrity on 25 December 2023. The authors tested fourteen tools with 54 test cases across six categories of text.
No tool scored above 80% accuracy. Only five tools scored above 70% accuracy. Here are the results broken down by type of text:
| Text type | Average accuracy across 14 tools |
|---|---|
| Human-written | 96% |
| AI-generated, untouched | 74% |
| Machine-translated | 76% |
| AI text edited by hand | 42% |
| AI text run through a paraphraser | 26% |
The false positive rates for the tools range from 0% for Turnitin to 50% for GPTZero. The false negative rates range from 8% for GPTZero to 100% for Content at Scale. The tools at the extremes of both metrics are the same two tools: one says "AI" too often, the other says "human" too often and neither is accurate.
The authors' own summary of the results is that the tools are "neither accurate nor reliable" and have "a main bias towards classifying the output as human-written rather than detecting AI-generated text".
OpenAI reached a similar conclusion about its own product first. OpenAI's AI Text Classifier launched on 31 January 2023 with published figures of a 26% true positive rate and a 9% false positive rate. However, OpenAI withdrew the classifier on 20 July 2023 due to its "low rate of accuracy". So OpenAI, with direct access to how its own models write, could not build an accurate classifier for them and stated as much within six months.
The June 2026 study is the one worth reading
Van Vlasselaer, Van Droogenbroeck and Spruyt published in the same journal on 29 June 2026. The authors tested GPTZero, Pangram, Copyleaks and Turnitin on 160 synthetic documents (40 in each of four categories) as well as 1,163 real master's theses.
There are two different scoring schemes that are relevant to these questions. One scoring scheme is for "strict" accuracy, which only counts true positives or true negatives as correct calls. The other is for "inclusive" accuracy, which also gives points for partially correct calls.
| Category | Pangram | Turnitin | Copyleaks | GPTZero |
|---|---|---|---|---|
| Fully human | 100% | 100% | 100% | ~97.5% |
| Fully AI (strict) | 65% | 0% | 0% | 0% |
| Fully AI (inclusive) | 97.5% | 0% | 25% | 30% |
| Hybrid human + AI | 92.5% | 60% | 30% | 0% |
| Rewritten to sound human | 92.5% | 50% | 22.5% | 2.5% |
The clean human documents are where every tool looks good and that is the row vendors are quoting. Move one column right and three of the four tools stop working.
The "rewritten to sound human" row is worth reading carefully, as it was not produced by a humanizer product. Instead, the researchers provided GPT-4o with a prompt that asked it to increase perplexity and burstiness in the rewritten text without changing the meaning of the text or its references. This prompt resulted in Copyleaks's accuracy dropping from 25% to 22.5% and GPTZero's accuracy dropping from 30% to 2.5%.
On the 1,163 real theses, Pangram detected AI writing in 45.5% of the theses, with an average AI share of 34%. So the authors conclude that "detection tools should not be used as sole evidence in high-stakes decision-making".
The false positives are not spread evenly
A 1% false positive rate sounds survivable until you ask which 1%.
Liang and colleagues at Stanford tested seven detectors on 91 TOEFL essays written by non-native English speakers. The average false positive rate for the seven detectors was 61.22%. Also, 89 of the 91 essays were flagged by at least one of the detectors. However, 18 of the 91 essays (19.78%) were flagged as AI-generated by all seven detectors. By contrast, the same detectors tested on 88 essays from US eighth graders had an average false positive rate of 5.19%.
This is not a mysterious process. The detectors calculate the level of predictability of the writing sample's wording and writers with smaller working vocabularies produce more predictable wording. The detectors are reading constrained English and reporting it as machine English. We wrote about this phenomenon in why an AI detector says your essay is AI when it isn't and it is the reason for the existence of pages dedicated to false positives and detector bias against ESL writers.
Turnitin's sentence-level number deserves the same treatment. The company's own post puts it plainly:
There is a 4% likelihood that a specific sentence highlighted as AI-written might be human-written.
Run that across a 40-sentence essay and you'll likely find a highlighted sentence or two within your own work. Turnitin adds that 54% of its false positive sentences sit next to genuine AI text within an essay, which creates clusters of highlighted sentences, making an otherwise clean essay look worse than it truly is.
What the percentage on your screen actually means
Since July 2024, Turnitin has not surfaced AI scores below 20%. Scores between 1% and 19% are represented by an asterisk, with no score and no highlighted text, and the model requires at least 300 words to produce any score at all.
This answers the question that people constantly ask. Is 20% AI detection bad? 20% is not a rating of the severity of the AI detection. Instead, it is the lowest number that Turnitin is willing to show users for their documents. Anything below 20% was determined to be too unreliable for Turnitin to display to users.
The "less than 1%" document-level false positive rate only applies to documents that score 20% AI or above. There is no published figure for the range of scores that the product now hides.
Why the two sets of numbers disagree
Three reasons, and none of them requires anyone to be dishonest.
Test sets are chosen. A detector trained and measured on unedited ChatGPT output scores brilliantly on unedited ChatGPT output. The percentage of correctly detected instances drops to 26% in Weber-Wulff's paraphrased category, and to zero for three out of four tools in the 2026 study's hybrid and rewritten categories.
Adversarial text breaks them. The RAID benchmark, published by Dugan and colleagues in May 2024, is six million generations across 11 models, 8 domains and 11 adversarial attacks. The authors found that current systems are "easily fooled by adversarial attacks, variations in sampling strategies, repetition penalties and unseen generative models". GPTZero's 95.7% comes from that same benchmark, which tells you a headline figure can sit on top of a paper whose conclusion is that the category is fragile.
Base rates do the rest. A 0.03% false positive rate is excellent but still means that students will be wrongly flagged. If you multiply this percentage by the number of students who submit their papers each year, the number of wrongly flagged students is no longer a rounding error.
If you want to see how the scoring system works, we broke it down in how AI detectors work. You can also run a passage through our AI detector and compare the results to those of your institution's AI detector. This is a faster way to understand the instability of these scores than reading any of the accuracy pages that have been published.
The short version
The detectors are accurate on clean, unedited AI text and on clean, native-speaker human text. But they fail on all other texts. Vendors claim 99% detection rates but these are only based on their own benchmarks. No independent tests have ever been able to replicate these results. In a 2026 test of unedited AI papers, the best tool achieved only 65% strict accuracy, while three tools did not detect any AI-written papers. Non-native English speakers are also more likely to be incorrectly identified as AI writers by these tools, with one study finding over 60% of these writers were misidentified. No published statistic should be used to prove guilt in an academic integrity case, according to the conclusions of two independent studies into AI detection tools.
If you want the longer treatment
- How AI detectors work: perplexity, classifiers and watermarks — the mechanism behind every number above.
- What AI detector do colleges actually use? — which institutions have switched detection off, and why.
- Why an AI detector says your essay is AI when it isn't — what to do if it happens to you.
- AI detector comparison: Turnitin vs GPTZero vs Grammarly — the three tools students meet most often.
- Is GPTZero accurate? What its 2026 numbers actually mean — one tool, four independent measurements, three orders of magnitude apart.
- How Turnitin AI detection works — the scoring and the thresholds in detail.
Keep reading
- AI Detector Comparison: Turnitin vs GPTZero vs Grammarly (2026)
We tested Turnitin, GPTZero, and Grammarly's AI detectors head-to-head to find which is most accurate, which generates false positives, and how to bypass each one.
- How AI Detectors Work: Perplexity, Classifiers and Watermarks
GPTZero stopped using perplexity and burstiness in autumn 2023, which is awkward, because that is still how almost every explainer says AI detectors work. Here are the three generations, and the one that actually scores your writing.
- What AI Detector Do Colleges Actually Use? (2026)
Turnitin ran its AI writing indicator across 33 million US college submissions between October 2025 and April 2026. A growing list of universities, including Vanderbilt, Pittsburgh, Johns Hopkins, Illinois and Curtin, has switched the feature off.
Humanize AI Text and Improve Your Writing Right Now
Rewrite for clarity, flow, and readability while keeping your original meaning and writing style.

