Key takeaways
- A 2023 Stanford study found seven GPT detectors misclassified more than half of TOEFL essays by non-native English writers as AI-generated, while classifying native-speaker essays almost perfectly.
- The bias is structural: detectors score linguistic predictability, and second-language writing is typically more regular in vocabulary and sentence construction.
- In the same study, rewriting the essays with richer vocabulary sharply reduced false flags — evidence that the detectors were reading proficiency, not authorship.
- This lands hardest on international students, who often face the highest stakes and the least institutional standing to contest an accusation.
- Improving your English is a legitimate response. Concealing authorship is not — and the distinction matters both ethically and practically.
What the research found
The central study is Liang, Yuksekgonul, Mao, Wu and Zou, 'GPT detectors are biased against non-native English writers', published in the Cell Press journal Patterns in July 2023.
The design was simple. The researchers took TOEFL essays written by non-native English speakers and essays written by US eighth-graders — native speakers, and younger writers at that — and ran both sets through seven widely used GPT detectors.
| Essay set | Author | Outcome |
|---|---|---|
| US eighth-grade essays | Native English speakers | Classified as human-written almost perfectly |
| TOEFL essays | Non-native English speakers | More than half misclassified as AI-generated |
| TOEFL essays, vocabulary enriched | Same non-native writers | Misclassification dropped sharply |
That third row is the one that settles the question. The authorship never changed. The only variable was how linguistically varied the text was. When the researchers prompted a model to rewrite the same essays with more sophisticated language, the detectors largely stopped flagging them.
A detector whose output changes that much based on vocabulary richness, while authorship stays constant, is not measuring authorship.
Why this study carries weight
It is peer-reviewed, it tested multiple commercial detectors rather than one, and it included a manipulation that isolated the mechanism. The comparison group being younger native speakers makes the result harder to explain away as a writing-quality effect.
Why second-language writing trips the signal
Detectors are built around two measurable properties, and second-language writing tends to score low on both for reasons that have nothing to do with machines.
A narrower working vocabulary
Writing in a second language usually means drawing on a smaller set of reliably known words. You choose the word you are confident is correct rather than the unusual one that might be slightly off.
That is good judgement. It also means the next word in your sentence is more often the statistically likely one — which is the definition of low perplexity, the primary thing detectors measure.
More consistent sentence construction
Second-language writers frequently rely on sentence patterns they have been explicitly taught and know to be correct. The result is even, well-formed, structurally consistent prose.
Detectors read that evenness as low burstiness. Native speakers, writing intuitively, tend to produce more erratic rhythm — fragments, asides, sentences that run long because the thought did.
The proofreading trap
There is a second-order problem that specifically affects ESL writers. If you are less confident in your English, you are more likely to run your work through a grammar checker or translation tool before submitting it.
Those tools move text toward conventional, expected phrasing — lowering perplexity further. So the writer most likely to be flagged is also the writer most motivated to use the tools that increase the chance of being flagged.
Worth being clear about
Using a grammar checker on your own writing is not academic misconduct at any institution we are aware of. But it can move your text toward the statistical profile a detector treats as machine-like, and it is better to know that in advance.
Who carries the cost
The consequences fall unevenly, and they fall hardest on people least able to absorb them.
- International students. Often paying substantially higher fees, sometimes on visas contingent on good academic standing, and frequently least familiar with the institution's appeals process. An accusation carries more risk and is harder to contest.
- Researchers publishing in English. Most academic publishing happens in English regardless of the author's first language. Detector screening at journals introduces a bias that has nothing to do with the quality of the research.
- Professionals writing for global employers. Reports, proposals and internal documents increasingly pass through AI-detection tooling. Being flagged in a workplace has no appeals process at all.
- Job applicants. Some hiring platforms screen cover letters for AI use. A non-native applicant writing carefully in English can be filtered out before a human reads a word.
Institutions have started responding
Some universities have stepped back from automated detection over reliability concerns — Vanderbilt disabled Turnitin's AI detector in August 2023 and published its reasoning. Turnitin's own guidance frames its score as the start of a conversation rather than a determination of misconduct. If you are contesting a flag, your institution's own policy language is often the most useful document you have.
What you can actually do about it
There is a legitimate response and an illegitimate one, and the line is worth stating plainly before the practical advice.
Improving the range and rhythm of English that you wrote is ordinary editing — the same thing a copy editor or a writing centre would help you do. Passing off work you did not do is misconduct regardless of what tooling is involved. Everything below assumes the first case.
Practical steps:
- Protect the record by default. Write in a tool that keeps revision history — Google Docs, Word, Notion. Do not draft elsewhere and paste in a finished block. A visible history of incremental work is the strongest evidence of authorship you can have, and it costs nothing to create.
- Deliberately vary sentence length. Follow two long sentences with a short one. This raises burstiness, which is the property detectors measure — and it genuinely improves readability for any reader.
- Add specifics only you would know. A detail from a source you actually read, an example from your own context, a particular number. Generic writing is what reads as machine-generated; specificity is both better writing and a stronger signal of a real author.
- Keep some of your own phrasing. Resist the urge to smooth every sentence into the most conventional possible form. Slight unconventionality is a human signal. Correct is not the same as maximally expected.
- Check before you submit, not after. Knowing how your text scores lets you decide what to do while you still have options. After an accusation, your leverage is much lower.
- If accused, argue process rather than score. Draft history, notes, sources and a willingness to discuss the content are far more persuasive than a competing detector reading. Our guide on false positives covers the full sequence.
- What to do if you are wrongly accused — the evidence checklist and escalation sequence
- AI humanizer for ESL writers — the product page, if that is what you came for
Where a tool like ours fits — and where it does not
We build an English rewriting tool, so treat this section as interested rather than neutral. It is worth being specific about what it does.
Our humanizer takes English text you paste and rewrites it with more varied vocabulary and sentence rhythm while keeping your meaning. It returns several options, and you choose the one that still sounds like you. There are no tone or strength settings — you paste text and get alternatives back.
For ESL writers, the honest framing is that this is an editing aid for English you already wrote. It is the same category of help as a writing centre appointment or a native-speaker friend reading your draft. It cannot verify authorship, and no tool should claim otherwise.
Two real limits
The model is optimized for English; results in other languages may be inaccurate, so this is not a translation workflow. And detector scores shift as detectors are retrained — anyone promising a permanent guarantee against a specific detector is overselling.
Paste English you wrote to see the alternatives. Free plan covers 500 words per month, 300 per request.
Sources
- Liang et al., 'GPT detectors are biased against non-native English writers' — Patterns (Cell Press), 2023
- Guidance on AI detection and why we're disabling Turnitin's AI detector — Vanderbilt University, 2023
- Understanding false positives in AI writing detection — Turnitin