Skip to main content

Does Turnitin Detect AI From ChatGPT, Claude and Gemini?

Alex Halpin

Alex Halpin

9/27/2026

#turnitin#ai-detector#chatgpt#academic#students
Rows of wooden bookshelves filled with books in a university library

The website for Turnitin describes the model behind its AI writing detector. The description names four systems: GPT-3.5, GPT-4, GPT-4o and GPT-4o-mini. Each of these models are from OpenAI. Claude is not mentioned. Gemini is not mentioned. Neither is Llama, Grok or DeepSeek.

That is the honest starting point for anyone asking whether Turnitin catches a particular chatbot. The vendor documents what the model was built against and what it documents is one company's products.

It does not follow that text from Claude or Gemini will sail through. It follows that nobody, Turnitin included, has published anything on the results of Turnitin encountering them.

What Turnitin documents about its own model

Turnitin's AI writing detection model page was read on 25 September 2026.

Language modelWhat Turnitin says it detectsDated
Englishtext "likely generated with the GPT-3.5 and GPT-4 large language models"July 2024
Spanishtext "likely generated with the GPT-3.5 and GPT-4 large language models"September 2024
JapaneseGPT-4 (version 0613), GPT-4o (May and August 2024 releases) and GPT-4o-miniApril 2025

The same documentation sets two limits on how a score should be read. Submissions under 300 words may result in an AI writing score that is likely less accurate (the limit was previously 150 words until May 2023). Also, scores below 20% are not represented by a percentage within the report; instead, an asterisk is displayed (this has been the case since July 2024).

Turnitin's product pages claim more than the model documentation does. They say the tool flags AI-paraphrased or bypassed content produced by what the company calls AI bypassers, and they name ChatGPT. I could not find a page where Turnitin claims coverage of Claude or Gemini by name.

A detector is not a list of models

The classifier does not identify which system wrote a given passage. Instead, it measures, sentence by sentence, how predictable the writing of that passage is compared to the writing of other passages that were analyzed during the development of the classifier. More information about the process can be found in our article on how Turnitin's AI detection works.

That design has a real consequence in both directions. A classifier trained mostly on OpenAI output will still flag a Claude essay if it writes in the same flat register. Conversely, it will also miss an OpenAI essay that writes in short sentences. The model that produces the text matters less than how the text reads.

Which is why a question phrased as "does Turnitin detect Gemini" has no clean answer and why every blog post that gives you one with a percentage attached is making it up.

What happened when researchers pointed it at a current model

The most careful published test that I have read was performed by Marijke Van Vlasselaer, Filip Van Droogenbroeck and Bram Spruyt and was published in the International Journal for Educational Integrity on 29 June 2026. They ran four AI detectors, including Turnitin, against 160 synthetic documents in four categories of 40 and against 1,163 real master's theses.

Text categoryTurnitin's accuracy
Fully AI-generated0% strict — "Turnitin classified 100% of the Fully AI generated papers as False Negatives"
Hybrid human and AI60.0%
Re-prompted to sound human50.0%
Fully human100% correct, no human paper wrongly flagged

The fully AI-generated papers in that study were created using GPT-4o Deep Research. So this is not a result about Claude or Gemini. It is a result about the newest OpenAI models, the ones Turnitin's documentation is built around, and under strict scoring the detector caught none of them. The researchers concluded that Turnitin "significantly underestimated GenAI content, particularly for texts generated with the most advanced model".

Set that against the study everyone quotes in Turnitin's favour, published in the same journal on 25 December 2023 by Weber-Wulff and colleagues. They tested 14 detection tools against 54 different cases and determined that Turnitin had a false positive rate of zero, making it the most accurate of all the tools they tested. Both results are real. Two and a half years separate them, and in that time the models doing the writing moved on.

The numbers Turnitin publishes about itself

Turnitin's own figures, published on a blog dated 14 June 2023, show that the company has a document-level false positive rate of "less than 1%" for documents that score at least 20% in AI writing and a sentence-level false positive rate of "around 4%", showing that there is a 4% chance that the sentence in question was written by a human author.

Turnitin is more careful about this than its resellers are. For example, the same blog post states that the company "cannot mitigate the risk of false positives completely" and that the highlighted areas of interest are meant to initiate a conversation, not to draw a conclusion. According to the Turnitin solutions page, false positive rates are 0.014 for English language learners and 0.013 for native English writers on submissions of at least 300 words.

Four per cent per sentence is not a small number when a report highlights forty sentences. We collected the rest of the published accuracy figures for this category in our table of every number the detectors have put in print.

So what about Claude and Gemini

Here is the whole of what can be said with evidence. Turnitin's documentation includes only OpenAI models. I could find no peer-reviewed study that has tested Turnitin against Claude or Gemini output and published per-model results. The one thorough 2026 test used GPT-4 and GPT-4o and found that Turnitin failed to detect fully AI-generated text.

As a student, the useful conclusion is that the score is not about the chatbot you opened but about your writing, and that it can be wrong in both directions for any given paper. For more on these topics, we wrote about whether detectors can tell the models apart in can AI detectors detect Claude, ChatGPT and Gemini, and about the limits of what a Turnitin score proves in bypassing Turnitin.

For instructors, the Van Vlasselaer figures are the ones to sit with. A tool that never flags a human paper and also never flags a fully AI one is not giving you what the dashboard implies it is.

The short version

Turnitin's model documentation names GPT-3.5, GPT-4, GPT-4o and GPT-4o-mini and does not mention any Claude or Gemini model. That is a statement about training, not about immunity, because the classifier scores how predictable your writing is rather than identifying a vendor. A study conducted in 2026 with 160 documents determined that Turnitin did not catch any of the fully AI-generated papers under strict scoring, but caught 60% of the hybrid papers and 50% of the re-prompted ones, while flagging no human paper wrongly. Turnitin's published false positive rates for AI writing are less than 1% per document above the 20% threshold and around 4% per highlighted sentence. Anyone quoting you a per-model detection rate for Claude or Gemini is quoting a number that has not been measured.

If you want the longer treatment:

Humanize AI Text and Improve Your Writing Right Now

Rewrite for clarity, flow, and readability while keeping your original meaning and writing style.