Do AI Detectors Actually Work? I Tested Three

Do AI Detectors Actually Work? — The AI Cheat Sheet

Short answer: not reliably enough to accuse anyone of anything. I ran two passages through three free AI detectors — one written by ChatGPT, one written by people at NOAA years before ChatGPT existed. All three correctly caught the AI passage. Two of the three also flagged the human passage, one of them at 100%. If a detector has ever been used to judge your work, that is the number that should worry you.

The interesting failure with these tools is not that they miss AI text. It is that they accuse real writing, and they do it with a confident percentage attached.

How the test was set up

A fair test needs two passages that differ in origin and as little else as possible, so I matched them deliberately:

  • The human passage is 179 words from NOAA's JetStream weather education material, about thunderstorm frequency and lightning. It is public domain, written by government staff, and published well before current AI writing tools existed. Provenance is not in question.
  • The AI passage is 180 words I asked ChatGPT to write on the same subject, in the same plain factual register, with specific figures and no headings or bullets.

Same topic, same length, same register, same lack of personality. The only real variable is who wrote it. Then both went through ZeroGPT, QuillBot and Sapling, all three free and none requiring an account.

The results

DetectorChatGPT passage (180 words)NOAA passage (179 words)
ZeroGPT100% AI — correct100% AI — wrong
QuillBot95% AI — correct0% AI, 100% human — correct
Sapling100% “fake” — correct91.4% “fake” — wrong

Every detector identified the ChatGPT passage. That is the easy half of the job, and they all did it.

The other half went badly. Sapling rated the NOAA passage 91.4% “fake” and highlighted every sentence of it in red.

Sapling AI detector showing a NOAA weather passage entirely highlighted in red with a score of 91.4 percent fake
Public-domain NOAA text, rated 91.4% “fake”. Sapling, September 2026.

ZeroGPT was worse: 100% AI, on a passage written by government meteorologists. Its headline hedged — “your text contains mixed signals” — while the gauge underneath read 100%.

QuillBot was the one that got it right, returning 0% AI and 100% human-written on the same text. Here are the two verdicts on the identical passage, side by side.

ZeroGPT rating the NOAA passage 100 percent AI above QuillBot rating the same passage 0 percent AI and 100 percent human-written
The same 179 words of NOAA text, judged by two detectors. September 2026.

Two tools, one text, 100% and 0%. At least one of them is badly wrong, and from the outside there is nothing to tell you which.

Three tools are not always three opinions

While testing I opened Scribbr's free AI detector as a fourth check and found the QuillBot logo sitting inside the results panel. Scribbr's detector displays QuillBot branding, so checking your text there and at QuillBot is not two independent second opinions. It is worth looking at whose engine is actually behind a tool before you treat agreement between two of them as confirmation.

Why the false positives happen

Detectors do not recognise AI writing the way you recognise a friend's handwriting. Broadly, they measure how predictable the text is — whether each word is the obvious next word, and how much the sentence lengths and structures vary. AI writing tends to be predictable and even. The problem is that so is a lot of good human writing.

Technical documentation, government and safety material, legal writing, instructions, textbook prose, and anything written by someone taught to write plainly for a general audience all share exactly the features these tools score as artificial. The NOAA passage is a clean example: plain declarative sentences, consistent structure, no flourishes. It reads as machine-like because it was written carefully, and being careful is the thing that gets flagged.

This also means the people most likely to be falsely accused are not random. Plain, well-structured writing scores worst, and so does writing by people who learned English in a classroom rather than at home.

How to sanity-check a detector in five minutes

Before trusting any of these tools, or arguing with someone who does, calibrate it. The method is the one used above and it takes about five minutes:

  • Find text of certain human origin. Anything published well before late 2022 works. Government agency pages are ideal because they are public domain, plainly written, and easy to date — NOAA, the National Weather Service, the CDC, the census. Take 150 to 200 words.
  • Get a matched AI passage. Ask any chatbot to write on the same subject at the same length in the same register. Matching matters; comparing a breezy blog post against a technical passage tests topic and tone, not origin.
  • Run both, and watch the gap. A tool worth listening to should score them far apart. If it flags your known-human passage, you have learned what its number is worth, and you can show someone else that result in about the time it takes to explain the problem.

Keep whichever human passage you used. Being able to produce a known-human text that a given detector calls artificial is a more useful thing to have than any score it gives your own writing.

If you have been accused

A detector score is not evidence. It is one tool's estimate, with no error rate you can see and, as above, no agreement with the tool next to it. What actually shows authorship is the history of the work, and that is what to produce:

  • Version history. Google Docs and Word both keep it. A document that grew over days, with revisions and deletions and reordered paragraphs, looks nothing like one pasted in whole.
  • Your working material. Notes, outlines, sources, earlier drafts, the half-finished paragraph you gave up on.
  • The ability to talk about it. Someone who wrote a piece can explain why they cut a section or chose one example over another.

It is also fair to ask whoever ran the detector which tool it was, what the number was, and whether they have tested that tool against writing of known origin. Running a few paragraphs of undisputed human text through it, as I did here, tends to end the conversation quickly.

What I would and would not use them for

As a private smell test on your own work — a first draft you suspect is flat and generic — a detector is a rough signal. A high score often means the writing is bland, which is worth knowing even when it is entirely yours.

For any decision with a consequence attached — a grade, a job, a contract, a public accusation — the results above should be disqualifying. Two tools disagreed completely about a passage whose origin is beyond dispute.

What this test does not prove

One matched pair of passages, both under 200 words, both on the same technical subject, run once each. That is a demonstration, not a measured accuracy rate, and I would not claim from it that any of these tools has a particular false positive rate. Detector models also change without notice, so the same text next month may score differently.

What the test does establish is narrower and enough: on a real human passage, these tools produced 100%, 91.4% and 0%. Any process that treats one of those numbers as proof is a process that will eventually punish someone for writing clearly.

Related reading

Written by Mitch, a software analyst who tests software for a living. This guide comes from actually using the tool on a real task. Tested on ChatGPT Plus, Claude Max and Gemini (free). More about this site · Corrections: [email protected].

Get the next guide by email

A new tested how-to when there is one — usually a couple a month. No hype, unsubscribe any time.

© 2026 The AI Cheat Sheet  ·  About  ·  Contact  ·  Privacy Policy  ·  RSS
Tested in real accounts. No affiliate links.

Popular posts from this blog

Is It Safe to Paste Work Documents Into ChatGPT?

How to Use AI to Write and Tailor Your Resume and Cover Letter

How to Spot AI Scams and Deepfakes Before They Cost You