Type an essay, a cover letter, or a blog post into an AI detector, and you’ll get a confident-looking number: “97% AI-generated” or “100% Human.” That number feels like a verdict. It isn’t one — and the gap between how these tools are used and how well they actually perform is bigger than most people realize.

The company that built the technology gave up on its own tool

The most telling data point isn’t from a critic. It’s from OpenAI. In early 2023, OpenAI released its own AI text classifier, trained specifically to distinguish human writing from output by GPT-family models. By its own published numbers, the tool correctly identified AI-written text only 26% of the time, while incorrectly flagging real human writing as AI-generated 9% of the time. Six months later, OpenAI pulled it entirely, stating plainly that the classifier was being discontinued due to its low rate of accuracy.

If the company that built the underlying models couldn’t build a reliable detector for its own output, that’s a meaningful signal about how hard this problem actually is — not a fluke specific to one bad tool.

Independent tests tell a similar story

This isn’t just a 2023 problem that’s since been solved. Independent testing continues to find wide, inconsistent variation across the popular tools people actually use. One recent large-scale comparison found that accuracy across major detectors ranges anywhere from 65% to 90% depending on the tool, with false positives remaining a real and well-documented problem rather than a rare edge case.

A rigorous independent academic study out of the University of Chicago’s Booth School of Business — built on a corpus of nearly 2,000 pre-2020 human texts paired with AI-generated texts — is one of the more careful attempts to measure this properly. It found that most detectors become essentially useless once you require a strict false-positive cap, with true-positive (correct AI detection) rates collapsing toward zero when tools are tuned to avoid falsely accusing real human writers. Only one tool in the study, Pangram, managed to hold onto meaningful detection power under a strict 0.5% false-positive cap — most others simply couldn’t do both at once.

The bias problem is arguably worse than the accuracy problem

Overall accuracy percentages hide a more serious issue: detectors don’t fail evenly across writers. A now widely-cited Stanford study found that AI detectors misclassified more than 61% of essays written by non-native English speakers as AI-generated, while achieving near-perfect accuracy on essays from native English speakers. A more recent follow-up using TOEFL essays found an almost identical pattern: a 61.3% false-positive rate for Chinese students’ essays compared to just 5.1% for essays from US students using the same detection setup.

The likely mechanism is that non-native speakers, along with some neurodivergent writers, tend to produce writing with lower “perplexity” — more predictable word choices and more uniform sentence structure — which happens to be one of the exact statistical signatures detectors are trained to flag as AI-like. In other words, some of the most vulnerable groups to a false accusation are being flagged not because their writing resembles AI text by coincidence, but because the detectors’ core methodology conflates “formulaic” with “generated.”

Even the vendors’ own numbers show the gap

It’s worth pointing out that detector companies aren’t hiding this entirely — the fine print tends to tell a more modest story than the marketing headline. Turnitin, for instance, publicly advertises a 98% accuracy claim with under 1% false positives for documents with significant AI content. But independent testing on real submissions has found Turnitin’s detection accuracy on unedited AI text landing closer to 90–95%, with false-positive rates on harder cases — non-native English writing, heavily edited drafts, technical prose — climbing to 5–12%. When Vanderbilt University reviewed the math on its own submission volume, it noted that even at Turnitin’s own claimed 1% false-positive rate, its roughly 75,000 annual submissions implied about 750 students could be flagged incorrectly in a single year — which is why the university disabled the AI detection feature rather than rely on it for disciplinary decisions.

They’re also easy to defeat on purpose

Beyond the false-positive problem, there’s a false-negative problem working in the opposite direction: detectors are notably easy to fool if someone wants to get past them. Paraphrasing tools and “AI humanizers” reliably reduce detection accuracy, and testing has found that even simple rewriting can drop detection rates by 20% or more. Because detectors largely work by measuring statistical patterns like word predictability and sentence-length uniformity, a light paraphrase pass is often enough to push a genuinely AI-written text below the flagging threshold — while, confusingly, the same statistical quirks can push some human writing above it.

So do any of them work?

The honest answer is: partially, and unevenly. The best-performing tools in independent testing (Pangram was a standout in the Booth study, and Originality.ai and Copyleaks score well in other comparisons) do meaningfully outperform free, casual tools like the original GPTZero on long-form academic text. But “meaningfully better than the worst option” is a very different claim from “reliable enough to base a decision on,” and no major detector — including the tools built by the same labs that created the AI models in question — has demonstrated the kind of accuracy that would justify treating a single score as proof of anything.

The practical takeaway echoes what most of the researchers behind this work actually recommend: treat a detector’s output as a prompt to look closer — draft history, writing samples, a conversation with the person — not as a verdict in itself. That’s not a hedge to avoid controversy; it’s the same conclusion the tool-makers themselves have reached when they’ve been honest about their own numbers.


If you’re evaluating whether to trust a detector score on something that matters — a job applicant’s writing sample, a student’s essay — the research above suggests the single most useful step is running the same text through more than one tool and treating disagreement between them as a signal in itself.

Leave a comment

Leave a Reply