How Accurate Are AI Detectors, Really?

Vendors claim 99%. Independent testing tells a different story, especially for non-native writers.

What the marketing number leaves out

A detector claiming 99% accuracy is usually reporting how often it catches unedited model output on a curated test set. That is the easy case. The hard cases — lightly edited AI, human text in plain prose, mixed drafts — are where the number quietly collapses.

The two failure modes that matter

Why they behave this way

Detectors estimate statistical predictability — how unsurprising each next word is. Simple, careful, well-edited human writing is also unsurprising. There is no signal that separates "clear" from "generated," which is why the error is structural rather than a bug to be patched.

The moving-target problem

Every new model generation shifts the distribution the detector was tuned on. Scores on the same text change between vendor updates. Treat any single score as one weak signal with a timestamp, never as evidence.

How to use them sensibly

Run two or three, note the spread, and treat a high score as a prompt to reread rather than a verdict. If you are on the receiving end of an accusation, the version history of your draft is worth more than every detector combined.

Free: get every new prompt pack + 50% off launch

Follow the store on Gumroad (free, no purchase) and every new drop lands in your library automatically — plus grab the whole store 50% off with code FIRST50 at checkout.

→ Follow free on Gumroad

Earn 20% sharing these tools

Like this? Join the affiliate program and keep 20% on every sale you refer — including the $9/mo Prompt Club recurring. No approval, no quota.

→ Become an affiliate

Want it done for you? Our $15/$45/$120 packages humanize and SEO-optimize any draft — or grab the self-serve Humanize Pro Pack.

← Back to HumanizeAI Lab