A detector claiming 99% accuracy is usually reporting how often it catches unedited model output on a curated test set. That is the easy case. The hard cases — lightly edited AI, human text in plain prose, mixed drafts — are where the number quietly collapses.
Detectors estimate statistical predictability — how unsurprising each next word is. Simple, careful, well-edited human writing is also unsurprising. There is no signal that separates "clear" from "generated," which is why the error is structural rather than a bug to be patched.
Every new model generation shifts the distribution the detector was tuned on. Scores on the same text change between vendor updates. Treat any single score as one weak signal with a timestamp, never as evidence.
Run two or three, note the spread, and treat a high score as a prompt to reread rather than a verdict. If you are on the receiving end of an accusation, the version history of your draft is worth more than every detector combined.
Follow the store on Gumroad (free, no purchase) and every new drop lands in your library automatically — plus grab the whole store 50% off with code FIRST50 at checkout.
Like this? Join the affiliate program and keep 20% on every sale you refer — including the $9/mo Prompt Club recurring. No approval, no quota.
Want it done for you? Our $15/$45/$120 packages humanize and SEO-optimize any draft — or grab the self-serve Humanize Pro Pack.
← Back to HumanizeAI Lab