AI Detectors Explained
AI detectors don't know you used AI — they measure statistical fingerprints that model output happens to carry. Understand the fingerprints and three things follow: why honest writing gets flagged, what actually changes scores, and how to verify any fix. This hub covers all three, then links to every detector-specific guide on this site.
The fingerprint, concretely
- Low variance in sentence length — the "burstiness" GPTZero made famous.
- Highly probable word choices — models pick the statistically likely next word, every time.
- AI-typical lexicon — furthermore, moreover, delve, leverage, comprehensive, testament, pivotal.
- Template structure — transition-led paragraphs, repeated openers, tidy three-part everything.
Our open methodology documents exactly which signals we measure and how they're weighted.
Anti-detector methods, ranked by what actually works
| Method | Verdict |
|---|---|
| Vary sentence rhythm (short + long mix) | Works — strongest single signal |
| Remove AI vocabulary, use plain verbs | Works |
| Add specifics (numbers, names, your observations) | Works |
| Verify with before/after scores | Essential — the only proof |
| Synonym-spinning / paraphrase tools | Fails — fingerprint survives swaps |
| Random typos | Fails — humans notice, detectors don't |
| Homoglyphs / invisible characters | Backfires — detected as evasion |
| Double translation | Fails — grammar mangles |
The verify loop (works on any detector)
- Baseline. Free detection run: score + flagged terms + suspect lines.
- Rewrite. Humanize up to 1,500 words: rhythm varied, lexicon cleaned, meaning and citations intact.
- Verify. Automatic re-score — numbers, not vibes.
- Hand-finish. Edit the few surviving flagged lines yourself.
- Insure. Keep drafts; process evidence beats any score argument.
Three scenarios in detail: AI drafts on a deadline, your own falsely-flagged writing, and client content passing buyer scans — all in the 5-step workflow.
How accurate are detectors, really?
Accurate on obvious cases, shaky at the edges. Benchmarks advertise 98-99%+; classrooms deliver meaningfully worse. False positives concentrate on formal academic writing, non-native English and neurodivergent styles — and the same text can score 15% on one tool and 60% on another. Full analysis: are AI detectors accurate. If you've been flagged unfairly: the false-accusation playbook.
Detector guides, one by one
- Turnitin — the academic standard; similarity vs AI score, what percentage means
- GPTZero — perplexity and burstiness, published openly
- Copyleaks — claims 99%+ and mixed-text detection
- Originality.ai — strictest for content buyers
- ZeroGPT — the free student favorite
- Winston AI — educators' pick, sentence highlighting
Turnitin specifics (similarity colors, 25%, no-repository, appeals): the full Q&A cluster. Comparing tools against each other: detectors like Turnitin, compared.
Why detectors disagree with each other
Run the same essay through three detectors and you'll get three different scores — sometimes wildly different. The disagreement isn't a bug; each tool weighs the same signal families with different models, thresholds and training data. The practical consequences:
| Source of disagreement | What it looks like |
|---|---|
| Different signal weights | One tool keys on vocabulary, another on rhythm — your edit moves one score, not the other |
| Different training cutoffs | Newer model output flags on older detectors less reliably |
| Different thresholds | The same internal probability displays as 35% on one tool and 62% on another |
| Chunking differences | Long texts get split differently; scores average out differently |
This is why the verify loop matters more than any single number: when the machine fingerprints are actually gone, every detector's score drops. Chasing one tool's percentage is how people end up rewriting text into gibberish.
The detectors at a glance
| Detector | Who runs it | Temperament |
|---|---|---|
| Turnitin | Schools — the academic default | High stakes, similarity + AI in one report |
| GPTZero | Education + individuals | Publishes its two metrics; jumpy on formal text |
| Copyleaks | Schools & enterprises | Claims mixed-text detection; strict |
| Originality.ai | Content buyers & SEO | Strictest; scans for marketing clients |
| ZeroGPT | Free self-checks | Fast and free; noisiest on short texts |
| Winston AI | Educators | Sentence-level highlighting, OCR scans |
If you've been flagged and you didn't use AI
False positives are the category's worst-kept secret — formal writing, non-native English and neurodivergent styles score high on exactly the signals detectors trust. The route that works:
- Don't confess to something you didn't do. An AI score is an indicator, never proof — the vendors' own docs say so.
- Assemble process evidence. Document version history, drafts, notes — Google Docs and Word timestamp every edit.
- Get an explainable second opinion. Our free detector shows which signals fired, so you can argue specifics instead of vibes.
- De-stiffen the flagged lines only. Vary rhythm, cut formal transitions, add specifics — then re-check.
The detailed playbook, including how reviews actually proceed: flagged but innocent · why human writing gets flagged.
What schools actually see
Worth knowing before finals week: instructors see the reports inside their own Turnitin (or equivalent) assignment — submissions, scores, timestamps, resubmission history. They do not see what you did on independent websites, and independent tools like ours store nothing that could appear in any institutional system. Drafts you pre-check privately stay private; what enters the university's repository depends on their submission settings (see no-repository mode and safe pre-checking).
A short history, for context
AI detection went mainstream in early 2023 when GPTZero's launch made "perplexity and burstiness" dinner-table words, and Turnitin shipped its AI indicator months later amid a wave of academic panic. Since then the story has been an arms race with a twist: detectors improved, but so did ordinary writing tools — grammar checkers, autocomplete, translation — all quietly nudging everyone's prose toward the statistical middle. That's the backdrop nobody markets: the detectors and the "clean human writing" they model are both moving targets, and the writers caught in between are usually innocent. It's also why we advocate process evidence (drafts, version history) over any score, ours included.
Common misconceptions, quickly
- "0% AI means safe." No number means safe — scores are indicators. An empty flagged-lines list is the robust signal.
- "Detectors can prove I used AI." They can't; every vendor's docs admit it. Scores open conversations, they don't close them.
- "Paraphrasing tools remove flags." Word-swaps leave the fingerprint intact — structure-level rewriting is what moves scores.
- "One tool's verdict is the verdict." Same text, three tools, three numbers. Fix the signals and all three drop together.
Reading any score honestly
A detector score answers one narrow question — "how much does this statistically resemble model output?" — and nothing else. It can't see your process, your sources or your intent. So read scores the way you'd read a smoke detector: a beep means check the kitchen, not evacuate your degree. Low score with flagged lines remaining? Fix the lines. High score on text you wrote yourself? The false-positive guides are your section. Score moved after editing? That's the signal working as intended. The writers who navigate this era best treat every number — ours included — as the start of a conversation with their own draft.
Measure it yourself — free, unlimited, no login
See the score and the exact flagged lines behind it, then humanize up to 1,500 words per run and re-score in one click.