Methodology: How the Detector and Humanizer Work
This page documents exactly how our AI detector and humanizer work, what they measure, what they don't, and where they can be wrong. We publish it because a score without a method is just a number — and because you deserve to evaluate the tool before trusting it with your writing.
How the detector works
Our detector is an explainable heuristic engine — not a neural classifier. It analyzes eight families of statistical signals that correlate with machine-generated writing, then blends them into a 0-100 score:
| Signal | What it measures | Why it matters |
|---|---|---|
| Sentence-rhythm variance ("burstiness") | Coefficient of variation of sentence lengths | Human writing fluctuates wildly; models stay uniform |
| AI-lexicon density | Frequency of ~130 over-represented words/phrases (furthermore, delve, leverage, testament…) | Models overuse a stable academic register |
| Transition density | However/therefore/moreover-style connectors per 100 words | Template-led paragraph structure |
| Sentence-length uniformity | Share of sentences within ±3 words of the mean | Uniformity is the core machine fingerprint |
| Punctuation variety | Diversity of ! ? ; : ( ) … " | Humans scatter; models standardize |
| Contraction rate | it's / don't / they're per 100 words | Formal model output rarely contracts |
| Em-dash density | — and -- per 100 words | Current-generation models overuse them |
| Opener repetition | Max share of sentences sharing a two-word opener | "The… The… The…" patterns |
Every result page shows the real values behind your score: lexicon density per 100 words, your burstiness value, average sentence length, the flagged terms, and the five most AI-sounding sentences. You can audit the arithmetic yourself.
What we deliberately do NOT claim
- No third-party equivalence. Our engine is ours. Turnitin, GPTZero, Copyleaks and the rest use different (mostly neural) methods with different weights. Our score correlates with what they measure — the same signal families — but it is not a prediction of any specific detector's output.
- No infallibility. Formal academic writing, non-native English and neurodivergent styles score higher on these signals — a documented bias across the whole category (see are AI detectors accurate). A high score is a reason to look at the flagged lines, never a verdict.
- No hidden benchmark. When we publish accuracy comparisons, they will live at a public URL with test sets, dates and downloadable data. Until then we claim nothing we don't show.
How the humanizer works
The humanizer rewrites text through a hosted language model under a strict editing contract: keep every fact, number, name, quotation and citation byte-identical; vary sentence rhythm; remove AI-typical vocabulary; preserve the original language and register (academic stays academic). After rewriting, the output is re-scored by the same detection engine, and we show you the measured before/after — including a note when a rewrite did not lower the score, because hiding that would be dishonest.
Verification workflow and editing tips: the 5-step guide. For who tends to be falsely flagged: our false-positive analysis.
Privacy under this method
Detection runs per request in memory; we store no text, no account, no profile (see Privacy Policy). The only counters we keep are anonymous usage tallies (how many detections/humanizations ran) that contain no user content.
Changes to this method
When signals or weights change materially, we update this page and note the date. The engine's calibration targets are simple: near-zero scores for informal human writing, high scores for raw model output, honest middles for everything else.
Why heuristics instead of a neural classifier
Most detectors are neural networks: powerful, but opaque — they produce a number and offer no way to check its reasoning. We made the opposite trade on purpose. A heuristic engine is slightly less accurate on adversarial edge cases, but every score is auditable: you can see each signal's value, look up its weight in the table above, and reproduce the arithmetic by hand. For a tool whose results get quoted in conversations with professors and editors, "show your work" beats "trust the black box." It also means when a detector-style bias hurts a particular group of writers — as documented across this category — we can see exactly which signal caused it and say so, rather than shrug at a neural net.
The trade-off we accept: heuristics can be gamed once someone reads this page, and they lag behind brand-new model writing styles until we update the word lists. We think that's fair — an arms race we decline to hype.
How the engine was calibrated
Calibration is the boring discipline behind every score. We tune weights against labeled sample sets and re-run them on every engine change. Current calibration targets, from the latest run:
| Sample type | Expected score | Latest measured |
|---|---|---|
| Raw AI output (academic register) | 85-90 (high) | 86-90 ✓ |
| Raw AI output (blog/marketing register) | 85-90 (high) | 90 ✓ |
| Casual human writing (forum, personal) | Below 25 | 4-20 ✓ |
| Human forum/technical writing | Below 25 | 20 ✓ |
| Human academic writing (careful student) | Below 25 | 9 ✓ |
| Mixed text (human + one AI paragraph) | 40-70 (middle) | 76 — known over-flag |
That last row is a real limitation we choose to publish: heavily formatted "mixed" text currently skews high. If your score lands mid-range, trust the flagged-lines list over the headline number — that's exactly what it's for.
Signal weights, published
| Signal | Weight |
|---|---|
| AI-lexicon density | 0.28 |
| Sentence-rhythm variance (burstiness) | 0.22 |
| Transition-word density | 0.14 |
| Sentence-length uniformity | 0.10 |
| Punctuation variety | 0.08 |
| Contraction rate | 0.07 |
| Em-dash density | 0.06 |
| Opener repetition | 0.05 |
Vocabulary and rhythm together carry half the score — which matches what matters most when editing: if you only fix two things, fix rhythm and word choice. Short texts (under ~40 words) are additionally pulled toward the middle, because statistics on tiny samples lie.
How the humanizer's contract is enforced
The rewrite model operates under explicit instructions — keep every fact, number, name, quotation and citation byte-identical; vary sentence rhythm; remove AI-typical vocabulary; preserve language and register — and two mechanical guardrails back the instructions: outputs shorter than 35% of the input are rejected and retried (catches summary-collapse), and every finished run is re-scored by this same engine so the before/after card you see is a measurement, not a promise. When a rewrite fails to lower the score, the UI says so instead of hiding it. Failures keep your original text intact rather than returning a mangled version.
Related reading: how detectors work · detector accuracy in practice · why human writing gets flagged · our editorial policy.
Reading the weights in plain words
The distribution tells a story. Vocabulary and rhythm — the two signals a writer can consciously change — carry exactly half the weight, which is why editing advice works: fix word choice and sentence-length variety and the score moves more than anything else you could do. The remaining half spreads across structural habits (transitions, uniformity, openers) that rewrite themselves naturally once the first two improve. Punctuation, contractions and em-dashes barely register on their own; they're corroborating evidence, not verdicts — which is also why our engine never flags a text on formal vocabulary alone when its rhythm is genuinely human. If your score surprises you, read it as a question — "which of these eight habits is strongest here?" — rather than an answer.
See the method on your own text
Every score we show comes with its signal breakdown — flagged terms, rhythm statistics and suspect lines. Free, unlimited, no login.