How accurate is PureRank, really?
We scored 275 labeled websites — 185 clearly legitimate, 90 documented AI content farms — and published exactly what we found, including the bugs the test caught in our own engine and the calibration it drove. Here is where the AI Spam Score is trustworthy, and where it's only a hint.
A risk gradient, not a lie detector
Most "AI detectors" sell a yes/no verdict. We don't, and this benchmark is why. The AI Spam Score is a gradient: it is most reliable at the two ends and deliberately cautious in the middle, because the cost of falsely accusing an honest website is worse than missing a borderline one.
In plain terms: if PureRank scores your site 60 or above, take it seriously — nothing legitimate in our test reached that zone. If it scores under 25, you're almost certainly fine. The 30–55 middle is a prompt to look, not a verdict — it holds both cautious-legitimate sites and AI farms we didn't fully catch. The report always shows which signals fired, so you read the reasons, not just the number.
The dataset & method
275 labeled sites, scanned blind at a 25-page sample on the same engine build:
- 185 legitimate sites chosen to be hard — independent newsrooms and magazines (CT Mirror, Colorado Sun, Defector, Science News), developer docs and personal blogs (Astro, Tailwind, Simon Willison, Julia Evans, Martin Fowler), portfolios, small independent e-commerce, nonprofits, hobby forums and non-English sites. Not global mega-brands (those are out of scope) — the mid-sized, high-volume, sometimes-templated sites where a spam heuristic is most likely to trip.
- 90 documented AI content farms, each named by a reputable investigation — 404 Media, Nieman Lab, MIT Technology Review — and re-checked live before scanning: AI "news" rip-off sites, made-for-advertising health and finance farms, crypto slop, and mass-produced recipe sites.
We measured specificity (how often we wrongly flag an honest site), class separation (do farms score higher than legit sites?), and the trade-off across every threshold. We did not compute a single "accuracy %" — with a gradient that would hide more than it shows.
Both sides were picked to be adversarial: fluent modern AI slop that reads human, and diverse legitimate publishers that stress our heuristics. A random sample of ordinary small-business sites would show much higher specificity than the figures here. Treat this as a floor, not the typical scan.
How the score behaved, band by band
| Score band | What it means | How it behaved on the benchmark |
|---|---|---|
| ≥ 60 · High–Critical | Confidently matches the scaled-content pattern | 0 of 185 legitimate sites. Caught the most blatant farms (a documented MFA health farm, a 300+‑page recipe mill). When it fires here, it's right. |
| 50–59 · High | Elevated — worth a real review | 95% precision — of everything we flag here, nineteen in twenty are documented farms. Exactly one honest site reached it: a hobby forum whose user posts are saturated with pasted AI answers. |
| 25–49 · Moderate | Advisory — read the signals | The crowded middle: cautious-legitimate publishers and AI farms we under-scored both live here. The band to investigate, not to trust blindly. |
| < 25 · Low | Confidently clean | Dominated by real human sites — docs, personal blogs, edited publications. Reliable "all clear". |
Documented AI farms averaged meaningfully higher than legitimate sites, but the classes overlap in the middle — which is exactly why we ship a gradient with an explained breakdown instead of a single verdict.
What this benchmark caught — and we fixed
A dependency upgrade had introduced an HTML-parser incompatibility that could abort a scan on some legitimate sites (a nonprofit newsroom, a food magazine). The benchmark hit it immediately. Fixed with a resilient parser fallback — no single page can crash a scan again.
A gap in the Internet Archive's coverage of a long-lived site was being read as evidence the domain had died and been hijacked — flagging, among others, a well-known developer's own blog as an expired-domain-abuse case. Fixed: only a confirmed dead/parked period counts now; a bare archive gap on a site that kept its identity is treated as continuous ownership.
An ordinary phrase — "we always recommend…" — was matching our prompt-injection blocklist and flooring innocent sites (a flagship Python library, a leather-goods shop) to a "High" manipulation verdict. Fixed: only genuinely AI-directed phrasing counts now.
Our strongest single signal — site-wide saturation with AI-typical language — was designed to floor a site into the High band. A multi-agent audit of our own scoring code found that this floor appeared in reports but was never actually applied to the score. Fixing that one line nearly doubled how many documented farms reach the High band (13% → 24%) at 95% precision.
With a big labeled set we could finally measure each signal's real predictive power. One — our trust/E-E-A-T surface — turned out to be anti-predictive: it scored real independent newsrooms as "thinner" than the content farms. We re-derived every dimension's weight from the data, leaning on the signals that actually separate (site-wide AI-language markers, formatting patterns) and cutting the ones that don't. Result on this set: sharper separation and roughly a 7× drop in false alarms at the "High" line.
Every fix shipped to production before this page updated. That's the whole reason to benchmark in public: it turns "trust us" into "here's what we found, and here's what we changed."
Where the score is weak (on purpose and not)
Modern AI farms produce text that reads cleanly, and three in four documented farms still scored only "Moderate" in our test. Catching them without raising false alarms on honest sites is the hard, central problem — and our next engineering focus. We'd rather under-flag than falsely accuse, so today the top band is precise but not exhaustive.
A small independent shop with many near-identical product pages, or a site whose marketing copy leans on stock phrasing, can still edge into the High band. The re-weighting above cleared most of these, but it's the one legitimate category that occasionally reaches it — read the signal breakdown before acting.
The expired-domain check reads the live Internet Archive, which is sometimes slow or incomplete — so that one signal can vary between scans. We flag it as lower-confidence when the archive doesn't respond fully.
Read your own report — signals, not just a number
Every scan shows the 7-dimension breakdown and every signal that fired, so you can judge the score for yourself. See the full methodology.
Run a free scan →