Free SEO Audit Run free scan
Transparency report — we tested ourselves

How accurate is PureRank, really?

We scored 275 labeled websites — 185 clearly legitimate, 90 documented AI content farms — and published exactly what we found, including the bugs the test caught in our own engine and the calibration it drove. Here is where the AI Spam Score is trustworthy, and where it's only a hint.

≤ 25 · confidently clean 30–55 · advisory middle ≥ 60 · confidently flagged
The honest headline

A risk gradient, not a lie detector

Most "AI detectors" sell a yes/no verdict. We don't, and this benchmark is why. The AI Spam Score is a gradient: it is most reliable at the two ends and deliberately cautious in the middle, because the cost of falsely accusing an honest website is worse than missing a borderline one.

0 / 185
legitimate sites scored in the confidently-flagged zone (≥ 60). No honest site in our set was called out at the top.
95%
of the sites we flag "High" (≥ 50) are documented AI farms — exactly one false alarm across 166 honest sites scored.
1 in 4
documented AI farms reach the "High" band (24% recall). The other three still read fluently enough to land "Moderate" — our known gap and next focus.

In plain terms: if PureRank scores your site 60 or above, take it seriously — nothing legitimate in our test reached that zone. If it scores under 25, you're almost certainly fine. The 30–55 middle is a prompt to look, not a verdict — it holds both cautious-legitimate sites and AI farms we didn't fully catch. The report always shows which signals fired, so you read the reasons, not just the number.

How we tested

The dataset & method

275 labeled sites, scanned blind at a 25-page sample on the same engine build:

  • 185 legitimate sites chosen to be hard — independent newsrooms and magazines (CT Mirror, Colorado Sun, Defector, Science News), developer docs and personal blogs (Astro, Tailwind, Simon Willison, Julia Evans, Martin Fowler), portfolios, small independent e-commerce, nonprofits, hobby forums and non-English sites. Not global mega-brands (those are out of scope) — the mid-sized, high-volume, sometimes-templated sites where a spam heuristic is most likely to trip.
  • 90 documented AI content farms, each named by a reputable investigation — 404 Media, Nieman Lab, MIT Technology Review — and re-checked live before scanning: AI "news" rip-off sites, made-for-advertising health and finance farms, crypto slop, and mass-produced recipe sites.

We measured specificity (how often we wrongly flag an honest site), class separation (do farms score higher than legit sites?), and the trade-off across every threshold. We did not compute a single "accuracy %" — with a gradient that would hide more than it shows.

An honest caveat about the numbers.

Both sides were picked to be adversarial: fluent modern AI slop that reads human, and diverse legitimate publishers that stress our heuristics. A random sample of ordinary small-business sites would show much higher specificity than the figures here. Treat this as a floor, not the typical scan.

Results

How the score behaved, band by band

Score bandWhat it meansHow it behaved on the benchmark
≥ 60 · High–CriticalConfidently matches the scaled-content pattern0 of 185 legitimate sites. Caught the most blatant farms (a documented MFA health farm, a 300+‑page recipe mill). When it fires here, it's right.
50–59 · HighElevated — worth a real review95% precision — of everything we flag here, nineteen in twenty are documented farms. Exactly one honest site reached it: a hobby forum whose user posts are saturated with pasted AI answers.
25–49 · ModerateAdvisory — read the signalsThe crowded middle: cautious-legitimate publishers and AI farms we under-scored both live here. The band to investigate, not to trust blindly.
< 25 · LowConfidently cleanDominated by real human sites — docs, personal blogs, edited publications. Reliable "all clear".

Documented AI farms averaged meaningfully higher than legitimate sites, but the classes overlap in the middle — which is exactly why we ship a gradient with an explained breakdown instead of a single verdict.

The point of testing yourself

What this benchmark caught — and we fixed

A crash on certain valid pages.

A dependency upgrade had introduced an HTML-parser incompatibility that could abort a scan on some legitimate sites (a nonprofit newsroom, a food magazine). The benchmark hit it immediately. Fixed with a resilient parser fallback — no single page can crash a scan again.

A false "hijacked expired domain" verdict.

A gap in the Internet Archive's coverage of a long-lived site was being read as evidence the domain had died and been hijacked — flagging, among others, a well-known developer's own blog as an expired-domain-abuse case. Fixed: only a confirmed dead/parked period counts now; a bare archive gap on a site that kept its identity is treated as continuous ownership.

A false "prompt-injection" accusation.

An ordinary phrase — "we always recommend…" — was matching our prompt-injection blocklist and flooring innocent sites (a flagship Python library, a leather-goods shop) to a "High" manipulation verdict. Fixed: only genuinely AI-directed phrasing counts now.

A safety floor that existed only on paper.

Our strongest single signal — site-wide saturation with AI-typical language — was designed to floor a site into the High band. A multi-agent audit of our own scoring code found that this floor appeared in reports but was never actually applied to the score. Fixing that one line nearly doubled how many documented farms reach the High band (13% → 24%) at 95% precision.

A dimension that was scoring backwards.

With a big labeled set we could finally measure each signal's real predictive power. One — our trust/E-E-A-T surface — turned out to be anti-predictive: it scored real independent newsrooms as "thinner" than the content farms. We re-derived every dimension's weight from the data, leaning on the signals that actually separate (site-wide AI-language markers, formatting patterns) and cutting the ones that don't. Result on this set: sharper separation and roughly a 7× drop in false alarms at the "High" line.

Every fix shipped to production before this page updated. That's the whole reason to benchmark in public: it turns "trust us" into "here's what we found, and here's what we changed."

What we're still working on

Where the score is weak (on purpose and not)

Recall on fluent AI content.

Modern AI farms produce text that reads cleanly, and three in four documented farms still scored only "Moderate" in our test. Catching them without raising false alarms on honest sites is the hard, central problem — and our next engineering focus. We'd rather under-flag than falsely accuse, so today the top band is precise but not exhaustive.

Templated e-commerce & stock-phrase copy.

A small independent shop with many near-identical product pages, or a site whose marketing copy leans on stock phrasing, can still edge into the High band. The re-weighting above cleared most of these, but it's the one legitimate category that occasionally reaches it — read the signal breakdown before acting.

Archive-dependent signals vary.

The expired-domain check reads the live Internet Archive, which is sometimes slow or incomplete — so that one signal can vary between scans. We flag it as lower-confidence when the archive doesn't respond fully.

Read your own report — signals, not just a number

Every scan shows the 7-dimension breakdown and every signal that fired, so you can judge the score for yourself. See the full methodology.

Run a free scan →

Questions

Is the AI Spam Score a proven Google-penalty detector?

No, and our own benchmark says so. It is an explainable, site-level risk gradient — most trustworthy at the extremes (a score of 60+ flagged zero legitimate sites in our test; a score under 25 is confidently clean) and advisory in the middle. It estimates how strongly a site matches the scaled-content pattern Google's spam policies target; it does not read Google's systems and cannot prove a penalty.

What did the benchmark measure?

We scored 275 labeled sites — 185 clearly legitimate (independent newsrooms, developer docs and blogs, portfolios, small e-commerce, nonprofits, forums, international sites) and 90 documented AI content farms named by 404 Media, Nieman Lab, MIT Technology Review and others. We report specificity (how often we wrongly flag honest sites), where the score separates the two classes, and where it doesn't.

Why publish your weaknesses?

Because a tool that asks you to trust a risk score should show its work. The benchmark caught two real bugs in our own engine, which we fixed. It also shows where we still under-detect fluent AI content. Publishing that is the point — you can see exactly where the number is strong and where it is only a hint.