How much of the web is invisible to AI assistants?

Aggregate results from every page scanned with Citeable. Not a survey and not a sample of the whole web. It is a sample of pages people cared enough to check, which skews toward sites already trying. The real numbers are likely worse.

211pages checked
19returned nothing at all
86distinct sites
95median score /100
6%block an AI search crawler

Where sites fail

CheckFail rateWhy it matters
AI search crawlers in robots.txt6% Removed from the answer entirely
Server serves AI crawlers16% A CDN rule silently overriding robots.txt
Content in raw HTML3% Crawlers do not run JavaScript
Structured data28% Nothing machine-readable about the page
Machine-readable date90% Cannot be confirmed current

Two things about these figures, because a page arguing that other people publish numbers they cannot explain has to explain its own. The counts are approximate under concurrency: each scan increments a shared record, and simultaneous scans can overwrite one another, so busy periods undercount rather than over. The median comes from a histogram in ten-point bands, so it names the band the middle page falls in rather than an exact value. Both are honest limits of counting cheaply at the edge, and both err toward saying less than we measured.

The percentages above are over the 192 pages we could grade, not over everything we tried. A page that never loaded has no robots.txt verdict to fail and no structured data to be missing, so folding it into those rates would drag every one of them down and make the web look healthier than it is. Refusals are counted on their own line instead. They are the most severe result this instrument can produce and, until recently, the only one it did not record at all.

A scan is counted once per page per six hours, so reloading a report does not inflate anything here. A refusal is counted on the same six-hour terms.

Check your own