How much of the web is invisible to AI assistants?
Aggregate results from every page scanned with Citeable. Not a survey and not a sample of the whole web. It is a sample of pages people cared enough to check, which skews toward sites already trying. The real numbers are likely worse.
Where sites fail
| Check | Fail rate | Why it matters |
|---|---|---|
| AI search crawlers in robots.txt | 6% | Removed from the answer entirely |
| Server serves AI crawlers | 16% | A CDN rule silently overriding robots.txt |
| Content in raw HTML | 3% | Crawlers do not run JavaScript |
| Structured data | 28% | Nothing machine-readable about the page |
| Machine-readable date | 90% | Cannot be confirmed current |
Two things about these figures, because a page arguing that other people publish numbers they cannot explain has to explain its own. The counts are approximate under concurrency: each scan increments a shared record, and simultaneous scans can overwrite one another, so busy periods undercount rather than over. The median comes from a histogram in ten-point bands, so it names the band the middle page falls in rather than an exact value. Both are honest limits of counting cheaply at the edge, and both err toward saying less than we measured.
The percentages above are over the 192 pages we could grade, not over everything we tried. A page that never loaded has no robots.txt verdict to fail and no structured data to be missing, so folding it into those rates would drag every one of them down and make the web look healthier than it is. Refusals are counted on their own line instead. They are the most severe result this instrument can produce and, until recently, the only one it did not record at all.
A scan is counted once per page per six hours, so reloading a report does not inflate anything here. A refusal is counted on the same six-hour terms.