Key findings
- 1
Across 6 categories, the share of homepages that block at least one tracked AI crawler runs from 2% (Cybersecurity and DevTools & Cloud) to 18% (Consumer apps).
- 2
Bytespider is the most-blocked crawler, on 8 of 219 homepages (4%). OAI-SearchBot and PerplexityBot are the least-blocked, each on 1 (under 1%).
- 3
Of the 2190 crawler-and-homepage pairs, 2% are blocked, 80% are limited by an applicable Disallow rule and 18% are open.
01
Which crawlers are blocked
We read the robots.txt of every homepage that completed a scan and checked it against each of 10 tracked AI crawlers. A crawler counts as blocked when the file disallows the whole site for it.
Bytespider is blocked most often, on 8 of 219 homepages (4%).
Fig. 1
Source: Citehound scans, 29 July 2026 to 31 July 2026, n=219
02
Blocking by category
The categories differ in how many of their homepages block at least one tracked crawler. B2B categories are navy and consumer categories are gold.
Fig. 2
Source: Citehound scans, 29 July 2026 to 31 July 2026, n=219
03
Blocked, limited and open
Most pairs of crawler and homepage are neither blocked nor fully open. They are limited: at least one Disallow rule applies to that crawler. That is usually an ordinary path such as an admin area or a search page, not a block on the site.
1761 of the 2190 pairs are limited, 41 are blocked and 388 are open.
Fig. 3
Source: Citehound scans, 29 July 2026 to 31 July 2026, n=219
Methodology and limitations
What we scanned
219 homepages completed a scan in six categories: 125 B2B and 94 consumer. 59 more could not be reached and are excluded from every figure.
What we pulled
For each homepage, its robots.txt, retrieved live on its category’s scan date: CRM software 29 July 2026, Cybersecurity 31 July 2026, DevTools & Cloud 31 July 2026, DTC brands 29 July 2026, Consumer apps 31 July 2026 and Hospitality & travel 31 July 2026.
What this measures
What a site’s robots.txt says to each tracked AI crawler. It does not measure whether a crawler obeys the file, or whether any assistant names a brand in a live answer.
What this doesn’t tell you
- Anything beyond homepages. Each site was read once, at its homepage.
- How common blocking is across the web. The sites are hand-picked well-known names in each category, not a random sample.
- How the sites look today. Each category has one scan date. The October 2026 rescan of the same six lists moved category averages by one point or less (see the methodology changelog).
- Anything about the 59 sites that could not be reached, which are counted here and left out.
- Whether a limited result is harmful. It means an applicable Disallow rule, usually an ordinary path, and is not a block.
How to cite this
Andaç Üzel, “Which AI crawlers do sites block? 219 homepages, six categories,” Citehound Research, 5 October 2026. https://getcitehound.com/research/crawler-access-2026
If you want to see what your own robots.txt says to each of these crawlers, run the free scan.
Scan your site free