Research

Which AI crawlers do sites block? 219 homepages, six categories.

The share of homepages that block each tracked AI crawler, by category, and how many only limit it with a Disallow rule. Computed from the benchmark scans, no new scans.

219homepages scanned
6categories
10AI crawlers
Jul 2026scan window

Key findings

  1. 1

    Across 6 categories, the share of homepages that block at least one tracked AI crawler runs from 2% (Cybersecurity and DevTools & Cloud) to 18% (Consumer apps).

  2. 2

    Bytespider is the most-blocked crawler, on 8 of 219 homepages (4%). OAI-SearchBot and PerplexityBot are the least-blocked, each on 1 (under 1%).

  3. 3

    Of the 2190 crawler-and-homepage pairs, 2% are blocked, 80% are limited by an applicable Disallow rule and 18% are open.

01

Which crawlers are blocked

We read the robots.txt of every homepage that completed a scan and checked it against each of 10 tracked AI crawlers. A crawler counts as blocked when the file disallows the whole site for it.

Bytespider is blocked most often, on 8 of 219 homepages (4%).

Fig. 1

Share of homepages blocking each tracked AI crawler Bytespider 4% (8 of 219). CCBot 3% (7 of 219). GPTBot 3% (7 of 219). Google-Extended 2% (5 of 219). ClaudeBot 2% (4 of 219). anthropic-ai 2% (4 of 219). ChatGPT-User 1% (2 of 219). Perplexity-User 1% (2 of 219). OAI-SearchBot under 1% (1 of 219). PerplexityBot under 1% (1 of 219). Bytespider 4% CCBot 3% GPTBot 3% Google-Extended 2% ClaudeBot 2% anthropic-ai 2% ChatGPT-User 1% Perplexity-User 1% OAI-SearchBot <1% PerplexityBot <1% Filled = homepages whose robots.txt disallows the whole site for that crawler. n=219.
Blocked share by crawler, across all 6 categories. Each bar is a count of homepages out of 219.

Source: Citehound scans, 29 July 2026 to 31 July 2026, n=219

02

Blocking by category

The categories differ in how many of their homepages block at least one tracked crawler. B2B categories are navy and consumer categories are gold.

Fig. 2

Share of homepages blocking at least one tracked AI crawler, by category Consumer apps 18% of 40 homepages. CRM software 10% of 31 homepages. Hospitality & travel 5% of 21 homepages. DTC brands 3% of 33 homepages. Cybersecurity 2% of 46 homepages. DevTools & Cloud 2% of 48 homepages. B2B SaaS Consumer Consumer apps · n=40 18% CRM software · n=31 10% Hospitality & travel · n=21 5% DTC brands · n=33 3% Cybersecurity · n=46 2% DevTools & Cloud · n=48 2% Filled = share of homepages blocking at least one of the 10 tracked crawlers.
Share of each category’s homepages that block at least one of the 10 tracked crawlers.

Source: Citehound scans, 29 July 2026 to 31 July 2026, n=219

03

Blocked, limited and open

Most pairs of crawler and homepage are neither blocked nor fully open. They are limited: at least one Disallow rule applies to that crawler. That is usually an ordinary path such as an admin area or a search page, not a block on the site.

1761 of the 2190 pairs are limited, 41 are blocked and 388 are open.

Fig. 3

Blocked, limited and open, for each tracked AI crawler Bytespider: blocked 8, limited 177, open 34 of 219. CCBot: blocked 7, limited 175, open 37 of 219. GPTBot: blocked 7, limited 172, open 40 of 219. Google-Extended: blocked 5, limited 175, open 39 of 219. anthropic-ai: blocked 4, limited 177, open 38 of 219. ClaudeBot: blocked 4, limited 175, open 40 of 219. Perplexity-User: blocked 2, limited 179, open 38 of 219. ChatGPT-User: blocked 2, limited 176, open 41 of 219. OAI-SearchBot: blocked 1, limited 178, open 40 of 219. PerplexityBot: blocked 1, limited 177, open 41 of 219. Blocked Limited Open Bytespider blocked 8 · limited 177 · open 34 CCBot blocked 7 · limited 175 · open 37 GPTBot blocked 7 · limited 172 · open 40 Google-Extended blocked 5 · limited 175 · open 39 anthropic-ai blocked 4 · limited 177 · open 38 ClaudeBot blocked 4 · limited 175 · open 40 Perplexity-User blocked 2 · limited 179 · open 38 ChatGPT-User blocked 2 · limited 176 · open 41 OAI-SearchBot blocked 1 · limited 178 · open 40 PerplexityBot blocked 1 · limited 177 · open 41 Bars are shares of the 219 homepages. Limited means an applicable Disallow rule.
Blocked, limited and open for each tracked crawler. “Limited” means an applicable Disallow rule, usually an ordinary path, not a block.

Source: Citehound scans, 29 July 2026 to 31 July 2026, n=219

Methodology and limitations

What we scanned

219 homepages completed a scan in six categories: 125 B2B and 94 consumer. 59 more could not be reached and are excluded from every figure.

What we pulled

For each homepage, its robots.txt, retrieved live on its category’s scan date: CRM software 29 July 2026, Cybersecurity 31 July 2026, DevTools & Cloud 31 July 2026, DTC brands 29 July 2026, Consumer apps 31 July 2026 and Hospitality & travel 31 July 2026.

What this measures

What a site’s robots.txt says to each tracked AI crawler. It does not measure whether a crawler obeys the file, or whether any assistant names a brand in a live answer.

What this doesn’t tell you

  • Anything beyond homepages. Each site was read once, at its homepage.
  • How common blocking is across the web. The sites are hand-picked well-known names in each category, not a random sample.
  • How the sites look today. Each category has one scan date. The October 2026 rescan of the same six lists moved category averages by one point or less (see the methodology changelog).
  • Anything about the 59 sites that could not be reached, which are counted here and left out.
  • Whether a limited result is harmful. It means an applicable Disallow rule, usually an ordinary path, and is not a block.

How to cite this

Andaç Üzel, “Which AI crawlers do sites block? 219 homepages, six categories,” Citehound Research, 5 October 2026. https://getcitehound.com/research/crawler-access-2026

If you want to see what your own robots.txt says to each of these crawlers, run the free scan.

Scan your site free