Research

Who controls your brand's AI representation?

We scanned 64 sites and found blocking is rare: 10% of B2B sites and 3% of consumer brands restrict AI crawlers. The larger gap is what those crawlers find once they're inside.

64sites scanned
2categories
13checks
Jul 2026scan window

Key findings

  1. 1

    76% of consumer brands have an llms.txt file, but the two files we examined by hand were near-identical and said nothing specific about the brand.

  2. 2

    91% of consumer-brand homepages and 48% of B2B homepages are missing content schema, the markup that tells AI systems what type of content they're looking at.

  3. 3

    22 of 85 attempted scans could not complete, blocked by protection that may also stop AI crawlers — a question this report leaves open.

Before the specifics: both categories land in a similar place on the 100-point score. Consumer brands trail B2B by 7 points on average, and the two ranges overlap across most of their span.

Fig. 1

Score distribution by category B2B SaaS sites range from 53 to 91 out of 100, averaging 78, across 31 sites. Consumer brand sites range from 46 to 87, averaging 71, across 33 sites. B2B SaaS · n=31 53 91 78 Consumer brands · n=33 46 87 71 0 25 50 75 100 AI visibility score (0–100)
The range is the bar; the marker is the category average. The two categories overlap for most of that span — they differ less between each other than the spread inside each one.

Source: Citehound scan, 29 July 2026, n=64

01

The llms.txt illusion

llms.txt is a young standard, barely two years old. It tells AI systems what a site is about and why it should be trusted. On paper, consumer brands have adopted it fast: 76% of the DTC sites we scanned have one, against 55% of B2B sites.

We read two of those consumer files by hand: Magic Spoon and Allbirds. The structure, section order and phrasing were nearly identical. Only the brand name changed.

The reason is Shopify. The platform generates an llms.txt automatically for every store it hosts, which explains why adoption looks high across consumer brands without any of them writing a word themselves. In the two files we examined, each close to 1,000 words, there was no mention of what the brand sells, who it's for, or what sets it apart. Instead, the file gives an AI agent Shopify's own shopping-skill and commerce-protocol instructions.

An AI system reading one of these files learns about Shopify, not about the brand behind the store.

B2B tells a different story. 55% of the B2B sites we scanned have an llms.txt, and no single platform is generating it for them. Every one of those files reflects a choice someone made.

Fig. 2

llms.txt adoption by category 76% of consumer brands have an llms.txt file, generated automatically by their platform in the cases we examined. 55% of B2B sites have one, each authored individually. 76% Consumer 55% B2B SaaS PLATFORM-GENERATED BY DEFAULT INDIVIDUALLY AUTHORED
Adoption looks close between the two categories. The difference is who wrote the file: a platform default for most consumer brands, a deliberate choice for every B2B site that has one.

Source: Citehound scan, 29 July 2026, n=64

02

Structural gaps and the trust divide

While the industry focuses on llms.txt, older structural signals are being left unfixed.

Content schema, the markup that tells a machine whether a page is an article, a product or an FAQ, is missing on 91% of consumer-brand homepages and 48% of B2B homepages. Organization schema, the baseline declaration of who a business is, is missing on 24% of consumer-brand homepages and 26% of B2B homepages — on both sides, roughly a quarter of sites have never told a machine they're a company.

Author and about signals are where the two categories split. All 31 B2B sites we scanned passed this check. 27% of consumer brands failed it: businesses selling products to the public while staying anonymous to the systems reading their pages.

Fig. 3

Check failure rates by category, sorted by consumer-brand rate Percentage of sites failing each of 13 checks. Content schema on the homepage fails for 91% of consumer brands and 48% of B2B sites. Single H1 heading fails for 52% of consumer brands and 23% of B2B. Meta description fails for 27% of consumer brands and 29% of B2B. Author or about signals fail for 27% of consumer brands and 0% of B2B. llms.txt absent for 24% of consumer brands and 45% of B2B. Structured data JSON-LD fails for 24% of consumer brands and 16% of B2B. Organization schema fails for 24% of consumer brands and 26% of B2B. Page title fails for 18% of consumer brands and 6% of B2B. Open Graph tags fail for 15% of consumer brands and 6% of B2B. Subheading structure fails for 9% of consumer brands and 10% of B2B. Canonical tag fails for 6% of both. html lang attribute fails for 3% of consumer brands and 0% of B2B. Contact signals fail for 0% of consumer brands and 6% of B2B. Consumer brands B2B SaaS Content schema (homepage) 91 48 Single H1 heading 52 23 Meta description 27 29 Author / about signals 27 0 llms.txt absent 24 45 Structured data (JSON-LD) 24 16 Organization schema 24 26 Page title 18 6 Open Graph tags 15 6 Subheading structure (H2) 9 10 Canonical tag 6 6 html lang attribute 3 0 Contact signals 0 6 Failure rate, % of sites in category. Sorted by consumer-brand rate.
Content schema is the largest single gap for consumer brands. llms.txt is the one check where B2B fails more often than consumer brands — the reverse of the adoption story in Fig. 2.

Source: Citehound scan, 29 July 2026, n=64

The pillar scores make the split explicit. B2B and consumer brands score almost the same on discoverability, 27 out of 40 for both, and on technical foundation, 18 against 17 out of 20. The entire gap between the two categories sits in content and trust: 33 out of 40 for B2B against 28 for consumer brands.

Fig. 4

Pillar scores against maximum, by category Discoverability, maximum 40 points: B2B scores 27, consumer brands score 27. Technical foundation, maximum 20 points: B2B scores 18, consumer brands score 17. Content and trust, maximum 40 points: B2B scores 33, consumer brands score 28. Discoverability · max 40 B2B 27/40 DTC 27/40 Technical foundation · max 20 B2B 18/20 DTC 17/20 Content & trust · max 40 B2B 33/40 DTC 28/40 Filled = points earned. Unfilled = gap to maximum.
Discoverability and technical foundation are close to identical between the two categories. The entire difference sits in content and trust.

Source: Citehound scan, 29 July 2026, n=64

Canonical tags, page titles and sitemaps are in reasonable shape across the board for both categories.

Both categories know how to satisfy a search engine. The signals that tell an AI system who to trust are thinner.

03

The unreachable sites

22 of the 85 sites we attempted could not be scanned: 8 of 39 B2B sites and 14 of 46 consumer-brand sites. These sites were reachable in a browser. They were not down.

What stopped the scan was bot protection that rejects traffic from cloud data centers. Our scanner runs on one. So do GPTBot, ClaudeBot and PerplexityBot.

Fig. 5

Sites that could not be scanned, by category 8 of 39 B2B sites attempted could not be scanned, 21%. 14 of 46 consumer-brand sites attempted could not be scanned, 30%. B2B SaaS — 8 of 39 attempted 21% Consumer brands — 14 of 46 attempted 30% Filled = could not be scanned. Unfilled = completed.
Consumer brands were unreachable at a higher rate than B2B, 30% against 21%, out of sites we attempted to scan.

Source: Citehound scan, 29 July 2026, n=85 attempted

We don't know whether the protection blocking our scanner also blocks AI crawlers. That's an open question, not a finding, and it's the one we're least able to answer from this data set alone.

Methodology and limitations

What we scanned

64 domains completed a scan, out of 85 attempted: 31 B2B CRM and software sites, 33 consumer and DTC brands.

What we pulled

For each domain: robots.txt, sitemap.xml, llms.txt and the homepage, retrieved live on 29 July 2026.

What we scored

13 checks across three categories: AI discoverability, technical foundation, and trust signals.

What this measures

Whether AI systems can reach a site, and whether what they find is structured well enough to establish trust. It does not measure whether a named assistant currently mentions a brand in a live answer. Model responses vary by prompt, session and training data, and no scan of public files can predict them.

What this doesn't tell you

  • Whether the 22 unreachable sites block AI crawlers specifically, or only generic cloud traffic. We could not distinguish between the two from outside the site.
  • Whether the Shopify-generated llms.txt pattern holds beyond the two files we read by hand. We inspected Magic Spoon and Allbirds; we did not audit every Shopify-detected file individually.
  • Whether any of these signals correlate with being cited in an AI answer. This report measures inputs, not outcomes.

How to cite this

Andaç Üzel, “Who Controls Your Brand's AI Representation?,” Citehound Research, 29 July 2026. https://getcitehound.com/research/llms-txt-adoption-2026

This report will be updated as we scan more categories. If you want to see where your own site lands against these numbers, run the free scan.

Scan your site free