Key findings
- 1
76% of consumer brands have an llms.txt file, but the two files we examined by hand were near-identical and said nothing specific about the brand.
- 2
91% of consumer-brand homepages and 48% of B2B homepages are missing content schema, the markup that tells AI systems what type of content they're looking at.
- 3
22 of 85 attempted scans could not complete, blocked by protection that may also stop AI crawlers — a question this report leaves open.
Before the specifics: both categories land in a similar place on the 100-point score. Consumer brands trail B2B by 7 points on average, and the two ranges overlap across most of their span.
Fig. 1
Source: Citehound scan, 29 July 2026, n=64
01
The llms.txt illusion
llms.txt is a young standard, barely two years old. It tells AI systems what a site is about and why it should be trusted. On paper, consumer brands have adopted it fast: 76% of the DTC sites we scanned have one, against 55% of B2B sites.
We read two of those consumer files by hand: Magic Spoon and Allbirds. The structure, section order and phrasing were nearly identical. Only the brand name changed.
The reason is Shopify. The platform generates an llms.txt automatically for every store it hosts, which explains why adoption looks high across consumer brands without any of them writing a word themselves. In the two files we examined, each close to 1,000 words, there was no mention of what the brand sells, who it's for, or what sets it apart. Instead, the file gives an AI agent Shopify's own shopping-skill and commerce-protocol instructions.
An AI system reading one of these files learns about Shopify, not about the brand behind the store.
B2B tells a different story. 55% of the B2B sites we scanned have an llms.txt, and no single platform is generating it for them. Every one of those files reflects a choice someone made.
Fig. 2
Source: Citehound scan, 29 July 2026, n=64
02
Structural gaps and the trust divide
While the industry focuses on llms.txt, older structural signals are being left unfixed.
Content schema, the markup that tells a machine whether a page is an article, a product or an FAQ, is missing on 91% of consumer-brand homepages and 48% of B2B homepages. Organization schema, the baseline declaration of who a business is, is missing on 24% of consumer-brand homepages and 26% of B2B homepages — on both sides, roughly a quarter of sites have never told a machine they're a company.
Author and about signals are where the two categories split. All 31 B2B sites we scanned passed this check. 27% of consumer brands failed it: businesses selling products to the public while staying anonymous to the systems reading their pages.
Fig. 3
Source: Citehound scan, 29 July 2026, n=64
The pillar scores make the split explicit. B2B and consumer brands score almost the same on discoverability, 27 out of 40 for both, and on technical foundation, 18 against 17 out of 20. The entire gap between the two categories sits in content and trust: 33 out of 40 for B2B against 28 for consumer brands.
Fig. 4
Source: Citehound scan, 29 July 2026, n=64
Canonical tags, page titles and sitemaps are in reasonable shape across the board for both categories.
Both categories know how to satisfy a search engine. The signals that tell an AI system who to trust are thinner.
03
The unreachable sites
22 of the 85 sites we attempted could not be scanned: 8 of 39 B2B sites and 14 of 46 consumer-brand sites. These sites were reachable in a browser. They were not down.
What stopped the scan was bot protection that rejects traffic from cloud data centers. Our scanner runs on one. So do GPTBot, ClaudeBot and PerplexityBot.
Fig. 5
Source: Citehound scan, 29 July 2026, n=85 attempted
We don't know whether the protection blocking our scanner also blocks AI crawlers. That's an open question, not a finding, and it's the one we're least able to answer from this data set alone.
Methodology and limitations
What we scanned
64 domains completed a scan, out of 85 attempted: 31 B2B CRM and software sites, 33 consumer and DTC brands.
What we pulled
For each domain: robots.txt, sitemap.xml, llms.txt and the homepage, retrieved live on 29 July 2026.
What we scored
13 checks across three categories: AI discoverability, technical foundation, and trust signals.
What this measures
Whether AI systems can reach a site, and whether what they find is structured well enough to establish trust. It does not measure whether a named assistant currently mentions a brand in a live answer. Model responses vary by prompt, session and training data, and no scan of public files can predict them.
What this doesn't tell you
- Whether the 22 unreachable sites block AI crawlers specifically, or only generic cloud traffic. We could not distinguish between the two from outside the site.
- Whether the Shopify-generated llms.txt pattern holds beyond the two files we read by hand. We inspected Magic Spoon and Allbirds; we did not audit every Shopify-detected file individually.
- Whether any of these signals correlate with being cited in an AI answer. This report measures inputs, not outcomes.
How to cite this
Andaç Üzel, “Who Controls Your Brand's AI Representation?,” Citehound Research, 29 July 2026. https://getcitehound.com/research/llms-txt-adoption-2026
This report will be updated as we scan more categories. If you want to see where your own site lands against these numbers, run the free scan.
Scan your site free