Research · Launch Edition 2026

State of AI Readiness

How ready are 219 well-known websites for AI crawlers and assistants? A benchmark across six categories, scanned 29 to 31 July 2026.

Findings

Four findings.

From 219 scored sites. Each figure is a count or an average of our own scan, on the dates shown. The score measures AI readiness: crawler access and on-page signals. It does not measure whether any assistant names or cites a site.

  1. Scanned 29 to 31 July 2026

    The typical site scores 73 out of 100.

    The mean is 73 and the median 75. The middle half of the 219 scored sites sit between 67 and 82, and the range runs from 30 to 100.

  2. Scanned 29 to 31 July 2026

    B2B software scored higher than consumer categories, 76 against 69.

    In this sample the three B2B software categories (125 sites) average 76.4 and the three consumer-facing ones (94 sites) 68.6, a gap of 7.8 points. Differences of a point or two between single categories are not worth reading into.

  3. Scanned 29 to 31 July 2026

    The most common gap is homepage content schema, missing from 76.7% of sites.

    No Article, FAQ, Product or similar type on the homepage. The scan reads only the homepage, so a missing type there is not by itself a defect. It is where the most points were left on the table.

  4. Scanned 29 to 31 July 2026

    14 of 219 sites (6.4%) block at least one of ten tracked AI crawlers outright.

    One site blocks all ten. For most sites, an ordinary Disallow rule on a path is what applies to AI crawlers, which is a limit, not a block.

How we measured it

  • 219 hand-picked, well-known websites in six categories: CRM software, cybersecurity, DevTools and cloud, DTC brands, consumer apps, and hospitality and travel.
  • One scan of each homepage and its public files (robots.txt, llms.txt, the sitemap), with the 16 checks of the public methodology: discoverability 40 points, technical 20, content and trust 40.
  • CRM software and DTC brands were scanned on 29 July 2026, the other four categories on 31 July. 278 sites were attempted; the 59 that could not be scored (21.2%) are left out of every figure.

What it does not show

  • Presence in AI answers. The scan measures readiness, not what an assistant says. No number links to citations, answers or traffic.
  • The web, or an industry. The sites are hand-picked, so "sites in our sample" is the widest claim the numbers support. There are no confidence intervals or significance claims.
  • Cause and effect. A check is worth some points on our scale. That says nothing about what fixing it changes outside the score.
  • A trend. Each site was scanned once, on three days in July. Nothing here is rising or falling.
  • Sites that could not be read. 59 of 278 are missing from every figure, 43 because robots.txt could not be read. The averages describe sites our scanner could read.

The scoring is public: the methodology lists all 16 checks and their points.

Free report

Get the full report by email.

Eight pages: the typical score, the six categories, where the points go, AI crawler access, signals by category, the sites we could not scan, and the limits. Free. We send a link to the PDF.

What we store: your address, the time you asked and whether you ticked the second box. We keep it for up to 12 months, to email you the report link and so that you can ask again, then delete it. The link in the email works for 30 days and contains no address. We send at most three report emails a day to one address. Every email has a link that deletes your address at once. The email carries no tracking pixel and no open or click tracking, and it is sent through Resend. The second box is separate and off unless you tick it. See the privacy page.

Want your own site’s numbers instead? Run the free scan. It applies the same 16 checks to your homepage.