Methodology

How the score is built.

Every point value below is exactly what runs in the live scan. Nothing here is rounded, hidden or simplified for marketing.

The three pillars

The 100-point score splits into three weighted pillars. The weighting is deliberate, not arbitrary.

Discoverability — 40 points

Carries the most weight because it gates everything else. A page an AI crawler cannot read cannot be evaluated on any other signal — access is the precondition, not one factor among many. AI crawler access alone is worth 26 of these 40 points: blocking even one major crawler removes you from that platform's answers entirely, which is a bigger loss than any single on-page fix can offset.

robots.txt present4 pts
llms.txt present4 pts
Sitemap declared6 pts
AI crawler access (10 crawlers)26 pts

Technical foundation — 20 points

Baseline machine-readability hygiene. These checks rarely make or break a citation on their own, which is why they carry the least weight, but a page failing most of them signals a site nobody has ever prepared for machine readers.

Canonical tag4 pts
html lang attribute3 pts
Page title3 pts
Meta description4 pts
Open Graph tags3 pts
Structured data (JSON-LD)3 pts

Content & trust — 40 points

Equal weight to Discoverability, deliberately. Access without credibility just means a model can read you without trusting you enough to cite you. These checks are the ones that tell a model who you are and whether your claims are verifiable.

Single H1 heading6 pts
Subheading structure (H2)5 pts
Organization / WebSite schema8 pts
Content schema (Article, FAQ…)5 pts
Author / about signals8 pts
Contact signals8 pts

Commerce only · scored separately

Commerce checks

These checks run only when a scan treats the site as a store. They have their own sub-score and are never added to the 100, so a store and a non-store stay comparable and the category benchmarks stay valid.

Detecting a store

A scan looks for five signals. Two or more means it treats the site as a store. The report lists which fired.

Product or Offer schema on the homepagesignal
Store platform markers in the page, its scripts or its headerssignal
Product or collection URLs in the sitemapsignal
A valid UCP merchant profile at /.well-known/ucpsignal
A cart or checkout link on the homepagesignal

The sitemap signal needs a real share of the file, not one matching URL: at least 5 product or collection URLs making up 20% of a sitemap, or a child sitemap named for products. Many software sites have a /products/ section of marketing pages, and a single match would have called them stores.

Platform markers cover Shopify, WooCommerce, BigCommerce, Magento and Salesforce Commerce Cloud. We checked Shopify and Salesforce Commerce Cloud against live stores. The other three use their documented asset paths and have not been checked against a live store of their own.

Commerce readiness — 8 points

UCP endpoint4 pts
Product schema: name1 pt
Product schema: price1 pt
Product schema: availability1 pt
Product schema: image1 pt

The UCP check passes when /.well-known/ucp returns JSON with a ucp block that has a version and declares services or capabilities. The report lists the versions and capabilities it finds. A URL that answers with a web page, or with JSON of another shape, does not pass. A failed request is an absent endpoint, never a failed scan.

The product check opens one page: the first /products/ or /product/ URL in the sitemap, following one product sitemap if the sitemap is an index, or else the first product link on the homepage. It reads Product or ProductGroup schema, using the first variant to fill gaps in a ProductGroup, and the report shows the URL it checked. If no product page can be found, if robots.txt disallows the path for CitehoundBot, or if the page cannot be read, the check is reported as not assessed and left out of the total. It is never counted as a failure.

The sub-score is shown as points earned out of the points that could be assessed, so a store whose product page could not be read shows a total out of 4.

Not scored: llms.txt authorship

For a store, the report also shows what the llms.txt checker finds: whether the file looks like a platform default, custom or unclear. It is a pattern match on the text and is never scored.

How a scan treats your site

The product page is the only page beyond the homepage that a scan reads. It is one request, made last and after a pause, with a user agent that names Citehound (CitehoundBot), and only where robots.txt allows the path. Sites that show no store signal cost no extra requests.

What this detection misses

Detection is a pattern match. On 2 October 2026 we ran the rules over the 156 reachable sites in our benchmark lists. They detected 30 of 34 DTC stores and none of the other 122 sites. The four misses had no platform markers, no product URLs in the sitemap and no cart link on the homepage. The sitemap thresholds were set after two software sites tripped the first version of the rule, so treat that result as a calibration, not an independent test.

What this scan does not measure

A readiness score, not a presence score.

Citehound measures AI readiness: whether AI crawlers can access your content, and whether your on-page signals give a model a reason to trust and cite you. It does not measure whether ChatGPT, Claude, Perplexity or any other assistant currently mentions your brand in a live answer.

That is a different, harder problem. Model responses vary by prompt, by session, and by training snapshot — no scanner can observe a black box or promise a citation. What we can measure, and verify with certainty, is whether the door is open. A perfect score does not guarantee an answer engine will cite you. A blocked crawler guarantees it cannot.

The crawler-access check gives half credit for a bot that has any Disallow rule applying to it, including ordinary paths such as /admin/. A site can lose points here while being open to all of its public content. Citehound’s own site scores 13 of 26 on this check for that reason.

We built the tool that measures the certain thing precisely, instead of the uncertain thing vaguely. That precision is the point.

How the data is collected

Every scan fetches these directly from your domain, live, at scan time: public robots.txt, llms.txt, the UCP profile at /.well-known/ucp, sitemap.xml (if not already declared in robots.txt), and the homepage. If the scan treats the site as a store, it may also read your declared sitemap or one product sitemap, and one product page. All requests happen inside a single Vercel serverless function — nothing routes through a third-party proxy.

Nothing is stored. There is no database behind the scan. The result exists in your browser for as long as the page stays open, and re-running the scan simply fetches everything again.

If robots.txt cannot be read at all — a network failure, not a genuine 404 — the scan stops and returns an error rather than guessing. A missing robots.txt (a real 404) is scored as default-open access, which is the true state of that site; a failed fetch is not, and reporting a guessed score in that case would be a fabricated result. We do not fabricate results.

Changelog

  1. 2026-10-03Fixed how the scanner reads a meta description. Until this date it cut the text at the first straight apostrophe, so a description such as “Scan your site’s AI visibility…” written with a plain ' was read as 14 characters and failed the length check. In the July scans, 7 of 219 sites failed that check with 1 to 49 characters captured, so up to 7 results (four points each) may be false failures. That is an upper bound of about 0.13 points on the overall average and under one point in any single category. The same cut could also have let a description longer than 170 characters pass; we have not measured that direction. The published benchmark data stays at the original scans. Fixed going forward.
  2. 2026-10-03The crawler-access rule, which gives half credit for a bot with any Disallow rule applying to it, is under review for a later scoring version. No scoring has changed.
  3. 2026-10-03Recorded a rescan check. A rescan of the same six lists on 2 October 2026 moved the category averages as follows: CRM software −1, DTC brands −1, Cybersecurity +1, Consumer apps +1, Hospitality 0, DevTools & Cloud +1. About a third of individual sites changed by 3 to 15 points between runs, and the most common cause was the Single H1 heading check. The raw rescan output was not retained in the repository, and the published benchmark data was kept at the original scan.
  4. 2026-10-02Added two commerce-only checks, structured product data and the UCP endpoint, plus store detection. They form a separate sub-score out of 8 and are never added to the 100, so the 16 checks, their point values and the six category benchmarks are unchanged. A scan of a store now reads at most one product page, after the other requests, with a user agent that names Citehound and only where robots.txt allows it.
  5. 2026-10-02Fixed entity decoding in the page title and meta description parser. Encoded characters such as & were counted at full length, which pushed some titles and descriptions over the length limit and failed the Page title or Meta description check. Across the benchmark corpus, 8 check results on 7 sites flipped from fail to pass. The DTC brands average moved from 71 to 72. No other category average changed, and the overall average stays 73.
  6. 2026-07-29Added llms.txt detection as a new check, bringing the total from 15 to 16. Rebalanced Discoverability: robots.txt 4, llms.txt 4, Sitemap declared 6, AI crawler access 26 (down from 30). Discoverability, Technical foundation and Content & trust totals (40 / 20 / 40) are unchanged.

See exactly how your own site scores against this.

Scan your site free