Privacy

What we fetch, what we store, what we don't.

Last updated 30 July 2026.

What the scanner fetches

When you scan a domain, Citehound fetches public files from that domain: robots.txt, sitemap.xml, llms.txt, /.well-known/ucp, and the homepage. If the site looks like a store, a scan also reads one product sitemap and one product page, made with a user agent that names Citehound and only where robots.txt allows it. All of it is public by definition — nothing behind a login, nothing private. We fetch whatever domain you type in, whether or not it's yours.

The two free tools fetch less. The llms.txt checker reads one file, llms.txt, from the domain you enter. The schema generator reads that domain's homepage for its title, description and language. Neither keeps what it reads, and both share the scanner's rate limit described below.

What we store

Almost nothing, today. There are no accounts, no email collection, and no scan history. A scan result exists in your browser for as long as the page stays open.

The exception is a full-site crawl, the engine behind the planned Pro report. It is not part of the public scanner yet, but its address is open. A crawl stores the domain, the list of pages it chose, and each page’s result, under a random 128-bit id, for 90 days after its last step. Nothing about who started it is stored with it, and nobody can list or search crawls: you need the id.

The one exception: to stop abuse of the scanner, we keep a one-way hash of your IP address paired with a request counter, purely to enforce a rate limit. It expires automatically within a day (the hourly counter within an hour), isn't linked to which domains were scanned, and is never stored as a raw IP address.

The MCP server keeps short-lived counters keyed by the scanned domain, kept for an hour, to limit repeated fetches of one site. They hold no caller identity.

Separately, we count visits from known AI crawlers (GPTBot, ClaudeBot, PerplexityBot and similar) to our own site, so we can publish real data on how often they show up. This is two aggregate counts: visits per crawler per day, and how often each page of our site was requested by those crawlers. It is never a per-visit log, and nothing about human visitors is recorded at all.

Will change

This section will be updated when score history ships. Until then, nothing is retained after your browser tab closes.

What we don't collect

No personal data. No email addresses. No accounts. No advertising trackers.

Third parties

Hosting: Vercel. The site and the scan function both run on Vercel's infrastructure.

Executive summaries: when a full-site crawl finishes, the report’s executive summary is written from aggregated figures: scores, counts of pages and the names of the checks. Those figures are sent to a model provider, Google’s Gemini API, to be phrased. No page content, no page addresses, no domain name and no personal data are sent. Google’s terms for unpaid use of the API may allow it to use what it receives to improve its products. If the model is unavailable, or its reply fails our checks, the summary is written by rules here and nothing is sent.

Placeholder

[Analytics tool name and what it collects will be added here when analytics is implemented. Nothing is in place yet.]

Contact

Questions about this policy: hey@getcitehound.com.

See what a scan of your own site looks like.

Scan your site free