CiteScore

Crawlers

Three robots.txt tokens the audit actually reads

How the audit reads robots.txt for GPTBot, PerplexityBot, and Google-Extended: allow, disallow, or unspecified. What each token does not control.

3 October 2026 · 8 min read

A CiteScore audit requests /robots.txt on the host you paid for and records three policies: GPTBot, PerplexityBot, and Google-Extended. Each one is allow, disallow, or unspecified. That trio is the AI-crawler dimension, together with whether /llms.txt returned a non-empty body. We do not crawl the internet looking for who linked to you, and we do not pretend a robots.txt edit is a citation.

The tokens, and the ones we skip

TokenWhat the vendor says it isIn this audit
GPTBotOpenAI’s crawler for content that may be used in training. Separate from ChatGPT-User (a person clicked something that fetches) and OAI-SearchBot (search).Parsed. A disallow lowers the crawler dimension. When that dimension is poor, the ChatGPT row is Absent. ChatGPT-User and OAI-SearchBot are not parsed.
PerplexityBotPerplexity’s crawler. Perplexity-User is the separate agent for fetches a person triggered.Parsed. Perplexity-User is not. The Perplexity row is still mostly evidence and markup. The bot policy is the crawler dimension, not the footnote status by itself.
Google-ExtendedGoogle’s token for whether content helps improve Gemini Apps and Vertex AI generative APIs. Google says it does not change inclusion or ranking in Google Search.Parsed, and a disallow lowers the crawler dimension. The AI Overviews row does not read this token. It reads answer-shaped copy and FAQ / entity markup.

We also note whether User-agent: * blocks /. A blanket disallow is recorded, and we still fetch the URL you paid to audit — you asked us to — with a warning on the crawl. Googlebot, Bingbot, ClaudeBot, Applebot-Extended, and CCBot are not in the three-token list. If your only question is “does Googlebot index this?”, use Search Console.

Allow, disallow, unspecified

The parser is ordinary robots.txt, applied to the path /. An agent-specific group wins. If that agent has no group, we use User-agent: *. If neither exists, the result is unspecified. An empty Disallow: is treated as allow, which matches the robots rule that an empty disallow means “nothing is disallowed.”

A file that only names GPTBot, with no * group, leaves PerplexityBot and Google-Extended unspecified. Unspecified does not subtract the disallow penalty. It also does not earn the allow bonus. Silence is a third state, and the report should show it as silence.

A block that matches “we want to be fetched”

This is an illustration, not a requirement. Leave Googlebot’s rules alone while you edit these. If you already opt out of a training crawler on purpose, do not paste an Allow over that decision to chase a dimension score.

User-agent: GPTBot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Google-Extended
Allow: /

The own-goal we see in audits is the inverse, copied from a staging template and left in production:

User-agent: GPTBot
Disallow: /

User-agent: *
Disallow: /

The second group blocks everyone who falls through to *, including the two tokens you did not name. Fix the file, then re-read the copy. A crawler score cannot compensate for a homepage that never defines the product.

What else the crawl fetches

Besides robots.txt and llms.txt, we request /sitemap.xml and treat a response that contains a urlset as a freshness signal. We do not walk every URL in the sitemap. The page set is the URL you submitted plus a short list of same-host candidates (about, pricing, FAQ, and similar) capped so the audit stays a single report. Login walls and private docs are out of scope. Privacy covers what an optional language model is allowed to see: the public extract, not your passwords.

Questions about the tokens

What does unspecified mean?

We did not find a matching group for that agent, and there was no User-agent: * group to fall back to. A missing robots.txt also comes back unspecified. Unspecified is not the same as Disallow.

If User-agent: * allows the site, do the three tokens show allow?

Yes. When an agent has no group of its own, we apply the * group. An Allow, or a * group that does not block /, is recorded as allow. A * group with Disallow: / is recorded as disallow for each of the three tokens.

Will you tell me to block training bots?

No. Opting out of GPTBot or Google-Extended is a policy choice. The audit records it and, for a disallow, lowers the crawler dimension. If the block is intentional, ignore that slice of the score. If it is a leftover from a staging template, remove it before you read the ChatGPT row as a verdict on your copy.