GeoPack 0.1.0 · technical snapshot · not a ranking product

AI-crawler readiness pack

Target: https://www.cloudflare.com/
Generated 2026-08-16 16:00 CDT (2026-08-16 21:00 UTC)

Honest scope. This is a point-in-time technical snapshot of robots.txt rules, live user-agent probes, drafted llms.txt, and schema/sitemap gaps. It does not measure rankings, citations, or appearance in ChatGPT / Claude / Perplexity / Google AI Overviews. llms.txt does not affect Google Search or AI Overviews.
7/7training tokens allow /
6/6retrieval tokens allow /
32/33UA probes reached (no challenge)
10content pages fetched

Training vs retrieval

Blocking a training crawler (GPTBot, ClaudeBot, CCBot, Google-Extended) is not the same as blocking a retrieval agent (OAI-SearchBot, ChatGPT-User, PerplexityBot, Googlebot). Sites that paste a blanket “block all AI bots” file often disappear from answer-engine citations while still leaking training data — or the reverse. GeoPack splits the two. This is a policy observation, not a promise that anyone will cite you.

robots.txt matrix — training

Evaluated for path /. Longest matching group wins; * is fallback only (RFC 9309-style), not first-match.

TokenOperatorJob/Rule sourceNote
GPTBotOpenAItrainingallow /explicitTraining crawl for OpenAI foundation models.
ClaudeBotAnthropictrainingallow /via *Anthropic training crawler.
anthropic-aiAnthropictrainingallow /explicitLegacy Anthropic robots.txt token. Still appears on many sites.
Google-Extended control token — no HTTP crawlerGoogletrainingallow /explicitrobots.txt control token for Gemini / AI training use of already-crawled content. Googlebot is the HTTP crawler. No Google-Extended user-agent is sent.
Applebot-Extended control token — no HTTP crawlerAppletrainingallow /via *robots.txt control token for Apple Intelligence training. Applebot is the HTTP crawler. No Applebot-Extended user-agent is sent.
BytespiderByteDancetrainingallow /via *ByteDance training crawler. Public reports say it often ignores robots.txt.
CCBotCommon Crawltraining-corpusallow /explicitOpen web corpus. Many labs train on Common Crawl snapshots.

robots.txt matrix — retrieval / user-fetch

TokenOperatorJob/Rule sourceNote
OAI-SearchBotOpenAIsearchallow /via *Indexes pages for ChatGPT search results. Independent of GPTBot.
ChatGPT-UserOpenAIuser-fetchallow /explicitOn-demand fetch when a ChatGPT user asks about a page. May ignore robots.txt.
PerplexityBotPerplexitysearchallow /explicitSearch indexer for Perplexity answers. Not a foundation-model trainer.
Google-CloudVertexBotGooglegroundingallow /via *Fetches pages to ground Vertex AI / Gemini Enterprise agents.
AmazonbotAmazonsearch-and-assistantallow /via *Amazon search / assistant crawler.
GooglebotGooglesearchallow /via *Classic Google Search crawler. Shown for context, not an AI-training bot.

Context tokens

TokenOperator/Reason
*defaultallow /inherited from *

robots.txt: https://www.cloudflare.com/robots.txt → HTTP 200. Content-Signal: ai-train=yes, search=yes, ai-input=yes.

Live user-agent probes

Homepage plus a few important URLs, fetched as each crawler UA. Allowed even when robots.txt would block that bot — the question is whether the site (WAF/CDN) blocks them. Gentle: few URLs, timeouts, sleep between requests. Not a full crawl.

Reached without challenge: 32 · Challenge-like body/headers: 0 · Error or other 4xx/5xx: 1

URLUA tokenPurposeStatusServerWAF hintsTime
https://www.cloudflare.com/GPTBottraining200cloudflarecf-ray=a2c35f36abfc23a3-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}386 ms
https://www.cloudflare.com/OAI-SearchBotretrieval200cloudflarecf-ray=a2c35f3f79dcc3af-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}502 ms
https://www.cloudflare.com/ChatGPT-Userretrieval200cloudflarecf-ray=a2c35f48ef04b338-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}408 ms
https://www.cloudflare.com/ClaudeBottraining200cloudflarecf-ray=a2c35f51a986ff1a-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}499 ms
https://www.cloudflare.com/anthropic-aitraining200cloudflarecf-ray=a2c35f5b0c085ed8-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}547 ms
https://www.cloudflare.com/PerplexityBotretrieval200cloudflarecf-ray=a2c35f64bf64dd76-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}532 ms
https://www.cloudflare.com/Google-CloudVertexBotretrieval200cloudflarecf-ray=a2c35f6e4f46e102-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}493 ms
https://www.cloudflare.com/Bytespidertrainingerror SSLError: HTTPSConnectionPool(host='www.cloudflare.com', port=443): Max retries exceeded with url: / (Caused by SSLError(SSLError(1, '[SSL: WRONG_VERSION_NUMBER] wrong version number (_ssl.c:1029)')))12 ms
https://www.cloudflare.com/CCBottraining200cloudflarecf-ray=a2c35f7df90968f1-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}486 ms
https://www.cloudflare.com/Amazonbotretrieval200cloudflarecf-ray=a2c35f874a7787cd-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}425 ms
https://www.cloudflare.com/Googlebotretrieval200cloudflarecf-ray=a2c35f9029f1ef34-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}518 ms
https://www.cloudflare.com/productsGPTBottraining200cloudflarecf-ray=a2c35f99c91a5ef1-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}469 ms
https://www.cloudflare.com/productsOAI-SearchBotretrieval200cloudflarecf-ray=a2c35fa2ed871561-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}62 ms
https://www.cloudflare.com/productsChatGPT-Userretrieval200cloudflarecf-ray=a2c35fa99ca3de1d-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}66 ms
https://www.cloudflare.com/productsClaudeBottraining200cloudflarecf-ray=a2c35fb039cff3b4-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}64 ms
https://www.cloudflare.com/productsanthropic-aitraining200cloudflarecf-ray=a2c35fb6edd89b30-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}61 ms
https://www.cloudflare.com/productsPerplexityBotretrieval200cloudflarecf-ray=a2c35fbd8d8eabbf-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}59 ms
https://www.cloudflare.com/productsGoogle-CloudVertexBotretrieval200cloudflarecf-ray=a2c35fc42fc3ff13-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}61 ms
https://www.cloudflare.com/productsBytespidertraining200cloudflarecf-ray=a2c35fcac9dab024-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}64 ms
https://www.cloudflare.com/productsCCBottraining200cloudflarecf-ray=a2c35fd179b92095-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}63 ms
https://www.cloudflare.com/productsAmazonbotretrieval200cloudflarecf-ray=a2c35fd81b826c24-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}67 ms
https://www.cloudflare.com/productsGooglebotretrieval200cloudflarecf-ray=a2c35fdeba104518-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}151 ms
https://www.cloudflare.com/products/GPTBottraining200cloudflarecf-ray=a2c35fe5ddd68969-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}50 ms
https://www.cloudflare.com/products/OAI-SearchBotretrieval200cloudflarecf-ray=a2c35fec7dde9976-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}53 ms
https://www.cloudflare.com/products/ChatGPT-Userretrieval200cloudflarecf-ray=a2c35ff30a6f68f1-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}52 ms
https://www.cloudflare.com/products/ClaudeBottraining200cloudflarecf-ray=a2c35ff999339d6b-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}48 ms
https://www.cloudflare.com/products/anthropic-aitraining200cloudflarecf-ray=a2c360002b87ef34-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}63 ms
https://www.cloudflare.com/products/PerplexityBotretrieval200cloudflarecf-ray=a2c36006ccfd1f9b-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}49 ms
https://www.cloudflare.com/products/Google-CloudVertexBotretrieval200cloudflarecf-ray=a2c3600d5a306c24-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}49 ms
https://www.cloudflare.com/products/Bytespidertraining200cloudflarecf-ray=a2c36013ea97b7b9-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}48 ms
https://www.cloudflare.com/products/CCBottraining200cloudflarecf-ray=a2c3601a783b5f03-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}52 ms
https://www.cloudflare.com/products/Amazonbotretrieval200cloudflarecf-ray=a2c360210c6c1571-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}147 ms
https://www.cloudflare.com/products/Googlebotretrieval200cloudflarecf-ray=a2c360283bd5ad66-PDX, server=cloudflare, nel={"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}51 ms

Sitemap (capped)

Sitemaps used: https://www.cloudflare.com/sitemap.xml. Listed 10 HTML URLs (cap 10).

Pages fetched for titles / H1 / JSON-LD

Our own crawl uses the GeoPack user-agent and respects robots.txt. Cap is small on purpose.

URLTitleH1JSON-LD types
https://www.cloudflare.com/Cloudflare: Build for the agent eraEverything we learned from powering 20% of the Internet—yours by defaultOrganization, SearchAction, WebPage, WebSite
https://www.cloudflare.com/plans/PricingScale predictablyOrganization, SearchAction, WebPage, WebSite
https://www.cloudflare.com/products/Products | CloudflareBuild without boundariesOrganization, SearchAction, WebPage, WebSite
https://blog.cloudflare.com/Cloudflare BlogCloudflare BlogWebSite
https://www.cloudflare.com/case-studies/Customer Case Studies | CloudflareOrganization, SearchAction, WebPage, WebSite
https://www.cloudflare.com/learning/Just a moment...
https://www.cloudflare.com/careers/Cloudflare Careers | CloudflareBuild without boundariesOrganization, SearchAction, WebPage, WebSite
https://www.cloudflare.com/careers/jobs/Careers at Cloudflare — Open Positions | CloudflareHelp Us Build a Better InternetOrganization, SearchAction, WebPage, WebSite
https://www.cloudflare.com/partners/Partners | CloudflareBuild without boundariesOrganization, SearchAction, WebPage, WebSite
https://www.cloudflare.com/partners/cloud-and-platform/oracle/Oracle Cloud Infrastructure Partner | CloudflareBuild without boundariesOrganization, SearchAction, WebPage, WebSite

Schema + sitemap gaps

Types observed: Organization, SearchAction, WebPage, WebSite

Common types not seen:

Schema presence is a technical observation from a tiny crawl. It is not a ranking, rich-result, or citation prediction.

Drafted llms.txt

/llms.txt → HTTP 200; /llms-full.txt → HTTP 200

See llms.txt and llms-full.txt in this pack. Every generated link is marked DRAFTED and needs a human edit before you publish it. Publishing the file does not opt you out of training and does not change Google Search.

How to use this pack

  1. Read the training vs retrieval split. Decide a policy (example: block training, allow retrieval — or the reverse).
  2. Compare robots.txt rules to live probe status. A robots Allow plus a WAF challenge means the file is lying.
  3. Edit the drafted llms.txt. Do not ship DRAFTED lines as-is.
  4. Treat schema gaps as a checklist for your developer, not as SEO score.