AI crawlers and robots.txt, one bot per page
Eleven AI crawlers publish official robots.txt tokens across OpenAI, Perplexity, Anthropic and Google. They do one of three jobs: collect training data, index for search, or fetch a page because a user asked. Blocking the wrong group removes a store from AI shopping answers for nothing.
- Start here
Which AI crawlers to allow in robots.txt, and what each one does
The real AI user agents, what each does, and why blocking retrieval bots quietly removes you from AI answers. Includes the one everyone gets wrong.
Every AI crawler user agent, in one table
The exact robots.txt token, full user-agent string and published IP range for every documented AI crawler from OpenAI, Perplexity, Anthropic and Google.
Should a Shopify store block AI crawlers?
Most AI crawler blocklists were written for publishers and cost stores money. Which bots to allow, which are a values call, and the one setting to avoid.
GPTBot: the user agent, the robots.txt rule, and what blocking it costs
GPTBot's exact user-agent string and robots.txt token, what OpenAI uses it for, how it differs from OAI-SearchBot, and how to block it on Shopify.
OAI-SearchBot, GPTBot and ChatGPT-User: which OpenAI agent does what
OpenAI runs four crawlers with four different jobs. The exact tokens and user-agent strings for OAI-SearchBot, GPTBot, ChatGPT-User and OAI-AdsBot, and which to allow.
PerplexityBot and Perplexity-User: two agents, two different rules
PerplexityBot's exact robots.txt token and user-agent string, how it differs from Perplexity-User, published IP ranges, and how to set the rule on Shopify.
ClaudeBot, Claude-User and Claude-SearchBot: Anthropic's three agents
The three Anthropic crawler tokens, what each one does, the shared IP range file, and which to allow if you want Claude to be able to read your store.
Google-Extended: what blocking it actually does
Google-Extended controls Gemini training and grounding, not Search. Google states it is not a ranking signal. What the token does, and what it does not touch.
Cloudflare blocks AI crawlers by default on September 15. Is your store affected?
From 15 September 2026, Cloudflare blocks AI training and agent bots by default on new domains, and legacy Block AI bots settings start catching Googlebot. Who is affected and how to check.
A copy-paste robots.txt for Shopify stores that want AI visibility
A robots.txt.liquid template that keeps Shopify's defaults, allows the AI search and user-fetch agents, and leaves the training bots as an explicit choice. With the reasoning per line.
Are you accidentally invisible? Auditing CDN and firewall bot blocking
Robots.txt can say yes while a CDN says no. How to audit Cloudflare AI Crawl Control, WAF rules and bot protections for silent blocks on the crawlers AI visibility depends on.
