How to let AI crawlers in: robots.txt for GPTBot, ClaudeBot and PerplexityBot
A practical guide to the AI user agents that matter, what each one does, and a robots.txt you can copy that lets assistants cite you without giving away the whole site.
By Seoptist Team · · 4 min read
In our audits, roughly one site in four blocks at least one major AI crawler, and most of the owners do not know it. The usual cause is a Disallow: / rule added for a scraper years ago, or a security plugin that bundles "block AI bots" as a default. The effect is that ChatGPT, Claude and Perplexity cannot read your pages when a customer asks about you, so they answer from stale training data or from a directory that lists you incorrectly.
This guide explains the user agents, what each one is for, and gives you a robots.txt to copy.
The AI user agents that matter
Each provider now runs more than one bot, and they do different jobs. Blocking one does not block the others.
OpenAI
GPTBotcollects content for training future models.OAI-SearchBotindexes pages for ChatGPT search results and citations.ChatGPT-Userfetches a page on demand when a user asks ChatGPT about it.
Anthropic
ClaudeBotcrawls for training.Claude-SearchBotindexes for Claude search features.Claude-Userfetches a page on demand during a conversation.
Perplexity
PerplexityBotindexes pages for Perplexity answers.Perplexity-Userfetches on demand when a user asks about a specific page.
Googlebotis unchanged and feeds both Search and AI Overviews. You cannot opt out of AI Overviews without leaving Search.Google-Extendedis a control token, not a crawler. Disallowing it stops your content being used to train Gemini models but has no effect on Search or AI Overviews.
Others worth knowing: Bingbot (Copilot answers use Bing's index), Applebot-Extended (Apple Intelligence training), CCBot (Common Crawl, used by many research datasets), Amazonbot and Meta-ExternalAgent.
Training versus retrieval: decide what you actually want
There are two distinct questions and you can answer them separately.
Do you want your content used to train models? Blocking training bots (GPTBot, ClaudeBot, Google-Extended) is a legitimate choice for publishers who sell content. For most businesses it is a mistake: being in the training data is how a model learns that you exist, what you do and where you are.
Do you want assistants to retrieve and cite your pages when a user asks? For almost every business the answer is yes, and that means allowing the search and user-triggered bots (OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User).
Note that the on-demand "User" agents generally ignore robots.txt by design, because a human asked for that specific page. Blocking them has little effect; the point of listing them is to be explicit.
A robots.txt you can copy
This version allows everything that helps you be recommended, keeps private areas out, and points to the sitemap. Adjust the disallowed paths to your site.
# Default rules for all crawlers
User-agent: *
Allow: /
Disallow: /admin/
Disallow: /account/
Disallow: /cart/
Disallow: /search
# AI assistants: allow indexing and retrieval so they can cite us
User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
Allow: /
Disallow: /admin/
Disallow: /account/
Sitemap: https://www.example.co.uk/sitemap.xml
If you do want to opt out of training but stay citable, replace the block above with explicit rules: allow OAI-SearchBot, Claude-SearchBot and PerplexityBot, and disallow GPTBot, ClaudeBot and Google-Extended.
Things that silently undo your robots.txt
A permissive robots.txt is not enough if something upstream blocks the request.
- CDN and WAF bot rules. Cloudflare, Akamai and others offer "block AI bots" toggles. Check the firewall events log for 403 responses to the user agents above.
- Rate limiting. AI crawlers can be aggressive. Rather than blocking, set a
Crawl-delay(respected by some) or use your CDN to throttle. - JavaScript-only content. Most AI crawlers do not execute JavaScript. If your service descriptions render client-side, the bot sees an empty page. Server-side rendering or static generation fixes this.
- Meta robots and X-Robots-Tag. A
noindexornoaiheader on key pages overrides robots.txt for the bots that honour it. - Login walls and cookie interstitials that return a different page to bots.
How to verify
Fetch your own pages with the user agent string and check the status code and body. For example:
curl -A "PerplexityBot" -I https://www.example.co.uk/services/
curl -A "GPTBot" -s https://www.example.co.uk/ | grep -i "<title>"
A 200 response with your real HTML means you are reachable. A 403, a challenge page or an empty body means something is in the way.
Seoptist's audit checks all of this automatically: robots.txt rules per AI user agent, WAF blocks, JavaScript rendering and the presence of llms.txt. Each problem appears in your checklist with the exact lines to change.