AI crawlers and robots.txt: who to allow
There are two kinds of AI crawler. Answer bots fetch pages to cite them; training bots copy text to train models. Treat them separately, and never confuse either with Googlebot.
The short answer
Allow the answer bots if you want to appear in AI answers. Decide about the training bots on your own terms; blocking them keeps your text out of future models but does not remove you from answers. Leave Googlebot alone.
The bots, by what they do
| Crawler | Owner | Used for |
|---|---|---|
OAI-SearchBot |
OpenAI | ChatGPT search results and citations (answers) |
ChatGPT-User |
OpenAI | Pages a ChatGPT user asks about, fetched live (answers) |
GPTBot |
OpenAI | Training data (training) |
Claude-User |
Anthropic | Pages a Claude user asks about, fetched live (answers) |
ClaudeBot |
Anthropic | Training data and Claude web answers (training + answers) |
PerplexityBot |
Perplexity | Perplexity index and citations (answers) |
Google-Extended |
Gemini training and grounding, not Google Search (training) | |
Applebot-Extended |
Apple | Apple Intelligence training (training) |
The lines
To allow answer bots and block training bots:
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: Claude-User
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: GPTBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Applebot-Extended
Disallow: /
ClaudeBot is the awkward one: it serves both training and Claude’s web answers, so blocking it removes you from those answers too.
Check before you change
Read your robots.txt first: a User-agent: * rule already applies to every bot you have not named. A site that blocks everything from a staging rule will block the answer bots too, and nothing in Search Console will tell you.
MonoRanks reads the file, shows the effect per bot and gives you the exact lines for your choice; on WordPress it can write them for you. See GEO readiness.
Sources
Common questions
Does blocking GPTBot hurt my Google rankings?
No. GPTBot is OpenAI’s training crawler; Google Search uses Googlebot, which these rules do not touch.
What is Google-Extended?
A control for whether Google may use your pages for Gemini training and grounding. It does not affect Google Search or AI Overviews.
Do all bots obey robots.txt?
The named ones from OpenAI, Anthropic, Google, Apple and Perplexity say they do. Unknown scrapers may not; that is a firewall question, not a robots.txt one.
About one email a month
What changed in search and AI answers, from what we see in audits. No card, no pitch, one click to leave.