AI crawlers and robots.txt: who to allow

There are two kinds of AI crawler. Answer bots fetch pages to cite them; training bots copy text to train models. Treat them separately, and never confuse either with Googlebot.

The short answer

Allow the answer bots if you want to appear in AI answers. Decide about the training bots on your own terms; blocking them keeps your text out of future models but does not remove you from answers. Leave Googlebot alone.

The bots, by what they do

Crawler Owner Used for
OAI-SearchBot OpenAI ChatGPT search results and citations (answers)
ChatGPT-User OpenAI Pages a ChatGPT user asks about, fetched live (answers)
GPTBot OpenAI Training data (training)
Claude-User Anthropic Pages a Claude user asks about, fetched live (answers)
ClaudeBot Anthropic Training data and Claude web answers (training + answers)
PerplexityBot Perplexity Perplexity index and citations (answers)
Google-Extended Google Gemini training and grounding, not Google Search (training)
Applebot-Extended Apple Apple Intelligence training (training)

The lines

To allow answer bots and block training bots:

User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: Claude-User
Allow: /
User-agent: PerplexityBot
Allow: /

User-agent: GPTBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Applebot-Extended
Disallow: /

ClaudeBot is the awkward one: it serves both training and Claude’s web answers, so blocking it removes you from those answers too.

Check before you change

Read your robots.txt first: a User-agent: * rule already applies to every bot you have not named. A site that blocks everything from a staging rule will block the answer bots too, and nothing in Search Console will tell you.

MonoRanks reads the file, shows the effect per bot and gives you the exact lines for your choice; on WordPress it can write them for you. See GEO readiness.

Sources

Common questions

Does blocking GPTBot hurt my Google rankings?

No. GPTBot is OpenAI’s training crawler; Google Search uses Googlebot, which these rules do not touch.

What is Google-Extended?

A control for whether Google may use your pages for Gemini training and grounding. It does not affect Google Search or AI Overviews.

Do all bots obey robots.txt?

The named ones from OpenAI, Anthropic, Google, Apple and Perplexity say they do. Unknown scrapers may not; that is a firewall question, not a robots.txt one.

About one email a month

What changed in search and AI answers, from what we see in audits. No card, no pitch, one click to leave.