"We block AI bots" is too vague to diagnose anything. Several providers now run separate crawlers for separate jobs, and your robots.txt can treat each one differently. That's a feature: it lets you decide about training independently from search visibility.
#Three purposes
Search and answer indexing. Crawlers that build the index an assistant retrieves from. Blocking them can keep your pages out of AI search results and citations. Examples documented by providers include OpenAI's OAI-SearchBot, Perplexity's PerplexityBot and Anthropic's Claude-SearchBot. Google's AI features in Search rely on the same Google Search systems as ordinary results.
Training. Crawlers that collect content that may be used to train or improve models: OpenAI's GPTBot, Anthropic's ClaudeBot, and the Google-Extended robots token for Gemini training and grounding. Google-Extended isn't a separate crawler, and Google documents that it doesn't affect Search inclusion.
User-triggered fetches. Requests made because a person asked an assistant about a page: ChatGPT-User, Perplexity-User, Claude-User. Providers describe these differently from automatic crawling — check each provider's current documentation for how they treat robots.txt.
#Deciding
There's no universal right answer. A reasonable starting framework:
- Most businesses that want to be recommended allow search and user-fetch crawlers.
- Training is a values-and-contracts decision. Some publishers restrict it; others allow it for reach. Decide deliberately and revisit it.
- Private or paid content belongs behind authentication. robots.txt is advisory, not access control.
An example that allows search and user-fetch bots but blocks two training crawlers:
User-agent: OAI-SearchBot
User-agent: PerplexityBot
User-agent: Claude-SearchBot
User-agent: ChatGPT-User
User-agent: Perplexity-User
User-agent: Claude-User
Allow: /
User-agent: GPTBot
User-agent: ClaudeBot
Disallow: /
Multiple User-agent lines before a rule set share that rule set, which keeps the file readable.
#Common mistakes
- A blanket
User-agent: */Disallow: /left over from a staging environment. - Blocking by firewall instead of robots.txt. A security rule that returns 403 to anything containing "bot" is invisible in robots.txt and breaks search crawlers too. Verify a crawler's identity (providers publish IP ranges or verification methods) before allowing traffic through.
- Assuming access equals inclusion. Letting a crawler in is a precondition. Retrieval and citation are separate decisions the engine makes.
- Forgetting JavaScript. Many retrieval crawlers read the HTML a server returns. Content that only exists after scripts run may never be seen.
#Check yours
Our free AI crawler checker fetches your robots.txt and tests it against 20 AI-related user agents, grouped by purpose. It takes about five seconds, and nothing is stored. For the underlying rules, see RFC 9309 and each provider's crawler documentation, which change over time.