Before an AI system can use your content, it has to be able to fetch it. This guide covers the access layer: who's knocking, what you can say about it and the common ways sites accidentally shut the door.
#Who's knocking
Several providers publish separate user agents for separate jobs. Check each provider's current documentation, because names and behaviour change.
| Purpose | Examples |
|---|---|
| Search and answer indexing | OAI-SearchBot (OpenAI), PerplexityBot (Perplexity), Claude-SearchBot (Anthropic), Googlebot, Bingbot |
| Training | GPTBot (OpenAI), ClaudeBot (Anthropic), Google-Extended token, Applebot-Extended, CCBot |
| User-triggered fetches | ChatGPT-User (OpenAI), Perplexity-User, Claude-User |
A few clarifications:
Google-Extendedis a robots.txt token, not a separate crawler. Google documents that it doesn't affect inclusion in Google Search.- Allowing a search crawler doesn't opt you into training, and vice versa, for providers that separate them.
- User-triggered fetchers behave differently from background crawlers. Providers describe how, or whether, robots.txt applies; read their docs before assuming.
#Configure by purpose
robots.txt works per user agent. Group the bots you treat the same:
# Allow search and user-triggered access
User-agent: OAI-SearchBot
User-agent: PerplexityBot
User-agent: Claude-SearchBot
User-agent: ChatGPT-User
User-agent: Perplexity-User
User-agent: Claude-User
Allow: /
# Opt out of model training
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
Disallow: /
# Everyone else
User-agent: *
Disallow: /admin/
Sitemap: https://example.com/sitemap.xml
Rules are matched per RFC 9309: the most specific user-agent group applies, and within a group the longest matching path wins.
Whether to restrict training is a business decision. Make it deliberately and document it.
#robots.txt is not access control
It's an honour system for well-behaved crawlers. For private content use authentication. For abusive traffic use your firewall — but note that blanket WAF rules ("block anything with 'bot' in the user agent") can lock out the search crawlers you want. Verify a crawler's identity using the provider's published IP ranges or verification method before allowing it through.
#Rendering matters
Many retrieval systems read the HTML your server returns and don't execute JavaScript. Check that:
- The main content and key facts appear in the raw HTML.
- Important pages don't depend on client-side navigation to be reachable.
- Response codes are correct (200 for pages, 404 for missing ones, no soft-404s).
A quick test: curl the page and read the output. If the answer isn't there, a crawler may not see it either. The page-readiness audit automates that check.
#What about llms.txt?
llms.txt is a proposed convention: a Markdown file at /llms.txt summarising your site for language models. It's not a standard that all engines honour, and Google's guidance says special AI text files aren't required for its search experiences. If you maintain one, treat it as documentation. Generate one in a minute with our llms.txt generator.
#Common mistakes
- Staging rules (
Disallow: /) shipped to production. - Allowing bots in robots.txt but blocking them in a CDN or WAF.
- Blocking by path patterns that accidentally catch documentation or pricing pages.
- Redirect chains on robots.txt itself.
- Assuming a successful crawl means inclusion in answers. Access is necessary, not sufficient.
#Check your site
Run the free AI crawler checker. It tests your robots.txt against 20 AI-related user agents grouped by purpose and shows the matching rule for each.
#Sources
- IETF, RFC 9309: Robots Exclusion Protocol.
- OpenAI, Overview of OpenAI crawlers.
- Perplexity, Perplexity crawlers.
- Google Search Central, Introduction to robots.txt and Optimizing for generative AI features in Search.
- llms.txt proposal.