AI crawler access checker
Enter a domain and we fetch its robots.txt, then test the user agents used by major AI search, training and user-triggered fetching bots.
Questions about this tool.
Does allowing a crawler guarantee my pages appear in AI answers?
No. Access is a precondition, not a guarantee. An engine still has to retrieve, select and choose to cite your page.
Should I block training crawlers?
That's a business decision. Training and search controls are separate for several providers, so you can restrict training while allowing search crawlers — verify each provider's current documentation.
Why does Google-Extended show up here?
It's a robots.txt token rather than a separate crawler. Google documents it as a control over the use of content for Gemini training and grounding; it doesn't affect Google Search inclusion.
Is robots.txt a security measure?
No. It's advisory. Use authentication and your firewall for anything that must stay private.
Track this over time.
Citeroot records AI answers on a schedule, traces them to sources and shows what changed.