Skip to main content

AI Bot Access Tester

Check which AI crawlers your robots.txt blocks. Tests GPTBot, ClaudeBot, Google-Extended, PerplexityBot, CCBot and 28 more against a pasted robots.txt using RFC 9309 precedence, grouped by whether the bot trains models, powers AI search, or fetches on a user's behalf — and emits a ready-to-paste block for the ones still allowed.

Input

Paste the robots.txt you want to check — yours, or the one you are about to deploy.

A path like /blog/post, or a full URL (only its path + query is used).

Output

Per-crawler verdict
CrawlerOperatorPurposeAccessDeciding rule
No data yet

User-triggered fetchers run because a person pasted your link into a chatbot. Several operators state these ignore robots.txt, so treat a “Blocked” verdict there as a request rather than a guarantee — enforce it at the edge if it matters.

Block the rest

robots.txt block for every crawler still allowed
 
Was this helpful?

Guides

Paste a robots.txt and see, in one table, which AI crawlers it actually blocks. The tool tests 33 documented AI user-agents — GPTBot, ClaudeBot, Google-Extended, PerplexityBot, CCBot, Bytespider, Applebot-Extended and the rest — against a path you choose, using the same RFC 9309 precedence rules a real crawler applies.

How to use it

  1. Paste your robots.txt into the first box. It does not have to be live yet — checking a file before you deploy it is the whole point.
  2. Set the path you care about. / answers "is my whole site open to AI?"; /blog/some-post answers it for one page, which matters when your rules use wildcards.
  3. Optionally narrow the list to one purpose — training, AI search, or user-triggered.
  4. Read the Deciding rule column. It names the exact directive and line number that produced each verdict, so a surprising result is traceable rather than mysterious.
  5. Copy the generated block at the bottom to disallow every crawler that is still allowed.

The three kinds of AI crawler

They are listed separately because blocking them are three different decisions:

  • Model training — collects pages to train or fine-tune models (GPTBot, ClaudeBot, CCBot, Google-Extended). Blocking these is an opt-out from training corpora and costs you nothing in search.
  • AI search — indexes pages so an answer engine can cite them (OAI-SearchBot, PerplexityBot, Claude-SearchBot, DuckAssistBot). Blocking these removes you from AI answers — usually the opposite of what a site owner wants.
  • User-triggered — fetches a single page because a person pasted your link into a chatbot (ChatGPT-User, Claude-User, Perplexity-User). Blocking these breaks a link for a real visitor.

A blanket Disallow: / aimed at "AI" hits all three. That is why the tool separates them.

Why isn't Googlebot in the list?

Because blocking it removes you from Google Search, not just from AI answers. The same goes for bingbot. Google's AI-training opt-out is the separate Google-Extended token, which is in the list — that is the one to disallow if you want ordinary search but not model training.

Does blocking a bot in robots.txt actually stop it?

For the well-behaved crawlers, yes — the major operators publish their tokens precisely so you can opt out. For the user-triggered fetchers, several operators state outright that a person-initiated fetch is not a crawl and does not consult robots.txt. If you need enforcement rather than a request, block by user-agent or IP range at your CDN or web server.

Why paste the file instead of entering a domain?

So it runs entirely in your browser. Nothing is uploaded, there is no rate limit, and it works on a draft file that is not deployed anywhere yet.

Related tools

To test one specific crawler against one path with the full list of applicable rules, use the Robots.txt Tester. To build a robots.txt from scratch rather than audit one, use the Robots.txt Generator. If you are checking a site's crawl setup more broadly, the XML Sitemap Generator covers the other half of the file crawlers look for.

airobots.txtcrawlerseogptbotllm

Love the tools? Lose the ads.

One payment clears every ad from your account, for good. No subscription, no tracking.