Free · runs in your browser

AI crawler checker for robots.txt

Paste your robots.txt and see, crawler by crawler, whether ChatGPT, Claude, Perplexity, Gemini and Apple may read the page — and exactly which line decides it.

Blocking an AI search crawler removes you from the answers it produces. Search and user-fetch crawlers — OAI-SearchBot, Claude-SearchBot, PerplexityBot, ChatGPT-User — are what let an assistant open and cite your page today. Training crawlers such as GPTBot, ClaudeBot and Google-Extended are a separate decision and do not affect whether you are cited. This checker applies RFC 9309 matching: most specific User-agent group wins, then longest path, and Allow breaks a tie.

Check a robots.txt

Paste the whole file, comments included. Nothing you paste is uploaded.

Only used to build the link below. Nothing is fetched.

The example is fictional — it is not your site.

/ answers for the homepage. Try a real page path to catch a directory rule.

Paste a robots.txt to see the verdicts. If your site has no robots.txt at all, every crawler is allowed everywhere — which is a valid, if unmanaged, position.

A robots.txt block that allows every AI crawler

Merge this into your existing file — do not replace the file with it, or you will drop the Disallow rules you meant to keep.
robots.txt (excerpt)

How to check and fix AI crawler access

  1. 1Open your robots.txtVisit https://yourdomain.com/robots.txt in a browser. If it returns 404 the file does not exist, which means every crawler is allowed everywhere by default.
  2. 2Copy the whole fileSelect all of it, comments included. Comments change nothing, but a partial paste can hide the group that is actually deciding.
  3. 3Paste it here and set the pathLeave the path as / for the homepage answer, or enter a specific path such as /blog or /ar/products to see whether a directory rule blocks it.
  4. 4Read the deciding line, not just the verdictEach row names the User-agent group and the Allow or Disallow line that won, with its line number. That is the line to change.
  5. 5Fix the blocks you did not intendCopy the ready-made allow block below, merge it into your file, redeploy, and re-open the URL to confirm the change is live.

Frequently asked questions

Does blocking GPTBot stop my site appearing in ChatGPT?

Not on its own. GPTBot is OpenAI's training crawler. The crawler that builds the index ChatGPT search reads is OAI-SearchBot, and the one that opens a page because a user asked about it is ChatGPT-User. Blocking GPTBot while allowing the other two keeps you out of training data but still citable in answers.

What does Google-Extended actually control?

Google-Extended controls whether your content may be used in Gemini apps and Vertex AI grounding. It does not affect Google Search ranking, and it does not remove you from Google AI Overviews — those are served by Googlebot's normal index, which robots.txt governs through the Googlebot token instead.

How does robots.txt decide which rule applies?

RFC 9309 says only the most specific matching User-agent group applies: if a group names your crawler, the wildcard group is ignored entirely for it. Within that group the longest matching path pattern wins, and if an Allow and a Disallow tie on length, Allow wins. This checker applies exactly those rules.

My robots.txt has no rules for AI crawlers. Is that a problem?

It means access is unmanaged rather than blocked: with no matching group, the wildcard group decides, and if there is no wildcard group either, everything is allowed. That is usually fine, but it also means a future edit to the wildcard group silently changes AI access. Naming the crawlers you care about makes the decision explicit.

Should I allow every AI crawler?

Allow the search and user-fetch crawlers if you want to be cited in AI answers — that is the whole mechanism. Training crawlers are a separate business decision: blocking them protects content from model training but does not improve or harm your visibility in AI search. Decide the two questions separately.

Why can't this tool fetch my robots.txt for me?

A browser cannot read another site's robots.txt because of the same-origin policy, and we will not proxy it through a server, because then the tool would be sending your domain to us. Pasting keeps the promise that nothing on this page leaves your browser.

Does robots.txt actually stop these crawlers?

The major operators publish their tokens and state that they honour robots.txt, and our own audits observe them doing so. It is a request, not an enforcement mechanism: a crawler that ignores robots.txt has to be blocked at the edge or the firewall instead.

Is anything I paste here sent to Search Genie?

No. The parsing and the verdicts happen in JavaScript in your browser. Nothing is uploaded, nothing is logged, and closing the tab discards it.

Want this checked against your live site, not a paste?

The free brand audit fetches your real robots.txt, applies the same rules, and reports it next to your Google rankings, AI Overview citations, structured data and real-user speed.

What the free audit checks