Use Cases AI AgentsBeta Pricing Embeds AI Skill MCP Server Sign in Get Started
Free tool AI Crawler Visibility Checker

AI crawler checker: is your site blocked from ChatGPT, Perplexity and Claude?

Paste a URL to see which AI crawlers your robots.txt allows or blocks, grouped into AI search engines and training bots, plus whether you have an llms.txt and whether your content is readable without JavaScript. Free, no signup.

5 free checks a day as a guest · unlimited with a free account

AI visibility report for

Checked:

Signals

Crawler access

Crawler Site root / This path Deciding robots.txt line

Allowed / blocked follows robots.txt only. CDN or firewall rules that challenge bots are not visible here. Crawler list last reviewed .

Fix it in robots.txt

Allow AI search engines

Add these groups to your robots.txt so ChatGPT search, Perplexity, Claude and the others can read and cite your pages. Explicit groups also override a User-agent: * block.


              
            

Optional: block AI training crawlers

Keeps your content out of training sets while leaving the search crawlers above untouched. Google-Extended and Applebot-Extended only affect training, never Search or Siri.


              
            

Snippets are generated from the same crawler list as the table. Put them above any User-agent: * group and re-run the check to confirm.

How it works

  1. Paste a page URL

    The homepage or any deep page. Crawlers are evaluated for the site root and for the exact path you enter.

  2. We read robots.txt, llms.txt and the page

    robots.txt is parsed the way Google and OpenAI document it; the page is fetched without JavaScript and its robots meta tags, X-Robots-Tag header and visible text are measured.

  3. Get a verdict and the fix

    A yes / partial / no answer for AI search visibility, the exact line that decides each crawler, and copy-ready robots.txt snippets to allow AI search engines or block training crawlers.

The AI crawlers this tool checks

20 crawler tokens in three groups, kept in one maintained list. Last reviewed .

AI search & answer engines

  • OAI-SearchBot · OpenAI Builds the index behind ChatGPT search results and citations. Docs
  • ChatGPT-User · OpenAI Fetches a page when a ChatGPT user asks about it or clicks a citation. Docs
  • PerplexityBot · Perplexity Indexes pages so they can be cited in Perplexity answers. Docs
  • Perplexity-User · Perplexity Fetches a page on behalf of a Perplexity user during a query. Docs
  • Claude-SearchBot · Anthropic Indexes pages that Claude can surface as search results. Docs
  • Claude-User · Anthropic Fetches a page when a Claude user asks about it. Docs
  • DuckAssistBot · DuckDuckGo Reads pages to generate DuckAssist answers in DuckDuckGo results. Docs
  • MistralAI-User · Mistral AI Fetches a page on behalf of a Le Chat user during a query. Docs

AI training crawlers

  • GPTBot · OpenAI Collects pages that may be used to train OpenAI models. Docs
  • ClaudeBot · Anthropic Collects pages that may be used to train Anthropic models. Docs
  • anthropic-ai · Anthropic Older Anthropic token still honoured; many sites list it next to ClaudeBot. Docs
  • Google-Extended · Google Controls whether Googlebot's crawl may train Gemini. Does not affect Search or AI Overviews. Docs
  • Applebot-Extended · Apple Controls whether Applebot's crawl may train Apple models. Does not affect Siri or Spotlight. Docs
  • CCBot · Common Crawl Builds the open Common Crawl corpus that many model training sets start from. Docs
  • Bytespider · ByteDance ByteDance (TikTok) crawler reported to gather training data; no public documentation.
  • Meta-ExternalAgent · Meta Collects pages for training Meta AI models and improving its products. Docs
  • Amazonbot · Amazon Indexes pages for Alexa answers and Amazon AI products. Docs
  • cohere-ai · Cohere Token widely listed for Cohere training crawls; no public documentation.

Classic search (for reference)

  • Googlebot · Google Google Search; Google AI Overviews and AI Mode draw on this index. Docs
  • Bingbot · Microsoft Bing Search; Microsoft Copilot answers draw on this index. Docs

Frequently asked questions

What is GPTBot and why does it matter?

GPTBot is the crawler OpenAI uses to collect pages that may train its models. It is separate from OAI-SearchBot, which builds the index behind ChatGPT search, and from ChatGPT-User, which fetches a page when someone asks ChatGPT about it. Blocking GPTBot keeps your content out of training sets without removing you from ChatGPT answers; blocking OAI-SearchBot or ChatGPT-User does remove you from them.

What is the difference between AI training crawlers and AI search crawlers?

Training crawlers (GPTBot, ClaudeBot, Google-Extended, CCBot, Bytespider, Meta-ExternalAgent and others) collect text to train models. AI search and answer engines (OAI-SearchBot, ChatGPT-User, PerplexityBot, Claude-SearchBot, DuckAssistBot and others) fetch pages to answer a question right now and cite the source. Most vendors publish separate tokens, so you can block training and still be visible in AI search. The checker groups the two so you can see which is which.

How do I check if my website is blocked from ChatGPT?

Paste a page URL above. The checker reads your robots.txt and applies the standard matching rules for OAI-SearchBot and ChatGPT-User: the group naming the crawler wins over the * group, the longest matching path rule wins, and Allow beats Disallow on a tie. A missing robots.txt (404) allows everything, but one that answers with a server error counts as unreachable and crawlers must treat the whole site as blocked until it recovers. It also fetches the page itself to confirm the server returns readable HTML, because a page that renders only in JavaScript is invisible to most AI crawlers even when robots.txt allows it.

What is llms.txt?

llms.txt is a proposed convention (llmstxt.org): a Markdown file at /llms.txt that names the site, describes it in one line and links its most useful pages, so language models and AI agents can find the right content quickly. It is not a standard and no vendor has committed to reading it, but it costs nothing to publish and the checker reports whether yours exists, how large it is and what its first heading says. Our llms.txt generator drafts one from your sitemap.

What do the noai and noimageai directives do?

They are robots meta tags or X-Robots-Tag headers (for example content="noai, noimageai") that ask AI systems not to use the page or its images. They were introduced by DeviantArt and are honoured by some vendors, not all. Unlike robots.txt, which stops a crawler from fetching the page, they are read after the fetch. The checker reports them alongside noindex and nosnippet because all four change what an AI answer engine may do with a page it has already read.

Why does the checker say my content is not readable without JavaScript?

It counts the visible words in the HTML your server sends, with scripts and styles removed, and looks for an H1. A page under 30 words is reported as empty and under 120 words as thin. AI crawlers mostly do not execute JavaScript, so a single-page app that fills the page client-side looks blank to them. Server-side rendering, prerendering or static HTML fixes it. The check reads server HTML only; it does not render the page in a browser.

Does the checker see Cloudflare or WAF bot blocks?

No. It reports what robots.txt and the HTML say. A CDN or firewall rule that challenges or blocks AI user agents is invisible here, so a crawler can be allowed in robots.txt and still fail at the edge. Check your CDN bot-management settings separately if a vendor reports it cannot reach you.

Which URLs can I check?

Any public page over http or https; a bare domain works too. Addresses on private networks, localhost, internal hostnames and unusual ports are refused. Guests get 5 checks a day; a free Helpdesky account removes the limit. Nothing is stored.

More free tools

No signup and no credit card. Each one runs in seconds.

A help center AI search can actually read

Helpdesky help centers are server-rendered with sitemaps and llms.txt built in, so ChatGPT, Perplexity and Claude can find, read and cite your answers from day one. Free plan, no credit card.

Create a free help center