Paste a URL to see which AI crawlers your robots.txt allows or blocks, grouped into AI search engines and training bots, plus whether you have an llms.txt and whether your content is readable without JavaScript. Free, no signup.
5 free checks a day as a guest · unlimited with a free account
| Crawler | Site root / |
This path | Deciding robots.txt line |
|---|
Allowed / blocked follows robots.txt only. CDN or firewall rules that challenge bots are not visible here. Crawler list last reviewed .
Add these groups to your robots.txt so ChatGPT search, Perplexity, Claude and the others can read and cite your pages. Explicit groups also override a User-agent: * block.
Keeps your content out of training sets while leaving the search crawlers above untouched. Google-Extended and Applebot-Extended only affect training, never Search or Siri.
Snippets are generated from the same crawler list as the table. Put them above any User-agent: * group and re-run the check to confirm.
The homepage or any deep page. Crawlers are evaluated for the site root and for the exact path you enter.
robots.txt is parsed the way Google and OpenAI document it; the page is fetched without JavaScript and its robots meta tags, X-Robots-Tag header and visible text are measured.
A yes / partial / no answer for AI search visibility, the exact line that decides each crawler, and copy-ready robots.txt snippets to allow AI search engines or block training crawlers.
20 crawler tokens in three groups, kept in one maintained list. Last reviewed .
GPTBot is the crawler OpenAI uses to collect pages that may train its models. It is separate from OAI-SearchBot, which builds the index behind ChatGPT search, and from ChatGPT-User, which fetches a page when someone asks ChatGPT about it. Blocking GPTBot keeps your content out of training sets without removing you from ChatGPT answers; blocking OAI-SearchBot or ChatGPT-User does remove you from them.
Training crawlers (GPTBot, ClaudeBot, Google-Extended, CCBot, Bytespider, Meta-ExternalAgent and others) collect text to train models. AI search and answer engines (OAI-SearchBot, ChatGPT-User, PerplexityBot, Claude-SearchBot, DuckAssistBot and others) fetch pages to answer a question right now and cite the source. Most vendors publish separate tokens, so you can block training and still be visible in AI search. The checker groups the two so you can see which is which.
Paste a page URL above. The checker reads your robots.txt and applies the standard matching rules for OAI-SearchBot and ChatGPT-User: the group naming the crawler wins over the * group, the longest matching path rule wins, and Allow beats Disallow on a tie. A missing robots.txt (404) allows everything, but one that answers with a server error counts as unreachable and crawlers must treat the whole site as blocked until it recovers. It also fetches the page itself to confirm the server returns readable HTML, because a page that renders only in JavaScript is invisible to most AI crawlers even when robots.txt allows it.
llms.txt is a proposed convention (llmstxt.org): a Markdown file at /llms.txt that names the site, describes it in one line and links its most useful pages, so language models and AI agents can find the right content quickly. It is not a standard and no vendor has committed to reading it, but it costs nothing to publish and the checker reports whether yours exists, how large it is and what its first heading says. Our llms.txt generator drafts one from your sitemap.
They are robots meta tags or X-Robots-Tag headers (for example content="noai, noimageai") that ask AI systems not to use the page or its images. They were introduced by DeviantArt and are honoured by some vendors, not all. Unlike robots.txt, which stops a crawler from fetching the page, they are read after the fetch. The checker reports them alongside noindex and nosnippet because all four change what an AI answer engine may do with a page it has already read.
It counts the visible words in the HTML your server sends, with scripts and styles removed, and looks for an H1. A page under 30 words is reported as empty and under 120 words as thin. AI crawlers mostly do not execute JavaScript, so a single-page app that fills the page client-side looks blank to them. Server-side rendering, prerendering or static HTML fixes it. The check reads server HTML only; it does not render the page in a browser.
No. It reports what robots.txt and the HTML say. A CDN or firewall rule that challenges or blocks AI user agents is invisible here, so a crawler can be allowed in robots.txt and still fail at the edge. Check your CDN bot-management settings separately if a vendor reports it cannot reach you.
Any public page over http or https; a bare domain works too. Addresses on private networks, localhost, internal hostnames and unusual ports are refused. Guests get 5 checks a day; a free Helpdesky account removes the limit. Nothing is stored.
No signup and no credit card. Each one runs in seconds.
Enter your help center or docs URL and get a scorecard for discoverability, content, structure, freshness and technical health.
Open toolPaste a site URL and get a ready-to-edit llms.txt built from its sitemap, page titles and descriptions.
Open toolPaste any website URL and get its color palette as hex codes, plus the logo and favicon.
Open toolAudit any URL for on-page SEO issues: meta tags, headings, structured data, links and performance.
Open toolEvery free tool in one place, with more on the way. All of them run without an account.
Browse the hubHelpdesky help centers are server-rendered with sitemaps and llms.txt built in, so ChatGPT, Perplexity and Claude can find, read and cite your answers from day one. Free plan, no credit card.
Create a free help center