Answer engine crawlers, one by one
This is the list our engine actually audits on your sites. We review it every quarter against the operators’ own documentation, and this page follows without anyone touching it.
- 19 AI crawlers audited measured
- 11 operators covered measured
The crawlers we audit, one by one
The filter works without JavaScript. One selection at a time: most crossings between purpose and robots.txt behaviour match no crawler at all.
| robots.txt token | Operator | Purpose | Honours robots.txt | User agent |
|---|---|---|---|---|
| GPTBot | OpenAI | Training | yes | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.3; +https://openai.com/gptbot |
| OAI-SearchBot | OpenAI | Search | yes | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-SearchBot/1.3; +https://openai.com/searchbot |
| ChatGPT-User | OpenAI | User request | yes | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot |
| ClaudeBot | Anthropic | Training | yes | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ClaudeBot/1.0; +claudebot@anthropic.com) |
| Claude-SearchBot | Anthropic | Search | yes | Not published by the operator. |
| Claude-User | Anthropic | User request | yes | Not published by the operator. |
| PerplexityBot | Perplexity | Search | yes | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot) |
| Perplexity-User | Perplexity | User request | no | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user) |
| Google-Extended | Control token | yes | Not published by the operator. | |
| Applebot-Extended | Apple | Control token | yes | Not published by the operator. |
| CCBot | Common Crawl | Training | yes | CCBot/2.0 (https://commoncrawl.org/faq/) |
| Bytespider | ByteDance | Training | partially | Mozilla/5.0 (Linux; Android 5.0) AppleWebKit/537.36 (KHTML, like Gecko) Mobile Safari/537.36 (compatible; Bytespider; spider-feedback@bytedance.com) |
| meta-externalagent | Meta | Training | yes | meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler) |
| meta-externalfetcher | Meta | User request | no | meta-externalfetcher/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler) |
| Amazonbot | Amazon | Training | yes | Mozilla/5.0 (Macintosh; Intel Mac OS X 10_10_1) AppleWebKit/600.2.5 (KHTML, like Gecko) Version/8.0.2 Safari/600.2.5 (Amazonbot/0.1; +https://developer.amazon.com/support/amazonbot) |
| Amzn-SearchBot | Amazon | Search | yes | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amzn-SearchBot/0.1) Chrome/120.0.0.0 Safari/537.36 |
| Amzn-User | Amazon | User request | yes | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amzn-User/0.1) Chrome/120.0.0.0 Safari/537.36 |
| DuckAssistBot | DuckDuckGo | Search | yes | Mozilla/5.0 (compatible; DuckAssistBot/1.2; +http://duckduckgo.com/duckassistbot.html) |
| MistralAI-User | Mistral | User request | yes | Mozilla/5.0 (compatible; MistralAI-User/1.0; +https://docs.mistral.ai/robots) |
Four purposes, and why they call for different answers
- Search
- This crawler feeds the answers the engine shows its users, with their sources. Blocking it removes your site from citations: it is the most expensive decision on this page, and often the one taken by default.
- User request
- This crawler fetches one specific page because a user just asked for it. Traffic is low and deliberate; blocking it means declining to answer someone who is looking for you.
- Training
- This crawler collects content to train models. Blocking it is a perfectly legitimate editorial choice with no direct effect on your citations, the one case where saying no costs you no visibility.
- Control token
- This is not a crawler but a token the operator recognises in robots.txt to govern one specific use. No request ever carries this name: probing it by user agent proves nothing, only reading robots.txt counts.
Find out which ones actually reach your site
The directory tells you what each crawler does. An audit tells you which ones get through: we test whether your pages are actually reachable, crawler by crawler, and record whether your brand is cited.
Access opens in waves: we email you when yours is ready.