Skip to content
NessFlow
Menu

Answer engine crawlers, one by one

This is the list our engine actually audits on your sites. We review it every quarter against the operators’ own documentation, and this page follows without anyone touching it.

  • 19 AI crawlers audited measured
  • 11 operators covered measured

The crawlers we audit, one by one

Filter the list

The filter works without JavaScript. One selection at a time: most crossings between purpose and robots.txt behaviour match no crawler at all.

robots.txt token Operator Purpose Honours robots.txt User agent
GPTBot OpenAI Training yes Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.3; +https://openai.com/gptbot
OAI-SearchBot OpenAI Search yes Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-SearchBot/1.3; +https://openai.com/searchbot
ChatGPT-User OpenAI User request yes Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot
ClaudeBot Anthropic Training yes Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ClaudeBot/1.0; +claudebot@anthropic.com)
Claude-SearchBot Anthropic Search yes Not published by the operator.
Claude-User Anthropic User request yes Not published by the operator.
PerplexityBot Perplexity Search yes Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)
Perplexity-User Perplexity User request no Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user)
Google-Extended Google Control token yes Not published by the operator.
Applebot-Extended Apple Control token yes Not published by the operator.
CCBot Common Crawl Training yes CCBot/2.0 (https://commoncrawl.org/faq/)
Bytespider ByteDance Training partially Mozilla/5.0 (Linux; Android 5.0) AppleWebKit/537.36 (KHTML, like Gecko) Mobile Safari/537.36 (compatible; Bytespider; spider-feedback@bytedance.com)
meta-externalagent Meta Training yes meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler)
meta-externalfetcher Meta User request no meta-externalfetcher/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler)
Amazonbot Amazon Training yes Mozilla/5.0 (Macintosh; Intel Mac OS X 10_10_1) AppleWebKit/600.2.5 (KHTML, like Gecko) Version/8.0.2 Safari/600.2.5 (Amazonbot/0.1; +https://developer.amazon.com/support/amazonbot)
Amzn-SearchBot Amazon Search yes Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amzn-SearchBot/0.1) Chrome/120.0.0.0 Safari/537.36
Amzn-User Amazon User request yes Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amzn-User/0.1) Chrome/120.0.0.0 Safari/537.36
DuckAssistBot DuckDuckGo Search yes Mozilla/5.0 (compatible; DuckAssistBot/1.2; +http://duckduckgo.com/duckassistbot.html)
MistralAI-User Mistral User request yes Mozilla/5.0 (compatible; MistralAI-User/1.0; +https://docs.mistral.ai/robots)

Four purposes, and why they call for different answers

Search
This crawler feeds the answers the engine shows its users, with their sources. Blocking it removes your site from citations: it is the most expensive decision on this page, and often the one taken by default.
User request
This crawler fetches one specific page because a user just asked for it. Traffic is low and deliberate; blocking it means declining to answer someone who is looking for you.
Training
This crawler collects content to train models. Blocking it is a perfectly legitimate editorial choice with no direct effect on your citations, the one case where saying no costs you no visibility.
Control token
This is not a crawler but a token the operator recognises in robots.txt to govern one specific use. No request ever carries this name: probing it by user agent proves nothing, only reading robots.txt counts.

Find out which ones actually reach your site

The directory tells you what each crawler does. An audit tells you which ones get through: we test whether your pages are actually reachable, crawler by crawler, and record whether your brand is cited.

Access opens in waves: we email you when yours is ready.