Guide

AI crawlers list: every major AI bot user agent and what it does

46 AI and search bot user agents, each looked up on its vendor's own crawler page where one exists: what the bot collects, whether it follows robots.txt, and how to block it.

By · Updated 15 Sept 2026 · 12 min read

Five small outlined shapes in a column, four joined by dotted blue lines to an empty browser window and one stopped at a closed gate marked with an orange dot, showing AI bots allowed or blocked from a site

An AI crawler is a bot that fetches web pages for an AI company, and it names itself with a user-agent token, such as GPTBot, that you can address in robots.txt. There are three kinds: training crawlers that collect pages to build AI models, search crawlers that index pages so AI answers can cite them, and user-triggered fetchers that open a page when a person asks an assistant about it.

robots.txt is the plain text file at yourdomain.com/robots.txt that tells bots which pages they may fetch. Some vendors also publish control tokens, such as Google-Extended, which change how your content is used without any bot visiting under that name.

The table lists each bot by vendor. We checked every row against the vendor's own documentation on 15 Sep 2026, and where a vendor publishes nothing, the row says so. The sections after the table explain each vendor and give the robots.txt lines to use.

AI crawler user agents by vendor, checked 15 Sep 2026
User-agent tokenVendorPurposeFollows robots.txtSource
GPTBotOpenAITraining: collects content to train OpenAI's foundation modelsYesOpenAI docs
OAI-SearchBotOpenAISearch: finds pages to show in ChatGPT search answersYesOpenAI docs
ChatGPT-UserOpenAIUser-triggered: visits a page when someone asks ChatGPT or a custom GPT about itMay not apply, says OpenAIOpenAI docs
OAI-AdsBotOpenAIChecks the landing pages of ads submitted to ChatGPT; not used for trainingNot statedOpenAI docs
GooglebotGoogleSearch: crawls for Google Search, including AI Overviews and AI ModeYesGoogle AI features
Google-ExtendedGoogleControl token, no crawler: covers Gemini model training and grounding in Gemini Apps and Vertex AIToken onlyGoogle crawlers
GoogleOtherGoogleGeneric crawler for Google product teams; its rules affect no specific productYesGoogle crawlers
ClaudeBotAnthropicTraining: collects web content for Anthropic's AI modelsYesAnthropic docs
Claude-SearchBotAnthropicSearch: crawls to improve search results in ClaudeYesAnthropic docs
Claude-UserAnthropicUser-triggered: fetches a page when someone asks Claude a questionYesAnthropic docs
anthropic-aiAnthropicOlder token, not on Anthropic's current crawler pageVendor docs not foundVendor docs not found
PerplexityBotPerplexitySearch: surfaces and links sites in Perplexity; not used for model trainingYesPerplexity docs
Perplexity-UserPerplexityUser-triggered: visits a page to answer a user's questionGenerally ignores it, says PerplexityPerplexity docs
Meta-ExternalAgentMetaTraining and products: training foundation AI models or improving productsYesMeta docs
Meta-WebIndexerMetaSearch: improves Meta AI search resultsYesMeta docs
Meta-ExternalFetcherMetaUser-triggered: fetches single links at a user's requestMay bypass it, says MetaMeta docs
FacebookBotMetaOlder token, not on Meta's current crawler pageVendor docs not foundVendor docs not found
ApplebotAppleSearch: powers search in Spotlight, Siri and SafariYes, and follows Googlebot rules when none name ApplebotApple docs
Applebot-ExtendedAppleControl token, no crawler: opts content out of training Apple's foundation modelsToken onlyApple docs
AmazonbotAmazonImproves Amazon products and services; may be used to train Amazon AI modelsYesAmazon docs
Amzn-SearchBotAmazonSearch: search in Amazon products; not used for model trainingYesAmazon docs
Amzn-UserAmazonUser-triggered: for example Alexa questions that need current informationMay not follow every rule, says AmazonAmazon docs
bingbotMicrosoftSearch: crawls for BingYesBing Webmaster blog
msnbotMicrosoftOlder Bing token; bingbot reads a msnbot group when it has no group of its ownRead by bingbot as a fallbackBing Webmaster blog
CCBotCommon CrawlBuilds Common Crawl's free, open archive of web crawl dataYesCommon Crawl
DuckAssistBotDuckDuckGoSearch: fetches pages in real time for AI-assisted answers; not used for trainingYesDuckDuckGo help
MistralAI-UserMistralUser-triggered: visits a page when a user asks Mistral's assistant a questionYesMistral docs
MistralAI-IndexMistralSearch: crawls for indexing only; not used for trainingYesMistral docs
MistralAI-TrainingMistralTraining: builds datasets for Mistral AI modelsYesMistral docs
AI2BotAllen Institute for AITraining: content used to train open language modelsNot statedAi2 crawler page
AhrefsBotAhrefsSearch and SEO index: powers the Yep search engine and the Ahrefs platformYesAhrefs docs
cohere-aiCohereCohere says it runs no crawlers for model training today and lists no bot namesVendor lists no botsCohere docs
cohere-training-data-crawlerCohereNot listed by Cohere, as aboveVendor lists no botsCohere docs
DiffbotDiffbotCrawls the web to build Diffbot's general search engineYes by default; can be overridden by agreementDiffbot docs
PetalBotHuaweiSearch: indexes sites for Petal SearchYesPetalBot page
ImagesiftBotHive (ImageSift)Collects public images, alt text and page text so ImageSift products can find similar imagesYes, and follows Googlebot rules when none name ImagesiftBotImageSift by Hive
SemrushBot-OCOBSemrushCrawls pages for Semrush's Content ToolkitYesSemrush bots
OmgilibotWebz.ioOlder Webz.io crawler, replaced by Webzio and Webzio-extendedNot stated for OmgilibotWebz.io bots
BytespiderByteDanceVendor docs not foundVendor docs not foundVendor docs not found
TikTokSpiderByteDanceVendor docs not foundVendor docs not foundVendor docs not found
DeepSeekBotDeepSeekVendor docs not foundVendor docs not foundVendor docs not found
Kangaroo BotKangaroo LLMVendor docs not foundVendor docs not foundVendor docs not found
TimpibotTimpiVendor docs not foundVendor docs not foundVendor docs not found
xAI-BotxAIVendor docs not foundVendor docs not foundVendor docs not found
Grok-UserxAIVendor docs not foundVendor docs not foundVendor docs not found
YouBotYou.comVendor docs not foundVendor docs not foundVendor docs not found
"Vendor docs not found" means we found no crawler page from the company itself, so we make no claim about that bot. robots.txt matches user-agent names without regard to letter case. Depra's AI crawler checker covers 41 tokens: every row here except OAI-AdsBot, Meta-WebIndexer, Amzn-SearchBot, Amzn-User, MistralAI-Index and MistralAI-Training, plus Scrapy, an open-source scraping tool. AhrefsBot, SemrushBot-OCOB, PetalBot and Diffbot are search or SEO crawlers; they are listed because Depra's checker covers them.

OpenAI crawlers: GPTBot, OAI-SearchBot and ChatGPT-User

OpenAI runs separate bots for training and for search, and its documentation says each robots.txt setting is independent of the others. GPTBot collects content that may be used to train OpenAI's generative AI foundation models. OAI-SearchBot finds pages for ChatGPT search, and OpenAI says sites that block it will not be shown in ChatGPT search answers.

ChatGPT-User visits a page when a person asks ChatGPT or a custom GPT about it. It does not crawl the web automatically, and OpenAI says robots.txt rules may not apply to it because a user started the request. OAI-AdsBot checks the landing pages of ads submitted to ChatGPT.

OpenAI says search changes can take about 24 hours to apply after you edit robots.txt. If you allow both GPTBot and OAI-SearchBot, OpenAI may use one crawl for both jobs.

To stay in ChatGPT search while opting out of training, add a group with the line User-agent: GPTBot followed by Disallow: /, and leave OAI-SearchBot allowed. A blocked OAI-SearchBot is one of the causes covered in why ChatGPT doesn't mention your brand.

Google: Googlebot for AI Overviews, Google-Extended and GoogleOther

Google AI Overviews and AI Mode use Googlebot, the same crawler as Google Search. An AI Overview is the AI summary at the top of some Google results. Google's guide to AI features says robots.txt rules for Googlebot are the control for Search, and a page must be indexed and eligible for a snippet to appear as a supporting link. Blocking Googlebot removes you from Search and AI Overviews together.

Google-Extended is a control token. No bot visits under that name, because Google crawls with its existing user agents. A group with User-agent: Google-Extended and Disallow: / tells Google not to use your content to train future Gemini models or for grounding in Gemini Apps and Vertex AI. Grounding means giving Gemini pages from the Google Search index while it writes an answer. Google says Google-Extended does not affect inclusion or ranking in Google Search, so it does not remove you from AI Overviews.

Since grounding in Gemini Apps is part of what Google-Extended covers, blocking it can change how Gemini Apps use your pages at answer time. That is a different trade from blocking a bot that only collects training data, so decide on it separately.

GoogleOther is a generic crawler that Google product teams use to fetch public pages, and Google says rules addressed to it do not affect any specific product. Google says Googlebot and GoogleOther always obey robots.txt when crawling automatically.

Anthropic crawlers: ClaudeBot, Claude-SearchBot and Claude-User

Anthropic documents three bots. ClaudeBot collects web content for its AI models. Claude-SearchBot crawls to improve search results in Claude. Claude-User fetches a page when a person asks Claude a question.

Anthropic says its bots honour robots.txt. Because each has its own name, you can block training with User-agent: ClaudeBot and Disallow: / and leave Claude-SearchBot and Claude-User allowed. The older anthropic-ai token does not appear on Anthropic's current crawler page.

Perplexity crawlers: PerplexityBot and Perplexity-User

PerplexityBot finds and links websites in Perplexity search results. Perplexity says it is not used to crawl content for AI foundation models, and it follows robots.txt. Perplexity recommends allowing it in your robots.txt file.

Perplexity-User visits a page when a user's question needs it. Perplexity says that because a user asked for the fetch, it generally ignores robots.txt rules. A robots.txt line will not reliably stop it; a firewall or CDN rule on its user agent will.

Meta crawlers: Meta-ExternalAgent, Meta-WebIndexer and Meta-ExternalFetcher

Meta-ExternalAgent crawls for uses such as training foundation AI models and improving products. Meta-WebIndexer crawls to improve Meta AI search results. Meta-ExternalFetcher fetches single links at a user's request, and Meta says it may bypass robots.txt for that reason.

To opt out of training while staying findable in Meta AI search, disallow Meta-ExternalAgent and leave Meta-WebIndexer allowed. FacebookBot, an older Meta token, is not on Meta's current crawler page.

Apple crawlers: Applebot and Applebot-Extended

Applebot crawls for search in Spotlight, Siri and Safari. It follows robots.txt rules that name Applebot, and when your file names Googlebot but not Applebot, it follows the Googlebot rules.

Applebot-Extended does not crawl. Apple uses it only to decide how pages Applebot fetched may be used. Disallowing it opts your content out of training Apple's foundation models, and Apple says those pages stay discoverable in Spotlight, Siri and Safari.

Amazon, Microsoft and Common Crawl

Amazon lists three agents. Amazonbot improves Amazon products and services and may be used to train Amazon AI models. Amzn-SearchBot supports search in Amazon products and is not used for model training. Amzn-User acts for users, for example on Alexa questions that need current information, and Amazon says it may not follow every robots.txt rule.

bingbot crawls for Bing. Bing says bingbot follows one group only: its own if it exists, then a msnbot group for backwards compatibility, then the catch-all User-agent: * group.

CCBot builds the free, open archive of web crawl data that Common Crawl publishes. It follows robots.txt, and Common Crawl warns that some crawlers falsely call themselves CCBot, so it suggests checking the address a request came from.

Other AI crawlers: Mistral, DuckDuckGo, Ai2, Cohere and more

Mistral documents three agents: MistralAI-User for pages a user asks its assistant about, MistralAI-Index for search indexing, and MistralAI-Training for training datasets. Only MistralAI-Training collects content to train Mistral's models.

DuckAssistBot fetches pages in real time for DuckDuckGo AI-assisted answers. DuckDuckGo says the data is not used to train AI models, a robots.txt change takes effect after 72 hours, and opting out does not change organic rankings on DuckDuckGo.

AI2Bot, from the Allen Institute for AI (Ai2), collects content to train open language models, and its page does not say whether it follows robots.txt. PetalBot indexes sites for Huawei's Petal Search. Diffbot follows robots.txt by default. AhrefsBot builds the index behind Ahrefs and the Yep search engine, and SemrushBot-OCOB crawls pages for Semrush's Content Toolkit.

Cohere says it runs no crawlers for model training today and lists no bot names. The cohere-ai and cohere-training-data-crawler tokens therefore have no vendor page behind them.

Which AI bots have no vendor documentation?

We found no crawler page from the company itself for Bytespider and TikTokSpider (ByteDance), DeepSeekBot (DeepSeek), xAI-Bot and Grok-User (xAI), YouBot (You.com), Timpibot (Timpi) and Kangaroo Bot (Kangaroo LLM). Third-party bot directories describe them, but without a vendor source we make no claim about what they collect or whether they follow robots.txt.

If one of these bots puts load on your server or you do not want it, add a robots.txt group for it and also block its user agent at your firewall or CDN, since robots.txt only works for bots that choose to read it.

How to allow AI search bots but block training bots

Put the training bots in one group that ends with Disallow: /, then let every other bot fall through to the catch-all group. A bot follows only the group that names it, and a bot with no group of its own reads the User-agent: * rules. The example below opts out of training where the vendor offers a separate training bot or token, and keeps ChatGPT search, Claude search, Perplexity, Google and Bing able to read the site.

Example robots.txt that allows AI search bots and blocks training bots
Line in robots.txtWhat it does
User-agent: GPTBotStarts a group for OpenAI's training crawler
User-agent: ClaudeBotAdds Anthropic's training crawler to the same group
User-agent: Meta-ExternalAgentAdds Meta's training and product crawler
User-agent: MistralAI-TrainingAdds Mistral's training crawler
User-agent: Applebot-ExtendedAdds Apple's training control token
User-agent: CCBotAdds Common Crawl's archive crawler
Disallow: /Closes the whole site to every name in this group
User-agent: *Starts the group for every other bot, including OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot and bingbot
Allow: /Lets those bots read the whole site
Leave a blank line between the two groups. Google-Extended is left out on purpose because it also covers grounding in Gemini Apps; add it only if you accept that trade. If your file already has a User-agent: * group with Disallow lines for private folders, keep those lines in that group when you add Allow: /. Replacing the whole group with this example would open those folders to every bot.

How to check which AI bots can crawl your site

You can check any site by hand in four steps.

  • Open the file. Go to yourdomain.com/robots.txt. A 404 means there are no rules, so every bot that follows robots.txt may crawl.
  • Find each bot's group. Search the file for the token, for example OAI-SearchBot. If a User-agent line names it, only that group applies. If no line names it, the User-agent: * group applies.
  • Read the rules. Disallow: / closes the whole site. A longer Allow line, such as Allow: /blog/, reopens that folder, because the most specific matching rule wins under the robots.txt standard, RFC 9309.
  • Check the server. robots.txt can allow a bot while a firewall or CDN bot setting answers it with a 403 error. Search your access logs for the bot name and look at the status codes.

Faster: check robots.txt for 41 bots on your site and your competitors' sites

Doing this by hand for 41 bots and several competitor sites takes a while. Depra's AI crawler checker reads robots.txt for your site and each competitor you track, shows which User-agent group decides each verdict, tests single URLs and reports AI bot visits from your server logs. It comes with every plan, with 7 days free and no card.

Crawler access only makes you eligible to be cited. To see whether AI engines then name your brand, compare generative engine optimization tools, and read how to get ChatGPT to recommend your brand for what to publish next. Terms such as grounding and AI Overview are defined in the AI search glossary.

Frequently asked

What is GPTBot?

GPTBot is OpenAI's web crawler for model training. It collects public pages that may be used to train OpenAI's generative AI foundation models, and it follows robots.txt, so User-agent: GPTBot with Disallow: / opts you out. OpenAI says this setting is independent of OAI-SearchBot, so blocking GPTBot does not remove you from ChatGPT search.

What is OAI-SearchBot?

OAI-SearchBot is the OpenAI crawler that finds pages for ChatGPT search. OpenAI says sites that block it will not be shown in ChatGPT search answers. It follows robots.txt, and OpenAI says a change can take about 24 hours to apply.

What is the difference between GPTBot and OAI-SearchBot?

GPTBot collects content to train OpenAI models. OAI-SearchBot finds pages to show and cite in ChatGPT search. They are separate robots.txt settings, so you can block GPTBot and allow OAI-SearchBot to stay visible in ChatGPT search while keeping your pages out of training.

Does Google-Extended affect AI Overviews?

No. AI Overviews use Googlebot, and Google says robots.txt rules for Googlebot are the control for its AI features in Search. Google-Extended covers training future Gemini models and grounding in Gemini Apps and Vertex AI, and Google says it does not affect inclusion in Google Search.

What is PerplexityBot?

PerplexityBot is Perplexity's search crawler. It finds and links websites in Perplexity search results, follows robots.txt, and Perplexity says it is not used to crawl content for AI foundation models. Its sibling Perplexity-User fetches pages for a user's question and generally ignores robots.txt.

Check your robots.txt against 41 AI bots

7 days free on any plan, no card. Your first scan starts at signup.

Keep reading