Skip to content
Depra AI

Docs

AI crawler check: which AI bots can read your site

What the Crawlability report reads, the four bot types, the URL Tester, llms.txt, AI bot visits from your logs, and how to allow or block a bot.

Updated

The AI crawler check shows which AI bots your robots.txt file lets in, for your site and your competitors' sites. robots.txt is the plain text file at yourdomain.com/robots.txt that tells bots which pages they may fetch. Use the check when AI answers never cite your pages, or after anyone edits that file.

What does the check read?

When you open the Crawlability report or test a URL in it, DepraBot reads robots.txt. It reads the file for your own site and for the competitor sites you entered. DepraBot is the Depra AI web reader, described on about DepraBot.

The report checks the live robots.txt against 41 AI crawlers from 26 vendors and gives each bot a verdict.

Verdicts in the Crawlability report
VerdictMeaning
AllowedThe bot may read the site
PartialThe site root is closed to that bot and an Allow line reopens part of it
BlockedThe rules that apply to that bot close the site

The report sorts the bots under two headings and lets you open the rule behind each verdict.

  • Explicitly mentioned: bots your file names in their own User-agent line. A User-agent line says which bot the rules under it are for.
  • Following global rules: bots that fall back to the catch-all User-agent: * group.
  • The rule behind a verdict: click a verdict to open the robots.txt file at the User-agent group that applies to that bot.
  • Other sites: use the brand picker to switch between your own domain and a tracked competitor that has a website address.

Note: If robots.txt returns a 404, no rule applies and every bot may crawl. If the file cannot be read, Depra AI shows the error in place of verdicts.

Note: The check comes with every plan. There is no public version without a login.

What are the four bot types?

Search bots, training bots, user-triggered fetchers and control tokens. Depra AI tags each of the 41 bots with one of these types, because blocking each type costs you something different.

The four bot types and what blocking each one costs
Bot typeWhat it doesExamplesIf you block it
Search botsIndex pages so an AI answer can cite themOAI-SearchBot, PerplexityBot, Claude-SearchBot, Googlebot, bingbotYou can drop out of that engine's answers
Training botsCollect pages to build future AI modelsGPTBot, ClaudeBotYour pages stay out of that training
User-triggered fetchersOpen one page when a person asks an assistant about itChatGPT-User, Perplexity-UserSome vendors say these fetchers may not follow robots.txt
Control tokensNames in robots.txt that no bot visits under. They tell a vendor how it may use fetched pagesGoogle-Extended, Applebot-ExtendedThe vendor is told not to use your pages for that purpose, such as training
What each vendor says its bots do is in the AI crawlers list, checked against vendor pages on 7 Oct 2026.

OpenAI says its GPTBot and OAI-SearchBot settings are independent. It says sites opted out of OAI-SearchBot are not shown in ChatGPT search answers but can still appear as navigational links (OpenAI crawler documentation, checked 7 Oct 2026).

How do I test one URL?

  1. 1

    Open the URL Tester tab

    It is a tab in the Crawlability report.

  2. 2

    Paste a page address

    Use its exact address.

  3. 3

    Read the result

    Depra AI checks that address against the site's robots.txt for all 41 bots. It shows Allowed or Blocked for each one, next to its bot type and platform.

What about llms.txt?

llms.txt is a proposed plain text file at yourdomain.com/llms.txt. It points AI tools to the most useful pages on a site.

To see whether a site has one, open yourdomain.com/llms.txt in a browser. A 404 means the site has none.

The file is still a proposal, so fix crawler access first. More is in the llms.txt guide.

How do I see which AI bots visit my pages?

robots.txt shows what bots may do. Your access logs show what they did. Depra AI reads logs from your server or CDN and keeps only AI bot visits. A CDN is the service that delivers your pages. Setup is manual.

  1. 1

    Get the private address

    Depra AI gives each log source a private web address.

  2. 2

    Send your logs

    Set your server, CDN or a scheduled job to send log lines to that address. You can also send an access log file.

  3. 3

    Read the visits

    You see total bot visits, requests over time, visits by bot type and the pages bots read most.

AI bot visits: formats, storage and limits
ItemDetail
Formats readCloudFront logs, Cloudflare and Vercel JSON log lines, JSON arrays, and Apache or nginx access logs in the standard combined format
What is storedOnly requests from the 41 AI bots, matched by user agent, the name a bot sends with each request. Human visits are dropped
Date rangeFrom the last 7 days to the last 365 days
ExportThe table of pages bots read most downloads as CSV

Warning: A scraper can copy a real bot's user agent, so the visit counts describe what your logs report.

How do I allow or block a bot?

Edit robots.txt at the root of your site. Give the bot its own group: a User-agent line with the bot's name, then its rules. A bot follows the group that names it. A bot with no group of its own follows the User-agent: * group (RFC 9309, checked 7 Oct 2026).

text
# Block one bot: close the whole site to OpenAI's training crawler
User-agent: GPTBot
Disallow: /

# Allow one bot: open the whole site to OpenAI's search crawler
User-agent: OAI-SearchBot
Allow: /
Two example groups for robots.txt.
  • Mind your private folders: a bot with its own group stops following the User-agent: * group. Copy your Disallow lines for private folders into the new group.
  • Check the result: after you publish the file, open the Crawlability report again or test a page in the URL Tester tab.

Warning: robots.txt is one layer. A firewall or CDN bot setting can refuse a bot with a 403 error while robots.txt allows it. Check the status codes for that bot in your access logs.

Note: Perplexity says Perplexity-User generally ignores robots.txt rules (Perplexity crawler documentation, checked 7 Oct 2026). Depra AI does not show Blocked for that bot, because the file cannot enforce it. To stop it, block its user agent in your firewall or CDN.

The action plan can also flag an AI crawler block on pages it has read.

Keep reading