The AI crawler check shows which AI bots your robots.txt file lets in, for your site and your competitors' sites. robots.txt is the plain text file at yourdomain.com/robots.txt that tells bots which pages they may fetch. Use the check when AI answers never cite your pages, or after anyone edits that file.
What does the check read?
When you open the Crawlability report or test a URL in it, DepraBot reads robots.txt. It reads the file for your own site and for the competitor sites you entered. DepraBot is the Depra AI web reader, described on about DepraBot.
The report checks the live robots.txt against 41 AI crawlers from 26 vendors and gives each bot a verdict.
| Verdict | Meaning |
|---|---|
| Allowed | The bot may read the site |
| Partial | The site root is closed to that bot and an Allow line reopens part of it |
| Blocked | The rules that apply to that bot close the site |
The report sorts the bots under two headings and lets you open the rule behind each verdict.
- Explicitly mentioned: bots your file names in their own User-agent line. A User-agent line says which bot the rules under it are for.
- Following global rules: bots that fall back to the catch-all User-agent: * group.
- The rule behind a verdict: click a verdict to open the robots.txt file at the User-agent group that applies to that bot.
- Other sites: use the brand picker to switch between your own domain and a tracked competitor that has a website address.
Note: If robots.txt returns a 404, no rule applies and every bot may crawl. If the file cannot be read, Depra AI shows the error in place of verdicts.
Note: The check comes with every plan. There is no public version without a login.
What are the four bot types?
Search bots, training bots, user-triggered fetchers and control tokens. Depra AI tags each of the 41 bots with one of these types, because blocking each type costs you something different.
| Bot type | What it does | Examples | If you block it |
|---|---|---|---|
| Search bots | Index pages so an AI answer can cite them | OAI-SearchBot, PerplexityBot, Claude-SearchBot, Googlebot, bingbot | You can drop out of that engine's answers |
| Training bots | Collect pages to build future AI models | GPTBot, ClaudeBot | Your pages stay out of that training |
| User-triggered fetchers | Open one page when a person asks an assistant about it | ChatGPT-User, Perplexity-User | Some vendors say these fetchers may not follow robots.txt |
| Control tokens | Names in robots.txt that no bot visits under. They tell a vendor how it may use fetched pages | Google-Extended, Applebot-Extended | The vendor is told not to use your pages for that purpose, such as training |
OpenAI says its GPTBot and OAI-SearchBot settings are independent. It says sites opted out of OAI-SearchBot are not shown in ChatGPT search answers but can still appear as navigational links (OpenAI crawler documentation, checked 7 Oct 2026).
How do I test one URL?
- 1
Open the URL Tester tab
It is a tab in the Crawlability report.
- 2
Paste a page address
Use its exact address.
- 3
Read the result
Depra AI checks that address against the site's robots.txt for all 41 bots. It shows Allowed or Blocked for each one, next to its bot type and platform.
What about llms.txt?
llms.txt is a proposed plain text file at yourdomain.com/llms.txt. It points AI tools to the most useful pages on a site.
To see whether a site has one, open yourdomain.com/llms.txt in a browser. A 404 means the site has none.
The file is still a proposal, so fix crawler access first. More is in the llms.txt guide.
How do I see which AI bots visit my pages?
robots.txt shows what bots may do. Your access logs show what they did. Depra AI reads logs from your server or CDN and keeps only AI bot visits. A CDN is the service that delivers your pages. Setup is manual.
- 1
Get the private address
Depra AI gives each log source a private web address.
- 2
Send your logs
Set your server, CDN or a scheduled job to send log lines to that address. You can also send an access log file.
- 3
Read the visits
You see total bot visits, requests over time, visits by bot type and the pages bots read most.
| Item | Detail |
|---|---|
| Formats read | CloudFront logs, Cloudflare and Vercel JSON log lines, JSON arrays, and Apache or nginx access logs in the standard combined format |
| What is stored | Only requests from the 41 AI bots, matched by user agent, the name a bot sends with each request. Human visits are dropped |
| Date range | From the last 7 days to the last 365 days |
| Export | The table of pages bots read most downloads as CSV |
Warning: A scraper can copy a real bot's user agent, so the visit counts describe what your logs report.
How do I allow or block a bot?
Edit robots.txt at the root of your site. Give the bot its own group: a User-agent line with the bot's name, then its rules. A bot follows the group that names it. A bot with no group of its own follows the User-agent: * group (RFC 9309, checked 7 Oct 2026).
# Block one bot: close the whole site to OpenAI's training crawler
User-agent: GPTBot
Disallow: /
# Allow one bot: open the whole site to OpenAI's search crawler
User-agent: OAI-SearchBot
Allow: /- Mind your private folders: a bot with its own group stops following the User-agent: * group. Copy your Disallow lines for private folders into the new group.
- Check the result: after you publish the file, open the Crawlability report again or test a page in the URL Tester tab.
Warning: robots.txt is one layer. A firewall or CDN bot setting can refuse a bot with a 403 error while robots.txt allows it. Check the status codes for that bot in your access logs.
Note: Perplexity says Perplexity-User generally ignores robots.txt rules (Perplexity crawler documentation, checked 7 Oct 2026). Depra AI does not show Blocked for that bot, because the file cannot enforce it. To stop it, block its user agent in your firewall or CDN.
The action plan can also flag an AI crawler block on pages it has read.