Article · published 18 September 2026 · DImato
Which AI bots should you allow in robots.txt?
Allow the search and fetch bots: OAI-SearchBot and ChatGPT-User for ChatGPT, Claude-SearchBot and Claude-User for Claude, PerplexityBot for Perplexity, Googlebot for Google AI Overviews and Gemini, Bingbot for Copilot. Blocking any of them removes your site from that engine's answers. Training bots, GPTBot, ClaudeBot and Google-Extended, are a separate decision: blocking them does not change whether you are cited. This page lists every bot, what it does and what blocking it changes, from the vendors' own documentation.
What is the difference between a search bot and a training bot?
Each AI vendor runs two or three crawlers with different jobs. A search bot builds the index the engine searches when a user asks a question. A fetch bot opens a page live during a conversation. A training bot collects text to train future models. Only the first two decide whether your page can be cited.
Blocking a training bot is a policy choice. Blocking a search bot is a visibility choice. Most sites that "blocked AI" in 2024 blocked both and became uncitable by accident.
Every AI bot and what blocking it changes
Two facts from the vendor documents matter most. Google-Extended does not affect Search, AI Overviews or AI Mode; those use Googlebot. And Claude's engine is Brave-backed, so Brave's crawler must reach your site for Claude to find it, though Anthropic has not published that dependency itself.
A robots.txt you can copy
This is the file on dimato.lt. It allows everything and says why. Change the training-bot block to Disallow: / if your policy is to opt out of training; leave the search bots alone.
What robots.txt does not cover
- The CDN or firewall. Cloudflare, Akamai and similar services ship "block AI bots" rules that stop search bots as well as training bots. Cloudflare's own data from 2025 shows about 80% of AI crawl volume is training and under 5% is search and user fetches; a blanket rule trades a small bandwidth saving for invisibility. Test with
curl -A "OAI-SearchBot" https://yoursite/for each bot and expect 200. - JavaScript. No AI crawler executes JavaScript except Googlebot. Across more than 500 million GPTBot fetches, zero ran JavaScript; GPTBot downloads script files 11.5% of the time and ClaudeBot 23.8%, and neither runs them. If your product text is rendered in the browser, the other engines see an empty page. Test with
curland look for the text in the raw HTML. - Server logs. A bot that is allowed but never arrives cannot cite you either. Count hits per user agent monthly. Zero hits from OAI-SearchBot or PerplexityBot on a site that wants AI traffic is a finding, not a comfort.
- New bot names. Vendors add crawlers. Check the four vendor pages below every quarter.
Should you block the training bots?
It is a policy decision with no citation cost either way. Companies that want their own product descriptions to be what a model remembers tend to allow training; companies with licensed content or editorial concerns tend to block it. Record the decision and apply it consistently across robots.txt, the CDN and any noai meta tags, so it can be audited later.
Sources
- OpenAI, Overview of OpenAI crawlers.
- Anthropic, Does Anthropic crawl data from the web.
- Perplexity, Perplexity crawlers.
- Google, Google common crawlers, updated July 2026, and AI features and your website.
- Vercel and MERJ, The rise of the AI crawler, December 2024, for the JavaScript figures.
- Cloudflare Radar, AI crawler traffic share, August 2025.