• English
  • Article · published 18 September 2026 · DImato

    Which AI bots should you allow in robots.txt?

    Allow the search and fetch bots: OAI-SearchBot and ChatGPT-User for ChatGPT, Claude-SearchBot and Claude-User for Claude, PerplexityBot for Perplexity, Googlebot for Google AI Overviews and Gemini, Bingbot for Copilot. Blocking any of them removes your site from that engine's answers. Training bots, GPTBot, ClaudeBot and Google-Extended, are a separate decision: blocking them does not change whether you are cited. This page lists every bot, what it does and what blocking it changes, from the vendors' own documentation.

    What is the difference between a search bot and a training bot?

    Each AI vendor runs two or three crawlers with different jobs. A search bot builds the index the engine searches when a user asks a question. A fetch bot opens a page live during a conversation. A training bot collects text to train future models. Only the first two decide whether your page can be cited.

    Blocking a training bot is a policy choice. Blocking a search bot is a visibility choice. Most sites that "blocked AI" in 2024 blocked both and became uncitable by accident.

    Every AI bot and what blocking it changes

    User agentVendorJobRespects robots.txtIf you block it
    OAI-SearchBotOpenAISearch index for ChatGPT searchYes, about 24 h lagRemoved from ChatGPT search answers
    ChatGPT-UserOpenAILive fetch when a user asksYesChatGPT cannot read your page during a conversation
    GPTBotOpenAITraining onlyYesNo effect on ChatGPT search citations
    Claude-SearchBotAnthropicSearch index qualityYes, including Crawl-delayReduced visibility in Claude answers
    Claude-UserAnthropicLive fetch for a user queryYesClaude cannot read your page for a user
    ClaudeBotAnthropicTrainingYesNo effect on Claude search
    PerplexityBotPerplexitySearch index, no trainingYesRemoved from Perplexity results
    Perplexity-UserPerplexityLive fetch for a userGenerally ignores it, because a user askedLittle effect; block at the firewall if you must
    GooglebotGoogleSearch index, also AI Overviews, AI Mode and Gemini groundingYesRemoved from Google Search, AI Overviews and AI Mode
    Google-ExtendedGoogleA control token, not a crawler: Gemini training and Gemini app groundingYesNo effect on Search, AI Overviews or AI Mode
    BingbotMicrosoftBing index, Copilot, part of ChatGPT groundingYesRemoved from Bing and Copilot, weaker ChatGPT grounding
    CCBotCommon CrawlOpen dataset used for training by manyYesNo effect on any engine's citations

    Two facts from the vendor documents matter most. Google-Extended does not affect Search, AI Overviews or AI Mode; those use Googlebot. And Claude's engine is Brave-backed, so Brave's crawler must reach your site for Claude to find it, though Anthropic has not published that dependency itself.

    A robots.txt you can copy

    This is the file on dimato.lt. It allows everything and says why. Change the training-bot block to Disallow: / if your policy is to opt out of training; leave the search bots alone.

    # Search and fetch bots: allowed, they make the site citable.
    User-agent: Googlebot
    Allow: /
    User-agent: Bingbot
    Allow: /
    User-agent: OAI-SearchBot
    Allow: /
    User-agent: ChatGPT-User
    Allow: /
    User-agent: Claude-SearchBot
    Allow: /
    User-agent: Claude-User
    Allow: /
    User-agent: PerplexityBot
    Allow: /
    
    # Training bots: a policy choice. Allow or Disallow, no effect on citations.
    User-agent: GPTBot
    Allow: /
    User-agent: ClaudeBot
    Allow: /
    User-agent: Google-Extended
    Allow: /
    User-agent: CCBot
    Allow: /
    
    User-agent: *
    Allow: /
    
    Sitemap: https://example.com/sitemap.xml

    What robots.txt does not cover

    • The CDN or firewall. Cloudflare, Akamai and similar services ship "block AI bots" rules that stop search bots as well as training bots. Cloudflare's own data from 2025 shows about 80% of AI crawl volume is training and under 5% is search and user fetches; a blanket rule trades a small bandwidth saving for invisibility. Test with curl -A "OAI-SearchBot" https://yoursite/ for each bot and expect 200.
    • JavaScript. No AI crawler executes JavaScript except Googlebot. Across more than 500 million GPTBot fetches, zero ran JavaScript; GPTBot downloads script files 11.5% of the time and ClaudeBot 23.8%, and neither runs them. If your product text is rendered in the browser, the other engines see an empty page. Test with curl and look for the text in the raw HTML.
    • Server logs. A bot that is allowed but never arrives cannot cite you either. Count hits per user agent monthly. Zero hits from OAI-SearchBot or PerplexityBot on a site that wants AI traffic is a finding, not a comfort.
    • New bot names. Vendors add crawlers. Check the four vendor pages below every quarter.

    Should you block the training bots?

    It is a policy decision with no citation cost either way. Companies that want their own product descriptions to be what a model remembers tend to allow training; companies with licensed content or editorial concerns tend to block it. Record the decision and apply it consistently across robots.txt, the CDN and any noai meta tags, so it can be audited later.

    Sources