AI Crawlers and robots.txt: Which AI Bots to Allow, Block or Limit
What GPTBot, OAI-SearchBot, ClaudeBot, Google-Extended and other AI bots do, what blocking each changes, and robots.txt rules by business type.
Quick answer
AI companies run three kinds of bots: training crawlers (GPTBot, ClaudeBot, Google-Extended, CCBot) that collect content for model training, search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot, plus Googlebot and Bingbot) that index pages so AI answers can cite them, and user-triggered fetchers (ChatGPT-User, Claude-User, Perplexity-User) that load a page because someone asked. Most businesses should allow search and user-triggered bots, decide about training bots on principle, and limit expensive URLs rather than blocking everything. Blocking a training bot does not remove you from AI search.
The AI bots that matter in 2026
The table summarizes each vendor's documentation as of October 2026 (OpenAI's crawler overview, Anthropic's support article on its crawlers, Google's list of common crawlers, Perplexity's crawler page and Common Crawl's FAQ). User agent names occasionally change, so check the vendor page before relying on an old list.
| User agent | Operator | Purpose | If you block it |
|---|---|---|---|
| GPTBot | OpenAI | Training data for OpenAI models | Signals your content should not be used for training; no effect on ChatGPT search |
| OAI-SearchBot | OpenAI | Index for ChatGPT search | Your pages are not shown in ChatGPT search answers (may still appear as navigational links) |
| ChatGPT-User | OpenAI | Fetches a page when a user's request needs it | User actions may still fetch pages; OpenAI notes robots.txt may not apply |
| ClaudeBot | Anthropic | Training data for Claude models | Future content excluded from training |
| Claude-SearchBot | Anthropic | Index to improve Claude's search results | May reduce visibility in Claude's search answers |
| Claude-User | Anthropic | Fetches pages for user questions | Claude cannot retrieve your pages when users ask |
| Google-Extended | Control token for Gemini model training and grounding | No effect on Google Search, AI Overviews or AI Mode (per Google) | |
| Googlebot | Google Search, including AI Overviews and AI Mode | Removes you from Google Search entirely | |
| Bingbot | Microsoft | Bing index, used by Copilot | Removes you from Bing and weakens Copilot visibility |
| PerplexityBot | Perplexity | Index for Perplexity answers | Not surfaced in Perplexity search results |
| CCBot | Common Crawl | Open web archive widely used to train models | Excluded from future Common Crawl snapshots |
Training, search and user-triggered bots are separate decisions
The most common mistake is treating "AI bots" as one thing. Many sites added a block for GPTBot in 2023 and assumed it would stop ChatGPT from using their content in answers. It never did that: ChatGPT search uses OAI-SearchBot, and user-requested fetches use ChatGPT-User. The opposite mistake is equally common: blocking every AI-related user agent and unknowingly disappearing from AI search.
Think about each category on its own terms:
- Training bots: a policy choice about whether your content may train models. It has no documented effect on being cited in AI search. Blocking applies only to future crawling
- Search bots: a visibility choice. Blocking them means AI answers cannot cite or link to you
- User-triggered fetchers: behave like a person's browser acting on request. Blocking them breaks assistants that a customer is actively using to look at your site
Worth noting
Google-Extended is a robots.txt product token, not a separate crawler. Google crawls with Googlebot and uses the token to decide whether content can be used for Gemini model training and grounding. Google says it does not affect Search inclusion or ranking.
What should your business allow?
There is no universal answer, but the trade-offs are predictable by business model.
| Business type | Training bots | Search bots | User-triggered | Reasoning |
|---|---|---|---|---|
| Service business or agency | Your choice; many allow | Allow | Allow | Being cited in answers brings qualified enquiries |
| SaaS | Your choice; docs often allowed | Allow | Allow | Buyers and developers research through assistants |
| Ecommerce | Your choice | Allow; limit cart, search, filter URLs | Allow | Product discovery increasingly starts in AI assistants |
| Publisher with ads | Often block | Usually allow | Allow | Visibility drives traffic; training use is the contested part |
| Paywalled or licensed content | Block | Consider partial access | Case by case | Content is the product; licensing deals may apply |
| Internal tools or staging sites | Block | Block | Block | Nothing public to gain; also protect with authentication |
robots.txt examples
robots.txt rules are grouped by user agent. A crawler follows the most specific group that matches its name, so a named group overrides the wildcard. These examples follow the Robots Exclusion Protocol (RFC 9309); test them in Search Console's robots.txt report before deploying.
# Training crawlers: opt out
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
Disallow: /
# Search and user-triggered AI bots: allow, minus costly paths
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
Disallow: /cart
Disallow: /checkout
Disallow: /account
Disallow: /search
# Everyone else, including Googlebot and Bingbot
User-agent: *
Disallow: /cart
Disallow: /checkout
Disallow: /account
Disallow: /search
Sitemap: https://www.example.com/sitemap.xmlPro tip
Because a named group replaces the wildcard group for that bot, repeat important disallow rules (cart, checkout, account, internal search) in every group. Leaving them out of a named group allows that bot into those paths.
When robots.txt is not enough
robots.txt is a public request, not a lock. Reputable crawlers follow it; scrapers that ignore it will not be stopped by it, and some disguise themselves with browser user agents. For enforcement you need controls at the CDN or firewall: verify claimed crawlers by IP range or reverse DNS (OpenAI, Google and others publish their ranges), rate-limit aggressive clients and challenge unverified automation on sensitive paths.
Be careful with blanket bot blocking. Bot-protection products that challenge every non-browser client can block search crawlers and legitimate AI agents shopping on a customer's behalf. Newer approaches let agents prove who they are with cryptographic signatures; our guide to verifying AI agent traffic explains how that works.
Controlling what appears, not just who crawls
Sometimes the question is not whether a bot can visit but what may be shown. For Google, the existing snippet controls apply to AI features too: `nosnippet`, `data-nosnippet` on specific elements, `max-snippet` and `noindex`. They limit what Search, including AI Overviews and AI Mode, can display from your pages. Use them precisely; a site-wide `nosnippet` also removes ordinary search snippets and makes you ineligible for AI features.
How to audit your current setup
- Fetch your live robots.txt and list every AI-related user agent and rule
- Check that rules match your intent for each category (training, search, user-triggered)
- Review CDN and firewall bot settings for blanket AI blocking or challenges
- Search server logs for each user agent to see what is actually crawling and how often
- Verify high-volume bots against published IP ranges before trusting the user agent
- Disallow expensive dynamic URLs (internal search, filters, cart) instead of whole bots
- Re-check quarterly; vendors add and rename bots
Need a crawler policy that matches your business?
ZSpace Labs reviews robots.txt, CDN bot rules and server logs and sets up crawler access that protects your infrastructure without hiding you from AI search. See website development services.
Conclusion
Treat AI crawlers as three groups with three decisions. Allow the search and user-triggered bots that let customers find you through AI assistants, make an explicit choice about training bots, and protect your infrastructure with targeted disallow rules and rate limits instead of blanket blocks. Then confirm the result in your logs. For the wider picture, see how to make your website discoverable in AI search.
Common questions.
GPTBot collects content that may be used to train OpenAI's models. OAI-SearchBot indexes pages so they can appear in ChatGPT search answers. They are controlled separately in robots.txt, so you can block one and allow the other.