Skip to content
Web Development7 min read

AI Crawlers and robots.txt: Which AI Bots to Allow, Block or Limit

What GPTBot, OAI-SearchBot, ClaudeBot, Google-Extended and other AI bots do, what blocking each changes, and robots.txt rules by business type.

01

Quick answer

AI companies run three kinds of bots: training crawlers (GPTBot, ClaudeBot, Google-Extended, CCBot) that collect content for model training, search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot, plus Googlebot and Bingbot) that index pages so AI answers can cite them, and user-triggered fetchers (ChatGPT-User, Claude-User, Perplexity-User) that load a page because someone asked. Most businesses should allow search and user-triggered bots, decide about training bots on principle, and limit expensive URLs rather than blocking everything. Blocking a training bot does not remove you from AI search.

02

The AI bots that matter in 2026

The table summarizes each vendor's documentation as of October 2026 (OpenAI's crawler overview, Anthropic's support article on its crawlers, Google's list of common crawlers, Perplexity's crawler page and Common Crawl's FAQ). User agent names occasionally change, so check the vendor page before relying on an old list.

User agentOperatorPurposeIf you block it
GPTBotOpenAITraining data for OpenAI modelsSignals your content should not be used for training; no effect on ChatGPT search
OAI-SearchBotOpenAIIndex for ChatGPT searchYour pages are not shown in ChatGPT search answers (may still appear as navigational links)
ChatGPT-UserOpenAIFetches a page when a user's request needs itUser actions may still fetch pages; OpenAI notes robots.txt may not apply
ClaudeBotAnthropicTraining data for Claude modelsFuture content excluded from training
Claude-SearchBotAnthropicIndex to improve Claude's search resultsMay reduce visibility in Claude's search answers
Claude-UserAnthropicFetches pages for user questionsClaude cannot retrieve your pages when users ask
Google-ExtendedGoogleControl token for Gemini model training and groundingNo effect on Google Search, AI Overviews or AI Mode (per Google)
GooglebotGoogleGoogle Search, including AI Overviews and AI ModeRemoves you from Google Search entirely
BingbotMicrosoftBing index, used by CopilotRemoves you from Bing and weakens Copilot visibility
PerplexityBotPerplexityIndex for Perplexity answersNot surfaced in Perplexity search results
CCBotCommon CrawlOpen web archive widely used to train modelsExcluded from future Common Crawl snapshots
03

Training, search and user-triggered bots are separate decisions

The most common mistake is treating "AI bots" as one thing. Many sites added a block for GPTBot in 2023 and assumed it would stop ChatGPT from using their content in answers. It never did that: ChatGPT search uses OAI-SearchBot, and user-requested fetches use ChatGPT-User. The opposite mistake is equally common: blocking every AI-related user agent and unknowingly disappearing from AI search.

Think about each category on its own terms:

  • Training bots: a policy choice about whether your content may train models. It has no documented effect on being cited in AI search. Blocking applies only to future crawling
  • Search bots: a visibility choice. Blocking them means AI answers cannot cite or link to you
  • User-triggered fetchers: behave like a person's browser acting on request. Blocking them breaks assistants that a customer is actively using to look at your site

Worth noting

Google-Extended is a robots.txt product token, not a separate crawler. Google crawls with Googlebot and uses the token to decide whether content can be used for Gemini model training and grounding. Google says it does not affect Search inclusion or ranking.

04

What should your business allow?

There is no universal answer, but the trade-offs are predictable by business model.

Business typeTraining botsSearch botsUser-triggeredReasoning
Service business or agencyYour choice; many allowAllowAllowBeing cited in answers brings qualified enquiries
SaaSYour choice; docs often allowedAllowAllowBuyers and developers research through assistants
EcommerceYour choiceAllow; limit cart, search, filter URLsAllowProduct discovery increasingly starts in AI assistants
Publisher with adsOften blockUsually allowAllowVisibility drives traffic; training use is the contested part
Paywalled or licensed contentBlockConsider partial accessCase by caseContent is the product; licensing deals may apply
Internal tools or staging sitesBlockBlockBlockNothing public to gain; also protect with authentication
05

robots.txt examples

robots.txt rules are grouped by user agent. A crawler follows the most specific group that matches its name, so a named group overrides the wildcard. These examples follow the Robots Exclusion Protocol (RFC 9309); test them in Search Console's robots.txt report before deploying.

Allow AI search, opt out of AI training, protect expensive URLs
# Training crawlers: opt out
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
Disallow: /

# Search and user-triggered AI bots: allow, minus costly paths
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
Disallow: /cart
Disallow: /checkout
Disallow: /account
Disallow: /search

# Everyone else, including Googlebot and Bingbot
User-agent: *
Disallow: /cart
Disallow: /checkout
Disallow: /account
Disallow: /search

Sitemap: https://www.example.com/sitemap.xml

Pro tip

Because a named group replaces the wildcard group for that bot, repeat important disallow rules (cart, checkout, account, internal search) in every group. Leaving them out of a named group allows that bot into those paths.

06

When robots.txt is not enough

robots.txt is a public request, not a lock. Reputable crawlers follow it; scrapers that ignore it will not be stopped by it, and some disguise themselves with browser user agents. For enforcement you need controls at the CDN or firewall: verify claimed crawlers by IP range or reverse DNS (OpenAI, Google and others publish their ranges), rate-limit aggressive clients and challenge unverified automation on sensitive paths.

Be careful with blanket bot blocking. Bot-protection products that challenge every non-browser client can block search crawlers and legitimate AI agents shopping on a customer's behalf. Newer approaches let agents prove who they are with cryptographic signatures; our guide to verifying AI agent traffic explains how that works.

07

Controlling what appears, not just who crawls

Sometimes the question is not whether a bot can visit but what may be shown. For Google, the existing snippet controls apply to AI features too: `nosnippet`, `data-nosnippet` on specific elements, `max-snippet` and `noindex`. They limit what Search, including AI Overviews and AI Mode, can display from your pages. Use them precisely; a site-wide `nosnippet` also removes ordinary search snippets and makes you ineligible for AI features.

08

How to audit your current setup

  • Fetch your live robots.txt and list every AI-related user agent and rule
  • Check that rules match your intent for each category (training, search, user-triggered)
  • Review CDN and firewall bot settings for blanket AI blocking or challenges
  • Search server logs for each user agent to see what is actually crawling and how often
  • Verify high-volume bots against published IP ranges before trusting the user agent
  • Disallow expensive dynamic URLs (internal search, filters, cart) instead of whole bots
  • Re-check quarterly; vendors add and rename bots

Need a crawler policy that matches your business?

ZSpace Labs reviews robots.txt, CDN bot rules and server logs and sets up crawler access that protects your infrastructure without hiding you from AI search. See website development services.

Start a Project
09

Conclusion

Treat AI crawlers as three groups with three decisions. Allow the search and user-triggered bots that let customers find you through AI assistants, make an explicit choice about training bots, and protect your infrastructure with targeted disallow rules and rate limits instead of blanket blocks. Then confirm the result in your logs. For the wider picture, see how to make your website discoverable in AI search.

FAQ

Common questions.

GPTBot collects content that may be used to train OpenAI's models. OAI-SearchBot indexes pages so they can appear in ChatGPT search answers. They are controlled separately in robots.txt, so you can block one and allow the other.

Get in touch

Have a project in mind?

Whether you're building a new digital product, improving an existing website, or looking to automate part of your business — let's talk.