Short answer: You control AI crawlers in robots.txt by adding a User-agent group for each bot name followed by Allow or Disallow rules. The important decision is which type of bot to block: training crawlers such as GPTBot or CCBot collect data for models, while search crawlers such as OAI-SearchBot or PerplexityBot fetch pages that can be cited in answers. Blocking training bots while allowing search bots is a common middle path that keeps you visible in AI search.
Three kinds of AI bots
“AI crawler” is not one thing. The bots that AI companies run fall into three broad groups, and each has different consequences when you block it.
- Training crawlers collect large amounts of public web content to train or improve language models. Blocking them keeps your future content out of new training datasets. It does not remove anything already collected.
- Search and retrieval crawlers build an index that an AI assistant searches when it answers questions. Blocking them usually means your pages cannot be retrieved or cited by that assistant’s search feature.
- User-triggered fetchers visit a specific page because a person asked the assistant to read it or because an answer needs a live look. Several companies state that these fetchers act on behalf of a user and may not treat robots.txt the same way as automated crawlers.
Understanding this split is the whole game. A site that blocks every AI-related name in one go often removes itself from AI search results without intending to.
Common AI user agents and what they do
Companies document their crawler names in their own help pages, and the list changes over time, so check the official documentation before relying on any list. The table below covers the names you will most often see in logs.
| User agent / token | Operator | Main purpose |
|---|---|---|
| GPTBot | OpenAI | Collecting content for model training |
| OAI-SearchBot | OpenAI | Indexing pages for search results in ChatGPT |
| ChatGPT-User | OpenAI | Fetching pages on a user’s request |
| ClaudeBot | Anthropic | Collecting content for model training |
| Claude-SearchBot / Claude-User | Anthropic | Search indexing and user-requested fetches |
| PerplexityBot | Perplexity | Indexing pages for Perplexity answers |
| Google-Extended | A control token for use in Gemini models, not a separate crawler | |
| Applebot-Extended | Apple | A control token for use in Apple’s AI training |
| CCBot | Common Crawl | Open web archive widely used to train models |
| Meta-ExternalAgent, Bytespider, Amazonbot | Meta, ByteDance, Amazon | Crawling for AI and other products |
Two entries deserve a note. Google-Extended and Applebot-Extended are not bots that fetch pages. Googlebot and Applebot do the crawling, and the “Extended” token only tells the company whether it may use the content for AI model purposes. Blocking Google-Extended does not remove your site from Google Search, and Google states it does not affect AI Overviews, which are part of Search.
How robots.txt rules work
The Robots Exclusion Protocol is documented as RFC 9309. A few rules explain almost every AI crawler question:
- Groups. Rules are grouped under one or more
User-agentlines. A crawler follows the most specific group that matches its name. - Specific beats general. If you have a
User-agent: GPTBotgroup, GPTBot ignores theUser-agent: *group entirely. That means rules you put under the wildcard do not apply to it. - Longest match wins. Within a group, the rule with the longest matching path takes priority, and
Allowwins a tie. - No rule means allowed. If no group matches a bot, or the group has no matching rule, the bot may crawl.
- It is voluntary. Reputable crawlers respect robots.txt, but it is a request, not a lock. Misbehaving bots need to be blocked at the server or firewall.
The second point causes the most confusion. Many sites carefully disallow /wp-admin/ and search pages under the wildcard, then add an AI-specific group with only Allow: /, accidentally opening those paths to that bot.
Copy-ready patterns
Here are three common policies. Adjust paths to your site and keep your existing rules for other bots.
Policy A: allow everything (visible everywhere). You need no AI-specific lines at all. If you want to be explicit, make sure no group blocks these names and that the wildcard group does not disallow /.
Policy B: block training, allow AI search.
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /
User-agent: OAI-SearchBot
User-agent: PerplexityBot
User-agent: Claude-SearchBot
Disallow: /wp-admin/
Disallow: /?s=
User-agent: *
Disallow: /wp-admin/
Disallow: /?s=
Note that the search bots get their own copy of the private-path rules, because they no longer read the wildcard group.
Policy C: block all AI use. List every training and search bot under one group with Disallow: /. Understand the trade-off: your pages will generally not be cited by those assistants, and user-triggered fetchers may still read a page when a user pastes your link.
Deciding which policy fits your business
There is no universally right answer. The question is what you gain from each kind of exposure.
- Businesses that sell products or services usually benefit from being cited, because an answer that mentions your brand is marketing. Allowing search bots is generally sensible, and many also allow training bots so that models learn accurate facts about them.
- Publishers whose content is the product have more reason to block training crawlers, since answers can substitute for a visit. Many still allow search crawlers because citations bring some traffic.
- Sites with sensitive or licensed material may prefer to block everything and handle access through agreements.
Whatever you choose, write it down. Robots.txt policies drift when different people edit the file for different reasons, and a documented decision makes audits much faster.
It also helps to separate the decision by section of the site. You might be happy for AI assistants to quote your product pages, help centre and pricing, while keeping a members-only resource library or paid reports out of reach. Robots.txt handles this well: allow the bot generally and disallow the specific folders. Just remember that anything truly private should sit behind a login, because a robots.txt rule also advertises the path to anyone who reads the file.
Finally, revisit the choice once a year. The balance between training, search and referral traffic is shifting quickly, and a policy that made sense when you wrote it may no longer match how your customers find you.
Blocking you did not choose: CDNs, firewalls and plugins
A clean robots.txt is not enough if something in front of your site turns bots away. Common culprits include:
- CDN bot management. Several CDNs offer a switch to block AI crawlers across all sites on an account, and some have enabled such protections by default for new domains.
- Security plugins and web application firewalls. Rules that block “unknown bots” or rate-limit aggressively can return 403 or 429 responses to AI crawlers.
- Hosting-level blocks. Some hosts block high-volume bots to protect shared servers.
- Challenge pages. JavaScript challenges and CAPTCHAs are effectively walls for crawlers.
The only reliable check is your server or CDN logs: filter by user agent and look at the status codes each AI bot receives. A bot that is allowed in robots.txt but gets 403 responses is blocked in practice.
Verifying and maintaining your rules
- Load your robots.txt in a browser and confirm it returns status 200 and plain text. A robots.txt that returns a server error can cause crawlers to pause crawling.
- Test specific URLs against specific user agents with a robots.txt testing tool, especially after editing groups.
- Check logs monthly for AI user agents, their status codes and how often they visit.
- Verify identity. Anyone can fake a user-agent string. Major operators publish IP ranges so you can confirm that a “GPTBot” request really comes from OpenAI.
- Review the list of bot names a couple of times a year. New crawlers appear and existing ones get renamed or split.
How Site SEO AI Audit helps
Site SEO AI Audit reads your robots.txt the way a crawler does and reports which AI crawlers are allowed and which are blocked, as part of the AI visibility area of the report. It also flags pages whose content depends on JavaScript, which many AI bots cannot run. On WordPress sites the fix steps point you to the exact place in wp-admin or your SEO plugin where robots.txt is managed. Plans with regular re-audits, listed on the pricing page, help catch a rule that changes after a plugin update.
Related reading
- Robots.txt for SEO: What to Block and What to Leave Open
- What Is Generative Engine Optimization? A Plain Guide
- llms.txt Explained: What It Is and Whether You Need One
The bottom line
AI crawler control is a business decision expressed in a few lines of robots.txt. Separate training bots from search bots, remember that a specific group replaces the wildcard group for that bot, and check that your CDN and firewall match your intent. Then verify with logs, because the file only states your policy; the logs show what actually happens.
KKK
Does blocking GPTBot remove my site from ChatGPT search?
Not by itself. OpenAI documents GPTBot as its training crawler and OAI-SearchBot as the crawler for search features, and the two are controlled separately. To stay citable in ChatGPT search, keep OAI-SearchBot allowed.
Will blocking Google-Extended hurt my Google rankings?
No. Google-Extended is only a control token for use in Gemini models, and Google states it does not affect Search rankings or inclusion in Search features. Googlebot continues to crawl as normal.
Does blocking AI crawlers remove content they already collected?
No. Robots.txt only affects future crawling. Content already in a training dataset or index is not deleted by changing the file, although search indexes typically refresh over time.
Do all AI bots respect robots.txt?
Major operators say their automated crawlers follow robots.txt, but compliance is voluntary. User-triggered fetchers and unknown bots may behave differently, so use server or firewall rules when you need enforcement.
Where do I edit robots.txt on WordPress?
WordPress generates a virtual robots.txt by default, and most SEO plugins include an editor for it. If a physical robots.txt file exists in the site root, it takes priority over the virtual one.


