Short answer: Yes, it is quite possible. CDNs, web application firewalls, security plugins and hosting providers can block or challenge AI crawlers even when your robots.txt allows them, and some CDNs now offer AI blocking as a one-click setting or a default for new domains. To check, review the bot settings at each layer, then filter your server or CDN logs by AI user agents and look for 403, 429 or challenge responses. Align every layer with a single, deliberate policy.
Why robots.txt is not the whole story
Robots.txt is a set of instructions that polite crawlers read before fetching pages. It does not enforce anything. The layers in front of your website, on the other hand, enforce rules directly: they decide whether a request reaches your server at all, and what response it gets if it does not.
A typical request from an AI crawler passes through several layers:
- DNS and CDN edge. The CDN receives the request first and applies bot management, rate limits and security rules.
- Web application firewall. A WAF, either at the CDN or at the host, inspects the request for suspicious patterns.
- Hosting platform. Many hosts run their own protection against aggressive bots to keep shared servers stable.
- Web server and application. Security plugins and server rules can block user agents or IP ranges.
A block at any layer overrides your robots.txt policy in practice. The crawler never sees your content, and your site quietly drops out of the AI answers that rely on that crawler.
The CDN layer
AI crawling became a major topic for CDNs in 2024 and 2025. Several providers introduced dedicated controls for AI bots, and Cloudflare announced in 2025 that new domains on its network would block known AI crawlers by default, with options for site owners to allow them. Settings you may find at your CDN include:
- A single “block AI bots” or “AI scrapers” toggle that blocks a list of known AI user agents, often including both training and search crawlers.
- Per-crawler controls that let you allow or block individual bots.
- Managed robots.txt features that add AI-related rules to your robots.txt automatically.
- Bot fight or bot protection modes that challenge any automated traffic not on a verified list.
- Rate limiting rules that apply to all clients, including crawlers.
Check who set these options and why. On agency-managed or inherited accounts, settings are often changed at the account level and apply to every site, including sites that want to be cited in AI answers.
If you are unsure which setting applies, the CDN’s security event log usually names the rule that acted on each request. Filter it by a crawler’s user agent and you will see exactly which feature blocked or challenged it.
The firewall and hosting layer
Web application firewalls block requests based on rules and reputation. AI crawlers can trigger them for ordinary reasons: they make many requests in a short time, come from cloud data centres rather than home connections, and do not run JavaScript challenges. Common causes of blocking include:
- Rules that block or challenge traffic from data centre IP ranges.
- Country-based blocking that happens to include the regions crawlers operate from.
- Rate limits set low to protect a small server.
- JavaScript or CAPTCHA challenges that no crawler can pass.
- Host-level bot protection, sometimes enabled without notice to the customer.
If you do not manage the firewall yourself, ask your host or IT provider directly: “Do you block or rate-limit GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot or bingbot?” The answer is often surprising.
The plugin and application layer
On WordPress and similar platforms, security and performance plugins can also block bots. Examples include bad-bot blocklists that contain AI user agents, login and brute-force protection rules that are too broad, and “block fake crawlers” features that fail to verify legitimate AI bots correctly. Custom rules in .htaccess or Nginx configuration, added years ago to stop a scraper, can also catch newer crawlers with similar names.
Search your server configuration and plugin settings for user-agent names such as GPTBot, ClaudeBot, CCBot, Bytespider and “bot” in general, and review each rule you find.
How to test whether you are blocked
Settings screens tell you what should happen. Tests and logs tell you what does happen. Use both:
| Method | What to do | What it shows |
|---|---|---|
| Log filtering | Filter CDN or server logs by AI user agents and group by status code | Real responses to real crawler requests |
| User-agent test | Request a page with a crawler’s user-agent string, for example with curl | Rules based on user agent alone |
| CDN analytics | Check bot and security event reports for AI crawler names | Blocks and challenges at the edge |
| Webmaster tools | Review crawl errors in Search Console and Bing Webmaster Tools | Blocks affecting Googlebot and bingbot |
| Response size check | Compare response sizes for bots with normal page sizes | Challenge pages returned with status 200 |
Note the limitation of user-agent tests: many protections look at IP addresses and behaviour as well as user agents, so a request from your laptop pretending to be GPTBot may be treated differently from a real one. Logs of genuine crawler traffic are the most reliable evidence.
Reading the results
- 403 Forbidden: the request was explicitly blocked by a rule.
- 429 Too Many Requests: rate limiting is active; occasional 429s are acceptable, constant ones are not.
- 503 Service Unavailable: sometimes used by protection systems or overloaded servers.
- 200 with a tiny response: often a JavaScript challenge or interstitial page instead of your content.
- No requests at all: either the crawler does not visit, or it is blocked before your logging point, which is why CDN logs matter.
Special cases that catch people out
Some blocks are not security rules at all but side effects of other settings. They are easy to miss because the site looks normal to visitors:
- Maintenance or coming-soon modes left active for bots after launch, often through a plugin that shows the real site only to logged-in users or known IP addresses.
- Geo-blocking introduced to reduce spam or for licensing reasons. Many crawlers operate from data centres in the United States, so blocking that region can block them too.
- Cookie or consent walls that replace the page content until a choice is made. Crawlers do not click buttons.
- Staging protections copied to production, such as HTTP authentication or IP allow lists, during a migration.
- Hotlink or referrer rules that refuse requests without a referrer, which is exactly how crawlers request pages.
After any launch, migration or change of hosting, request a few key pages the way a crawler would, without cookies, without a referrer and from outside your office network, and confirm that you receive the full content.
Fixing it without opening the door to abuse
You do not have to choose between blocking everything and allowing everything. A balanced setup:
- Decide your policy per crawler type. For example: allow search engines and AI search crawlers, decide separately on training crawlers, block unverified bots.
- Use verified bot lists. Many CDNs maintain lists of verified crawlers checked by IP range. Allow verified bots rather than trusting user-agent strings, which anyone can fake.
- Set reasonable rate limits that stop abuse without cutting off crawlers that fetch a few pages per second.
- Exempt robots.txt, sitemaps and llms.txt from challenges, so crawlers can at least read your instructions.
- Cache aggressively so that crawler traffic is served from the edge and does not load your server.
- Keep robots.txt and edge rules consistent, and write the policy down.
- Re-test after changes and check logs again a week later.
How Site SEO AI Audit helps
Site SEO AI Audit reads your robots.txt and reports which AI crawlers are allowed and which are blocked, in the AI visibility area of the report. When the audit’s own crawler, SEOAuditBot, receives errors or blocked responses on your pages, those show up as crawl issues too, which is often the first hint that a firewall or CDN is stricter than you thought. The crawler follows robots.txt and makes about three requests per second. You can run a free audit to see what a crawler experiences on your site.
Related reading
- How to Allow or Block AI Crawlers in robots.txt
- How to Find and Verify AI Crawlers in Your Server Logs
- 12 AI Search Mistakes That Make Your Site Invisible
The bottom line
Your robots.txt states a policy, but your CDN, firewall, host and plugins enforce it. Check each layer for AI bot settings, test with logs rather than assumptions, and align everything with one deliberate policy that allows the crawlers you want and blocks the rest. Then re-check after every infrastructure change, because these settings drift silently.
DUK
Can my CDN block AI crawlers even if robots.txt allows them?
Yes. CDNs enforce their own bot and security rules before requests reach your server. If a CDN setting blocks AI bots, robots.txt permissions make no difference.
How do I know if AI crawlers get blocked?
Filter your CDN or server logs by AI user agents and look at the status codes. Frequent 403, 429 or suspiciously small 200 responses indicate blocking or challenges.
Is it safe to allow AI crawlers through my firewall?
It is safest to allow verified crawlers, checked against published IP ranges or your CDN’s verified bot list, while still blocking unverified traffic that only claims to be a known bot.
Do JavaScript challenges block crawlers?
Usually yes. Crawlers generally cannot solve challenge pages or CAPTCHAs, so they receive the challenge instead of your content. Exempt verified crawlers from challenges.
Should I block AI crawlers to reduce server load?
Consider caching and sensible rate limits first. Blocking AI search crawlers removes your pages from the answers they power, which usually costs more than the bandwidth saved.


