Short answer: For years the only tools for AI crawlers were allow or block in robots.txt. Newer options add a middle ground: CDN features that charge AI crawlers per request or require a deal before access, machine-readable licensing terms, and emerging standards for declaring how content may be used by AI. These options are real but young, and most only work with crawlers that choose to take part. For most businesses the practical steps today are a clear robots.txt policy, verifying what your CDN already does, and measuring AI crawler activity before deciding whether charging is worth it.
Why “allow or block” is no longer the whole story
Robots.txt gives a binary answer per crawler and path: you may crawl, or you may not. That worked when crawlers were search engines that sent visitors back. AI crawlers changed the balance. Some collect content for training models, some fetch pages to answer questions in real time, and many send far fewer visitors per request than traditional search.
Publishers and content businesses started asking a different question: not “should this bot get in?” but “on what terms?”. That question led to three families of solutions:
- Paid access at the network level, where the CDN or a gateway asks crawlers to pay per request.
- Licensing terms and deals, from bilateral contracts to machine-readable licence files.
- Usage preferences, standard signals saying which uses (training, search, summarisation) you allow.
The broader decision about training use is covered in our guide should you let AI train on your content.
Pay-per-crawl: how it works
Pay-per-crawl moves the decision from a text file to the network. The idea, introduced by large CDN providers such as Cloudflare in 2025, works roughly like this:
- The site owner sets a price per request for AI crawlers, or chooses to allow or block them.
- When a participating crawler requests a page, the network checks whether it has agreed to pay.
- If not, the crawler receives an HTTP
402 Payment Requiredresponse with pricing information instead of the content. - If it has agreed, it receives the page and the charge is recorded and settled through the provider.
The 402 status code has existed in the HTTP standard for decades, reserved for future use (see RFC 9110). Pay-per-crawl gives it a practical role.
Important limits: these schemes depend on crawlers identifying themselves honestly and signing up. They work through the provider’s network, so they are only available if your site uses that provider. Availability, pricing models and participating crawlers change quickly, so check the current status with your CDN before planning around it.
Default blocking at the CDN
A related and more immediate change: some CDNs and hosting platforms now block or challenge known AI crawlers by default, or offer a one-click switch. Many site owners do not know which setting applies to them.
This matters in both directions. If you want AI search tools to read and cite your pages, a default block can make you invisible to them without any change in robots.txt. If you want to block training crawlers, the CDN setting may already do more than your robots.txt. Check both, and make sure they tell the same story. Our guide on whether your CDN or firewall blocks AI crawlers shows how to test it.
Licensing: contracts and machine-readable terms
Licensing is the oldest route. Large publishers and platforms have signed direct agreements with AI companies, allowing use of their content in exchange for payment, attribution or technology. For most businesses, direct deals are out of reach: AI companies negotiate with owners of large, valuable archives.
Machine-readable licensing tries to make terms available to everyone. Initiatives such as RSL (Really Simple Licensing), announced in 2025 by a group of publishers and platforms, propose a standard file that states the licence terms for AI use, referenced from robots.txt, with collective organisations handling payment. The concept is promising, but its effect depends on whether AI companies adopt and honour it.
If your content is a real product, such as research, premium articles or specialised data, keeping an eye on these standards is worthwhile. For a typical business website, licensing income is unlikely to be significant.
Usage preferences and emerging standards
A third approach separates access from use. Instead of blocking a crawler, you declare which uses you allow: for example search and citation yes, model training no. Several signals exist or are in development:
- Crawler-specific user agents, where companies separate training crawlers from search or user-triggered fetchers, so robots.txt can treat them differently.
- Control tokens such as Google-Extended, which controls use for certain AI models without affecting Google Search. See Google-Extended explained.
- Standardisation work, including an IETF working group on AI usage preferences, which aims to create a common vocabulary that could be expressed in robots.txt or HTTP headers.
- Informal meta tags such as noai, which are not widely supported.
The general principle: signals only work when the receiving side honours them. Named user agents with published policies are currently the most reliable lever. Our guide to allowing or blocking AI crawlers in robots.txt lists how to write the rules.
Comparing your options
| Option | What it does | Maturity | Best for |
|---|---|---|---|
| Allow all | Maximises chance of being read and cited | Established | Most businesses that want visibility |
| Selective robots.txt | Allows search and answer bots, blocks training crawlers | Established, depends on bot honesty | Sites that want citations but limit training |
| CDN block or challenge | Enforces blocking at the network level | Established, provider-specific | Sites under heavy bot load or with valuable content |
| Pay-per-crawl | Charges participating crawlers per request | Early, provider-specific | Publishers with content AI companies want |
| Licensing terms or deals | Sets conditions and payment for use | Deals established; open standards early | Large archives, premium or specialised content |
How to decide what fits your site
Work through these questions before changing anything:
- What do you sell? If you sell products or services and content is marketing, visibility in AI answers is usually worth more than any crawl fee. If content is the product, protection and payment matter more.
- How much AI crawling happens? Measure it first. Our guide on finding and verifying AI crawlers in server logs shows how.
- What do AI tools send back? Check referral traffic from AI assistants in your analytics. If it is meaningful, blocking or charging the crawlers behind it has a cost.
- What does your CDN already do? A default setting may already have made the decision for you.
- What is behind paywalls or logins? Content that should never be reused belongs behind access control, not only behind a robots.txt rule. See paywalls and gated content in AI search.
Many sites end up with a mixed policy: marketing pages open to everyone, premium or proprietary sections blocked or priced, and training crawlers treated differently from AI search fetchers.
Risks and open questions
Before committing to charging or licensing, weigh the uncertainties:
- Participation is voluntary. Crawlers that do not identify themselves, or ignore the rules, are not affected by prices or licence files. Only network-level blocking stops them, and even that relies on detection.
- Visibility trade-offs. If an AI search tool cannot read a page, it cannot cite it. A fee that no crawler pays is, in practice, a block.
- Provider lock-in. Pay-per-crawl programmes are tied to a specific network provider. Changing hosting or CDN may mean losing the arrangement.
- Legal questions. How copyright, licensing and data protection rules apply to AI training is still being debated and litigated in several countries. Contracts and terms should be reviewed by a lawyer.
- Fast change. Crawler names, company policies and standards evolve. Whatever you choose, review it at least twice a year.
Keeping your policy consistent
The most common real-world problem is not the choice of strategy but contradictions: robots.txt allows a crawler that the CDN blocks, an SEO plugin adds rules nobody remembers, or a staging-era block was never removed. Site SEO AI Audit’s AI visibility area checks which AI crawlers your robots.txt allows or blocks, whether you have llms.txt and whether content can be read without JavaScript, so you can confirm that the live site matches the policy you chose. The first audit is free.
Related reading
- How to track AI referral traffic in Google Analytics 4
- noai and noimageai meta tags: do they actually work?
- llms.txt explained: what it is and whether you need one
The bottom line
AI crawler control is moving beyond allow and block, towards paid access, licensing terms and standard usage preferences. These tools are promising but still early and depend on crawlers taking part. Most businesses should first measure AI crawling and AI referrals, align robots.txt and CDN settings, and protect truly valuable content with access control. Watch pay-per-crawl and licensing standards if your content is your product.
KKK
What is pay-per-crawl?
It is a model in which AI crawlers pay per page request to access a site. Crawlers that have not agreed to pay receive an HTTP 402 response instead of the content. It is currently offered through specific CDN providers.
Can a small website earn money from AI crawlers?
Possibly a little, but for most small sites the amounts are likely to be small, and blocking or charging may reduce visibility in AI answers. Measure crawler activity and AI referrals before deciding.
Does robots.txt support licensing terms?
Standard robots.txt only expresses allow and disallow rules. Proposals such as RSL reference licence files from robots.txt, but they only work with crawlers that support them.
Will blocking AI training crawlers remove me from AI search?
Not necessarily. Many companies use separate crawlers for training and for live answers. Blocking training crawlers while allowing search and answer bots is a common mixed policy.
Is my CDN already blocking AI crawlers?
It might be. Some providers block or challenge known AI crawlers by default or offer a simple switch. Check your CDN dashboard and test requests with the relevant user agents.


