Site SEO AI Auditот Internet Solutions

Pay-Per-Crawl and AI Licensing: Options for Site Owners

1 октября 2026 г.Время чтения: 8 минAI-поиск
Pay-Per-Crawl and AI Licensing: Options for Site Owners

Short answer: For years the only tools for AI crawlers were allow or block in robots.txt. Newer options add a middle ground: CDN features that charge AI crawlers per request or require a deal before access, machine-readable licensing terms, and emerging standards for declaring how content may be used by AI. These options are real but young, and most only work with crawlers that choose to take part. For most businesses the practical steps today are a clear robots.txt policy, verifying what your CDN already does, and measuring AI crawler activity before deciding whether charging is worth it.

Why “allow or block” is no longer the whole story

Robots.txt gives a binary answer per crawler and path: you may crawl, or you may not. That worked when crawlers were search engines that sent visitors back. AI crawlers changed the balance. Some collect content for training models, some fetch pages to answer questions in real time, and many send far fewer visitors per request than traditional search.

Publishers and content businesses started asking a different question: not “should this bot get in?” but “on what terms?”. That question led to three families of solutions:

  1. Paid access at the network level, where the CDN or a gateway asks crawlers to pay per request.
  2. Licensing terms and deals, from bilateral contracts to machine-readable licence files.
  3. Usage preferences, standard signals saying which uses (training, search, summarisation) you allow.

The broader decision about training use is covered in our guide should you let AI train on your content.

Pay-per-crawl: how it works

Pay-per-crawl moves the decision from a text file to the network. The idea, introduced by large CDN providers such as Cloudflare in 2025, works roughly like this:

The 402 status code has existed in the HTTP standard for decades, reserved for future use (see RFC 9110). Pay-per-crawl gives it a practical role.

Important limits: these schemes depend on crawlers identifying themselves honestly and signing up. They work through the provider’s network, so they are only available if your site uses that provider. Availability, pricing models and participating crawlers change quickly, so check the current status with your CDN before planning around it.

Default blocking at the CDN

A related and more immediate change: some CDNs and hosting platforms now block or challenge known AI crawlers by default, or offer a one-click switch. Many site owners do not know which setting applies to them.

This matters in both directions. If you want AI search tools to read and cite your pages, a default block can make you invisible to them without any change in robots.txt. If you want to block training crawlers, the CDN setting may already do more than your robots.txt. Check both, and make sure they tell the same story. Our guide on whether your CDN or firewall blocks AI crawlers shows how to test it.

Licensing: contracts and machine-readable terms

Licensing is the oldest route. Large publishers and platforms have signed direct agreements with AI companies, allowing use of their content in exchange for payment, attribution or technology. For most businesses, direct deals are out of reach: AI companies negotiate with owners of large, valuable archives.

Machine-readable licensing tries to make terms available to everyone. Initiatives such as RSL (Really Simple Licensing), announced in 2025 by a group of publishers and platforms, propose a standard file that states the licence terms for AI use, referenced from robots.txt, with collective organisations handling payment. The concept is promising, but its effect depends on whether AI companies adopt and honour it.

If your content is a real product, such as research, premium articles or specialised data, keeping an eye on these standards is worthwhile. For a typical business website, licensing income is unlikely to be significant.

Usage preferences and emerging standards

A third approach separates access from use. Instead of blocking a crawler, you declare which uses you allow: for example search and citation yes, model training no. Several signals exist or are in development:

The general principle: signals only work when the receiving side honours them. Named user agents with published policies are currently the most reliable lever. Our guide to allowing or blocking AI crawlers in robots.txt lists how to write the rules.

Comparing your options

Option What it does Maturity Best for
Allow all Maximises chance of being read and cited Established Most businesses that want visibility
Selective robots.txt Allows search and answer bots, blocks training crawlers Established, depends on bot honesty Sites that want citations but limit training
CDN block or challenge Enforces blocking at the network level Established, provider-specific Sites under heavy bot load or with valuable content
Pay-per-crawl Charges participating crawlers per request Early, provider-specific Publishers with content AI companies want
Licensing terms or deals Sets conditions and payment for use Deals established; open standards early Large archives, premium or specialised content

How to decide what fits your site

Work through these questions before changing anything:

  1. What do you sell? If you sell products or services and content is marketing, visibility in AI answers is usually worth more than any crawl fee. If content is the product, protection and payment matter more.
  2. How much AI crawling happens? Measure it first. Our guide on finding and verifying AI crawlers in server logs shows how.
  3. What do AI tools send back? Check referral traffic from AI assistants in your analytics. If it is meaningful, blocking or charging the crawlers behind it has a cost.
  4. What does your CDN already do? A default setting may already have made the decision for you.
  5. What is behind paywalls or logins? Content that should never be reused belongs behind access control, not only behind a robots.txt rule. See paywalls and gated content in AI search.

Many sites end up with a mixed policy: marketing pages open to everyone, premium or proprietary sections blocked or priced, and training crawlers treated differently from AI search fetchers.

Risks and open questions

Before committing to charging or licensing, weigh the uncertainties:

Keeping your policy consistent

The most common real-world problem is not the choice of strategy but contradictions: robots.txt allows a crawler that the CDN blocks, an SEO plugin adds rules nobody remembers, or a staging-era block was never removed. Site SEO AI Audit’s AI visibility area checks which AI crawlers your robots.txt allows or blocks, whether you have llms.txt and whether content can be read without JavaScript, so you can confirm that the live site matches the policy you chose. The first audit is free.

Related reading

The bottom line

AI crawler control is moving beyond allow and block, towards paid access, licensing terms and standard usage preferences. These tools are promising but still early and depend on crawlers taking part. Most businesses should first measure AI crawling and AI referrals, align robots.txt and CDN settings, and protect truly valuable content with access control. Watch pay-per-crawl and licensing standards if your content is your product.

FAQ

What is pay-per-crawl?

It is a model in which AI crawlers pay per page request to access a site. Crawlers that have not agreed to pay receive an HTTP 402 response instead of the content. It is currently offered through specific CDN providers.

Can a small website earn money from AI crawlers?

Possibly a little, but for most small sites the amounts are likely to be small, and blocking or charging may reduce visibility in AI answers. Measure crawler activity and AI referrals before deciding.

Does robots.txt support licensing terms?

Standard robots.txt only expresses allow and disallow rules. Proposals such as RSL reference licence files from robots.txt, but they only work with crawlers that support them.

Will blocking AI training crawlers remove me from AI search?

Not necessarily. Many companies use separate crawlers for training and for live answers. Blocking training crawlers while allowing search and answer bots is a common mixed policy.

Is my CDN already blocking AI crawlers?

It might be. Some providers block or challenge known AI crawlers by default or offer a simple switch. Check your CDN dashboard and test requests with the relevant user agents.

#AI crawlers#AI search#Generative engine optimization#robots.txt
Проверьте свой сайт — бесплатно.Все SEO-проблемы вашего сайта — и как именно их исправить.
Начать бесплатно

Ещё из блога

Все статьи →
Internet Solutions

Другие продукты нашей команды

Сделано Internet Solutions. Попробуйте и другие наши продукты — каждый экономит время по-своему.

internet-solutions.net ↗
01Автопостинг в соцсети
PostRSS

Новые записи из вашего RSS-фида автоматически публикуются в Facebook, X, LinkedIn, Telegram и ещё 60+ сетях.

Бесплатный тариф · с 2014Перейти →
02AI-чат для сайтов
Talkmio

Ваш сайт отвечает посетителям 24/7 на основе вашего контента и на их языке.

Бесплатный тариф · без картыПерейти →
03AI-ассистент
Ask Mio

Чат, код, дизайн, тексты и исследования. Mio подбирает лучшую модель для каждой задачи.

Бесплатный тарифПерейти →
04AI-автопилот для блога и соцсетей
AI Blog Autopilot

AI пишет SEO-статьи на 2000–3000 слов и публикует каждую в 58+ соцсетях.

Первые 3 статьи бесплатноПерейти →
05Проверка здоровья сайта
Site AI Audit

SEO, скорость, SSL, безопасность и настройка почты в одном отчёте — по порядку, что исправлять первым.

Первый аудит бесплатноПерейти →
06RSS и товарные фиды
RSS Feed Creator

Создавайте RSS из любой веб-страницы, а также товарные фиды для Google и Meta, которые обновляются сами.

Бесплатный тарифПерейти →
07Разработка сайтов и SEO
Internet Solutions

Сайты, интернет-магазины и индивидуальные системы — проектирует, создаёт и сопровождает наша команда.

С 2011Перейти →
Site SEO AI Audit
Обзор конфиденциальности

Этот сайт использует cookie, чтобы мы могли обеспечить вам наилучший пользовательский опыт. Информация cookie хранится в вашем браузере и выполняет такие функции, как узнавание вас при повторном посещении сайта, а также помогает нашей команде понять, какие разделы сайта вам наиболее интересны и полезны.