Site SEO AI Auditod Internet Solutions

Crawl Budget Explained: When It Matters and How to Save It

7 sierpnia 2026Czas czytania: 8 minSEO techniczne
Crawl Budget Explained: When It Matters and How to Save It

Short answer: crawl budget is the number of URLs a search engine is able and willing to crawl on your site in a given time. It depends on how much your server can handle and how much the search engine wants your content. Sites with a few thousand pages rarely need to worry about it; large shops, sites with faceted filters and sites that generate URLs automatically do, and the fix is almost always to remove low-value URLs and speed up server responses.

What crawl budget actually means

Search engines do not crawl every URL on the web all the time. For each site they balance two things, which Google describes in its guide to managing crawl budget for large sites:

Crawl budget is roughly the combination of the two: the URLs the search engine wants to crawl, limited by what your server can handle. It is not a fixed number you can look up, and it changes over time.

Does your site need to care?

Honestly, most sites do not. If your site has a few hundred or a few thousand pages, and new pages usually get crawled within a few days of publishing, crawl budget is not your problem. Spending time on it would be better spent on content and internal links.

It starts to matter when one or more of these is true:

Where crawl budget gets wasted

Waste comes from URLs that cost a request but bring nothing to search. The usual suspects:

How to see how your site is being crawled

You do not have to guess. Three sources together give a clear picture:

  1. Search Console Crawl stats report. Found under Settings, it shows total crawl requests, average response time and a breakdown by response code, file type and purpose (discovery versus refresh). A rising response time or a high share of errors is a warning sign.
  2. Server access logs. Filter requests by verified search engine user agents and count which URL patterns they hit. If half of the requests go to ?sort= URLs, you have found your leak.
  3. A full site crawl. Crawling the site yourself shows how many unique URLs your internal links expose. If a site with 5,000 products produces 300,000 crawlable URLs, the difference is the waste.

How to reduce crawl waste

Fixes fall into two groups: stop exposing useless URLs, and make the useful ones clearly more important.

How to increase crawl capacity

The other half is making each request cheaper for your server:

If a crawler is overloading your server, the right short-term tool is returning 503 or 429 briefly, not blocking it in robots.txt, which has longer-lasting effects.

A quarterly crawl check for large sites

For a site where crawling really matters, a short routine every few months keeps problems from building up unnoticed:

  1. Compare three numbers: the pages you want indexed (from your database or CMS), the URLs a full crawl finds through internal links, and the pages Search Console reports as indexed. Large gaps between them show where to look.
  2. Review the Crawl stats trend. Check whether total requests, average response time and the share of 5xx errors have changed since the last review, and link any change to releases or hosting changes.
  3. Sample the logs. Take one week of verified crawler requests and group them by URL pattern. Any pattern that takes a large share of requests but has no search value is a candidate for blocking or removal.
  4. Test new features before launch. New filters, sorting options, search features and tracking parameters are the most common source of new crawl waste. Check whether they produce crawlable URLs before they go live.
  5. Check the time to crawl new pages. Note when a batch of new products or articles was published and when they were first crawled. If that delay grows, investigate before it becomes a traffic problem.

This routine takes an hour or two and usually catches issues months before they show up as lost traffic.

Crawl budget myths

A few ideas circulate that do not hold up:

Myth Reality
Every site should optimise crawl budget Small and medium sites are usually crawled fully; it matters mainly for large or URL-heavy sites
Noindex saves crawl budget Noindex pages still have to be crawled to see the tag; only reducing links or blocking saves requests
Nofollow on internal links saves crawl budget The URLs may still be discovered elsewhere; fix the URLs themselves
More crawling means better rankings Crawling is a precondition, not a ranking factor
You can set crawl rate in robots.txt for Google Googlebot ignores the crawl-delay rule

How Site SEO AI Audit helps with crawl efficiency

An audit crawl shows the same structure a search engine sees. SEOAuditBot follows your links and your sitemap, respects robots.txt, and reports the issues that waste crawling: redirect chains, broken links, soft errors, duplicate and parameter URLs, pages that are deep in the click structure or orphaned, and server response times on every page. Each issue shows how many pages it affects, so you can see whether you have a handful of bad URLs or a systemic leak. Larger sites can be crawled up to their plan limit; see plan details.

Related reading

The bottom line

Crawl budget is the balance between what your server can handle and what search engines want to crawl. If your site is small and new pages are crawled quickly, ignore it. If it is large or generates URLs automatically, find the patterns that waste requests, block or remove them, and make your server respond fast and without errors.

FAQ

How do I know if I have a crawl budget problem?

Typical signs are new pages taking weeks to be crawled, many URLs marked “Discovered – currently not indexed” in Search Console, and server logs showing crawlers spending most requests on parameter or junk URLs. Small sites rarely show these signs.

Does site speed affect crawl budget?

Yes. Faster, error-free server responses allow search engines to crawl more pages without overloading your server. Slow responses and 5xx errors cause them to reduce crawling.

Should I block faceted navigation in robots.txt?

Block the filter combinations that have no search value and create huge numbers of URLs. Keep valuable filtered pages, such as a brand or major attribute within a category, crawlable with clean URLs if people search for them.

Does noindex reduce crawling?

Not immediately. A noindexed page must still be crawled for the directive to be seen. Over time such pages tend to be crawled less often, but blocking or removing links is the direct way to save requests.

Can I ask Google to crawl my site more?

There is no setting to increase crawling. You can improve server speed, fix errors, publish content people want and keep sitemaps accurate. Google adjusts crawling based on those signals.

#Crawling#E-commerce SEO#Site architecture#Technical SEO
Sprawdź swoją stronę — za darmo.Każdy problem SEO na Twojej stronie — i dokładnie, jak go naprawić.
Zacznij za darmo
Internet Solutions

Więcej od naszego zespołu

Stworzone przez Internet Solutions. Wypróbuj nasze pozostałe produkty — każdy oszczędza czas na swój sposób.

internet-solutions.net ↗
Site SEO AI Audit
Przegląd prywatności

Ta strona używa plików cookie, abyśmy mogli zapewnić Ci jak najlepsze wrażenia. Informacje z plików cookie są przechowywane w Twojej przeglądarce i pełnią funkcje takie jak rozpoznawanie Cię po powrocie na stronę oraz pomagają naszemu zespołowi zrozumieć, które sekcje strony są dla Ciebie najciekawsze i najbardziej przydatne.