Site SEO AI AuditInternet Solutions ürünü

Noindex vs Disallow: How to Keep Pages Out of Search

6 Ağustos 20267 dk okumaTeknik SEO
Noindex vs Disallow: How to Keep Pages Out of Search

Short answer: use noindex when a page may be crawled but should not appear in search results, and use a robots.txt Disallow when you want crawlers to stop requesting a group of URLs at all. Noindex removes pages from the index; Disallow only stops crawling, and a blocked URL can still be indexed from links. Never put both on the same URL, because a crawler that is blocked cannot see the noindex.

Two tools for two different jobs

Search engines work in stages. First they crawl, meaning they request a URL and download what it returns. Then they index, meaning they decide whether to store the page and make it eligible for results. Noindex and Disallow act at different stages, which is why they behave so differently:

Everything else follows from this order. Google explains the directive in its guide to blocking indexing with noindex.

How noindex works

There are two ways to set noindex. In the page’s HTML head:

<meta name="robots" content="noindex">

Or as an HTTP response header, which also works for PDFs, images and other files that have no HTML head:

X-Robots-Tag: noindex

Key points about noindex:

How robots.txt Disallow works

A Disallow rule in robots.txt tells compliant crawlers not to request matching URLs:

User-agent: *
Disallow: /search/
Disallow: /*?sort=

What it does well:

What it does not do:

Noindex vs Disallow side by side

Question noindex robots.txt Disallow
Stops crawling? No, the page must be crawled Yes, for compliant crawlers
Removes from search results? Yes, after the next crawl No, the URL can still be indexed
Works on non-HTML files? Yes, via X-Robots-Tag header Yes
Scales to URL patterns? Per page or per template Yes, with wildcards
Saves crawl resources? Only slowly, over time Yes, immediately
Keeps content private? No No
Typical use Thank-you pages, thin archives, filtered pages you still link to Internal search, carts, endless parameters, crawl traps

Why you should not combine them on the same URL

A very common pattern in audits: a site owner wants pages gone from Google, adds noindex, and to be thorough also blocks them in robots.txt. The result is the opposite of what they wanted. Because the crawler is blocked, it never sees the noindex, so pages that were already indexed stay indexed, often with a “No information is available for this page” snippet.

The correct order for removing a group of indexed pages is:

  1. Make sure the pages are not blocked in robots.txt.
  2. Add noindex to the pages, or return 404 or 410 if the pages should no longer exist at all.
  3. Wait until the pages drop out of the index. Check with the page indexing report and URL Inspection.
  4. Only then, if the URLs keep wasting crawl time, consider adding a Disallow rule.

For an urgent case, such as private data exposed in results, Search Console’s Removals tool hides URLs temporarily while the permanent fix takes effect.

Which one to use: common scenarios

Most real decisions fall into a handful of situations:

Mistakes that remove the wrong pages

Noindex is powerful, and a mistake in a template can remove whole sections. Watch for these:

  1. “Discourage search engines” left on after launch. WordPress’s reading setting adds noindex to every page. It is useful on a development copy and disastrous on a live site.
  2. Noindex on paginated category pages, which over time weakens the crawl path to products or posts listed deeper.
  3. SEO plugin settings that noindex a whole post type, such as products or portfolio items, after an update or import.
  4. A noindex in the HTTP header set by a server rule or CDN, which does not show up when you look at the page source.
  5. Noindex pages still listed in the XML sitemap, sending contradictory signals.
  6. Canonical pointing to a noindexed page, which leaves search engines with no clear page to index.

After any theme, plugin or server change, check a few key templates, home, category, product and article, for both the meta robots tag and the X-Robots-Tag header.

How to check what is actually happening

Because a directive can come from several places, verify on the live page rather than in settings:

How Site SEO AI Audit reports indexing directives

The crawl and index area of the audit covers robots.txt, noindex and canonicals together. SEOAuditBot follows robots.txt like a search engine, records noindex from both meta tags and headers, and flags contradictions such as noindex pages listed in your sitemap. Each issue lists the affected pages and is weighted by how much of the site it touches, so a template-wide noindex lands at the top of the fix list. WordPress sites get the exact steps in wp-admin and the SEO plugin. Start with a free audit to see which pages crawlers can index.

Related reading

The bottom line

Noindex keeps pages out of results but requires crawling; Disallow stops crawling but does not keep URLs out of results. Choose based on the job, never block a page whose noindex you need search engines to see, and use authentication for anything that must stay private.

SSS

Can a page blocked by robots.txt appear in Google?

Yes. If other pages link to it, Google can index the URL without crawling it, usually showing it without a description. To keep it out, allow crawling and add a noindex directive.

Does noindex also mean nofollow?

No. A noindex page’s links can still be followed unless you add nofollow too. Over a long time, though, noindexed pages tend to be crawled less often, so do not rely on them as the main route to important pages.

How long does noindex take to remove a page?

It takes effect the next time the page is crawled, which can be days for popular pages and weeks for rarely crawled ones. You can request a recrawl of individual URLs with URL Inspection in Search Console.

Is noindex in robots.txt supported?

No. Google stopped supporting a noindex rule inside robots.txt in 2019. Use a meta robots tag or an X-Robots-Tag HTTP header instead.

What should I use for a staging site?

Use password protection or an IP restriction. Robots.txt and noindex are both easy to copy to the live site by accident, and neither actually keeps people or links away from the staging content.

#Crawling#Indexing#robots.txt#Technical SEO
Kendi web sitenizi kontrol edin — ücretsiz.Sitenizdeki her SEO sorunu — ve tam olarak nasıl düzeltileceği.
Ücretsiz başla

Blogdan daha fazlası

Tüm makaleler →
Internet Solutions

Ekibimizden diğer ürünler

Internet Solutions tarafından geliştirildi. Diğer ürünlerimizi de deneyin — her biri size farklı bir şekilde zaman kazandırır.

internet-solutions.net ↗
Site SEO AI Audit
Gizlilik özeti

Bu web sitesi, size mümkün olan en iyi kullanıcı deneyimini sunabilmek için çerez kullanır. Çerez bilgileri tarayıcınızda saklanır ve sitemize geri döndüğünüzde sizi tanımak, ekibimizin sitenin hangi bölümlerini en ilginç ve faydalı bulduğunuzu anlamasına yardımcı olmak gibi işlevler görür.