Site SEO AI Auditby Internet Solutions

Noindex vs Disallow: How to Keep Pages Out of Search

6 tháng 8, 20267 phút đọcSEO kỹ thuật
Noindex vs Disallow: How to Keep Pages Out of Search

Short answer: use noindex when a page may be crawled but should not appear in search results, and use a robots.txt Disallow when you want crawlers to stop requesting a group of URLs at all. Noindex removes pages from the index; Disallow only stops crawling, and a blocked URL can still be indexed from links. Never put both on the same URL, because a crawler that is blocked cannot see the noindex.

Two tools for two different jobs

Search engines work in stages. First they crawl, meaning they request a URL and download what it returns. Then they index, meaning they decide whether to store the page and make it eligible for results. Noindex and Disallow act at different stages, which is why they behave so differently:

Everything else follows from this order. Google explains the directive in its guide to blocking indexing with noindex.

How noindex works

There are two ways to set noindex. In the page’s HTML head:

<meta name="robots" content="noindex">

Or as an HTTP response header, which also works for PDFs, images and other files that have no HTML head:

X-Robots-Tag: noindex

Key points about noindex:

How robots.txt Disallow works

A Disallow rule in robots.txt tells compliant crawlers not to request matching URLs:

User-agent: *
Disallow: /search/
Disallow: /*?sort=

What it does well:

What it does not do:

Noindex vs Disallow side by side

Question noindex robots.txt Disallow
Stops crawling? No, the page must be crawled Yes, for compliant crawlers
Removes from search results? Yes, after the next crawl No, the URL can still be indexed
Works on non-HTML files? Yes, via X-Robots-Tag header Yes
Scales to URL patterns? Per page or per template Yes, with wildcards
Saves crawl resources? Only slowly, over time Yes, immediately
Keeps content private? No No
Typical use Thank-you pages, thin archives, filtered pages you still link to Internal search, carts, endless parameters, crawl traps

Why you should not combine them on the same URL

A very common pattern in audits: a site owner wants pages gone from Google, adds noindex, and to be thorough also blocks them in robots.txt. The result is the opposite of what they wanted. Because the crawler is blocked, it never sees the noindex, so pages that were already indexed stay indexed, often with a “No information is available for this page” snippet.

The correct order for removing a group of indexed pages is:

  1. Make sure the pages are not blocked in robots.txt.
  2. Add noindex to the pages, or return 404 or 410 if the pages should no longer exist at all.
  3. Wait until the pages drop out of the index. Check with the page indexing report and URL Inspection.
  4. Only then, if the URLs keep wasting crawl time, consider adding a Disallow rule.

For an urgent case, such as private data exposed in results, Search Console’s Removals tool hides URLs temporarily while the permanent fix takes effect.

Which one to use: common scenarios

Most real decisions fall into a handful of situations:

Mistakes that remove the wrong pages

Noindex is powerful, and a mistake in a template can remove whole sections. Watch for these:

  1. “Discourage search engines” left on after launch. WordPress’s reading setting adds noindex to every page. It is useful on a development copy and disastrous on a live site.
  2. Noindex on paginated category pages, which over time weakens the crawl path to products or posts listed deeper.
  3. SEO plugin settings that noindex a whole post type, such as products or portfolio items, after an update or import.
  4. A noindex in the HTTP header set by a server rule or CDN, which does not show up when you look at the page source.
  5. Noindex pages still listed in the XML sitemap, sending contradictory signals.
  6. Canonical pointing to a noindexed page, which leaves search engines with no clear page to index.

After any theme, plugin or server change, check a few key templates, home, category, product and article, for both the meta robots tag and the X-Robots-Tag header.

How to check what is actually happening

Because a directive can come from several places, verify on the live page rather than in settings:

How Site SEO AI Audit reports indexing directives

The crawl and index area of the audit covers robots.txt, noindex and canonicals together. SEOAuditBot follows robots.txt like a search engine, records noindex from both meta tags and headers, and flags contradictions such as noindex pages listed in your sitemap. Each issue lists the affected pages and is weighted by how much of the site it touches, so a template-wide noindex lands at the top of the fix list. WordPress sites get the exact steps in wp-admin and the SEO plugin. Start with a free audit to see which pages crawlers can index.

Related reading

The bottom line

Noindex keeps pages out of results but requires crawling; Disallow stops crawling but does not keep URLs out of results. Choose based on the job, never block a page whose noindex you need search engines to see, and use authentication for anything that must stay private.

FAQ

Can a page blocked by robots.txt appear in Google?

Yes. If other pages link to it, Google can index the URL without crawling it, usually showing it without a description. To keep it out, allow crawling and add a noindex directive.

Does noindex also mean nofollow?

No. A noindex page’s links can still be followed unless you add nofollow too. Over a long time, though, noindexed pages tend to be crawled less often, so do not rely on them as the main route to important pages.

How long does noindex take to remove a page?

It takes effect the next time the page is crawled, which can be days for popular pages and weeks for rarely crawled ones. You can request a recrawl of individual URLs with URL Inspection in Search Console.

Is noindex in robots.txt supported?

No. Google stopped supporting a noindex rule inside robots.txt in 2019. Use a meta robots tag or an X-Robots-Tag HTTP header instead.

What should I use for a staging site?

Use password protection or an IP restriction. Robots.txt and noindex are both easy to copy to the live site by accident, and neither actually keeps people or links away from the staging content.

#Crawling#Indexing#robots.txt#Technical SEO
Kiểm tra website của bạn — miễn phí.Mọi lỗi SEO trên website của bạn — và cách sửa chính xác.
Bắt đầu miễn phí
Internet Solutions

Sản phẩm khác từ đội ngũ chúng tôi

Do Internet Solutions phát triển. Hãy thử các sản phẩm khác của chúng tôi — mỗi sản phẩm giúp bạn tiết kiệm thời gian theo một cách riêng.

internet-solutions.net ↗
01Tự động đăng mạng xã hội
PostRSS

Bài mới từ nguồn cấp RSS của bạn được tự động đăng lên Facebook, X, LinkedIn, Telegram và hơn 60 mạng khác.

Gói miễn phí · từ 2014Truy cập →
02Chat trực tuyến AI cho website
Talkmio

Website của bạn trả lời khách truy cập 24/7 từ chính nội dung của bạn, bằng ngôn ngữ của họ.

Gói miễn phí · không cần thẻTruy cập →
03Trợ lý AI
Ask Mio

Trò chuyện, viết code, thiết kế, viết bài và nghiên cứu. Mio chọn mô hình tốt nhất cho từng việc.

Gói miễn phíTruy cập →
04Lái tự động AI cho blog và mạng xã hội
AI Blog Autopilot

AI viết bài SEO dài 2.000–3.000 từ và chia sẻ từng bài lên hơn 58 mạng xã hội.

3 bài đầu tiên miễn phíTruy cập →
05Kiểm tra sức khỏe website
Site AI Audit

SEO, tốc độ, SSL, bảo mật và cấu hình email trong một báo cáo, sắp xếp theo việc cần sửa trước.

Lần kiểm tra đầu tiên miễn phíTruy cập →
06Nguồn cấp RSS và sản phẩm
RSS Feed Creator

Tạo RSS từ bất kỳ trang web nào, cùng nguồn cấp sản phẩm cho Google và Meta tự động cập nhật.

Gói miễn phíTruy cập →
07Phát triển website và SEO
Internet Solutions

Website, cửa hàng trực tuyến và hệ thống theo yêu cầu, do đội ngũ của chúng tôi thiết kế, xây dựng và vận hành.

Từ 2011Truy cập →
Site SEO AI Audit
Tổng quan quyền riêng tư

Website này dùng cookie để mang lại trải nghiệm người dùng tốt nhất có thể. Thông tin cookie được lưu trong trình duyệt của bạn và thực hiện các chức năng như nhận ra bạn khi bạn quay lại, giúp đội ngũ chúng tôi hiểu phần nào của website bạn thấy thú vị và hữu ích nhất.