Site SEO AI Auditde la Internet Solutions

X-Robots-Tag: How to Control Indexing With HTTP Headers

7 septembrie 20268 min de cititSEO tehnic
X-Robots-Tag: How to Control Indexing With HTTP Headers

Short answer: X-Robots-Tag is an HTTP response header that carries the same indexing directives as the meta robots tag, such as noindex and nofollow. Because it lives in the header rather than the HTML, it works for PDFs, images, videos, feeds and any other file type, and it can be applied to whole folders or file types with one server rule. It is powerful and invisible in the page source, so check headers directly after every server or CDN change.

What X-Robots-Tag is

Most people know the meta robots tag:

<meta name="robots" content="noindex">

It works well for HTML pages, but many files have no HTML head: PDFs, Word documents, images, videos, JSON and XML files. For those, search engines read directives from the HTTP response instead:

HTTP/1.1 200 OK
Content-Type: application/pdf
X-Robots-Tag: noindex

Google documents the supported directives and syntax in its guide to robots meta tag and X-Robots-Tag specifications. Other major search engines support the core directives as well.

Supported directives

Directive What it does
noindex Do not show this URL in search results
nofollow Do not follow links on this resource
none Same as noindex, nofollow
nosnippet Do not show a text snippet or video preview
max-snippet:[n] Limit snippet length to n characters
max-image-preview:setting Limit image preview size: none, standard or large
noimageindex Do not index images on this page
unavailable_after:[date] Stop showing the URL after a given date
indexifembedded Allow indexing when embedded in another page despite noindex

Directives can be combined with commas, and you can target a specific crawler by prefixing its name, for example X-Robots-Tag: googlebot: noindex. Without a crawler name, the directive applies to all crawlers that support it.

When to use the header instead of the meta tag

For ordinary HTML pages managed in a CMS, the meta tag is easier to maintain and easier to see. Use one method per URL; if both are present and conflict, search engines apply the most restrictive directive.

How to set X-Robots-Tag

Apache (in the virtual host configuration or .htaccess, with mod_headers enabled):

<FilesMatch "\.pdf$">
  Header set X-Robots-Tag "noindex"
</FilesMatch>

nginx (inside the relevant server block):

location ~* \.pdf$ {
  add_header X-Robots-Tag "noindex";
}

Application code: most languages can send a header before output, for example in PHP with header('X-Robots-Tag: noindex');. On WordPress, SEO plugins handle meta robots for pages and posts; for media files and other paths, server rules are usually the cleanest method.

CDNs: many CDNs let you add response headers with rules based on path or file type. That is convenient, but it also means a header can be added without anyone seeing it in the server configuration.

Important rules for it to work

  1. The URL must be crawlable. If robots.txt blocks it, search engines never fetch the response and never see the header.
  2. It must be on the final response. A header on a redirect response does not apply to the target.
  3. It takes effect on the next crawl, which for rarely visited files can take weeks.
  4. Test the exact file type and path; nginx location matching and Apache file matching are easy to get subtly wrong.
  5. Watch caching. CDNs may cache responses with or without the header; purge after changes.

The dangerous side: invisible noindex

Because the header does not appear in the page source, a mistaken X-Robots-Tag can go unnoticed for a long time. Real-world ways it happens:

The result looks exactly like a site-wide meta noindex: pages drop out as they are recrawled. Checking the source of the page shows nothing wrong, which sends people looking in the wrong place.

The safest habit is to include a header check in every deployment and infrastructure change. After a server move, CDN change, new security plugin or framework upgrade, request the home page, a category or service page, an article and a PDF, and read the headers. It takes a minute and catches the problem before search engines do. For larger sites, a scheduled crawl that alerts on new noindex pages provides the same protection automatically.

How to check X-Robots-Tag headers

Test several file types: HTML pages, a PDF, an image and the sitemap. A rule meant for PDFs should not appear on HTML pages, and nothing should noindex your sitemap in a way that confuses reporting.

Common use cases, done right

Snippet and preview controls

Not every directive is about keeping pages out of search. Several control how a page appears when it is shown:

These directives can be set in the meta robots tag or in the header, and the same rules apply: they must be on the final, crawlable response. Because they influence how content appears in results, change them deliberately and check the effect in Search Console performance data rather than setting them once and forgetting them.

How Site SEO AI Audit checks it

SEOAuditBot records robots directives from both meta tags and HTTP headers on every page it crawls. The crawl and index area reports noindex pages, including those set by X-Robots-Tag, and highlights contradictions such as noindex URLs listed in your sitemap. A header rule that accidentally matches a whole template appears at the top of the fix list, because it affects many pages. You can run a free audit to see every directive your server sends.

Related reading

The bottom line

X-Robots-Tag brings indexing directives to any file type and lets you apply them with a single server rule. Use it for PDFs, media and paths you cannot edit, keep those URLs crawlable so the header is seen, and check response headers after every server, CDN or plugin change, because an unseen noindex header can remove a whole site from search.

FAQ

What is the difference between meta robots and X-Robots-Tag?

They support the same directives. The meta tag sits in the HTML head of a page, while X-Robots-Tag is an HTTP header, so it also works for PDFs, images and other files and can be applied to many URLs with one server rule.

Can I use X-Robots-Tag to noindex PDFs?

Yes. It is the standard way to keep PDFs and other non-HTML files out of search results. Make sure the files are not blocked in robots.txt, so crawlers can see the header.

How do I see if a page has an X-Robots-Tag?

Check the HTTP response headers with curl, browser developer tools or a crawler. The header does not appear in the page source, which is why it is easy to miss.

What happens if meta robots and X-Robots-Tag conflict?

Search engines combine the directives and apply the most restrictive one. If either says noindex, the page will not be indexed.

Can X-Robots-Tag target only Google?

Yes. Prefix the directive with the crawler name, for example googlebot: noindex. Without a name, it applies to all crawlers that support the header.

#Indexing#Status codes#Technical SEO
Verifică-ți propriul site — gratuit.Fiecare problemă SEO a site-ului tău — și exact cum o rezolvi.
Începe gratuit

Mai multe de pe blog

Toate articolele →
Internet Solutions

Mai multe de la echipa noastră

Create de Internet Solutions. Încearcă și celelalte produse ale noastre — fiecare îți economisește timp în alt fel.

internet-solutions.net ↗
Site SEO AI Audit
Prezentare generală a confidențialității

Acest site folosește cookie-uri pentru a-ți oferi cea mai bună experiență posibilă. Informațiile din cookie-uri sunt stocate în browserul tău și îndeplinesc funcții precum recunoașterea ta când revii pe site și ajutarea echipei noastre să înțeleagă ce secțiuni ale site-ului găsești cele mai interesante și utile.