Short answer: X-Robots-Tag is an HTTP response header that carries the same indexing directives as the meta robots tag, such as noindex and nofollow. Because it lives in the header rather than the HTML, it works for PDFs, images, videos, feeds and any other file type, and it can be applied to whole folders or file types with one server rule. It is powerful and invisible in the page source, so check headers directly after every server or CDN change.
What X-Robots-Tag is
Most people know the meta robots tag:
<meta name="robots" content="noindex">
It works well for HTML pages, but many files have no HTML head: PDFs, Word documents, images, videos, JSON and XML files. For those, search engines read directives from the HTTP response instead:
HTTP/1.1 200 OK
Content-Type: application/pdf
X-Robots-Tag: noindex
Google documents the supported directives and syntax in its guide to robots meta tag and X-Robots-Tag specifications. Other major search engines support the core directives as well.
Supported directives
| Directive | What it does |
|---|---|
noindex |
Do not show this URL in search results |
nofollow |
Do not follow links on this resource |
none |
Same as noindex, nofollow |
nosnippet |
Do not show a text snippet or video preview |
max-snippet:[n] |
Limit snippet length to n characters |
max-image-preview:setting |
Limit image preview size: none, standard or large |
noimageindex |
Do not index images on this page |
unavailable_after:[date] |
Stop showing the URL after a given date |
indexifembedded |
Allow indexing when embedded in another page despite noindex |
Directives can be combined with commas, and you can target a specific crawler by prefixing its name, for example X-Robots-Tag: googlebot: noindex. Without a crawler name, the directive applies to all crawlers that support it.
When to use the header instead of the meta tag
- Non-HTML files. PDFs that duplicate web pages, internal price lists, old brochures, image files you do not want in image search.
- Whole folders or file types. One server rule can noindex every PDF in a downloads folder, instead of editing each file.
- Pages where you cannot edit the HTML, for example output from a third-party application served under your domain.
- Staging environments as an extra layer, applied to every response by the server configuration, next to password protection.
- Feeds and API responses that sometimes get indexed and appear as odd search results.
For ordinary HTML pages managed in a CMS, the meta tag is easier to maintain and easier to see. Use one method per URL; if both are present and conflict, search engines apply the most restrictive directive.
How to set X-Robots-Tag
Apache (in the virtual host configuration or .htaccess, with mod_headers enabled):
<FilesMatch "\.pdf$">
Header set X-Robots-Tag "noindex"
</FilesMatch>
nginx (inside the relevant server block):
location ~* \.pdf$ {
add_header X-Robots-Tag "noindex";
}
Application code: most languages can send a header before output, for example in PHP with header('X-Robots-Tag: noindex');. On WordPress, SEO plugins handle meta robots for pages and posts; for media files and other paths, server rules are usually the cleanest method.
CDNs: many CDNs let you add response headers with rules based on path or file type. That is convenient, but it also means a header can be added without anyone seeing it in the server configuration.
Important rules for it to work
- The URL must be crawlable. If robots.txt blocks it, search engines never fetch the response and never see the header.
- It must be on the final response. A header on a redirect response does not apply to the target.
- It takes effect on the next crawl, which for rarely visited files can take weeks.
- Test the exact file type and path; nginx location matching and Apache file matching are easy to get subtly wrong.
- Watch caching. CDNs may cache responses with or without the header; purge after changes.
The dangerous side: invisible noindex
Because the header does not appear in the page source, a mistaken X-Robots-Tag can go unnoticed for a long time. Real-world ways it happens:
- A staging configuration that adds
noindexto all responses is copied to production during a server move. - A CDN rule meant for one folder is written with a broader pattern and matches the whole site.
- A security or performance plugin adds headers that include robots directives.
- An application framework adds
noindexto error responses, and a bug makes normal pages go through the error path.
The result looks exactly like a site-wide meta noindex: pages drop out as they are recrawled. Checking the source of the page shows nothing wrong, which sends people looking in the wrong place.
The safest habit is to include a header check in every deployment and infrastructure change. After a server move, CDN change, new security plugin or framework upgrade, request the home page, a category or service page, an article and a PDF, and read the headers. It takes a minute and catches the problem before search engines do. For larger sites, a scheduled crawl that alerts on new noindex pages provides the same protection automatically.
How to check X-Robots-Tag headers
- Command line:
curl -sI https://www.example.com/page/shows all response headers. Look forx-robots-tagin any capitalisation. - Browser DevTools: in the Network panel, select the document request and read the response headers.
- URL Inspection in Search Console reports whether indexing is allowed and, if not, that a noindex was detected, including one sent in the header.
- A crawl that records response headers finds every URL with the header, which is the only practical way to spot a rule that matches more than intended.
Test several file types: HTML pages, a PDF, an image and the sitemap. A rule meant for PDFs should not appear on HTML pages, and nothing should noindex your sitemap in a way that confuses reporting.
Common use cases, done right
- PDFs that duplicate web pages: add
noindexfor the PDF path or file type, and keep the HTML pages indexable, so searchers land on the page with navigation and context rather than a bare document. - Internal documents: use password protection. The header can be an extra layer, but it does not keep people out.
- Product images appearing without context:
noimageindexon the page or a header on the image path, used sparingly, because image search can also bring visitors. - An offer page that expires on a set date:
unavailable_afterwith the date, then remove or redirect the page when the offer ends. - Staging environments: a password plus a site-wide noindex header in the staging server configuration, which is never deployed to the live site.
Snippet and preview controls
Not every directive is about keeping pages out of search. Several control how a page appears when it is shown:
nosnippetremoves the text snippet under the result. It is rarely a good idea for normal pages, because the snippet is what persuades people to click, and it can also limit how the content is used in search features.max-snippetsets a maximum snippet length in characters. A value of-1means no limit, and0behaves like nosnippet.max-image-preview:largeallows large image previews, which can make results more attractive in some search surfaces. Many sites set it deliberately on article templates.max-video-previewlimits the length of video previews in seconds.
These directives can be set in the meta robots tag or in the header, and the same rules apply: they must be on the final, crawlable response. Because they influence how content appears in results, change them deliberately and check the effect in Search Console performance data rather than setting them once and forgetting them.
How Site SEO AI Audit checks it
SEOAuditBot records robots directives from both meta tags and HTTP headers on every page it crawls. The crawl and index area reports noindex pages, including those set by X-Robots-Tag, and highlights contradictions such as noindex URLs listed in your sitemap. A header rule that accidentally matches a whole template appears at the top of the fix list, because it affects many pages. You can run a free audit to see every directive your server sends.
Related reading
- Noindex vs Disallow: how to keep pages out of search
- Staging site indexed by Google? How to fix and prevent it
- HTTP status codes for SEO: the ones that actually matter
The bottom line
X-Robots-Tag brings indexing directives to any file type and lets you apply them with a single server rule. Use it for PDFs, media and paths you cannot edit, keep those URLs crawlable so the header is seen, and check response headers after every server, CDN or plugin change, because an unseen noindex header can remove a whole site from search.
BUJ
What is the difference between meta robots and X-Robots-Tag?
They support the same directives. The meta tag sits in the HTML head of a page, while X-Robots-Tag is an HTTP header, so it also works for PDFs, images and other files and can be applied to many URLs with one server rule.
Can I use X-Robots-Tag to noindex PDFs?
Yes. It is the standard way to keep PDFs and other non-HTML files out of search results. Make sure the files are not blocked in robots.txt, so crawlers can see the header.
How do I see if a page has an X-Robots-Tag?
Check the HTTP response headers with curl, browser developer tools or a crawler. The header does not appear in the page source, which is why it is easy to miss.
What happens if meta robots and X-Robots-Tag conflict?
Search engines combine the directives and apply the most restrictive one. If either says noindex, the page will not be indexed.
Can X-Robots-Tag target only Google?
Yes. Prefix the directive with the crawler name, for example googlebot: noindex. Without a name, it applies to all crawlers that support the header.


