Short answer: hreflang connects pages that search engines can crawl, index and show. If an alternate is set to noindex, blocked in robots.txt, redirects elsewhere or returns an error, it cannot serve as a working member of the cluster: search engines either cannot read its return links or will not show it. Every hreflang target should return status 200, be allowed in robots.txt, carry no noindex and be its own canonical. When a page should stay out of search, remove it from the hreflang clusters instead.
Why hreflang depends on indexability
hreflang is a way of telling search engines: “instead of this page, show that equivalent page to searchers in that language or country.” For that swap to happen, the equivalent page must be something the search engine can show. A page that is excluded from the index cannot be swapped in.
There is also the return-link requirement. Search engines confirm hreflang relationships by reading the annotations on both pages. If one page cannot be crawled, its annotations cannot be read, and the relationship cannot be confirmed from that side.
So hreflang is only as strong as the indexability of every page in the cluster. Conflicts with noindex, robots.txt, redirects and errors are among the most common reasons hreflang silently stops working.
Conflict 1: noindex on an alternate
A language version carries <meta name="robots" content="noindex"> or an equivalent HTTP header, while other versions list it as an alternate. Common causes:
- A new language was built on staging with noindex, and the setting was carried over to production.
- A multilingual plugin or SEO plugin sets noindex on untranslated or “incomplete” pages.
- Someone noindexed a language temporarily during a review and forgot to remove it.
The noindexed page will not appear in search, so searchers in its language get another version or nothing. If the noindex is intended, remove that page from all hreflang clusters. If it is not intended, remove the noindex, and the cluster starts working as designed once the page is recrawled.
Conflict 2: alternates blocked in robots.txt
A robots.txt rule blocks a language folder, a subdomain or a URL pattern that includes some alternates. Blocked pages cannot be crawled, so search engines cannot read their content, their canonical or their hreflang. The result is that return links from those pages are invisible.
This often happens by accident:
- A rule meant to block a staging path such as
/dev/also matches a language folder such as/de/because the pattern was written too broadly. - Parameter rules block
?lang=URLs on sites that still use parameters for some languages. - A separate robots.txt on a language subdomain or country domain is left from development and blocks everything.
Check robots.txt on every host involved: each subdomain and each country domain has its own file. The Robots Exclusion Protocol is host-specific, so the main domain’s file does not apply to de.example.com.
Conflict 3: redirected alternates
An hreflang link points to a URL that redirects, for example an old slug, a URL without a trailing slash, an http version or a version that redirects visitors by location. Search engines may follow the redirect for crawling, but the relationship is declared with a URL that is not the final page, so it does not match the return link on the destination page.
Fix these by pointing every hreflang link directly at the final URL, exactly as that page’s canonical declares it. On sites with many redirected alternates, the root cause is usually that hreflang is generated from a different URL source than canonicals, or that URL changes are not propagated to all versions.
Conflict 4: errors and missing pages
An hreflang link points to a page that returns 404, 410 or a server error. This typically happens when:
- A translation was deleted, but the other versions still list it.
- Templates output hreflang for every configured language, whether or not a translation exists.
- A product was discontinued in one market but kept in others.
Remove links to missing pages from every version’s cluster. A cluster that lists only existing, working versions is always better than a larger cluster with broken members.
Soft errors deserve a mention too. Some sites return status 200 for pages that are effectively missing, showing a “page not found” message or an empty template instead of a proper 404. Search engines may classify these as soft 404s and drop them from the index. In an hreflang cluster, a soft 404 behaves like a missing page, even though a status check shows 200. Make sure missing translations return a real 404 or 410, and that they are removed from clusters.
Conflict 5: canonicals pointing elsewhere
A page’s canonical tag points to a different URL, often to another language or a parameter-free version. Search engines treat the canonical target as the main page and may disregard the hreflang on the non-canonical one. The fix is to give each language version a self-referencing canonical and to point hreflang only at canonical URLs.
Special case: content you keep out of one market on purpose
Sometimes a business deliberately hides a page in one market while keeping it in others. A product may not be approved for sale in one country, a promotion may be legally restricted, or a service may not be offered there yet. The instinct is often to noindex or block the page in that market’s version while leaving the hreflang templates untouched. That produces exactly the conflicts described above.
A cleaner approach is to treat the page as not existing in that market:
- Do not publish the page in the restricted market’s folder or domain at all, if possible.
- Remove that market from the page’s hreflang cluster in every other version, so no version points to a page that is hidden.
- Keep the language switcher honest, hiding or disabling the restricted market for that page.
- Consider x-default for searchers in the restricted market; they will be shown whatever fallback you define, which may be appropriate or may need a clear notice about availability.
If the page must exist in the restricted market for legal or practical reasons, for example to explain that the product is not available there, it can remain as a normal, indexable page with that message and stay in the cluster. What should be avoided is a page that is listed as an alternate but hidden from search at the same time.
Document these exceptions. On sites with many markets, a simple list of which pages are excluded from which markets, and why, prevents someone from “fixing” the missing hreflang later and reintroducing the conflict.
Quick reference: what to do in each case
| Situation | Intended? | Action |
|---|---|---|
| Alternate has noindex | Yes | Remove it from all hreflang clusters |
| Alternate has noindex | No | Remove noindex; recrawl |
| Alternate blocked in robots.txt | Yes | Remove it from clusters; consider noindex instead of blocking if it must be crawled |
| Alternate blocked in robots.txt | No | Fix the rule; check every host’s robots.txt |
| Alternate redirects | Either | Point hreflang at the final URL |
| Alternate returns 404 or 410 | Either | Remove it from clusters, or restore the page |
| Alternate canonicalises elsewhere | Rarely | Self-referencing canonical, or remove from clusters |
Finding conflicts before they cost traffic
A quick manual test for any single cluster takes a few minutes. Take the list of hreflang URLs from one page, request each one and note the status code, then view each page’s source and look for a robots meta tag and the canonical. Finally, open the robots.txt of each host involved and check whether any rule matches the paths. If every URL returns 200, is allowed, has no noindex and canonicalises to itself, the cluster is technically sound.
These conflicts are invisible to visitors: the pages load, the switcher works and nothing looks wrong. They show up only when you check status codes, robots rules, meta robots and canonicals for every hreflang target, which is exactly what a crawler does well. Check after every release that touches templates, robots.txt, plugins or URL structure.
Site SEO AI Audit crawls a site like a search engine and checks status codes, robots.txt, noindex and canonicals in its crawl and indexing area, and hreflang return links and broken language versions in its Languages area. Seeing both in one report makes conflicts between them easy to spot, and each issue is weighted by the number of pages it affects. The first audit is free.
Related reading
- Noindex vs disallow: how to keep pages out of search
- hreflang return links: how to fix “no return tags” errors
- Canonical tags and hreflang: how to use them together
The bottom line
hreflang only works between pages that are crawlable, indexable, final and canonical. noindex, robots.txt blocks, redirects, errors and cross-pointing canonicals all break clusters silently. Decide whether each exclusion is intended; if it is, remove the page from hreflang, and if it is not, remove the exclusion. Then recrawl to confirm that every alternate returns 200 and confirms the relationship.
BUJ
Can a noindex page be part of an hreflang cluster?
It should not be. A noindex page cannot be shown in search, so it cannot be swapped in for its language. Remove it from the cluster or remove the noindex.
What happens if one language folder is blocked in robots.txt?
Search engines cannot crawl those pages or read their hreflang, so return links from them are missing and the cluster is weakened for every language.
Is it a problem if hreflang points to a redirect?
Yes. hreflang should point to final URLs that return 200. A redirected alternate does not match the return link on the destination page.
Does each subdomain need its own robots.txt check?
Yes. robots.txt applies per host, so each subdomain and each country domain has its own file that must allow the language versions to be crawled.
Should I noindex untranslated pages?
It is better not to create translated URLs until a translation exists. If such pages exist, noindex them and keep them out of all hreflang clusters.


