Site SEO AI Auditpar Internet Solutions

Google-Extended Explained: What Blocking It Does and Doesn’t

13 août 20267 min de lectureRecherche IA
Google-Extended Explained: What Blocking It Does and Doesn’t

Short answer: Google-Extended is a product token you can use in robots.txt to tell Google whether content it crawls from your site may be used to train Gemini models and to ground answers in Gemini apps and certain developer products. It is not a separate crawler, and blocking it does not affect Google Search, your rankings or AI Overviews, which follow normal Googlebot and snippet rules. Use it if you want to limit Gemini use of your content while staying fully visible in Search.

What Google-Extended is

Google introduced Google-Extended in 2023 as a way for publishers to manage how their content is used for Google’s generative AI products. It appears in Google’s documentation of its crawlers as a “standalone product token”. That wording matters. Most entries in robots.txt refer to an actual bot that makes requests. Google-Extended does not make requests. Googlebot and Google’s other crawlers do the fetching, and Google then checks the Google-Extended rules to decide what it may do with the content.

According to Google’s documentation, the token manages whether content crawled from your site may be used for training future generations of Gemini models, and for grounding, which means supplying content from the Search index to a model at the moment it answers a question in Gemini apps and in Google’s grounding services for developers. The exact list of products covered is set out in Google’s overview of Google crawlers, which is the reference to check if anything changes.

Because it is a token rather than a crawler, you will never see “Google-Extended” in your server logs. If you are looking for evidence that the rule works, logs are the wrong place.

What blocking it does

When you disallow Google-Extended, you ask Google not to use the affected content for the purposes the token covers. In practical terms:

These effects are policy commitments from Google. You cannot directly observe them, just as you cannot directly observe how any company uses crawled data. What you can verify is that your robots.txt is correctly written and served.

What blocking it does not do

This is where most confusion lies. Blocking Google-Extended does not:

In short, Google-Extended separates the Gemini use of your content from its Search use. It does not separate Search from AI features inside Search.

How the Google controls compare

Control Where you set it What it affects Side effects
Disallow Googlebot robots.txt Crawling for Search, including AI features Pages drop out of Search over time
noindex Meta robots tag or HTTP header Indexing in Search Page removed from all Search results
nosnippet / max-snippet Meta robots tag or HTTP header Text shown in snippets and AI Overviews Less attractive regular results
data-nosnippet HTML attribute on an element Specific passages in snippets and AI features Only the marked text is excluded
Google-Extended robots.txt Gemini training and grounding None in Search

The table shows the trade-off clearly. There is no control that keeps you in regular Search results but removes your content from AI Overviews without also affecting how your snippets appear. Google-Extended is the only control that has no Search side effect, but it only covers Gemini products outside Search.

How to set it up

Adding the rule takes a minute. To block Gemini use of the whole site, add this group to robots.txt:

User-agent: Google-Extended
Disallow: /

To block only part of the site, for example a paid research section, disallow that path instead:

User-agent: Google-Extended
Disallow: /research/

A few practical points:

  1. Keep your Googlebot rules separate. Do not put Google-Extended in the same group as Googlebot unless you intend both to have the same rules.
  2. Check the served file. On WordPress, a physical robots.txt file overrides the virtual one generated by WordPress and SEO plugins. Edit the one that is actually served.
  3. Allow time. Google caches robots.txt, typically for up to a day, so changes are not instant.
  4. Document the decision. Note why you chose to allow or block, so the next person editing the file does not undo it by accident.

Should you block Google-Extended?

It depends on what you want from Gemini. There are sensible reasons both ways.

Whatever you decide, apply the same thinking to other AI companies’ crawlers. A policy that blocks Google-Extended but allows every other training crawler, or the reverse, is usually an accident rather than a strategy.

Applebot-Extended and similar tokens

Google is not the only company to use this pattern. Apple offers Applebot-Extended, which lets publishers opt out of their content being used to train Apple’s AI models while Applebot continues to crawl for features such as Siri and Spotlight suggestions. The logic is the same: the token controls a use, not a crawl, so it will not appear in logs.

Other operators take a different approach and run separate crawlers for training and for search, such as GPTBot and OAI-SearchBot. In those cases you control each bot by name. When you review robots.txt, it helps to sort entries into “tokens that control use” and “bots that fetch pages” so that you understand what each line actually changes.

Common mistakes

Google-Extended is simple, but audits regularly find the same errors around it:

After any change, reload the live robots.txt in a browser, read the groups line by line, and confirm that Googlebot’s rules are exactly what they were before.

How Site SEO AI Audit helps

Site SEO AI Audit reads your robots.txt and reports which AI crawlers and tokens are allowed or blocked in the AI visibility area of its report, next to checks for llms.txt, JavaScript-dependent content and quotable page structure. Because the same crawl also checks noindex tags and snippet-related issues, it helps you spot the more damaging mistake of restricting Googlebot or snippets when you only meant to limit Gemini. You can run a free audit to see your current settings.

Related reading

The bottom line

Google-Extended is a narrow, safe control. It lets you decide whether Google may use your content for Gemini training and grounding, with no effect on Search, rankings or AI Overviews. Block it if your content is your product, allow it if exposure in Gemini helps your business, and remember that controlling AI features inside Search requires snippet controls instead.

FAQ

Is Google-Extended a crawler?

No. It is a product token used only in robots.txt. Google’s existing crawlers fetch the pages, and the token tells Google whether that content may be used for Gemini training and grounding.

Will blocking Google-Extended remove me from AI Overviews?

No. AI Overviews are part of Google Search and follow Googlebot and snippet controls. Google-Extended only covers Gemini apps and related products outside Search.

Does blocking Google-Extended affect my rankings?

No. Google states that Google-Extended does not affect inclusion or ranking in Google Search. Your pages continue to be crawled and indexed normally.

Why can’t I see Google-Extended in my server logs?

Because it never makes requests. Googlebot and Google’s other crawlers do the fetching, so their user agents appear in logs instead.

Does blocking it remove content already used for training?

No. Robots.txt changes apply going forward. Content that was already used to train earlier models is not removed by adding the rule.

#AI crawlers#AI Overviews#robots.txt
Vérifiez votre propre site — gratuitement.Chaque problème SEO de votre site — et comment le corriger précisément.
Essai gratuit

Plus d’articles du blog

Tous les articles →
Internet Solutions

Plus de notre équipe

Conçus par Internet Solutions. Découvrez nos autres produits — chacun vous fait gagner du temps à sa manière.

internet-solutions.net ↗
Site SEO AI Audit
Aperçu de la confidentialité

Ce site utilise des cookies afin de vous offrir la meilleure expérience utilisateur possible. Les informations des cookies sont stockées dans votre navigateur et remplissent des fonctions telles que vous reconnaître lorsque vous revenez sur notre site et aider notre équipe à comprendre quelles sections du site vous trouvez les plus intéressantes et utiles.