Google-Extended

Google-Extended: a robots.txt token, not a crawler.

Who runs it, what it controls, why it never shows in your logs, and how Googlebot differs from it in Search.

Short answer

What is this token?

Before you write a Google-Extended rule, note that Googlebot Smartphone and Googlebot Desktop obey the same robots.txt token, so you cannot target either one separately. Google notes that blocking Googlebot from crawling a page does not keep its URL out of search results.

In brief

Four things to know first.

  • A token, not a bot

    Read Google's crawler documentation before you add a rule for Google-Extended. Google says some crawlers have more than one user agent token, and you need to match only one for a rule to apply.

  • It scopes Gemini use

    Googlebot is a separate crawler, and Google's documentation ties it to Google Search.

  • Nothing shows in your logs

    Most Googlebot crawl requests come from the mobile crawler, and a minority come from the desktop crawler. Both crawler types obey the same robots.txt token.

  • Block it without leaving Search

    Googlebot is the generic name for the two crawlers Google Search uses. Googlebot Image, Video and News have their own tokens, but they also follow a Googlebot rule.

What it is

Who runs it, and what it does.

Google runs this token. It is a product token you can name in robots.txt, in the same file where you address Googlebot. It has no crawler of its own.

According to Google, Googlebot Image, Video and News each have their own token. They also respond to the generic Googlebot token. One matching token is enough for a rule to apply.

The tokenGooglebot
What it isA robots.txt tokenA crawler Google runs
ControlsGemini training and groundingCrawling for Google Search
In your logsRarely appearsAppears as Googlebot
If you block itSearch is unaffectedYou leave Google Search

Recognise it

Why you won't see it in logs.

Other crawlers often fake the Googlebot user agent. To verify a request, run a reverse DNS lookup on the source IP, or match it against Google's published Googlebot IP ranges.

Spoofing is common. Google says the Googlebot user-agent header is often faked, so check the source IP, not the string.

  • Reverse DNS

    Google's common crawlers resolve to hostnames matching crawl-***-***-***-***.googlebot.com or geo-crawl-***-***-***-***.geo.googlebot.com.

  • Published IP ranges

    Match the source IP against the ranges Google lists in common-crawlers.json.

  • Wildcard the version

    The Chrome/W.X.Y.Z string in Google user agents changes over time. Use wildcards in log filters.

Control it

Block Google-Extended in robots.txt.

Add a rule that names the token. Google's documentation says its common crawlers obey robots.txt on automatic crawls, so the rule is honored; check the current wording in Google's own documentation before you ship it.

  1. Open your robots.txt

    Edit the file at the root of your domain, for example example.com/robots.txt.
  2. Name the token

    Start a rule group with a User-agent line naming the token you want to target.
  3. Set the rule

    Add Disallow: / beneath it to opt your whole site out of this use of your content.
  4. Leave Googlebot alone

    Do not add a Disallow for Googlebot. Per Google's documentation, blocking Googlebot affects Google Search, Discover, Images, Video and News.
  5. Test and monitor

    Fetch robots.txt to confirm the rule is live. Google's common crawlers obey robots.txt rules when they crawl automatically.

Unsure what your robots.txt blocks?

Our technical GEO work checks crawler access, so the pages you want cited can be reached and read.

The trade-off

Should you block it?

Per Google's documentation, blocking Googlebot affects Google Search (including Discover and all Search features) as well as Google Images, Video and News. Check Google's current documentation for exactly what a rule for this token covers.

Googlebot remains the crawler that matters for being found in Search. Blocking a page from crawling also does not stop its URL from appearing in results.

What each robots.txt choice affects
ChoiceEffectRisk to Search
Block this tokenCheck Google's documentation for what it coversNone, per Google's brief
Block GooglebotStops Google Search crawlingYou leave Search
Allow bothDefault behaviourNone

FAQ

Google-Extended: common questions.

Is this token a crawler?

Google's common crawlers obey robots.txt rules when they crawl automatically. Googlebot Smartphone and Googlebot Desktop share one token, so one rule covers both.

Does blocking this token affect my Google rankings?

Search crawling depends on Googlebot, not on this token. Per Google's documentation, blocking Googlebot affects Google Search (including Discover and all Search features) as well as Google Images, Video and News.

How do I add this token to robots.txt?

Google says you need to match only one crawler token for a rule to apply. Googlebot Smartphone and Googlebot Desktop obey the same token, so robots.txt cannot target either one alone.

How can I verify real Google crawlers in my logs?

Verify Google's crawlers with a reverse DNS lookup on the source IP, or match it against Google's published IP ranges. Google warns the Googlebot user-agent header is often spoofed, so the string alone proves nothing.

Will blocking this token hide my site from Gemini answers?

Track your AI answers before and after you change the rule. AI mention tracking measures whether brands, URLs and sources appear inside generated answers.

Find out what AI engines can read on your site.

We run your buyers' questions through ChatGPT, Claude, Perplexity and Gemini and show you the gap.