Crawler guide
Googlebot: Google's Search crawler, explained.
Google Search uses two crawler types under one robots.txt token, so you cannot block mobile or desktop alone. Blocking crawling does not keep your URLs out of results.
Short answer
How do Google's AI features find your pages?
Googlebot is Google's crawler for Search. Because other crawlers often spoof its user agent, you verify it with a reverse DNS lookup or Google's published IP ranges. Blocking it in robots.txt stops crawling, but a blocked URL can still appear in search results.
In brief
Four things to know first.
One name, two crawlers
Google's main crawler is the generic name for the smartphone and desktop crawlers, both run by Google Search.
The user agent can lie
Other crawlers often spoof Google's crawler user agent. Check the source IP, not the header.
One robots.txt token
Both crawler types obey the same token, so you cannot target the smartphone or desktop crawler alone.
Blocking costs you Google
Blocking it affects Search, Discover, Images, Video and News.
Who runs it
What is Googlebot?
Google runs this crawler for Google Search. Google's documentation says it is the generic name for two crawler types: the smartphone crawler and the desktop crawler.
For most sites Google primarily indexes the mobile version. So most of its requests come from the mobile crawler, and a minority from the desktop one.
Purpose
Fetches pages so Google Search can index them. It is not an AI training crawler.
Crawl rate
For most sites, Google's crawler shouldn't hit you more than once every few seconds on average.
Size limit
It fetches the first 2MB of a supported file type and the first 64MB of a PDF.
Timezone
From US IP addresses, its timezone is Pacific Time.
Recognise it
Googlebot user agent and verification
Its user agent contains a Chrome/W.X.Y.Z string. That is a version placeholder that rises over time, so use wildcards for the version in log searches.
Google warns that other crawlers often spoof its crawler's user-agent header. Verify each request instead.
Find the request
Find the source IP of each request whose user agent says Googlebot. Google warns that other crawlers often spoof this header. Check the IP with a reverse DNS lookup or Google's published Googlebot IP ranges.Run a reverse DNS lookup
Google says its crawler hostnames start with crawl- or geo-crawl-, followed by the IP digits, on Google's own crawler domain.Or match the IP ranges
Compare the source IP against the IP ranges Google publishes for this crawler. Common crawlers are listed in common-crawlers.json.Block what fails
Treat requests that fail both checks as spoofed, whatever their user agent says.
Control it
Googlebot robots.txt rules
It obeys robots.txt. Google's common crawlers always obey robots.txt rules when crawling automatically.
Google's Image, Video and News crawlers also respond to the generic token for its main crawler. You need to match only one token for a rule to apply.
| Goal | How | Note |
|---|---|---|
| Stop crawling | robots.txt, token for Google's main crawler | URL can still appear in results |
| Stop indexing | noindex on the page | Page must stay crawlable to be read |
| Block all access | Password protection | Keeps out crawlers and visitors |
Not sure what your robots.txt blocks?
We check crawler access, rendering and schema so the pages you want read and cited are reachable.
The trade-off
Should you block Googlebot?
Rarely. Blocking it affects Google Search, including Discover and all Search features, plus Google Images, Video and News.
It also does not fully hide a page. Google says blocking crawling doesn't stop the URL appearing in search results. Use noindex to block indexing.
| Allow Google's crawler | Block Google's crawler | |
|---|---|---|
| Google Search | Pages can be crawled and indexed | Pages leave normal crawling |
| Other Google products | Discover, Images, Video, News work | All are affected |
| Hiding a page | Use noindex or a password | URL can still appear in results |
| Server load | Reduce the crawl rate if needed | Not the right tool for load |
FAQ
Googlebot: common questions.
What is Google's main crawler?
Google's main web crawler serves Search. Google's documentation calls it the generic name for two crawler types, the smartphone crawler and the desktop crawler. It fetches pages so Google can consider them for indexing in Search.
How do I verify a request from Google's crawler?
Requests from Google's main crawler are verified by reverse DNS lookup on the source IP, or by matching the IP against Google's published crawler IP ranges. Google says the user-agent header is often spoofed by other crawlers, so the header alone proves nothing.
Can I block only the smartphone crawler?
The smartphone crawler cannot be blocked alone in robots.txt. Google says both crawler types obey the same product token, so you cannot selectively target either one using robots.txt. A rule for that token applies to both crawlers.
Does blocking Google's crawler remove a page from Google?
Blocking Google's main crawler in robots.txt does not reliably remove a page. Google says it doesn't prevent the URL from appearing in search results. Use noindex to block indexing, or password protection to block all access.
How often does Google's crawler visit my site?
For most sites, Google's crawler shouldn't access your site more than once every few seconds on average, according to Google. Short-term rates can look slightly higher. Google says you can reduce the crawl rate if your site struggles.
Does Google's crawler read my whole page?
Google's crawler reads only part of large files. For Search, it fetches the first 2MB of a supported file type and the first 64MB of a PDF. The limit applies to uncompressed data and to each referenced resource, such as CSS and JavaScript.
Make sure crawlers can reach the pages that matter.
We run your buyers' questions through ChatGPT, Claude, Perplexity and Gemini and show you the gap.