How to Verify Googlebot: We Found 698 Fakes in 30 Days
Anyone can send Googlebot's user agent. On our site the fakes outnumbered the real crawler 12 to 1. Here is how to count only verified Googlebot, and why it matters.
Contents
We were trying to answer a simple question about one of our sites: is Googlebot still crawling it? Filtering our request logs for the Googlebot user agent said yes, busily. Filtering for Google’s verified crawler said it had made 58 requests in 30 days.
The difference was 698 requests from fake Googlebots — and they were not confused, they were hostile. They were probing for /.env, /Dockerfile and similar files.
Had we trusted the user-agent count, we would have concluded the opposite of the truth about our own site. Here is how to avoid that.
Why the user agent proves nothing
The User-Agent header is a string the client chooses. There is no authentication in it. Sending
Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)
requires one flag in curl. Scanners do it constantly, for two reasons: some sites whitelist Googlebot past rate limits and bot protection, and some operators simply do not want their scans in the “suspicious” bucket.
So on any site with public exposure, “Googlebot traffic” measured by user agent is Googlebot plus whoever is currently probing you. The ratio is not a rounding error. Ours was twelve fakes for every real request.
Method 1: reverse DNS, which is what Google documents
Google’s own documented verification is a two-step DNS check on the requesting IP address:
- Run a reverse DNS lookup on the IP. The result must end in
googlebot.comorgoogle.com. - Run a forward lookup on that hostname. It must resolve back to the original IP.
Both steps matter. Step one alone is spoofable by anyone who controls a DNS zone; the round trip is what makes it sound.
# 1. reverse lookup
host 66.249.66.1
# → 1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com.
# 2. forward lookup on that hostname
host crawl-66-249-66-1.googlebot.com
# → crawl-66-249-66-1.googlebot.com has address 66.249.66.1
If the forward lookup returns a different address, or the reverse lookup returns anything not under googlebot.com or google.com, it is not Googlebot.
Google also publishes its crawler IP ranges as JSON, which is faster for bulk log processing than a DNS lookup per line:
https://developers.google.com/static/search/apis/ipranges/googlebot.json— the crawlershttps://developers.google.com/static/search/apis/ipranges/special-crawlers.json— special caseshttps://developers.google.com/static/search/apis/ipranges/user-triggered-fetchers.json— fetches caused by a user action
Match the request IP against those ranges. Refresh the file periodically; the ranges change.
Method 2: let your CDN classify it
If your traffic passes through a CDN, it has already done this work. On Cloudflare the analytics API exposes a verified-bot classification, and querying by that field instead of by user agent is the difference between our two numbers.
The distinction is worth naming precisely, because it is the whole point:
| What you filter on | What you get |
|---|---|
userAgent contains "Googlebot" | real Googlebot plus every impostor |
| verified-bot category | only crawlers the CDN authenticated |
In a GraphQL analytics query against Cloudflare, the useful dimensions are the verified-bot flag and the detected bot category, not the raw user-agent string. Group by response status as well — you want to know what the crawler received, not only that it came.
What we actually learned from counting properly
Once the fakes were out, the picture was unambiguous and not what the raw logs suggested.
| Verified crawler | Requests, 30 days | Requests to articles |
|---|---|---|
| Googlebot | 58 | 3 |
| bingbot | 124 | — |
| YandexBot | 285 | — |
Three article requests in a month, on a site with over 130 articles. At that rate a full pass takes roughly three and a half years. That single number reframed everything we thought about the site: we had been asking why Google refused our content, when the answer was that Google had largely stopped looking at it.
The fakes, meanwhile, told their own story. They were not crawling — they were scanning for secrets. The paths they requested were /.env, /Dockerfile and similar configuration files, which is a credential hunt, not an indexing attempt.
Three practical consequences
Your crawl-rate dashboards may be fiction. Any tool that counts crawler hits from raw logs without verification inherits this problem. If your “Googlebot activity” graph looks healthy while your indexing does not, verify before you trust the graph.
Never whitelist by user agent. If bot protection, rate limits or a WAF rule let “Googlebot” through based on the header alone, you have built a bypass and published the password. Whitelist verified bots, or by IP range.
Check what the crawler received, not just that it arrived. Group crawler requests by response status. We found verified Googlebot still requesting /feed/, /comments/feed/, /page-sitemap.xml and /category-sitemap.xml — all 404 — four months after a migration removed them. The crawler was following a map of a site that no longer existed. That is a different problem from not being crawled, and you only see it if you look at status codes.
The quick version
- The Googlebot user agent is a claim, not an identity.
- Verify by reverse DNS round trip, by Google’s published IP ranges, or by your CDN’s verified-bot classification.
- Then look at what the verified crawler actually fetched and what status it got back.
On our site those three steps turned a comfortable-looking number into 58 requests and 3 article fetches — which was the real story, and the one we needed.