Googlebot user agent on 34 homepages: 24 served the same bytes, 3 shut the door

A Googlebot user agent is a string anyone can write. On 34 homepages fetched twice, 24 returned byte-identical HTML to both a browser and a Googlebot user agent, 3 returned less, and 3 refused the fake Googlebot with a 403.

Crawling & Indexing6 min read2905 views
Googlebot user agent on 34 homepages: 24 served the same bytes, 3 shut the door

FIELD TEST · 2026-10-09 · 34 homepages · two user agents

Sample and method: the fixed 34-domain panel this blog has used for its header surveys since 2026-08-15, each homepage fetched twice on 2026-10-09 — once with a Chrome browser user agent, once with a Googlebot user agent string — from the same machine, no rendering, no retries. Four sites blocked both requests and are dropped, leaving 30 below.

The Googlebot user agent is a line of text you write into your own request; it proves nothing and needs no key. We sent it to 34 homepages alongside an ordinary browser user agent, and on most of them it changed nothing at all: 24 of the 30 that answered returned byte-identical HTML to both strings, 3 of them refused the fake Googlebot with a 403, and only Google's own homepage served meaningfully less. The lesson is short — a user agent is a label, not a key.

How we measured it

One HTTP GET per host per user agent, no retries, no JavaScript, no rendering. The panel is a convenience set of developer, news and SaaS homepages plus four sites we run ourselves, not a random sample of the web, so read every count as "among these 30". We recorded the HTTP status, the decoded byte length of the HTML, the page <title>, and the length of the visible text after stripping tags.

The two strings are the only thing that differs between the two fetches. The browser string is a normal Chrome 126 UA; the second string is the standard Googlebot mobile UA, which is public and copyable:

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1;
+http://www.google.com/bot.html) Chrome/126.0 Safari/537.36

One caveat belongs up front, because it shapes how to read the refusals below. Our requests came from a single datacenter IP, and a real Googlebot request comes from Google's published IP ranges and carries matching reverse DNS. Any site that verifies by IP will treat our "Googlebot" as an impostor, which is exactly what the verification is for. That means a 403 here is not evidence that the site blocks the real Googlebot; it is evidence that the site does not take the string on faith.

What the Googlebot user agent changed

Of the 34 homepages, 4 answered 403 to the browser string as well — nytimes.com, stackoverflow.com, medium.com and canva.com block automated fetches from datacenter addresses regardless of user agent, so they cannot be compared and are dropped. The remaining 30 answered the browser fetch, and this is how the Googlebot fetch compared:

ResultHomepagesWhat it means
Byte-identical HTML24The server does not look at the user agent at all
Different byte length3The server varies the response by user agent
403 to the fake Googlebot3The site refuses a self-declared Googlebot

Only six sites varied, and three of them are worth naming. Google's own homepage returned 66,485 bytes to the Googlebot string and 298,383 to the browser one — about 78% less HTML, from its own servers. slack.com returned roughly 12 KB less to the Googlebot string, and github.com differed by a single byte, which is a nonce or a timestamp rather than a decision. Then there are the three refusals.

SiteBrowserGooglebot UADelta
google.com200 · 298,383 B200 · 66,485 B−231,898 B
slack.com200 · 252,783 B200 · 252,764 B−19 B
github.com200 · 577,246 B200 · 577,245 B−1 B
notion.com200403refused
reddit.com200403refused
en.wikipedia.org200403refused

The remaining 24 homepages — from mozilla.org and bbc.com to stripe.com, figma.com and our own four sites — served byte-for-byte the same HTML to both strings. That is the normal case, and it is the right one: for a page that is the same for everyone, the user agent is not something a server needs to read.

What this means for you

Two habits follow from this, and one warning. The habit: never treat the word "Googlebot" in your logs as proof that the real crawler visited — a spoofed string looks identical there, and the only reliable check is reverse DNS against Google's published ranges, which is what how to verify Googlebot walks through. The warning is narrower but sharper: do not build anything that shows different content to a user agent string. The string is trivially forged, it is not the mechanism Google uses to identify itself, and varying the page by it is the shape of cloaking.

  • Log the user agent, but verify before you trust it — reverse DNS plus the published IP list.
  • Serve the same HTML to every user agent; if a page needs less markup for a crawler, that is a rendering question, not a user-agent one.
  • Do not gate content on a user agent string. It is self-declared and anyone can copy it.
  • Do not read a 403 to your own "Googlebot" fetch as a site blocking Google — from a datacenter IP, that is the site doing the right thing.

Common questions

How did you measure this?

We fetched each of 34 homepages twice on 2026-10-09, once with a Chrome browser user agent and once with the public Googlebot user agent string, from one machine, and compared the HTTP status, decoded byte length and <title>. No rendering, no retries. The raw counts are in the table above.

Does Googlebot get different content from a normal browser?

On this panel, usually not. 24 of the 30 comparable homepages returned the same bytes to both. The exceptions were Google's own homepage, which served a smaller page to the Googlebot string, and three sites that refused the string outright. There was no case of a site serving more, or richer, content to the Googlebot string.

Can I just set the user agent to Googlebot to see what Google sees?

No, and this is the point of the test. The user agent string is one of several signals, and the one Google actually relies on for verification is not a string at all — it is the IP address and its reverse DNS. A request from your laptop with a Googlebot string reaches the server looking exactly like an impostor, which is why three sites here answered it with a 403.

Why did four sites block even the browser fetch?

Because they block datacenter IPs, not user agents. nytimes.com, stackoverflow.com, medium.com and canva.com returned 403 to a plain Chrome user agent as well, from an address that is clearly not a residential visitor. That is ordinary bot management, and it is a reason to be careful about any "I fetched it and got 403" conclusion: the block is often about where you fetched from, not what you claimed to be.

The honest limit of this test: we sent a string, not a crawler. We cannot tell whether notion.com, reddit.com or en.wikipedia.org would serve different content to a request from Google's real IP ranges, because we cannot make one. What the data does show is narrower and still useful — the user-agent string on its own changes very little, and when it does change something, the change is a refusal or a smaller page, never a richer one. Front-loading the effort here is a poor trade; if your goal is that crawlers see your content, the accessibility question is the one that pays, which is what what AI crawlers actually see and the AI crawler accessibility check are for.

Googlebot user agent on 34 homepages: 24 served the same bytes, 3 shut the door