ScreenshotNeo

BlogEngineering

How Google Scrapes Websites: Inside Googlebot’s Crawl and Index

Learn how Googlebot discovers, fetches, renders, and indexes pages—and how to diagnose crawling, robots.txt, noindex, and sitemap problems.

By the ScreenshotNeo team30 September 20269 min read

How Google Scrapes Websites: Inside Googlebot’s Crawl and Index

When people say Google “scrapes” a website, they usually mean Google’s web crawling and indexing pipeline. Googlebot discovers URLs, requests the resources it is allowed to fetch, may render JavaScript, and sends the result through indexing systems. A successful fetch does not guarantee indexing, ranking, or appearance in Search.

The practical model is:

  1. Discovery: Google learns a URL from links, sitemaps, redirects, or other references.
  2. Crawl and fetch: Googlebot schedules a request and downloads the response and important resources.
  3. Render: For pages that need it, Google renders HTML, CSS, and JavaScript with a recent Chrome version.
  4. Index processing: Google evaluates content, metadata, duplicates, canonical signals, and whether the page is suitable for the index.
  5. Search presentation: An indexed page may or may not rank or appear for a particular query.

Google describes these as separate stages in its crawling and indexing documentation. Treating them separately makes SEO debugging much faster.

What Googlebot is and how it finds URLs

Googlebot is the name for Google’s automated crawlers. Googlebot Smartphone and Googlebot Desktop use the same robots.txt product token, and Google primarily indexes the mobile version for most sites. You cannot use robots.txt to permit one Googlebot subtype while blocking the other.

Google’s pipeline separates discovery, fetching, rendering, indexing, and Search presentation.
Google’s pipeline separates discovery, fetching, rendering, indexing, and Search presentation.

Google discovers URLs mainly by following crawlable links from pages it already knows. A sitemap is an additional discovery channel. It can list URLs and metadata such as last-modified dates, but it is a hint rather than a command to crawl or index every entry. Google’s sitemap guidance explicitly says that submitting a sitemap does not guarantee that Google will download it or use it for crawling URLs.

Make important URLs discoverable

  • Link important pages with ordinary HTML links whose destinations are real URLs.
  • Publish an accurate XML sitemap and keep its last-modified values honest.
  • Use canonical URLs consistently in links, redirects, and metadata.
  • Remove accidental URL variants caused by tracking parameters, session IDs, and unbounded filters.
  • Make navigation understandable without requiring a click handler that produces no crawlable destination.

A sitemap file is limited to 50 MB uncompressed or 50,000 URLs. Split larger inventories into multiple sitemap files and reference them from a sitemap index.

What happens during a Googlebot fetch

Google’s systems schedule requests according to crawl demand for your URLs and the capacity of your site. There is no universal fixed number of requests that every site should try to consume. Server errors, slow responses, connection failures, and repeated timeouts can cause Googlebot to reduce activity. Google says its crawlers try not to overload sites.

For each request, check the complete delivery chain:

Layer What to inspect Typical failure
DNS and TLS Resolution, certificate, protocol negotiation Intermittent DNS or certificate errors
HTTP response Status, redirects, content type, compression 5xx errors, redirect loops, soft 404s
HTML Title, links, canonical, robots directives Empty shell or accidental noindex
Resources CSS, JavaScript, images, fonts, APIs Robots rules or authentication block assets
Rendering Final DOM and visible content Content appears only after a failing script

Googlebot fetches referenced resources separately. Google’s documentation states that its fetch limit is the first 2 MB for supported file types and the first 64 MB for PDF files, measured on uncompressed data. Keep critical content and links available in the initial response when possible.

Rendering JavaScript pages

Google may render a fetched page with a recent version of Chrome. Rendering is not a license to hide essential content behind fragile client-side code. If JavaScript fails, takes too long, depends on a blocked API, or creates links only after an interaction Google cannot complete, the rendered view can be incomplete.

Use server-rendered HTML or progressive enhancement for primary text, headings, canonical information, and navigation. Then verify the rendered result in Search Console URL Inspection. Check that:

  • The main text exists in the rendered DOM.
  • Important links have stable, crawlable destinations.
  • CSS does not hide the content unintentionally.
  • APIs needed to build the page do not require a user session.
  • Robots rules allow critical scripts, styles, and images to load.

Crawled, indexed, and ranking are different states

A URL can be known but not yet crawled, crawled but not indexed, indexed but absent for a query, or indexed with a different canonical URL. “Crawled – currently not indexed” does not automatically mean that Google failed to read the page. It means the page did not enter the index at that point in processing.

Google evaluates usefulness, duplication, canonical relationships, content signals, and user demand. Google’s troubleshooting guidance says a page may not appear even when it has been crawled. Allowing crawling therefore cannot guarantee visibility.

Why a crawled page might not be indexed

  • It duplicates another page and Google selected a different canonical.
  • The response is thin, empty, soft-404-like, or mainly boilerplate.
  • Rendering failed and the useful content was absent from the final page.
  • A noindex directive was present when Google processed the URL.
  • The URL is new and has not completed indexing evaluation.
  • Internal links and sitemap signals do not establish that the page matters.

Do not infer a ranking penalty from one inspection result. Compare representative URLs, inspect server logs, and look for site-wide patterns.

robots.txt, noindex, and authentication

These controls affect different stages and are not interchangeable.

Control Purpose Can Google see the page? What it guarantees
robots.txt Disallow Limit crawling of paths or resources Not necessarily Requests may be blocked; the URL can still be known or surfaced
noindex meta or HTTP header Exclude a crawlable page from Search Yes, if fetch is allowed Google can process the exclusion instruction
Authentication Keep content private No, without credentials Unauthenticated users and crawlers cannot access the content

Google’s official wording is precise: “blocking Googlebot from crawling a page doesn’t prevent the URL of the page from appearing in search results.” If robots.txt blocks a page, Google may never see its noindex tag. Use robots.txt for crawl control, noindex when Google should fetch and then exclude a page, and authentication for genuinely private content.

Example robots.txt

User-agent: *
Disallow: /admin/
Disallow: /internal-search/
Sitemap: https://example.com/sitemap.xml

Place this file at the site root, such as https://example.com/robots.txt. Keep rules narrow. Blocking an entire site or a directory containing canonical pages is an easy deployment mistake.

Example noindex HTML and HTTP header

<meta name="robots" content="noindex,follow">
HTTP/1.1 200 OK
X-Robots-Tag: noindex

The crawler must be able to request the response containing the directive. Do not combine a robots.txt block with a noindex rule when your objective is reliable exclusion.

How to check whether Googlebot can access a page

  1. Open Search Console URL Inspection for the exact URL.
  2. Review whether the URL is known, crawled, indexed, or excluded.
  3. Inspect the reported canonical and compare it with your intended canonical.
  4. Check the rendered page and screenshot of the fetched result where available.
  5. Review Page Indexing and Crawl Stats reports for patterns across the site.
  6. Compare the inspection time with server logs and deployment events.

For a command-line first check, inspect status, redirects, and headers:

curl -I -L https://example.com/page

This does not impersonate Googlebot or prove what Google sees. It is useful for finding obvious HTTP and redirect failures. User-agent strings can be spoofed, so if logs contain a suspicious request claiming to be Googlebot, verify it with reverse DNS or compare the source address with Google’s published crawler IP guidance.

Sitemaps and crawl budget

For most sites, Google recommends keeping the sitemap current and monitoring the Page Indexing report. Crawl-budget work matters most for very large or frequently changing sites. Start with an inventory of URL variants and server health.

Useful actions include consolidating duplicate URLs, returning accurate status codes, limiting unbounded faceted navigation, and removing session identifiers from crawlable links. Repeatedly adding and removing robots.txt rules does not generally transfer crawl capacity to other URLs. A sitemap helps discovery; it does not force immediate crawling or indexing.

Or skip the browser setup

If you need a reliable visual record of what a public page looks like while auditing crawl and rendering issues, ScreenshotNeo provides a website screenshot API. It accepts one GET request and returns PNG, JPEG, WebP, or PDF. Cookie and consent banners, newsletter popups, and chat widgets are removed before capture, and each cleanup step can be turned off.

Only clean shots are billed. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

See the ScreenshotNeo API documentation for all options. The basic call is:

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo also supports full-page captures with lazy images loaded, CSS-selector element captures, dark mode, device presets and custom viewports, retina scale, custom CSS and JavaScript, click actions, selector waits, delays, network-idle waits, request blocking, custom headers and cookies, user agents, authorization, timezone and geolocation, transparent backgrounds, resizing, cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification. Parameter names used by other screenshot APIs also work, which can simplify migration.

The Free plan includes 1,000 screenshots each month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Create a free ScreenshotNeo account.

Troubleshooting checklist

“Blocked by robots.txt”

Cause: A matching Disallow rule prevents the request or an important resource. Fix: Test the exact path, including trailing slashes and query variants, then narrow or remove the rule. Do not expect a blocked URL to disappear from Search automatically.

A clean capture removes common overlays before recording the page.
A clean capture removes common overlays before recording the page.

“Excluded by noindex”

Cause: A meta tag or X-Robots-Tag was delivered. Fix: Remove it from templates, middleware, and preview environments if indexing is intended. Ensure the page is crawlable so Google can observe the change.

“Crawled – currently not indexed”

Cause: Google processed the URL but did not select it for the index at that time. Fix: Check canonicalization, duplication, rendered content, internal links, status codes, and sitemap accuracy. Avoid repeated request submissions without addressing evidence.

Pages appear blank

Cause: JavaScript errors, blocked APIs, slow client rendering, or content that requires interaction. Fix: inspect the rendered DOM, browser console, network requests, and server-rendered fallback. Permit critical resources.

Googlebot requests slow or stop

Cause: 5xx responses, timeouts, overloaded infrastructure, or unstable DNS. Fix: correlate Crawl Stats with origin and CDN logs, reduce latency, fix errors, and keep capacity available during launches.

Suspected fake Googlebot traffic

Cause: Any client can send a Googlebot user-agent string. Fix: perform reverse-DNS verification and compare the address with Google’s published crawler IP ranges before treating it as Google traffic.

Performance, reliability, and cost considerations

  • Performance: Keep HTML useful on the first response, reduce render-blocking work, and avoid generating unlimited URL combinations.
  • Reliability: Monitor DNS, TLS, origin status, redirects, resource availability, and deployment changes. A single successful manual fetch is not a site-wide guarantee.
  • Indexing expectations: Google says updates are checked and indexed in a reasonably timely way, but for most sites this can be three days or more. Treat that as guidance, not a service-level promise.
  • Measurement: Use URL Inspection for individual evidence and Page Indexing, Crawl Stats, logs, and sitemap reports for patterns.
  • Screenshot cost: If you use ScreenshotNeo for visual audits, failed loads, bot checks, blank pages, timeouts, and cache hits are not billed. Select a cache TTL, resize output when appropriate, and use bulk capture for up to 100 URLs per call.

FAQ

Does Google scrape every URL in my sitemap?

No. A sitemap is a discovery hint and does not guarantee downloading, crawling, or indexing.

Can robots.txt remove a page from Google?

No. It controls crawling. A known blocked URL can still appear based on external references. Use noindex on a crawlable page or authentication for private content.

Does Google run JavaScript?

Google may render pages with a recent Chrome version, but rendering can fail when scripts, APIs, or resources are inaccessible. Provide crawlable HTML for essential content.

Why is the mobile version important?

Google primarily indexes the mobile version for most sites. Ensure mobile HTML contains the same essential content and metadata as desktop.

How can I prove a request was from Googlebot?

Do not rely only on the user-agent string. Use reverse DNS or compare the source IP with Google’s published crawler ranges.

How quickly will a new page appear?

There is no guaranteed schedule. Google’s guidance says most sites may wait three days or more for updates, while highly time-sensitive content can differ.