ScreenshotNeo

BlogHow-to

How to Scrape Google Search Results in Python Without Getting Blocked

Learn a policy-aware way to collect Google results in Python, reduce blocks, handle errors, and choose an authorized API for production.

By the ScreenshotNeo team29 September 20269 min read

How to Scrape Google Search Results in Python Without Getting Blocked

Direct answer: do not build a production scraper that repeatedly sends raw automated queries to Google. Google classifies automated queries, including scraping results for rank checking without express permission, as machine-generated traffic. A direct Python scraper can work for a small, permitted experiment, but it is fragile and can trigger CAPTCHAs, 429 responses, IP blocks, or JavaScript challenges. For production, use an official API when your project qualifies or a contractually authorized hosted SERP API. Whichever route you choose, minimize requests, cache and deduplicate queries, avoid unnecessary pages, and follow Google’s terms and machine-readable instructions.

What “without getting blocked” really means

There is no Google-published universal requests-per-hour threshold that guarantees safe scraping. A vendor guide reports that raw scraping may work for about 50 requests before a CAPTCHA, IP block, or JavaScript challenge, but that is vendor experience, not a Google limit. Treat any fixed “safe rate” found online as unverified.

Google’s Terms of Service prohibit automated access that violates machine-readable instructions. Google Search Central describes machine-generated traffic as automated queries and specifically includes scraping results for rank checking or other automated access without express permission in its spam policies. A user-agent string does not grant permission or make a crawler safe.

robots.txt is a traffic-management signal, not authentication. Google explains that robots.txt instructions cannot enforce crawler behavior and that blocked URLs may still appear in Search. If you crawl a site linked from a result, inspect that site’s robots.txt and terms separately; Google’s robots rules concern the site that publishes them.

Choose an access method

Method Policy and permission fit Block exposure Control Maintenance Best use
Direct Python HTTP requests Only where expressly permitted High High over your own request and parser High; markup and challenges change Small, controlled experiments
Browser automation Still requires permission High; JavaScript challenges remain High, including rendered content Very high Testing a permitted workflow that needs a browser
Hosted SERP API Depends on provider contract and your use Provider handles much of it Geography, language, pagination and schema vary Lower; structured output reduces parser work Production systems needing predictable operations
Search Researcher Result API For eligible researchers and non-commercial use Quota-controlled Defined API response and rolling 24-hour limits Lower than HTML scraping Qualifying research projects

The Search Researcher Result API is a narrow option: eligibility, non-commercial program terms and rolling 24-hour request limits apply. Commercial applications need a separate, verified arrangement. Do not assume a research credential covers a commercial rank tracker.

A cautious collection pipeline validates permission, limits traffic and caches accepted responses.
A cautious collection pipeline validates permission, limits traffic and caches accepted responses.

Build a small, policy-aware Python collector

The following example is suitable for a permitted, low-volume experiment. It demonstrates a session, explicit timeout, retry handling for transient server errors, caching, deduplication and a parser that fails clearly when the expected markup is absent. It does not attempt to bypass CAPTCHAs or blocks.

1. Install dependencies

python -m pip install requests beautifulsoup4

2. Complete Python example

from __future__ import annotations

import hashlib
import json
import time
from pathlib import Path
from typing import Any
from urllib.parse import quote_plus

import requests
from bs4 import BeautifulSoup

CACHE_DIR = Path("serp-cache")
CACHE_DIR.mkdir(exist_ok=True)


def cache_path(query: str, page: int) -> Path:
    key = hashlib.sha256(f"{query}\0{page}".encode()).hexdigest()
    return CACHE_DIR / f"{key}.json"


def parse_result_page(html: str) -> list[dict[str, str]]:
    soup = BeautifulSoup(html, "html.parser")
    results = []
    # Selectors are examples only. Google markup changes; do not assume they
    # remain valid. Stop and review when the selector returns no results.
    for block in soup.select("div.MjjYud"):
        link = block.select_one("a[href]")
        heading = block.select_one("h3")
        if not link or not heading:
            continue
        results.append({"title": heading.get_text(" ", strip=True),
                        "url": link["href"]})
    return results


def fetch_permitted(query: str, page: int = 0) -> list[dict[str, str]]:
    path = cache_path(query, page)
    if path.exists():
        return json.loads(path.read_text())

    # Use only where you have express permission. Do not add proxy rotation,
    # CAPTCHA solving or identity spoofing to evade restrictions.
    url = "https://www.google.com/search"
    params = {"q": query, "start": page * 10, "hl": "en"}
    headers = {"User-Agent": "PermittedResearchClient/1.0 (contact: you@example.com)"}

    with requests.Session() as session:
        response = session.get(url, params=params, headers=headers, timeout=20)

    if response.status_code in (429, 403):
        raise RuntimeError(f"Access refused with HTTP {response.status_code}; stop and review permission and rate.")
    response.raise_for_status()

    text = response.text.lower()
    challenge_markers = ("captcha", "unusual traffic", "consent.google")
    if any(marker in text for marker in challenge_markers):
        raise RuntimeError("A challenge or consent page was returned; do not retry aggressively.")

    results = parse_result_page(response.text)
    if not results:
        raise RuntimeError("No result blocks found. Markup may have changed or a non-result page was returned.")

    path.write_text(json.dumps(results, ensure_ascii=False, indent=2))
    return results


if __name__ == "__main__":
    queries = list(dict.fromkeys(["python web scraping", "python web scraping"]))
    for index, query in enumerate(queries):
        if index:
            time.sleep(10)  # Conservative pacing; no universal safe rate exists.
        print(query, fetch_permitted(query, page=0))

Important limitations: the selector is not an API contract, result layouts vary by location and experiment, and a successful response today does not establish a safe long-term rate. Keep the collector small and stop when Google returns a challenge, refusal or unexpected page.

3. Add pagination only when necessary

Every extra page multiplies traffic and block exposure. Request only the pages your use case needs, record the query and page number, and never fetch the same pair twice. A simple key such as sha256(query + page) makes deduplication deterministic. Cache raw responses only for as long as your legal and contractual terms permit.

Request pacing, caching and data quality

  • Deduplicate first: normalize whitespace, case and tracking parameters before enqueueing a query.
  • Cache results: use a time-to-live appropriate to the decision you are making; daily reporting rarely needs minute-by-minute refreshes.
  • Space requests: use a conservative delay and a bounded worker count. Pacing advice is operational guidance, not a guaranteed Google threshold.
  • Stop on signals: 403, 429, CAPTCHA text, JavaScript challenges and sudden empty pages should pause the job for review.
  • Preserve provenance: save retrieval time, query, language, country, page, HTTP status and parser version.
  • Validate output: detect an unexpected result count or missing fields instead of silently writing bad rankings.

Do not rotate proxies, spoof crawler identities or solve CAPTCHAs as a general recipe for evading controls. Those techniques increase operational and policy risk and do not create permission.

cURL and Node.js patterns for an authorized API

If your provider gives you an authorized JSON endpoint, call that endpoint instead of scraping Google’s HTML. Set the endpoint and credentials from the provider’s current documentation; do not hard-code keys in source control.

export SERP_API_URL='https://your-authorized-provider.example/search'
export SERP_API_KEY='YOUR_API_KEY'
curl --fail-with-body -G "$SERP_API_URL" \
  -H "Authorization: Bearer $SERP_API_KEY" \
  --data-urlencode "q=python web scraping" \
  --data-urlencode "language=en" \
  --data-urlencode "page=1"
const endpoint = process.env.SERP_API_URL;
const key = process.env.SERP_API_KEY;
if (!endpoint || !key) throw new Error('Set SERP_API_URL and SERP_API_KEY');
const params = new URLSearchParams({ q: 'python web scraping', language: 'en', page: '1' });
const response = await fetch(`${endpoint}?${params}`, {
  headers: { Authorization: `Bearer ${key}`, Accept: 'application/json' },
  signal: AbortSignal.timeout(20_000)
});
if (!response.ok) throw new Error(`SERP API returned ${response.status}`);
const data = await response.json();
console.log(JSON.stringify(data, null, 2));

Compare providers on geography and language controls, schema stability, quotas, retention, terms, latency and total cost. A hosted API can reduce parser maintenance, but no provider should be described as permanently unblockable. Verify current commercial terms before committing.

Or skip the browser setup

If your workflow also needs a clean capture of a result page, documentation page or dashboard after collection, ScreenshotNeo provides a single-call website screenshot API. The API base is https://api.screenshotneo.com/v1/shot; see the API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, newsletter popups and chat widgets are removed before the shot. Bot checks, blank pages and failed loads are not billed; response headers identify the page verdict and whether it was billed. Its MCP server lets Claude, Cursor and other MCP clients call take_screenshot, get_page_info and capture_pdf. The Free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

Troubleshooting

Symptom Likely cause Fix
HTTP 429 Traffic was throttled Stop the queue, inspect permission, reduce volume and use an authorized API.
HTTP 403 Access denied, policy restriction or challenge Do not retry in a loop. Review terms and credentials; request permission or switch to an approved API.
CAPTCHA or “unusual traffic” HTML Google identified automated access End the run. Do not automate CAPTCHA solving or rotate identities to evade it.
Empty parser output Markup changed, consent page or challenge returned Save the response, inspect its title and status, update the parser only for a permitted source.
Requests hang No timeout or overloaded route Set connect/read timeouts, bound concurrency and record failed URLs for later review.
Different results by run Location, language, personalization or index changes Record locale parameters and retrieval time; avoid treating one page as permanent truth.
Duplicate rows Repeated queries or pagination overlap Normalize queries and deduplicate by canonical URL plus query and page.

Performance, reliability and cost

Direct scraping has a low apparent per-request cost but a high maintenance cost: challenges interrupt jobs, HTML changes break selectors and browser automation consumes more CPU and memory. More workers usually increase failure probability rather than producing linear throughput. Measure success rate, latency, challenge rate, parser-empty rate and cost per accepted result.

Consent banners and overlays can be removed before a clean page capture.
Consent banners and overlays can be removed before a clean page capture.

Hosted APIs generally charge by request or result and shift anti-bot handling and parsing maintenance to the provider. Review quotas, overage behavior, data retention and regional processing before sending sensitive queries. The official researcher API has rolling 24-hour limits and non-commercial terms, so it may be inexpensive for eligible research but unsuitable for a commercial product.

For reliable jobs, use a queue with bounded concurrency, idempotency keys, exponential backoff only for explicitly retryable provider errors, and a dead-letter list for responses requiring human review. Alert on sudden changes in status codes, response size, result count and parser version.

Operational checklist

  1. Define why you need Search data and confirm permission and applicable terms.
  2. Check whether an official or contractually authorized API fits.
  3. Deduplicate queries and limit pages before sending traffic.
  4. Cache responses with a documented retention period.
  5. Set timeouts, bounded concurrency and structured logs.
  6. Detect CAPTCHAs, refusals and empty pages; stop instead of escalating.
  7. Store locale, timestamp, query, page and source metadata.
  8. Review provider quotas, retention and commercial terms before launch.

FAQ

Can I make scraping safe by changing the User-Agent?

No. A User-Agent identifies a client; it does not provide permission or prevent throttling. Google warns that crawler User-Agent headers are often spoofed and recommends reverse-DNS or source-IP checks when verifying Googlebot, not when disguising your own traffic.

Does robots.txt allow me to scrape Google results?

robots.txt controls crawler guidance for the site that publishes it. It is not authentication and does not replace Google’s terms or an API agreement.

How many requests per hour are safe?

Google does not publish a universal safe threshold in the referenced documentation. Use the least aggressive design your permitted use case allows and stop on challenge or refusal signals.

Is browser automation more reliable than requests?

It can render JavaScript, but it still faces policy restrictions, CAPTCHAs, IP controls and higher resource usage. Rendering does not turn unauthorized access into authorized access.

When should I use the Search Researcher Result API?

Use it only if your project meets the eligibility and non-commercial program terms and can operate within its rolling 24-hour limits. Verify current documentation before building around it.