ScreenshotNeo

BlogHow-to

Gemini AI Web Scraping in Python: Fetch, Then Extract

Learn a reliable Python workflow that fetches web pages first, then uses Gemini to extract structured data—with URL Context limits and fixes.

By the ScreenshotNeo team29 September 20269 min read

Gemini AI Web Scraping in Python: Fetch, Then Extract

Use Gemini for the part of scraping it is good at: turning page content into structured information. Keep retrieval and extraction as two separate steps. Your Python code fetches a URL, checks the response, removes irrelevant markup, and sends a bounded piece of content to Gemini with a strict output schema. Gemini can also retrieve specific public URLs through URL Context, but that feature is a supplied-URL retrieval tool, not an unrestricted crawler.

This distinction makes a scraper easier to debug, cheaper to run, and safer to operate. You can tell whether a failure came from the target site, your parser, or the model. The examples below show both approaches and include limits, permissions, retries, output validation, and a production checklist.

1. Choose the retrieval approach

Approach Who fetches? Best fit Main limit
Fetch, then extract Your Python application Known URLs, authenticated requests, custom headers, crawl queues, deterministic retries You own HTTP, parsing, rate limits, and HTML cleanup
Gemini URL Context Gemini after you provide URLs Analysis of a small set of public pages without writing a fetcher Up to 20 URLs per request and 34 MB of retrieved content per URL; it does not follow nested links
Gemini CLI web_fetch Gemini CLI using URL Context Interactive command-line investigations It is a CLI interface, not a Python crawler library

Google describes URL Context as a way to provide URLs as additional context for a model. It can use indexed content first and fall back to a live fetch, and responses can contain URL citations and retrieval metadata. See the official URL Context documentation for current model and request details.

Do not use Google Search grounding as a programmatic source of crawl targets. The Gemini API Additional Terms, effective March 23, 2026, restrict collecting grounded results, suggestions, or links to identify destination pages for crawling or scraping.

2. Fetch a page in Python, then extract fields with Gemini

The following example is intentionally explicit. It fetches one page, rejects unsuccessful responses, strips scripts and styles, limits the text sent to the model, asks for JSON, and validates the result. Install the two ordinary Python dependencies first:

Keep retrieval, cleaning, and Gemini extraction as separate observable stages.
Keep retrieval, cleaning, and Gemini extraction as separate observable stages.
python -m pip install requests beautifulsoup4

Set GEMINI_API_KEY in your environment. The REST endpoint and model name can change, so confirm them in the current Gemini API documentation before production deployment.

import json
import os
import sys
from urllib.parse import urlparse

import requests
from bs4 import BeautifulSoup

URL = "https://example.com/article"
API_KEY = os.environ["GEMINI_API_KEY"]
MODEL = os.getenv("GEMINI_MODEL", "gemini-2.5-flash")


def fetch_text(url: str) -> tuple[str, str]:
    parsed = urlparse(url)
    if parsed.scheme not in {"http", "https"}:
        raise ValueError("Only http and https URLs are allowed")

    response = requests.get(
        url,
        headers={"User-Agent": "research-fetcher/1.0"},
        timeout=(10, 30),
        allow_redirects=True,
    )
    response.raise_for_status()

    content_type = response.headers.get("content-type", "")
    if "html" not in content_type:
        raise ValueError(f"Expected HTML, received {content_type!r}")

    soup = BeautifulSoup(response.text, "html.parser")
    for node in soup(["script", "style", "noscript", "svg", "template"]):
        node.decompose()

    title = soup.title.get_text(" ", strip=True) if soup.title else ""
    main = soup.find("main") or soup.body or soup
    text = " ".join(main.get_text(" ", strip=True).split())
    return title, text[:60_000]


def extract_with_gemini(title: str, text: str) -> dict:
    prompt = f"""Extract article metadata from the supplied page text.
Return JSON only with these keys:
- title: string
- summary: string of at most 40 words
- topics: array of strings
- facts: array of objects with claim and evidence

If a value is absent, use an empty string or empty array. Do not infer facts.
Page title: {title}
Page text:
{text}"""

    endpoint = (
        "https://generativelanguage.googleapis.com/v1beta/models/"
        f"{MODEL}:generateContent?key={API_KEY}"
    )
    payload = {
        "contents": [{"parts": [{"text": prompt}]}],
        "generationConfig": {
            "temperature": 0,
            "responseMimeType": "application/json",
        },
    }
    response = requests.post(endpoint, json=payload, timeout=90)
    response.raise_for_status()
    body = response.json()
    raw = body["candidates"][0]["content"]["parts"][0]["text"]
    result = json.loads(raw)

    required = {"title", "summary", "topics", "facts"}
    if set(result) != required:
        raise ValueError(f"Unexpected extraction keys: {set(result)}")
    return result


if __name__ == "__main__":
    try:
        page_title, page_text = fetch_text(URL)
        extracted = extract_with_gemini(page_title, page_text)
        print(json.dumps(extracted, indent=2, ensure_ascii=False))
    except (requests.RequestException, KeyError, ValueError, json.JSONDecodeError) as exc:
        print(f"scrape failed: {exc}", file=sys.stderr)
        raise SystemExit(1)

This is a complete pipeline, but it is not a universal crawler. JavaScript-rendered pages may return an empty shell to requests; login walls and bot checks may block the request; and a page can change between fetch and analysis. Add a browser renderer only when you have confirmed that ordinary HTTP is insufficient.

Why clean the HTML before sending it?

  • Scripts, navigation, cookie notices, and repeated footer text consume context without helping extraction.
  • A character limit prevents one unusually large page from exhausting the request.
  • A small, named schema makes downstream storage predictable.
  • temperature: 0 reduces variation for repeatable extraction; it does not guarantee factual correctness.

3. Use Gemini URL Context for known public URLs

URL Context is useful when your application already knows the pages to analyze. Supply complete URLs in the request and ask a focused question. Google documents a maximum of 20 URLs per request and a maximum retrieved content size of 34 MB per URL. Paywalled pages and some content types are unsupported. It does not discover links or traverse a site’s internal navigation for you.

The exact REST shape depends on the model and API version. Follow the current URL Context documentation when enabling the tool. Conceptually, your request contains a user prompt, the URLs, and the URL Context tool declaration:

payload = {
    "contents": [{
        "parts": [{
            "text": "Compare the pricing terms on these supplied pages. Return JSON with url, price, and evidence."
        }]
    }],
    "tools": [{"url_context": {}}]
}
# POST payload to the current Gemini generateContent endpoint.

Keep this path separate from a crawler queue. If you need pagination, sitemap traversal, authenticated cookies, or domain-specific throttling, fetch those URLs in Python and pass selected content to Gemini instead.

4. cURL, Python, and Node.js request patterns

For a direct Gemini API call, the cURL shape below illustrates the ordinary JSON request. Confirm the current model and tool fields in the official documentation before copying it into a deployment script.

curl -sS "https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash:generateContent?key=$GEMINI_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{"contents":[{"parts":[{"text":"Extract the title and dates from this supplied text: ..."}]}]}'

Python uses the same HTTP contract:

import os, requests

r = requests.post(
    "https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash:generateContent",
    params={"key": os.environ["GEMINI_API_KEY"]},
    json={"contents": [{"parts": [{"text": "Return JSON for: ..."}]}]},
    timeout=90,
)
r.raise_for_status()
print(r.json())

Node.js 18 or newer includes fetch:

const key = process.env.GEMINI_API_KEY;
const res = await fetch(
  `https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash:generateContent?key=${key}`,
  {
    method: 'POST',
    headers: { 'content-type': 'application/json' },
    body: JSON.stringify({
      contents: [{ parts: [{ text: 'Return JSON for: ...' }] }]
    })
  }
);
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
console.log(await res.json());

5. Make extraction reliable

Define fields and evidence

Ask for a fixed object rather than “summarize this page.” Include an evidence field containing a short phrase or source location. Reject objects with missing keys before writing them to a database. For lists, require an array and cap its length.

Separate fetching from model retries

Store the response status, final URL, content type, fetch timestamp, and a hash of the cleaned text. Retry transient HTTP 429 and 5xx responses with exponential backoff, but do not blindly retry 401, 403, or 404 responses. A model timeout should not trigger another page download if the fetched text is already stored.

Control concurrency

Use a bounded worker pool and per-domain rate limits. Cache successful fetches for a period appropriate to the site. Sending the same unchanged text repeatedly wastes model tokens; hash-based deduplication lets you skip extraction when nothing changed.

Respect access controls

Check the site’s robots.txt, access controls, terms, and the requirements that apply to your project and jurisdiction. Google describes robots.txt as a mechanism for site owners to allow or disallow crawler access. It is not, by itself, a complete legal determination for a particular project.

6. Or skip the browser setup

If your real goal is a clean visual capture rather than raw HTML, ScreenshotNeo makes one GET request and returns a PNG, JPEG, WebP, or PDF. Its capture pipeline accepts consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and whether it was billed.

A rendered capture can remove consent overlays and other visual clutter before analysis.
A rendered capture can remove consent overlays and other visual clutter before analysis.

See the ScreenshotNeo API documentation for all options. The basic call is:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

You can also capture a full page with lazy images loaded, select one element by CSS selector, set a device preset or custom viewport, use dark mode and retina scale, inject CSS or JavaScript, click before capture, wait for a selector, delay, or network idle, block ads or resource types, provide headers, cookies, a user agent, Authorization, timezone, or geolocation, make a transparent image, resize it, cache it with a chosen TTL, create signed public image links, submit async jobs with signed webhooks, capture up to 100 URLs per bulk call, and generate PDFs with paper size, margins, orientation, and page ranges. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

Free accounts include 1,000 screenshots each month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Create a free ScreenshotNeo account.

7. Troubleshooting

Symptom Likely cause Fix
403 or 429 from the site Access control or rate limiting Slow down, identify your client, honor robots.txt and terms, and do not rotate identities to evade controls.
Empty extracted text Content is rendered by JavaScript Use a permitted browser capture or an API provided by the site; confirm the final HTML before calling Gemini.
Gemini returns invalid JSON Prompt or model response is not schema-constrained Request JSON explicitly, use a response MIME type where supported, parse defensively, and retry only the model step.
Context or request too large Raw HTML includes navigation, scripts, or huge documents Extract the main content, truncate by a known limit, split sections, and merge validated results.
Wrong facts The page is ambiguous or the model inferred missing values Require evidence, say “use empty when absent,” and review low-confidence records.
URL Context cannot retrieve a page URL is private, paywalled, unsupported, or above the size limit Fetch it in your application if you are authorized, or provide a public supported URL.

8. Performance, reliability, and cost checklist

  • Set separate connect and read timeouts for page fetches.
  • Persist raw status metadata and cleaned text so model retries do not refetch.
  • Use conditional requests such as ETag or Last-Modified when the target supports them.
  • Batch only related pages; URL Context allows at most 20 URLs and 34 MB per URL.
  • Estimate model input size before submission and trim repeated boilerplate.
  • Log model, prompt version, response status, token usage when available, and parser errors.
  • Keep API keys in environment variables or a secret manager, never in page content or source control.
  • Run a small evaluation set of pages and compare extracted fields after prompt or model changes.

There is no single universal cost for a scrape: your total depends on HTTP infrastructure, model input and output usage, retries, and how often pages change. Caching and deduplication usually have the largest practical effect on spend and latency.

9. FAQ

Can Gemini crawl an entire website?

No. URL Context retrieves URLs you supply and does not follow nested links. Build discovery and crawl scheduling in your application, subject to the site’s controls and terms.

Should I send raw HTML to Gemini?

Usually no. Remove scripts and repeated chrome, select the main content, and preserve headings, tables, and links that are relevant to the fields you need.

Is Gemini Search grounding the same as fetching a known URL?

No. Fetching a known URL is a retrieval step for that URL. The Gemini terms restrict using Search grounding results or links to programmatically identify crawl destinations.

When should I use URL Context?

Use it when you already have a small list of public URLs and want Gemini to retrieve and analyze them. Use your own fetcher when you need authentication, custom headers, domain throttling, or link traversal.

Can screenshots replace structured scraping?

No. Screenshots are useful for visual records, QA, and rendered pages. Structured extraction should use accessible HTML, an API, or another authorized data interface.