ScreenshotNeo

BlogAI agents

How to Use the Gemini API for Web Data Extraction

Learn how to give Gemini web pages, extract validated JSON, discover sources with Search grounding, and preserve citations in production pipelines.

By the ScreenshotNeo team1 October 20268 min read

Short answer: give Gemini the URLs with the URL Context tool, state an extraction contract, request Structured Outputs with a JSON Schema, validate the result, and store the source URL and citation metadata with every record. Use Google Search grounding when Gemini must discover pages first. URL access, extraction, schema validation, and provenance are separate concerns.

This guide shows a complete workflow for extracting fields such as product names, prices, currencies, availability, article metadata, and key findings from public pages.

1. Choose the retrieval mode

Situation Use Why
You already know the pages URL Context Provide one or more public URLs for direct inspection.
You need Gemini to find pages Google Search grounding Gemini searches current public information and returns citation annotations.
You need machine-readable records Structured Outputs Constrain the final response to a JSON Schema.
Extraction triggers an application action Function calling Let Gemini request an application-owned function, then execute it in your code.

URL Context supports common public formats including HTML, JSON, plain text, XML, CSS, JavaScript, CSV, and RTF. Retrieval can still fail because of safety checks or URL limitations, so treat access as fallible.

2. Write an extraction contract

Before writing code, define:

  • Fields and their types.
  • Normalization rules, such as decimal prices and ISO currency codes.
  • What to do when a field is absent: usually null, never a guessed value.
  • Whether the model should quote source text or summarize it.
  • How to identify the source URL for each record.

A useful contract for product pages might be:

Extract the product name, price, currency, and availability from each URL.
Use the page's stated values. Normalize price to a number and currency to a
three-letter code when possible. If a field is missing or ambiguous, return null.
Do not infer values from similar products. Return one object per URL.

3. Define a JSON Schema

Structured Outputs makes the response syntactically conform to your schema, but it does not make the extracted values factually correct. Validate business rules after parsing. Gemini supports a subset of JSON Schema: string, number, integer, boolean, object, array, and null are the core types.

{
  "type": "object",
  "properties": {
    "items": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "url": {"type": "string"},
          "product_name": {"type": ["string", "null"]},
          "price": {"type": ["number", "null"]},
          "currency": {"type": ["string", "null"]},
          "availability": {"type": ["string", "null"]},
          "evidence": {"type": ["string", "null"]}
        },
        "required": ["url", "product_name", "price", "currency", "availability", "evidence"]
      }
    }
  },
  "required": ["items"]
}

4. Python: extract fields from known URLs

Install the current Google GenAI SDK and Pydantic, then set GEMINI_API_KEY. SDK method names and supported models can change, so check the current getting-started documentation when upgrading.

pip install -U google-genai pydantic
import json
import os
from typing import Optional
from pydantic import BaseModel
from google import genai

class Item(BaseModel):
    url: str
    product_name: Optional[str]
    price: Optional[float]
    currency: Optional[str]
    availability: Optional[str]
    evidence: Optional[str]

class Extraction(BaseModel):
    items: list[Item]

urls = [
    "https://example.com/product-a",
    "https://example.com/product-b",
]

client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])
prompt = f"""
Extract product_name, price, currency, availability, and a short evidence quote
from each URL below. Return one item per URL. Use null when a field is absent or
ambiguous. Never guess. URLs: {json.dumps(urls)}
"""

interaction = client.interactions.create(
    model="gemini-2.5-flash",
    input=prompt,
    tools=[{"type": "url_context"}],
    response_format={
        "type": "text",
        "mime_type": "application/json",
        "schema": Extraction.model_json_schema(),
    },
)

result = Extraction.model_validate_json(interaction.output_text)
for item in result.items:
    if item.price is not None and item.price < 0:
        raise ValueError(f"Invalid negative price for {item.url}")
    print(item.model_dump())

The model receives the URLs as input and URL Context fetches them. Keep the URL list bounded and reject URLs that your application should not fetch.

5. JavaScript: return validated objects

npm install @google/genai zod
import { GoogleGenAI } from "@google/genai";
import { z } from "zod";

const Item = z.object({
  url: z.string().url(),
  product_name: z.string().nullable(),
  price: z.number().nullable(),
  currency: z.string().nullable(),
  availability: z.string().nullable(),
  evidence: z.string().nullable()
});
const Extraction = z.object({ items: z.array(Item) });

const urls = ["https://example.com/product-a", "https://example.com/product-b"];
const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY });

const interaction = await ai.interactions.create({
  model: "gemini-2.5-flash",
  input: `Extract product_name, price, currency, availability, and a short evidence quote from each URL. Use null when absent or ambiguous; never guess. URLs: ${JSON.stringify(urls)}`,
  tools: [{ type: "url_context" }],
  response_format: {
    type: "text",
    mime_type: "application/json",
    schema: {
      type: "object",
      properties: {
        items: {
          type: "array",
          items: {
            type: "object",
            properties: {
              url: { type: "string" },
              product_name: { type: ["string", "null"] },
              price: { type: ["number", "null"] },
              currency: { type: ["string", "null"] },
              availability: { type: ["string", "null"] },
              evidence: { type: ["string", "null"] }
            },
            required: ["url", "product_name", "price", "currency", "availability", "evidence"]
          }
        }
      },
      required: ["items"]
    }
  }
});

const data = Extraction.parse(JSON.parse(interaction.output_text));
console.log(data);

6. REST and cURL pattern

The exact REST fields depend on the API surface and model version. The following illustrates the Generate Content shape: put URL Context in tools, request JSON with responseMimeType, and provide your schema in responseSchema. Check the current REST reference before production deployment.

curl "https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash:generateContent?key=$GEMINI_API_KEY" \\
  -H "Content-Type: application/json" \\
  -d '{
    "contents": [{
      "parts": [{
        "text": "Extract product name, price, currency and availability from https://example.com/product. Use null when absent; never guess."
      }]
    }],
    "tools": [{"url_context": {}}],
    "generationConfig": {
      "responseMimeType": "application/json",
      "responseSchema": {
        "type": "object",
        "properties": {
          "product_name": {"type": ["string", "null"]},
          "price": {"type": ["number", "null"]},
          "currency": {"type": ["string", "null"]},
          "availability": {"type": ["string", "null"]}
        },
        "required": ["product_name", "price", "currency", "availability"]
      }
    }
  }'

7. Discover pages with Google Search grounding

Use Search grounding when the URL is unknown or the information changes. Gemini can issue search queries and return inline URL citation annotations. Preserve those annotations, including the cited URL and title, with the extracted record. You can combine Search grounding with URL Context: search discovers candidate pages, then URL Context inspects URLs you explicitly provide.

from google import genai
import os

client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])
response = client.models.generate_content(
    model="gemini-2.5-flash",
    contents="Find the official product pages for Acme's current plans, then list each plan's monthly price.",
    config={"tools": [{"google_search": {}}]},
)
print(response.text)
# Store response.candidates[0].grounding_metadata and its web URI/title objects.

Search grounding is for discovery and evidence. It does not replace schema validation or post-processing.

8. Keep provenance with every record

Store at least:

  • The requested URL and the final fetched URL if available.
  • Retrieval timestamp.
  • Model name and SDK version.
  • Schema version.
  • Raw model output or a protected hash of it.
  • Grounding URL annotations or GroundingChunk web URI and title objects.
  • Validation errors and retry count.

For quotations, keep the quote short and attach it to the URL it came from. For summaries, retain the source URL even when no direct quote is requested.

9. Security and reliability checklist

  • Allow only http and https; block private IP ranges, localhost, cloud metadata endpoints, and internal hostnames.
  • Limit URL count, page size, extracted record count, and evidence length.
  • Treat page text as untrusted input. A page can contain prompt-injection instructions; your extraction contract remains authoritative.
  • Use explicit nulls for missing fields and reject unexpected extra fields.
  • Validate ranges, currencies, dates, IDs, and enumerations in application code.
  • Retry transient API failures with exponential backoff and an upper bound. Do not retry schema or safety failures indefinitely.
  • Make writes idempotent using a key such as normalized URL plus retrieval date.
  • Redact secrets before logging prompts and responses.

10. Performance and cost considerations

There is no universal latency or accuracy benchmark for this workflow. Retrieval time depends on the number and size of pages, model choice, and whether Search performs one or more queries. Reduce work by sending only required URLs, using a smaller schema, limiting evidence text, and caching unchanged pages in your own system.

Pricing, quotas, token counts, and model availability change. Consult the official pricing page before estimating cost. Search grounding can incur tool charges according to the model and billing rules; URL Context and retrieved content also affect usage.

11. Common errors and fixes

Error Likely cause Fix
URL cannot be retrieved Unsupported URL, robots restriction, authentication, safety block, or transient fetch failure. Confirm the page is public, test a smaller set, handle the error state, and provide a fallback URL or manual review.
Valid JSON, wrong values Schema controls shape, not factual accuracy. Require evidence, validate ranges and enums, compare critical values with deterministic parsers, and flag uncertainty.
Schema rejected Unsupported JSON Schema feature or excessive nesting. Use supported primitive, object, array, and null types; simplify deeply nested schemas.
Missing citations URL Context extraction was used without Search grounding, or annotations were discarded. Use Search grounding for discovery and persist citation annotations or GroundingChunk metadata.
Empty or truncated output Large pages, too many URLs, or output limits. Split batches, reduce evidence length, cap fields, and retry only transient failures.
Function called unexpectedly Function calling was used where a final JSON response was required. Use Structured Outputs for final formatting; reserve Function Calling for application actions.

12. Or skip the browser setup

If your workflow needs screenshots of rendered pages before extraction, ScreenshotNeo provides a website screenshot API and MCP server. It accepts one GET request and returns PNG, JPEG, WebP, or PDF. Cookie and consent banners, newsletter popups, and chat widgets are removed before capture; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. Responses identify the page verdict and billing status.

See the ScreenshotNeo API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo also has an MCP server so Claude, Cursor, and other MCP clients can take screenshots, inspect pages, and capture PDFs. One thousand screenshots each month are free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

13. FAQ

Can Gemini scrape any website?

No. URL Context is intended for supported public URLs and can fail safety checks, access restrictions, or documented URL limitations.

Should I ask for JSON in the prompt?

Ask for the fields in the prompt, but enforce the final shape with Structured Outputs and validate the parsed object.

When should I use Search instead of URL Context?

Use Search when Gemini must discover current pages. Use URL Context when you already know the exact pages to inspect.

Does Structured Outputs prove the extraction is correct?

No. It guarantees a supported response shape. Semantic checks, evidence review, and domain validation remain your application’s responsibility.

Can extraction trigger a database write?

Use Function Calling for an intermediate request to an application-owned function, then authorize and execute that function server-side.