ScreenshotNeo

BlogHow-to

How to Use LLMs for Web Scraping

Learn how to discover pages, retrieve and clean content, extract schema-shaped data with an LLM, and validate every result against its source.

By the ScreenshotNeo team4 October 202610 min read

To use an LLM for web scraping, separate the work into stages: define the fields you need, find or select pages, retrieve their content, clean and segment it, ask the model for schema-shaped data, then validate each result against the source. Search finds candidate pages; scraping retrieves a known page; crawling discovers and processes pages across a site. An LLM can help interpret messy content, but it does not replace retrieval, provenance, or validation.

This guide uses Python for a runnable example. It fetches a known, static HTML page, extracts readable text, asks an OpenAI model for structured output, and checks the result against the document. Adapt the retrieval step when a page requires JavaScript rendering or when you need to discover many URLs.

1. Define the extraction task before collecting pages

Write down the question and the exact fields you want before you fetch content. Give every field a type and decide whether it is required. Specify what the model should return when evidence is absent; for example, use null instead of a guess.

Decision Example
Data question What product name, listed price, and currency does this page state?
Fields and types product_name: string, price: number|null, currency: string|null
Required fields Product name is required; price and currency may be unknown.
Evidence rule Include a short source passage for each non-null value.
Missing evidence Return null; do not infer a value from context.

Keep the task narrow. Asking for a few named fields from a relevant section is easier to validate than asking for a broad summary of an entire site.

2. Choose search, scraping, or crawling

Method Use it when Typical output
Web search You need to discover candidate pages for a topic or query. URLs and search-result context; retrieve the pages separately if you need their contents.
Scraping a known URL You already know which page to retrieve. One page’s HTML, text, or rendered content.
Crawling You need to discover and process pages across a site or section. A set of page records, often with Markdown or structured output.

For a known, mostly static page, a basic HTTP client and HTML parser may be enough. JavaScript-rendered content may require a browser or a rendering service. Site-wide discovery is a crawler problem. Firecrawl describes crawling, rendering, and Markdown or structured JSON output in its Web Crawling API documentation; compare current capabilities and terms for your use case. OpenAI documents web search with sourced citations in its web search guide. Search results and page scraping are separate retrieval paths.

Keep the canonical URL, fetch time, and page title with every retrieved document. Those fields make later review and deduplication more practical.

3. Check access and retrieval needs

Before collection, review the site’s terms and crawler rules, use conservative request rates, and do not bypass authentication, CAPTCHAs, or other access barriers. Do not assume that every crawler follows the same rules. Anthropic describes its own bots as respecting robots.txt and anti-circumvention technology in its crawler FAQ.

Google explains that robots.txt is a crawler-access protocol, not a privacy mechanism or a guarantee that a URL will stay out of search. A blocked URL may still be indexed if discovered elsewhere; use authentication for access restriction or a noindex directive for search exclusion, as appropriate. See Google’s robots.txt introduction and robots.txt specification guidance. Rules apply within the relevant host, protocol, and port, and crawler implementations can differ.

Decide whether the target page is static or needs rendering. Do not retry a CAPTCHA or blocked page by changing identity or trying to evade the barrier. Record inaccessible pages as failures or exclusions.

4. Retrieve a known page and extract readable text with Python

Install the dependencies:

python -m pip install requests beautifulsoup4 openai pydantic

Set OPENAI_API_KEY in the environment, then save this as scrape_extract.py. It retrieves one URL, removes non-content elements, limits the text sent to the model, requests typed structured output, and performs basic checks.

import os
from datetime import datetime, timezone
from urllib.parse import urlparse

import requests
from bs4 import BeautifulSoup
from openai import OpenAI
from pydantic import BaseModel, Field

URL = "https://example.com/"

class ProductFacts(BaseModel):
    product_name: str = Field(description="Name stated on the page")
    price: float | None = Field(description="Numeric listed price, or null if absent")
    currency: str | None = Field(description="Currency stated on the page, or null if absent")
    evidence: list[str] = Field(description="Short exact passages supporting the values")


def retrieve_text(url: str) -> tuple[str, str]:
    parsed = urlparse(url)
    if parsed.scheme not in {"http", "https"} or not parsed.netloc:
        raise ValueError("Provide a complete http or https URL")

    response = requests.get(
        url,
        headers={"User-Agent": "ResearchExtractor/1.0"},
        timeout=(10, 30),
    )
    response.raise_for_status()
    content_type = response.headers.get("Content-Type", "")
    if "html" not in content_type.lower():
        raise ValueError(f"Expected HTML, got {content_type!r}")

    soup = BeautifulSoup(response.text, "html.parser")
    title = soup.title.get_text(" ", strip=True) if soup.title else ""
    for node in soup(["script", "style", "noscript", "svg", "nav", "footer"]):
        node.decompose()
    text = "\n".join(line.strip() for line in soup.get_text("\n").splitlines() if line.strip())
    if not text:
        raise ValueError("No readable page text found; the page may require JavaScript rendering")
    return title, text


def main() -> None:
    if not os.environ.get("OPENAI_API_KEY"):
        raise RuntimeError("Set OPENAI_API_KEY before running")

    title, page_text = retrieve_text(URL)
    # Keep the prompt bounded. For longer pages, select the relevant section instead.
    bounded_text = page_text[:18000]
    client = OpenAI()
    result = client.responses.parse(
        model="gpt-4o-mini",
        input=[
            {
                "role": "system",
                "content": (
                    "Extract only facts supported by the supplied page text. "
                    "Use null for missing price or currency. Do not infer. "
                    "Include short exact evidence passages."
                ),
            },
            {
                "role": "user",
                "content": f"URL: {URL}\nPage title: {title}\nPage text:\n{bounded_text}",
            },
        ],
        text_format=ProductFacts,
    )
    facts = result.output_parsed
    if facts is None:
        raise RuntimeError("The model did not return parsed structured output")

    # Basic validation: required value, evidence presence, and source support.
    if not facts.product_name.strip():
        raise ValueError("Required product_name is empty")
    for passage in facts.evidence:
        if passage not in bounded_text:
            raise ValueError(f"Evidence passage not found verbatim in page text: {passage!r}")

    record = {
        "url": URL,
        "fetched_at": datetime.now(timezone.utc).isoformat(),
        "page_title": title,
        "data": facts.model_dump(),
    }
    print(record)


if __name__ == "__main__":
    main()

The example uses the OpenAI Python SDK’s Responses API parsing interface and a model name shown in the API documentation. Check the current Structured Outputs guide for supported models and current SDK usage. The sample URL is a placeholder: replace it with a page you are permitted to retrieve. The extraction schema is illustrative, so change the fields and evidence rules to fit your data.

5. Clean and segment content before extraction

HTML includes navigation, scripts, cookie controls, footers, and repeated layout text. Remove irrelevant elements and normalize whitespace before prompting. Parsing with selectors tied to a page’s structure can be more stable than asking an LLM to interpret a full raw HTML dump, but page markup can change.

  • Keep sections intact: split at headings or article boundaries rather than cutting text at arbitrary character counts when possible.
  • Preserve context: include the heading and nearby units or labels that give a value its meaning.
  • Bound input: send only the sections relevant to the fields. Avoid whole-site dumps.
  • Retain provenance: store source URL and section or passage references alongside extracted values.

If a page is too long, process relevant sections independently and combine the resulting records in a later step. Define how to reconcile repeated or conflicting values before merging them.

6. Ask for schema-shaped output with evidence

Request JSON or another defined schema rather than prose. Describe each field, its type, whether it is required, and what to return when the page does not provide evidence. Ask for the source passage or section for each extracted value where feasible.

Structured output constrains the response format; it does not establish that the values are true. A well-formed JSON object can still contain unsupported or incorrectly interpreted data. Treat a value without source support as unverified.

7. Validate and store results

Run mechanical checks before accepting a record:

  • Parse the response and validate it against the expected schema and types.
  • Check required fields, empty strings, out-of-range values, and missing values.
  • Check that cited passages or section references exist in the retrieved source.
  • Normalize URLs and identify duplicate pages or duplicate records.
  • Sample-check extracted values against the source; keep uncertain fields unknown.
  • Store the original URL, canonical URL if available, fetch time, title, content reference, model response, and validation outcome.

For research answers, cite the underlying pages and distinguish directly extracted facts from model-generated summaries. OpenAI’s web-search documentation describes sourced citations for its search tool; preserve those citations when using that retrieval path.

8. Scale to multiple pages carefully

For a list of known URLs, process pages with bounded concurrency and per-request timeouts. For a site section or corpus, use a crawler that supports the discovery scope and output format you need. Firecrawl describes site crawling and Markdown or JSON outputs in its product documentation. Verify current service limits, pricing, and behavior directly before choosing a provider.

Track successes, empty pages, retrieval errors, parse failures, and model validation failures separately. Retry only transient failures with a limit and backoff; do not endlessly retry a structurally invalid result or an access barrier. Deduplicate before spending model calls on repeated content. Cache fetched content when permitted and when freshness requirements allow it.

9. Performance, reliability, and cost

  • Performance: retrieval, rendering, and model calls each add latency. Limit page text to relevant sections, avoid duplicate processing, and use bounded concurrency appropriate to the target site and provider limits.
  • Reliability: pages change, network requests fail, and schemas can evolve. Save source snapshots or durable content references when permitted, record timestamps, validate outputs, and make retries bounded.
  • Cost: total cost depends on retrieval or rendering, model choice, input size, output size, and retries. No comparative cost or accuracy benchmark is established here. Check current provider pricing and rate limits before estimating a production workload.
  • Quality: do not treat structured format as proof of correctness. Measure your own review sample for the specific pages and fields you use before relying on automated output.

10. Troubleshooting

Symptom Likely cause Fix
Text is empty or only contains a shell The page fills content with JavaScript, or the parser selected the wrong region. Inspect the response and page structure. Use a rendering-capable retrieval method when appropriate; do not evade access controls.
HTTP 403 or CAPTCHA The site denied automated access or requires an access check. Stop automated retries, review the site’s access rules, and request authorized access if needed. Do not bypass the barrier.
HTTP 429 Requests are too frequent or a provider limit was reached. Reduce concurrency, respect retry guidance, and use bounded backoff.
Timeout or connection error Network latency, a slow page, or an unavailable host. Set connect and read timeouts, retry transient failures a limited number of times, and record the failed URL.
Unexpected content type The URL returned a PDF, image, or other non-HTML response. Route by content type to a suitable parser or reject the response explicitly.
Model output fails schema parsing Unsupported structured output configuration, incomplete response, or incompatible model/SDK settings. Check the current SDK and model documentation, inspect the response status, and handle refusal or incomplete output explicitly.
Valid JSON contains wrong values Schema validity was mistaken for evidence or the page context was ambiguous. Require evidence passages, validate them against source text, tighten field definitions, and mark unsupported values unknown.
Repeated or conflicting records Duplicate URLs, canonical variants, or values found in multiple page sections. Normalize URLs, deduplicate, preserve section provenance, and define a deterministic conflict policy.

11. Screenshot pages when visual rendering matters

Some tasks need a visual record of a rendered page—for example, a review of page layout alongside extracted text. ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. One GET request can return a PNG, JPEG, WebP, or PDF. It is a rendering and capture step; it does not replace URL discovery, text extraction, schema validation, or source review.

Or skip the browser setup

For a screenshot of a known URL, use the API directly. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));

ScreenshotNeo accepts cookies or consent banners as a visitor would and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Every feature is on every plan. See ScreenshotNeo and the docs.

Sign up free for 1,000 screenshots a month with no card.

12. Frequently asked questions

Can an LLM scrape a website by itself?

An LLM can interpret retrieved content and return structured fields. A separate retrieval step is needed to fetch known pages or discover pages, unless the workflow uses a search or browsing tool that performs retrieval.

Should I send HTML or Markdown to the model?

Send the cleanest representation that preserves the context needed for your fields. Readable text or Markdown is often simpler to inspect; raw HTML can help when attributes or structure matter.

Does robots.txt give permission to scrape?

It is a crawler preference protocol, not a complete statement of permission, privacy, or search visibility. Review the site’s terms and applicable access rules separately.

Does structured output guarantee accurate extraction?

No. It constrains output shape. Validate values and evidence against the source page.

Sources