ScreenshotNeo

BlogHow-to

How to Use ChatGPT for Web Scraping

Use ChatGPT to design, debug and improve a reliable scraper, then run it locally with validation, retries and permission checks.

By the ScreenshotNeo team1 October 20269 min read

Yes, ChatGPT can help you scrape a website by planning the extraction, defining a schema, writing or reviewing Python code, and explaining errors. You still need to run the scraper in your own approved environment, verify the output, and confirm that the site permits the collection.

For a dependable workflow, use ChatGPT as a coding assistant rather than as a guarantee that any URL can be crawled. Give it a small HTML sample, specify the fields and pagination rules, ask for tests and error handling, then inspect the results against the live page.

What ChatGPT can and cannot do

Task What ChatGPT helps with What you must provide or verify
Static HTML Selectors, parsing code, normalization and CSV export A permitted page or HTML sample and correct selectors
JavaScript pages Diagnosis and browser-automation code A browser tool or API that can render the page
Pagination Loop design, stopping rules and deduplication Knowledge of the site’s page or cursor behavior
Authenticated data Request design and session handling advice Permission and a safe way to provide credentials; never paste passwords into chat
Validation Assertions, fixtures and comparison checks Known expected values or counts

ChatGPT’s supported site tools are a separate path. Where available, they can use the currently open page, its current state and your signed-in session, and can show tool activity in the conversation. Availability depends on the account and the website. Site-tool guidance warns about prompt injection and data-exfiltration risks; sensitive actions require confirmation. The fact that ChatGPT can reach a page does not grant permission to copy it.

Before writing code: define the scrape

  1. Permission: read the terms, robots.txt directives, API documentation and authentication rules. Prefer an official API or export.
  2. Row identity: decide what makes a record unique, such as a product URL or article ID.
  3. Fields: list required and optional columns, data types, units and acceptable missing values.
  4. Scope: specify start URLs, page limits, date filters and whether detail pages are included.
  5. Output: choose CSV, JSON or a database table, and decide whether raw HTML is retained.
  6. Failure policy: define retry counts, timeouts, rate limits and what happens when a selector is missing.

A prompt that produces better scraper code

You are reviewing a permitted web-scraping task.

Target: https://example.com/catalog
Fields per row:
- name: required string
- price: optional decimal in USD
- product_url: required absolute URL (unique key)

Rules:
- Follow numbered pagination until the next link is absent.
- Stop after 20 pages.
- Sleep at least 1.5 seconds between requests.
- Retry transient HTTP failures up to 3 times with exponential backoff.
- Deduplicate by product_url.
- Preserve raw HTML separately from cleaned CSV output.
- Treat missing prices as blank and log the URL.

Write a complete Python 3 script using requests and BeautifulSoup.
Include a requirements list, clear selectors marked for review, structured logging,
CSV output, a small fixture-based test, and an explanation of every assumption.
Do not bypass authentication, CAPTCHAs or access controls.

Supply a short, permitted HTML sample whenever possible. Ask ChatGPT to mark selectors that need confirmation instead of pretending it knows the site’s markup.

Complete Python and BeautifulSoup example

The following pattern extracts titles, prices and links, follows simple numbered pagination, retries temporary failures, deduplicates rows and writes a CSV. Replace the example selectors after inspecting the target page.

from __future__ import annotations

import csv
import logging
import time
from decimal import Decimal, InvalidOperation
from urllib.parse import urljoin

import requests
from bs4 import BeautifulSoup
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

START_URL = "https://example.com/products"
OUTPUT = "products.csv"
MAX_PAGES = 20
DELAY_SECONDS = 1.5

logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")


def make_session() -> requests.Session:
    retry = Retry(
        total=3,
        connect=3,
        read=3,
        backoff_factor=1,
        status_forcelist=(429, 500, 502, 503, 504),
        allowed_methods=frozenset({"GET"}),
        respect_retry_after_header=True,
    )
    session = requests.Session()
    session.headers.update({"User-Agent": "PermittedCatalogResearch/1.0"})
    adapter = HTTPAdapter(max_retries=retry)
    session.mount("https://", adapter)
    session.mount("http://", adapter)
    return session


def parse_price(value: str) -> str:
    cleaned = value.replace("$", "").replace(",", "").strip()
    try:
        return str(Decimal(cleaned))
    except InvalidOperation:
        logging.warning("Could not parse price %r", value)
        return ""


def parse_page(html: str, page_url: str) -> tuple[list[dict[str, str]], str | None]:
    soup = BeautifulSoup(html, "html.parser")
    rows: list[dict[str, str]] = []
    for card in soup.select("article.product-card"):  # review selector
        link = card.select_one("a.product-link")
        title = card.select_one("h2, h3")
        price = card.select_one(".price")
        if not link or not title:
            logging.warning("Skipped a card without the required title or link")
            continue
        href = link.get("href")
        if not href:
            continue
        rows.append({
            "name": title.get_text(" ", strip=True),
            "price": parse_price(price.get_text(" ", strip=True)) if price else "",
            "product_url": urljoin(page_url, href),
        })
    next_link = soup.select_one("a[rel='next']")
    next_url = urljoin(page_url, next_link["href"]) if next_link and next_link.get("href") else None
    return rows, next_url


def main() -> None:
    session = make_session()
    seen: set[str] = set()
    records: list[dict[str, str]] = []
    url: str | None = START_URL

    for page_number in range(1, MAX_PAGES + 1):
        if not url:
            break
        logging.info("Fetching page %s: %s", page_number, url)
        response = session.get(url, timeout=(10, 30))
        response.raise_for_status()
        page_rows, next_url = parse_page(response.text, response.url)
        for row in page_rows:
            if row["product_url"] not in seen:
                seen.add(row["product_url"])
                records.append(row)
        url = next_url
        if url:
            time.sleep(DELAY_SECONDS)

    with open(OUTPUT, "w", newline="", encoding="utf-8") as handle:
        writer = csv.DictWriter(handle, fieldnames=["name", "price", "product_url"])
        writer.writeheader()
        writer.writerows(records)
    logging.info("Wrote %s unique rows to %s", len(records), OUTPUT)


if __name__ == "__main__":
    main()

Install dependencies with python -m pip install requests beautifulsoup4. Keep a raw response archive when the data matters; it lets you diagnose selector changes without immediately re-fetching the site.

Exporting results to CSV safely

  • Write UTF-8 and include a header row.
  • Keep one logical value per column; do not put comma-separated lists into fields unless that is your documented schema.
  • Normalize prices, dates and URLs before deduplication.
  • Use a stable key such as an ID or canonical URL.
  • Save raw and cleaned outputs separately so a later parser fix does not destroy evidence.
  • Compare the number of rows with the number shown by the site or a known sample.

Scraping JavaScript-rendered pages

Requests and BeautifulSoup only see the HTML returned by the server. If the records appear after JavaScript runs, inspect the browser’s network requests for an official JSON endpoint first. If no permitted endpoint exists, use a browser automation tool such as Playwright in an approved environment, wait for a specific selector, and capture the rendered DOM. Infinite scroll needs an explicit stopping condition and duplicate tracking. A CAPTCHA or bot check is an access-control signal; do not attempt to bypass it.

Pagination, rate limits and repeat runs

  • Prefer cursor or next-link pagination over guessing page numbers.
  • Stop when the next link disappears, a cursor repeats, or the configured page limit is reached.
  • Honor documented rate limits, add delays and backoff on 429 responses, and cache pages during development.
  • Record retrieval timestamps, status codes and the parser version.
  • For scheduled jobs, add change detection, alerting and a failure queue before increasing volume.

Login-protected pages and sensitive data

Confirm that your account and the site’s rules allow automated collection. Do not send passwords, session cookies or API secrets to ChatGPT. Enter credentials directly in an approved browser flow or load them from a secret manager in your runtime. Limit collection to the fields you need, protect the resulting files and delete them on a defined schedule.

Common errors and fixes

Error Likely cause Fix
Zero rows Wrong selectors or content rendered by JavaScript Inspect saved HTML, confirm selectors and use a browser or API for rendered content.
403 or 429 Permission, authentication or rate limit Check site rules, slow down, authenticate through an approved method and stop if access is denied.
Timeout Slow server or oversized page Set connect/read timeouts, retry a small number of times and log the URL.
Duplicate rows Overlapping pages or repeated infinite-scroll batches Deduplicate on a stable ID or canonical URL.
Missing fields Optional markup or selector drift Use explicit defaults, log affected URLs and add fixture tests.
Broken characters Encoding guessed incorrectly Use the response encoding when declared and write CSV as UTF-8.
Stale ChatGPT code Generated selectors or libraries no longer match Pin dependencies, keep a fixture and ask ChatGPT to review the current error and HTML sample.

Reliability, performance and cost

Reliability comes from small batches, bounded retries, idempotent output and validation. Measure requests, rows, skipped records, duplicate count and elapsed time. Parallel requests can overload a site or trigger limits, so begin sequentially and increase concurrency only when the owner permits it. ChatGPT-generated code has no guaranteed accuracy: compare a sample with the page and fail loudly when required fields disappear.

ChatGPT assistance may reduce development time, but your runtime still pays for network traffic, browser execution, proxies or hosted jobs. An official API is usually easier to repeat and monitor than HTML parsing. A managed browser or scraping service can handle rendering and scheduling, but compare its permission model, rate limits, login support, extraction accuracy and total cost before adopting it.

ChatGPT site tools and web search are different

Site tools operate on supported websites and the current signed-in session; they are not a universal crawler. Search results and cached indexes are not equivalent to a complete live-site crawl. OpenAI’s crawler documentation also distinguishes OAI-SearchBot, used for search discovery, from GPTBot, which has separate robots.txt controls; robots.txt changes may take about 24 hours to propagate. These controls describe discovery and access behavior, not a license to scrape.

Or skip the browser setup

For a screenshot of a permitted page, ScreenshotNeo provides a single GET request that returns PNG, JPEG, WebP or PDF. See the API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, newsletter popups and chat widgets are removed before the shot. Bot checks, blank pages and failed loads are never billed, and response headers identify the page verdict and billing result. An MCP server lets AI agents use take_screenshot, get_page_info and capture_pdf. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

Checklist before scheduling a scraper

  • Permission and authentication rules are documented.
  • Fields, row identity and missing-value behavior are explicit.
  • Selectors have fixture tests and a change-detection check.
  • Pagination, retries, delays and maximum scope are bounded.
  • Raw and cleaned data are stored separately.
  • Logs and alerts identify failed pages and schema changes.
  • Secrets are outside prompts and source control.

FAQ

Can ChatGPT scrape a website directly?

Only through supported site tools and only where the website and account expose them. Otherwise it can generate and review code that you run.

Can ChatGPT scrape behind a login?

Supported site tools may use an existing signed-in session. For your own script, use an approved authentication flow and never paste passwords or session secrets into chat.

Is scraping a site allowed?

It depends on the site’s terms, robots directives, API rules, privacy obligations and applicable law. Check those conditions before collecting data.

How do I scrape a JavaScript website?

Find a permitted API first. If none exists, use browser automation that waits for the required content and respects access controls.

Why did my ChatGPT-generated scraper stop working?

Markup changes, selector drift, pagination changes and bot controls are common causes. Save a fixture, compare the current HTML and update selectors with a logged test.