ScreenshotNeo

BlogHow-to

How to Scrape Web Pages with ChatGPT

Use ChatGPT to inspect a page, work with supported browser tools, or build a repeatable scraper you run yourself. Learn how to validate results and handle common failures.

By the ScreenshotNeo team4 October 202610 min read

Short answer: ChatGPT can help you extract information from webpages in three ways: ask it to inspect a page using Search or an available browser feature, use a supported browser interaction for an interactive task, or ask it to write code that you run in your own environment. For repeatable collections, a scraper you run yourself gives you the most control over fetching, parsing, validation, and output. ChatGPT Data Analysis can help clean and analyze collected files, but it cannot fetch webpages from its Python environment.

This guide shows how to choose an approach, ask ChatGPT for auditable results, and build a small Python scraper for accessible HTML. It also covers dynamic pages, failure handling, validation, and how to pass the resulting data back to ChatGPT.

1. Choose what you mean by “scrape with ChatGPT”

Approach Best for Key limit Check
Search or ordinary page reading A few current facts or a one-off extraction Does not guarantee a complete structured capture Source links, missing fields, and current values
ChatGPT desktop site tools An interactive task on a supported webpage Availability depends on your account, model, and the tools the open page exposes Which tools are available and what actions they take
ChatGPT Work cloud browser A supported public or signed-in browser task Supported sites and actions vary, and a site may block access Correct site, access flow, and resulting records
Code written with ChatGPT and run externally Repeatable collection from accessible pages You need a runtime, permission, and maintenance when pages change Scope, selectors, failures, completeness, and changes
Official API or export Repeated or larger structured collections when offered Available fields and limits are set by the provider Provider documentation and allowed use

First check for an official API, download, or export. For a single page, ask ChatGPT for the exact fields and a source URL. For a recurring dataset, use code or an official structured source and keep a validation step in the workflow.

2. Extract a few fields from one page

  1. Give ChatGPT the exact page address and name the fields you want.
  2. Ask for one record per row, explicit column names, and a source URL for each row or group.
  3. Tell it to leave absent fields blank or mark them null, and to distinguish page facts from inference.
  4. Ask it to report the number of records found and any page sections or fields it could not access.
  5. Verify values against the page, especially dates, prices, identifiers, and totals.

Example prompt:

Inspect https://example.com/catalog and extract the product name, listed price, availability, and product URL for each visible product. Return one product per row with these exact column names. Include the source URL, leave missing values null, and do not infer values. Report how many rows you found and anything you could not inspect.

Use Search when you need current web research with source links. A page-specific tool or browser feature may be able to interact with a supported page, but access depends on the account and the website. Check the tool indicator in the desktop app when using site tools; follow the access flow for Work cloud browser. Neither method promises that every record on a page was captured.

See OpenAI’s documentation for site tools and cloud browser for their current availability and limits.

3. Build a repeatable scraper with ChatGPT

For a recurring collection, define the target, fields, scope, output format, and update frequency before asking ChatGPT to draft code. Provide a permitted sample of HTML or a saved page when you need help choosing selectors. Review the generated code and assumptions, run it in your own environment, and retain source URLs and retrieval dates.

A runnable Python example for accessible HTML

This example fetches a page, extracts elements marked up with article.product, and writes JSON. The selector is illustrative: replace it with selectors that actually match the permitted target page. It does not execute page JavaScript, sign in, or bypass access controls.

from __future__ import annotations

import json
import sys
from datetime import datetime, timezone
from urllib.parse import urljoin

import requests
from bs4 import BeautifulSoup

URL = "https://example.com/catalog"
TIMEOUT_SECONDS = 20


def clean_text(node):
    return node.get_text(" ", strip=True) if node else None


def main():
    try:
        response = requests.get(
            URL,
            headers={"User-Agent": "ExampleResearchBot/1.0 (contact: you@example.com)"},
            timeout=TIMEOUT_SECONDS,
        )
        response.raise_for_status()
    except requests.Timeout:
        sys.exit(f"Timed out fetching {URL}")
    except requests.HTTPError as exc:
        sys.exit(f"HTTP error fetching {URL}: {exc}")
    except requests.RequestException as exc:
        sys.exit(f"Request failed for {URL}: {exc}")

    soup = BeautifulSoup(response.text, "html.parser")
    records = []
    for card in soup.select("article.product"):
        link = card.select_one("a.product-link")
        price = card.select_one(".price")
        availability = card.select_one(".availability")
        records.append({
            "name": clean_text(card.select_one("h2")),
            "price": clean_text(price),
            "availability": clean_text(availability),
            "url": urljoin(response.url, link["href"]) if link and link.get("href") else None,
            "source_url": response.url,
            "retrieved_at": datetime.now(timezone.utc).isoformat(),
        })

    output = {"count": len(records), "records": records}
    print(json.dumps(output, indent=2, ensure_ascii=False))


if __name__ == "__main__":
    main()

Install the dependencies with python -m pip install requests beautifulsoup4, save the code as scrape.py, then run python scrape.py > results.json. The placeholder URL and selectors must be adapted to the page you are permitted to access. Check the page’s terms and access instructions; do not use the script to defeat authentication, CAPTCHA, paywalls, or anti-bot measures.

Fetch the page with cURL

For a quick inspection of a public page’s returned HTML, save the response and inspect it locally:

curl --fail --location --max-time 20 \
  --user-agent "ExampleResearchBot/1.0 (contact: you@example.com)" \
  "https://example.com/catalog" \
  --output page.html

cURL fetches the server response; it does not parse the records or run the page’s JavaScript. For repeatable scraping, use a parser in a language with suitable HTML tools, or use an official API/export if available.

Fetch the page with Node.js

This Node.js example fetches accessible HTML and saves it. Add a parser such as cheerio if you want to extract elements from the saved response.

import { writeFile } from 'node:fs/promises';

const url = 'https://example.com/catalog';
const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), 20_000);

try {
  const res = await fetch(url, {
    headers: { 'User-Agent': 'ExampleResearchBot/1.0 (contact: you@example.com)' },
    signal: controller.signal,
  });
  if (!res.ok) throw new Error(`HTTP ${res.status} ${res.statusText}`);
  const html = await res.text();
  await writeFile('page.html', html, 'utf8');
  console.log(`Saved ${html.length} characters from ${res.url}`);
} finally {
  clearTimeout(timer);
}

Save it as fetch.mjs and run node fetch.mjs with a current Node.js runtime that supports built-in fetch. This obtains the server response, not a browser-rendered page.

4. Make results auditable before asking ChatGPT to analyze them

  • Use descriptive field names and one record per row.
  • Include the source URL and retrieval date, and preserve the original value when you normalize a field.
  • Use an explicit missing-value convention such as JSON null or an empty CSV cell.
  • Check a sample of records against the page and compare counts or totals where possible.
  • Inspect generated code, outputs, and assumptions before using computed results.
  • Recheck selectors when the site changes its markup; keep a small validation sample for comparison.

After collection, upload a CSV, JSON, XML, text file, or another supported file for cleaning and analysis. OpenAI recommends descriptive headers and one record per row, and notes that complex, image-based, or scanned tables may not produce exact values reliably. Its Data Analysis documentation states: “The Python environment used for data analysis cannot make external web requests or API calls.” In other words, provide the collected file or an available connected source; Data Analysis is not a general-purpose web fetcher. See Data Analysis with ChatGPT.

5. Handle dynamic, interactive, and signed-in pages

A basic HTTP request receives the server’s response and does not execute client-side JavaScript. If the desired content appears only after scripts run or a user interacts with the page, first check whether the site provides an API or export. Otherwise, use a supported browser interaction for the task if the relevant account and site expose one. Site tools are page-specific, and cloud browser runs in its own session rather than reusing local browser cookies. A website can block automated access.

If access is blocked, use an allowed export/API or obtain the information through an authorized human workflow. Do not try to work around the site’s controls. Do not paste passwords or security codes into chat, and review the site and any consequential browser action before proceeding. See the official documentation for site tools and cloud browser.

6. Troubleshoot common scraping problems

Symptom Likely cause What to do
No records found The example selectors do not match the target markup, or the records are rendered by JavaScript. Inspect a permitted saved HTML sample and update selectors. If content requires rendering, use an available supported browser flow or an official source.
HTTP 403 or access denied The site declined the request or requires a supported access path. Check site access instructions and use an official API/export or authorized workflow. Do not attempt to bypass the restriction.
HTTP 404 The URL is wrong, the page moved, or the resource is no longer available. Check the address and the site’s current links; record the final URL after redirects.
HTTP 429 The service is rate limiting requests. Stop or slow collection, follow the site’s instructions, and use a documented API limit or export if offered.
Timeout or connection error Slow response, transient network issue, or unreachable host. Use a finite timeout, retry cautiously with a delay for transient failures, and log the URL and error. Do not retry indefinitely.
Values are missing or malformed Optional fields, changed markup, inconsistent formats, or an incorrect selector. Represent missing values explicitly, preserve source text, and validate normalization rules against examples.
Duplicate records Pagination overlap, repeated cards, or multiple links to the same item. Choose a stable record key such as the item URL, deduplicate on that key, and compare before/after counts.
ChatGPT cannot analyze the live URL in Data Analysis Its analysis Python environment cannot make external web requests or API calls. Fetch data in your own environment or use an available source connection, then upload the collected file.
Browser tool cannot complete the task The account, page, website, or requested action may not be supported, or the site may block access. Confirm available tools and the site access flow. Use an allowed API/export or authorized human workflow if unavailable.

7. Performance, reliability, and cost

For a small collection, one request at a time is easier to inspect and less likely to burden a site. For recurring work, limit request frequency, set finite timeouts, record failures, and avoid aggressive retries. A scraper depends on the target site’s availability and markup stability; an API or export can provide a more structured route when the site offers one. Keep a sample-based validation step so a successful HTTP response is not mistaken for a complete extraction.

ChatGPT feature availability can vary by plan, model, workspace settings, and website. Recheck the current official help pages for the account you use. The reviewed documentation does not establish a universal price or legal rule for scraping. Consider site terms, access instructions, data type, and applicable law; seek qualified legal advice when the stakes warrant it. OpenAI crawler settings for OAI-SearchBot, GPTBot, and ChatGPT-User describe OpenAI product behavior and do not grant blanket permission to unrelated scrapers. See OpenAI crawler documentation.

8. Or skip the browser setup

If your goal is to capture a page as an image or PDF for review, documentation, or a visual workflow, ScreenshotNeo is a website screenshot API and MCP server. A single GET request returns a PNG, JPEG, WebP, or PDF. Its capture flow accepts cookie/consent banners like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers say which page verdict and billing status applied. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. It is for visual capture, not a substitute for extracting arbitrary structured records from a site.

Example cURL request (replace the target URL and API key):

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status} ${res.statusText}`);
await Bun.write('shot.webp', res);

See the ScreenshotNeo API documentation for request options and setup. The API also supports full-page capture with lazy images loaded, element capture, device and viewport settings, retina scale, PDF options, custom CSS and JavaScript, waits, request blocking, headers and cookies, caching, signed links, asynchronous jobs, bulk capture, and more. Plans include every feature: 1,000 shots a month free with no card; paid plans start at $5 for 3,000 shots. Sign up for 1,000 free screenshots a month, with no card required.

9. FAQ

Can ChatGPT turn a webpage table into a CSV?

It can help extract or format a table when the page is accessible to the active tool, or help write code that produces a CSV. Check row counts and sample values against the source before relying on the file.

Can ChatGPT scrape a whole website?

There is no single universal ChatGPT scraping feature that guarantees a complete site-wide collection. For a larger or repeatable dataset, check for an official API/export and define scope, access, and validation before collecting.

Can I ask ChatGPT to write the scraper without giving it the page?

Yes, but selectors and page-specific assumptions need evidence. A permitted HTML sample or saved page makes it easier to design and review extraction logic.

Do OpenAI crawler settings determine whether I can scrape a site?

No. Those settings describe how OpenAI crawlers and user-triggered visits operate; they do not decide permission for unrelated scraping. Check the target site’s terms and access instructions.