How to Scrape Structured Responses from ChatGPT
Learn why ChatGPT website scraping is restricted and how to request validated JSON with OpenAI Structured Outputs instead.
There are two different tasks hidden in the phrase “scrape structured responses from ChatGPT”:
- Extracting messages from the consumer ChatGPT website with browser automation or page scraping.
- Calling a model from your application and asking it to return data that matches a JSON schema.
These workflows are not interchangeable. OpenAI’s individual Terms of Use, revised December 11, 2024, list “Automatically or programmatically extract data or Output” among prohibited acts. Business and organizational customers may be governed by separate agreements, including terms that permit extraction through the API. Check the agreement that applies to your account, organization, location and use case before automating exports.
For an application, use the OpenAI API with a JSON Schema response format and Structured Outputs when the selected model and endpoint support it. This gives your program a predictable shape. It does not make the model’s claims true or complete: validate business rules, handle refusals and incomplete responses, and add human review where the result matters.
Website scraping versus API responses
| Question | ChatGPT website extraction | OpenAI API Structured Outputs |
|---|---|---|
| Purpose | Read content rendered in a consumer-facing website | Build an application that requests model output |
| Official position in the cited sources | The cited individual terms prohibit automatic or programmatic extraction | The API reference documents response formats and JSON Schema |
| Structure | Depends on UI markup and can break when the interface changes | A supported schema constrains the returned JSON structure |
| Accuracy | Copied text can still be wrong | Schema adherence is formatting, not fact checking |
Do not build around DOM selectors, copied session cookies, reverse engineering, or attempts to bypass bot checks and other protective controls. If you have a permitted export workflow, use the method and agreement supplied for that account.
How to request structured JSON through the API
The API reference documents a response format with type: "json_schema" and a schema object. With strict mode enabled, supported models follow the defined schema subject to the supported subset of JSON Schema. The reference says JSON Schema is preferred where supported.
Example schema
This schema asks the model to classify a support message and return a short summary with extracted order IDs.
{
"type": "object",
"properties": {
"category": {
"type": "string",
"enum": ["billing", "technical", "shipping", "other"]
},
"summary": { "type": "string" },
"order_ids": {
"type": "array",
"items": { "type": "string" }
}
},
"required": ["category", "summary", "order_ids"],
"additionalProperties": false
}
JavaScript with the official SDK
The current developer quickstart shows the Responses API and reads generated text through response.output_text. Model names and feature support can change, so verify the current API reference before deployment.
import OpenAI from "openai";
const client = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });
const response = await client.responses.create({
model: "YOUR_SUPPORTED_MODEL",
input: [
{
role: "system",
content: "Return only data that matches the supplied schema. Do not invent order IDs."
},
{
role: "user",
content: "Order 1842 arrived late. Please refund the shipping charge."
}
],
text: {
format: {
type: "json_schema",
name: "support_ticket",
strict: true,
schema: {
type: "object",
properties: {
category: { type: "string", enum: ["billing", "technical", "shipping", "other"] },
summary: { type: "string" },
order_ids: { type: "array", items: { type: "string" } }
},
required: ["category", "summary", "order_ids"],
additionalProperties: false
}
}
}
});
if (!response.output_text) throw new Error("The model returned no text");
const data = JSON.parse(response.output_text);
if (!Array.isArray(data.order_ids)) throw new Error("Invalid order_ids");
console.log(data);
cURL
curl https://api.openai.com/v1/responses \\
-H "Authorization: Bearer $OPENAI_API_KEY" \\
-H "Content-Type: application/json" \\
-d '{
"model": "YOUR_SUPPORTED_MODEL",
"input": "Classify this message: Order 1842 arrived late. Please refund shipping.",
"text": {
"format": {
"type": "json_schema",
"name": "support_ticket",
"strict": true,
"schema": {
"type": "object",
"properties": {
"category": {"type": "string", "enum": ["billing", "technical", "shipping", "other"]},
"summary": {"type": "string"},
"order_ids": {"type": "array", "items": {"type": "string"}}
},
"required": ["category", "summary", "order_ids"],
"additionalProperties": false
}
}
}
}'
The response envelope and output text shape can vary by endpoint and SDK version. Inspect the current API reference rather than assuming that every response contains the same nesting.
Python
import json
import os
from openai import OpenAI
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
response = client.responses.create(
model="YOUR_SUPPORTED_MODEL",
input="Classify this message: Order 1842 arrived late. Please refund shipping.",
text={
"format": {
"type": "json_schema",
"name": "support_ticket",
"strict": True,
"schema": {
"type": "object",
"properties": {
"category": {"type": "string", "enum": ["billing", "technical", "shipping", "other"]},
"summary": {"type": "string"},
"order_ids": {"type": "array", "items": {"type": "string"}}
},
"required": ["category", "summary", "order_ids"],
"additionalProperties": False
}
}
}
)
if not response.output_text:
raise RuntimeError("The model returned no text")
data = json.loads(response.output_text)
if not isinstance(data.get("order_ids"), list):
raise ValueError("Invalid order_ids")
print(data)
JSON Schema mode versus JSON mode
json_object is the older JSON mode. It ensures valid JSON, but the API reference says your prompt still needs to instruct the model to generate JSON. It does not provide the same schema-adherence guarantee. Prefer json_schema for supported models and endpoints.
Neither mode guarantees that required facts exist, that values obey your business rules, or that the answer is true. After parsing, validate enums, ranges, identifiers, permissions and required relationships in application code.
Validation and failure handling
- Check the HTTP status and catch SDK exceptions.
- Handle refusals explicitly; a refusal may not match your requested object.
- Handle incomplete output caused by limits, cancellation or transport failure.
- Parse JSON only after confirming that text exists.
- Run a JSON Schema validator and then business-rule checks.
- Store the prompt, schema version, model identifier and validation result for debugging.
- Send uncertain or high-impact cases to a human reviewer.
Common errors
| Error | Likely cause | Fix |
|---|---|---|
| Unsupported parameter or format | The model or endpoint does not support Structured Outputs | Check the current API reference and choose a supported combination, or use JSON mode with weaker guarantees |
| 400 invalid schema | Unsupported JSON Schema features, missing required fields, or additional properties enabled | Reduce the schema to the supported subset and mark every required property |
| JSON parsing failure | Refusal, incomplete output, wrong response field, or non-JSON mode | Check refusal and completion state, inspect the SDK response, then parse |
| Valid JSON but bad values | Schema validation does not check truth or domain logic | Apply application validation and review |
| Website automation breaks | UI markup, login flow or terms changed | Do not bypass controls; use an authorized export or the API |
Performance, reliability and cost
Keep schemas focused: every required field increases the work the model must complete. Use bounded arrays and concise descriptions. Retry transient API failures with exponential backoff and an idempotency strategy appropriate to your operation. Set request timeouts, log request IDs when available, and cap retries so a queue cannot multiply cost. Cache results only when the input, schema version and model behavior make reuse acceptable. Estimate spend from your chosen model’s current pricing and token usage; Structured Outputs does not remove normal API charges.
Or skip the browser setup
If your real task is capturing the ChatGPT or documentation page as an image or PDF for review, ScreenshotNeo provides a website screenshot API and MCP server. One request returns a PNG, JPEG, WebP or PDF. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers report the page verdict and billing status.
See the ScreenshotNeo API documentation for all options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
It also offers an MCP server with take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients. Free accounts include 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
FAQ
Can I scrape ChatGPT responses?
The cited individual terms prohibit automatic or programmatic extraction of data or Output. Check the current agreement governing your account before any export or automation.
How do I get JSON from ChatGPT?
For an application, call the API and request json_schema Structured Outputs when your model and endpoint support it.
Is JSON mode enough?
It produces valid JSON, but it does not enforce your schema. Use JSON Schema where supported.
Does a schema make answers accurate?
No. Validate facts and business rules separately and use human review when appropriate.
Can I use the API and the ChatGPT website interchangeably?
No. They are separate workflows with different interfaces, capabilities and governing terms.


