ScreenshotNeo

BlogHow-to

How to Convert a Website to JSON

Extract existing JSON-LD or map website content into your own JSON schema. Includes runnable Python, Node.js, and cURL examples.

By the ScreenshotNeo team4 October 202610 min read

There are two different tasks people mean by “convert a website to JSON”:

  • Extract JSON the website already publishes: check for an API, feed, or JSON-LD embedded in the HTML.
  • Create your own JSON from page content: choose a schema, select the relevant page elements, and map their text or attributes into that schema.

Start with an official API or feed if one exists. Otherwise inspect the HTML for JSON-LD. If the fields you need are missing, extract the relevant elements and map them yourself. Use a browser-rendered page when the content is added by JavaScript and is absent from the initial HTML. The right path depends on the target site and the data it exposes.

1. Choose the conversion method

What you need Where to start What you get
Data the site publishes for developers Official API or downloadable feed The site’s intended structured response
Structured data already embedded in a page JSON-LD script tags, then other structured markup if needed Existing entities and properties, which may not cover every field you want
A custom object based on visible page content HTML selectors and a schema you define Your chosen fields, subject to the page’s markup and behavior
Content that appears after scripts run Browser rendering, followed by the same structured-data or selector extraction The rendered DOM, which can include client-loaded content

JSON-LD is JSON-based structured data embedded in a <script type="application/ld+json"> element. The W3C JSON-LD processing specification describes optional extraction of these scripts from HTML for processors that support it. Google also describes JSON-LD as a structured-data format embedded in a script tag. Those facts do not make arbitrary visible text a ready-made JSON object: for that, you still need to define the fields and extraction rules. W3C JSON-LD 1.1 Processing Algorithms and API; Google structured data introduction.

2. Inspect the page before writing a scraper

  1. Look for an official API, feed, export, or documented data endpoint.
  2. Fetch the page HTML and search for application/ld+json. A page may contain more than one JSON-LD script.
  3. Compare the fields in those scripts to the fields you need. JSON-LD can describe only part of a page.
  4. If needed content is missing, inspect the page’s HTML and identify stable selectors for the relevant elements.
  5. Check whether the initial response contains the content. If the page fills it in using JavaScript, use browser rendering and inspect the rendered DOM.
  6. Check the site’s access instructions, authentication requirements, rate limits, and terms before automating requests.

Use browser developer tools’ Elements panel to inspect rendered markup, and the Network panel to distinguish the original document from later API requests. Do not assume a visible element is present in the first HTML response.

3. Extract existing JSON-LD with Python

This example retrieves the HTML, parses every JSON-LD script it finds, and writes the results to a JSON file. It does not turn ordinary page text into structured data. Install the dependencies with python -m pip install requests beautifulsoup4.

import json
import requests
from bs4 import BeautifulSoup

url = "https://example.com/article"
response = requests.get(
    url,
    headers={"User-Agent": "Mozilla/5.0 (compatible; JsonDataExtractor/1.0)"},
    timeout=(5, 30),
)
response.raise_for_status()

soup = BeautifulSoup(response.text, "html.parser")
items = []
for script in soup.select('script[type="application/ld+json"]'):
    raw = script.string or script.get_text()
    if not raw.strip():
        continue
    try:
        items.append(json.loads(raw))
    except json.JSONDecodeError as exc:
        print(f"Skipping invalid JSON-LD script: {exc}")

with open("website-data.json", "w", encoding="utf-8") as file:
    json.dump(items, file, ensure_ascii=False, indent=2)

print(f"Saved {len(items)} JSON-LD script(s) to website-data.json")

Keeping each script’s value as a separate item avoids silently assuming that multiple blocks should be merged. A page may use arrays, objects, or a graph structure such as @graph. Inspect the result and define how your application should interpret those shapes.

4. Extract existing JSON-LD with Node.js

Using Node.js with Cheerio gives a compact HTML parsing workflow. Install dependencies with npm install cheerio. Save this as extract-jsonld.mjs and run it with node extract-jsonld.mjs.

import * as cheerio from "cheerio";
import { writeFile } from "node:fs/promises";

const url = "https://example.com/article";
const response = await fetch(url, {
  headers: { "user-agent": "Mozilla/5.0 (compatible; JsonDataExtractor/1.0)" },
  signal: AbortSignal.timeout(30000),
});
if (!response.ok) {
  throw new Error(`HTTP ${response.status} ${response.statusText}`);
}

const html = await response.text();
const $ = cheerio.load(html);
const items = [];

$('script[type="application/ld+json"]').each((_, element) => {
  const raw = $(element).text().trim();
  if (!raw) return;
  try {
    items.push(JSON.parse(raw));
  } catch (error) {
    console.warn("Skipping invalid JSON-LD script:", error.message);
  }
});

await writeFile("website-data.json", JSON.stringify(items, null, 2), "utf8");
console.log(`Saved ${items.length} JSON-LD script(s) to website-data.json`);

Node’s built-in fetch is available in current Node.js releases. If your runtime does not provide it, use a supported HTTP client or upgrade the runtime.

5. Check a page with cURL

cURL is useful for checking whether structured markup is present in the raw server response. Save the HTML, then inspect it for JSON-LD scripts:

curl --fail --location --max-time 30 \
  --user-agent "Mozilla/5.0 (compatible; JsonDataExtractor/1.0)" \
  "https://example.com/article" \
  --output page.html

rg -n 'application/ld\+json' page.html

This downloads HTML; it does not parse the JSON-LD or execute page JavaScript. Use one of the code examples above to parse scripts, or use a browser renderer if the data appears only after the page runs.

6. Map page elements into a custom JSON schema

If the site does not publish the fields you need, define the output shape first. For example, a product-card extractor could map a heading, price, and link into {"name":"…","price":"…","url":"…"}. The selector names below are examples: inspect the target page and replace them with selectors that actually match it.

import json
import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin

page_url = "https://example.com/products/widget"
response = requests.get(page_url, timeout=(5, 30))
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")

name_el = soup.select_one("h1")
price_el = soup.select_one(".price")
link_el = soup.select_one("a.primary-link")

record = {
    "name": name_el.get_text(" ", strip=True) if name_el else None,
    "price": price_el.get_text(" ", strip=True) if price_el else None,
    "url": urljoin(page_url, link_el["href"]) if link_el and link_el.has_attr("href") else None,
}

print(json.dumps(record, ensure_ascii=False, indent=2))

Prefer explicit missing values such as null when a field is absent, and decide whether a missing required field should fail the record. Normalize whitespace deliberately. Convert numbers and dates only when their formats are understood; preserve the original text when conversion would be ambiguous.

Design a schema that stays useful

  • Choose stable field names and document whether values are strings, numbers, booleans, arrays, objects, or null.
  • Use absolute URLs for links by resolving relative href values against the page URL.
  • Represent repeated elements as arrays, even when a particular page has only one.
  • Keep source information such as page URL and retrieval time if you need traceability.
  • Validate output before consuming it downstream. A syntactically valid JSON document can still have missing or incorrectly typed fields.

7. Render JavaScript-driven pages when necessary

When the raw response lacks the required content, a browser must load the page and allow its scripts to update the DOM before extraction. A browser-rendering service can accept a URL or HTML plus selectors; for example, Cloudflare documents a /scrape endpoint for selecting elements and returning details such as inner HTML. That is a vendor-specific option, not a guarantee that every target site will work with it. See the Cloudflare Browser Rendering documentation.

Whether using a local browser automation library or a hosted renderer, make the wait condition reflect the page. Waiting for a specific selector is usually more meaningful than an arbitrary short delay. A page may also require authentication, user interaction, or an API request that is not available to an anonymous visitor. Extract from the rendered DOM only after confirming the target element exists, and keep timeouts bounded.

8. Convert multiple pages safely

For a site-wide job, separate URL discovery from extraction:

  1. Get URLs from an official feed, sitemap, API, or allowed page links.
  2. Filter to the paths and page types your schema supports.
  3. Process a small sample and validate its output before scaling up.
  4. Limit concurrency and add a delay or backoff when the site signals that requests should slow down.
  5. Record each URL’s status and extraction errors so one malformed page does not erase the rest of the results.
  6. Use a stable record identifier or source URL to make retries idempotent and avoid duplicate output.

Do not assume a sitemap grants permission to crawl every listed URL, or that a successful single-page request implies the whole site has the same markup.

9. Access, reliability, and cost considerations

Respect access instructions

Review the target site’s terms and access guidance. Google explains that robots.txt tells crawlers which URLs they may access and is mainly used to manage crawler traffic; it is not a privacy or security mechanism, and a blocked URL may still appear in search results. Google’s robots.txt guide explains these limits. That guide does not resolve every site’s contractual or legal conditions.

Make the pipeline reliable

  • Set connect and read timeouts; do not let a stalled request hang a worker indefinitely.
  • Check HTTP status codes and content type before parsing. Error pages and access challenges can return HTML with a successful-looking body.
  • Retry transient network errors and server errors with capped exponential backoff; avoid rapid retries for access denials or rate limits.
  • Log the URL, status, parse error, and a small sanitized context. Avoid logging credentials, cookies, or personal data.
  • Keep extraction selectors in one place and monitor for empty or unexpectedly changed fields.
  • Validate the output against your schema before storing or passing it to another service.

Estimate the operating cost

For direct HTTP parsing, the main costs are your compute, bandwidth, storage, and maintenance. Browser rendering generally adds more resource use and operational complexity than parsing a static response. Hosted extraction or browser services may charge according to their own plans and request limits; check current vendor documentation before choosing one. No universal price or speed applies to every site.

10. Troubleshooting

Symptom Likely cause Fix
No JSON-LD scripts found The page does not include JSON-LD in its initial HTML, or JavaScript adds it later. Inspect the rendered DOM and try browser rendering. If structured data is absent, define selectors and a custom schema.
JSON parse error The script content is malformed, truncated, or not valid standalone JSON. Print the failing script safely, verify the response is complete, and skip or quarantine invalid blocks. Do not repair malformed content by blindly deleting characters.
Fields are missing or nested unexpectedly JSON-LD may use arrays, @graph, or different entity types; your target page may also vary. Inspect the complete object and write mapping logic for the structures you actually encounter.
Selector returns no element The selector is wrong, markup changed, or the content is loaded dynamically. Check the live markup, wait for the target selector in a browser, and add a clear missing-field policy.
HTTP 403 or an access challenge The site denied the request or requires a supported access path. Check the site’s access instructions and use its API or authorized method if available. Do not try to evade access controls.
HTTP 429 The server is limiting request frequency. Reduce concurrency, honor any retry guidance, and back off before retrying.
Request succeeds but output is an error page The server returned a challenge or fallback HTML rather than the expected page. Check status, final URL, content type, and a small response sample before parsing.
Duplicate records in a crawl Multiple URLs may resolve to the same content or a retry appended a record twice. Normalize URLs where appropriate and upsert using a stable page identifier or canonical URL.

11. Or skip the browser setup

If your JSON workflow needs a rendered screenshot or a visual record alongside extracted data, ScreenshotNeo is a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. It does not convert arbitrary page content into a custom JSON schema, so use the extraction methods above for that part.

For a screenshot, first create a free account and put your API key in place of YOUR_API_KEY. The API accepts the same parameter names other screenshot APIs use, which can make switching easier. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/article -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com/article"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/article' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed, and response headers say what happened. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up free for 1,000 screenshots a month, with no card required.

12. FAQ

Does converting a website to JSON preserve its design?

No. JSON represents data, not page layout or styling. Choose fields and values you want to retain.

Can I convert an entire website with one request?

Usually, a multi-page site requires URL discovery and one or more extraction requests per page. A site’s API or export may offer a more direct bulk path.

Is JSON-LD the same as all the page’s content?

No. It is structured data a site has chosen to publish, and it may describe only selected entities or properties.