Scrape Any Website to JSON with CSS Selectors
Map CSS selectors to a JSON schema, handle JavaScript-rendered pages, extract attributes safely, and choose between Scrapy and hosted APIs.
To scrape a website into JSON with CSS selectors, map every JSON key to a selector and an extraction rule. Use the selector to find an element, then read its text, an attribute such as href, or a typed value. Use nested rules for objects and repeated containers for arrays.
CSS selectors describe paths to elements in the DOM. Scrapy adds ::text and ::attr(name) shortcuts; .get() returns the first match, .getall() returns every match, and an unmatched selector returns None. See the Scrapy selector documentation and the W3C Selectors specification.
1. Define a JSON schema from CSS selectors
Start with the smallest useful schema. Each key should have one job and one selector.
{
"title": {"selector": "h1", "extract": "text"},
"canonical_url": {"selector": "link[rel=canonical]", "extract": "href", "type": "url"},
"description": {"selector": "meta[name=description]", "extract": "content"},
"products": {
"selector": ".product-card",
"multiple": true,
"fields": {
"name": {"selector": ".product-name", "extract": "text"},
"url": {"selector": "a.product-link", "extract": "href", "type": "url"},
"price": {"selector": ".price", "extract": "text", "type": "number"},
"image": {"selector": "img", "extract": "src", "type": "url"}
}
}
}
The outer .product-card selector defines an array. Child selectors run inside each card, so a page with ten cards produces ten objects. Missing fields should become null (or be omitted if your downstream contract requires that); do not silently shift values between records.
Text, attributes, and typed values
- Read visible text from an element such as
h1or.price. - Read attributes with selectors such as
a.next[href],img::attr(src)in Scrapy, or an explicitattr: "src"rule in a schema API. - Convert values deliberately: prices to numbers, dates to ISO strings, and links to absolute URLs.
- Trim whitespace and define whether embedded text nodes are joined with spaces.
- Return
nullfor a missing or invalid typed value so consumers can distinguish an empty result from a valid zero.
2. Inspect the DOM your scraper will actually receive
Open developer tools, inspect the target element, and copy a selector. Then simplify it. Prefer an ID, semantic class, data-* attribute, schema markup, or an ARIA attribute over a long positional path such as body > div:nth-child(3) > div:nth-child(2).
/* More stable */
article[data-product-id] .price
/* Fragile after a redesign */
main > div:nth-child(2) > section:nth-child(1) > div:nth-child(4) > span
Check the downloaded HTML as well as the browser inspector. A JavaScript application may return an almost empty HTML shell and fill the page only after scripts run. In that case, a plain HTTP client cannot see the content you selected.
3. A complete Python example with Scrapy
Scrapy provides CSS and XPath selectors, selector chaining, and JSON feed exports. Create a project, add this spider, and run it against a page containing repeated cards.
import scrapy
class ProductsSpider(scrapy.Spider):
name = "products"
start_urls = ["https://example.com/products"]
def parse(self, response):
for card in response.css(".product-card"):
price_text = card.css(".price::text").get()
price = None
if price_text:
cleaned = price_text.replace("$", "").replace(",", "").strip()
try:
price = float(cleaned)
except ValueError:
pass
href = card.css("a.product-link::attr(href)").get()
yield {
"name": card.css(".product-name::text").get(default="").strip() or None,
"url": response.urljoin(href) if href else None,
"price": price,
"image": response.urljoin(card.css("img::attr(src)").get())
if card.css("img::attr(src)").get() else None,
}
# Run from the project directory:
# scrapy crawl products -O products.json
Use .get() when one value is expected and .getall() for all matches:
title = response.css("h1::text").get()
tags = [text.strip() for text in response.css(".tag::text").getall()]
next_url = response.css("a.next::attr(href)").get()
Scrapy translates CSS queries to XPath internally, so XPath is available when CSS cannot express a relationship clearly. Its feed exports can write JSON, JSON Lines, CSV, or XML; -O products.json overwrites the output, while -o products.json appends according to the feed format.
4. A small Python extractor for one page
For a single static page, Requests plus Beautiful Soup is often enough. This example implements text, attributes, repeated records, URL resolution, and explicit nulls.
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup
url = "https://example.com/products"
response = requests.get(
url,
headers={"User-Agent": "Mozilla/5.0 (compatible; JsonExtractor/1.0)"},
timeout=30,
)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
def text_or_none(node):
return node.get_text(" ", strip=True) if node else None
def attr_or_none(node, name):
return node.get(name) if node and node.has_attr(name) else None
result = {
"title": text_or_none(soup.select_one("h1")),
"canonical_url": urljoin(url, attr_or_none(soup.select_one('link[rel="canonical"]'), "href"))
if attr_or_none(soup.select_one('link[rel="canonical"]'), "href") else None,
"products": [],
}
for card in soup.select(".product-card"):
href = attr_or_none(card.select_one("a.product-link"), "href")
image = attr_or_none(card.select_one("img"), "src")
result["products"].append({
"name": text_or_none(card.select_one(".product-name")),
"url": urljoin(url, href) if href else None,
"price": text_or_none(card.select_one(".price")),
"image": urljoin(url, image) if image else None,
})
import json
print(json.dumps(result, ensure_ascii=False, indent=2))
This fetches returned HTML only. It does not execute page JavaScript, click controls, wait for network activity, or maintain a browser session.
5. JavaScript-rendered pages and wait conditions
Use a rendered browser when the field appears only after JavaScript executes. A useful readiness strategy is, in order:
- Wait for a selector that proves the required data exists, such as
[data-loaded="true"]or.product-card. - Use network-idle waiting when no stable readiness selector exists.
- Add a short fixed delay only for pages with animations or delayed third-party widgets.
Cloudflare Browser Run documents gotoOptions.waitUntil values such as networkidle0 and networkidle2, plus waitForSelector. Browserless documents selectors against the fully rendered DOM. Microlink describes rules that run on a rendered page when needed. Rendering too early produces empty arrays; waiting forever makes jobs slow and unreliable.
// Generic browser pseudocode
await page.goto(url, {waitUntil: "networkidle2"});
await page.waitForSelector(".product-card", {timeout: 15000});
const data = await page.$$eval(".product-card", cards => cards.map(card => ({
name: card.querySelector(".product-name")?.textContent.trim() ?? null,
url: card.querySelector("a.product-link")?.href ?? null
})));
For infinite scroll, repeatedly scroll and collect records until the count stops increasing or a next-page control disappears. Set a maximum page count and record count so a broken page cannot run forever.
6. Selector patterns that cover common fields
| Need | Selector | Extraction |
|---|---|---|
| First heading | h1 |
text |
| Meta description | meta[name="description"] |
content attribute |
| Canonical URL | link[rel="canonical"] |
href attribute |
| All cards | .card |
repeat container |
| Links inside each card | a.card-link |
href, resolved against the page URL |
| Data attribute | [data-id] |
data-id attribute |
| Schema markup | script[type="application/ld+json"] |
text, then parse JSON |
| Fallback heading | h1, [role="heading"] |
first non-empty text |
Use fallback selectors only when they represent the same concept. If both selectors match different elements, validate and choose the first non-empty value rather than silently merging unrelated text.
7. cURL, Node.js, and hosted extraction choices
cURL is useful for checking the raw response and headers:
curl -L --fail --compressed \
-A 'Mozilla/5.0 (compatible; JsonExtractor/1.0)' \
'https://example.com/products' \
-o page.html
It cannot execute JavaScript or apply CSS selectors by itself. Pipe the saved document into a parser such as Scrapy, Beautiful Soup, Cheerio, or another DOM library.
import * as cheerio from "cheerio";
const response = await fetch("https://example.com/products");
if (!response.ok) throw new Error(`${response.status} ${response.statusText}`);
const html = await response.text();
const $ = cheerio.load(html);
const data = {
title: $("h1").first().text().trim() || null,
products: $(".product-card").map((_, card) => ({
name: $(card).find(".product-name").first().text().trim() || null,
url: $(card).find("a.product-link").attr("href") || null,
price: $(card).find(".price").first().text().trim() || null
})).get()
};
console.log(JSON.stringify(data, null, 2));
Hosted extraction APIs reduce browser and parser operations to one request. Compare them on JavaScript rendering, selector grammar, nested and repeated data, type conversion, null behavior, wait controls, authentication, proxy and session support, output formats, quotas, and operational ownership. Scrapy is a good fit when you need custom crawling, pipelines, retries, or on-premise execution. A hosted service is useful when you want fetch, render, and extraction in one request.
8. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. It can render a page before you inspect or archive it, with waits for a selector, delay, or network idle; custom headers, cookies, user agents and authorization; custom JavaScript and CSS; click and hide-selector actions; and request or resource blocking. Those controls help you capture the same rendered state that a selector-based scraper needs, while the extraction itself remains your JSON parser’s job. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/products -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com/products"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/products' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await Bun.write("shot.webp", bytes);
Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing; response headers report X-Page-Verdict and X-Billed. The MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
9. Reliability checklist
- Store the source URL, fetch time, HTTP status, and parser version with every record.
- Prefer semantic selectors, IDs, data attributes, and schema markup.
- Keep a fallback selector for known template variants.
- Measure null rates per field and alert when they change sharply.
- Validate arrays, numeric ranges, URLs, and required keys before writing output.
- Use bounded retries with backoff for transient network errors.
- Cache unchanged pages where your freshness requirements allow it.
- Respect terms of service, robots directives, rate limits, and applicable law.
- Keep authentication data and cookies out of logs.
10. Performance, cost, and scale
Raw HTTP fetching is usually faster and cheaper than launching a browser. Browser rendering is necessary when content is client-rendered, but it adds startup, JavaScript, network, and waiting time. Reduce work by selecting only required fields, avoiding unnecessary resources, reusing browser sessions where supported, and setting a maximum wait.
For large jobs, queue URLs, limit concurrency per host, deduplicate URLs, and write JSON Lines incrementally so one failed page does not lose the batch. Track cost as requests plus browser time or hosted API units, and distinguish cache hits from fresh captures. A successful HTTP response can still contain an empty field after a redesign, so correctness metrics matter as much as latency.
11. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Selector returns no result | Wrong selector or different DOM | Inspect saved HTML, simplify the selector, and test the rendered DOM. |
| HTML contains an app shell | Content is rendered by JavaScript | Use a browser renderer and wait for a known selector or network idle. |
| Only the first item is exported | Used .get() instead of .getall() or a repeated rule |
Select the container and iterate over matches. |
| Relative links are unusable | Href was copied without its base URL | Resolve with response.urljoin or urljoin. |
| Fields belong to the wrong card | Global child selectors were used | Run child selectors inside each repeated container. |
| Numbers fail validation | Currency symbols, locale separators, or hidden text | Normalize deliberately and retain the original string for auditing. |
| Timeouts on dynamic pages | Waiting for network idle never settles | Wait for a stable content selector and cap the total timeout. |
| HTTP 403 or CAPTCHA | Access controls or rate limits | Check permission, reduce concurrency, use supported authentication, and do not bypass access controls. |
| Output became empty after a redesign | Selector drift | Monitor null rates, add a tested fallback, and update the schema. |
12. FAQ
Can CSS selectors extract JSON directly?
Selectors identify and extract DOM values; your code or a hosted API maps those values into JSON. The schema defines names, nesting, repetition, types, and missing-value behavior.
What is the difference between text and HTML extraction?
Text extraction returns readable text nodes. HTML extraction preserves markup and is useful when formatting or embedded links must be retained. Choose one explicitly for each field.
How do I scrape pages behind login?
Use an authorized session, cookies, headers, or an authenticated browser context. Keep credentials secret and confirm that your use complies with the site’s rules.
Should I use Scrapy or a hosted API?
Choose Scrapy for control over crawling, pipelines, retries, and deployment. Choose a hosted API when browser rendering and extraction operations should be managed as a service. Compare rendering, schema support, wait controls, quotas, cost, and ownership before deciding.
How do I know a scraper still works?
Validate required fields, monitor null and error rates, save representative fixtures, and run selector checks against each important template variant.


