Preparing Web Pages for Data Extraction
A practical workflow for inspecting, rendering, selecting, validating, and safely extracting reliable data from modern web pages.
Reliable extraction starts before the parser runs. Define the fields you need, inspect a representative page, determine whether the data is in the initial HTML or added by JavaScript, choose an extractor that matches the page type, and validate the result against the rendered page. Article pages often suit Mozilla Readability; listings, tables, catalogs, and dashboards usually need selectors or structured data. If the required content is client-rendered, fetch it with a browser first.
1. Define the extraction target
Write down the smallest useful output before collecting a page. For an article, that might be title, author, published_at, and body_html. For a product listing, it could be name, price, currency, availability, and url.
- Specify required and optional fields.
- Choose output types: strings, numbers, dates, arrays, or nested objects.
- Define how missing values are represented, such as
null. - Decide whether you need text, sanitized HTML, links, images, or embedded structured data.
- Avoid collecting the entire page when a few fields are sufficient.
2. Classify the page
| Page shape | Preferred approach | Why |
|---|---|---|
| Article or documentation page | Readability plus metadata checks | Heuristics can isolate the main article from navigation and sidebars. |
| Repeated products, jobs, or records | CSS selectors, XPath, or JSON-LD | Repeated structure maps naturally to records and fields. |
| Tables | Table-specific parsing and header mapping | Rows and columns carry meaning that article heuristics may discard. |
| Dashboard or single-page app | Browser rendering, then DOM or network inspection | The initial response may contain little or none of the visible data. |
| Catalog with embedded data | Parse JSON-LD or application state when available | Structured data is often more stable than visual class names. |
Use semantic containers and meaningful attributes such as href, src, alt, aria-*, data-*, table headers, metadata, and structured data. Avoid selectors based only on presentation, generated class names, or a particular visual position.
3. Save a representative page
During development, save the response or rendered HTML locally. A fixture makes extraction reproducible, reduces repeated requests, and lets you compare code changes against the same input. Keep fixtures that cover normal pages, missing fields, unusually long content, redirects, consent banners, and error pages. Record the URL and retrieval time separately from the HTML.
4. Inspect the initial HTML
First determine whether the desired content exists before JavaScript runs. Inspect the response source, not only the browser’s Elements panel. Search for a distinctive value, a semantic container, JSON-LD, or a script-held application state. If the value is absent, an HTML parser cannot recover it from that response; render the page first.
Minimal cURL capture
curl -L --compressed --max-time 30 'https://example.com/page' -o page.html
rg -n 'article|price|application/ld\+json|target phrase' page.html
Use -L for redirects, a timeout to prevent hung jobs, and a saved file for repeatable inspection. Treat the response as untrusted input.
5. Extract article content with Mozilla Readability
Readability estimates the main content of article-like pages and can return a title and body from a DOM. It is a poor fit for product grids, price tables, dashboards, and pages whose content is missing from the initial HTML. In Node.js, provide a DOM with jsdom.
Node.js: fetch HTML and run Readability
import { JSDOM } from 'jsdom';
import { Readability } from '@mozilla/readability';
const url = 'https://example.com/article';
const res = await fetch(url, { headers: { 'User-Agent': 'ExtractionBot/1.0' } });
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const html = await res.text();
const dom = new JSDOM(html, { url });
const article = new Readability(dom.window.document).parse();
if (!article) throw new Error('Readability could not identify article content');
console.log(JSON.stringify({ title: article.title, text: article.textContent, html: article.content }, null, 2));
Check the returned title and body against the saved page. Readability can remove useful fields such as product prices or table columns; keep a selector-based fallback for those page types.
6. Extract repeated records with selectors
Selectors should describe stable meaning and relationships. Prefer a semantic item container, a label, and a value over a long chain of nested classes.
Python with BeautifulSoup
import json
import requests
from bs4 import BeautifulSoup
url = 'https://example.com/products'
r = requests.get(url, headers={'User-Agent': 'ExtractionBot/1.0'}, timeout=30)
r.raise_for_status()
soup = BeautifulSoup(r.text, 'html.parser')
records = []
for card in soup.select('[data-product], article.product, li.product'):
name = card.select_one('[data-name], .product-name, h2, h3')
price = card.select_one('[data-price], .price')
link = card.select_one('a[href]')
records.append({
'name': name.get_text(' ', strip=True) if name else None,
'price': price.get_text(' ', strip=True) if price else None,
'url': link.get('href') if link else None,
})
print(json.dumps(records, ensure_ascii=False, indent=2))
Normalize relative links with the page URL, parse numeric values with a locale-aware rule, and preserve the original text when normalization could lose meaning.
7. Parse structured data before visual markup
Many pages embed JSON-LD in <script type='application/ld+json'>. Parse it defensively: a page can contain multiple objects, an array, a graph, invalid JSON, or unrelated schema types.
import json
from bs4 import BeautifulSoup
scripts = soup.select("script[type='application/ld+json']")
objects = []
for script in scripts:
try:
value = json.loads(script.string or script.get_text())
objects.extend(value if isinstance(value, list) else [value])
except json.JSONDecodeError:
continue
products = [o for o in objects if isinstance(o, dict) and o.get('@type') in ('Product', ['Product'])]
Use structured data as an anchor and compare it with visible content. Do not assume it is complete, current, or intended to expose every field.
8. Render JavaScript pages before extraction
If the initial response lacks the required information, use a browser automation environment, wait for a meaningful condition, and then inspect the resulting DOM. Waiting for a fixed delay alone is fragile; prefer a selector, a network-idle condition, or an application-specific readiness signal.
Node.js with Playwright
import { chromium } from 'playwright';
const browser = await chromium.launch({ headless: true });
const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
try {
await page.goto('https://example.com/dashboard', { waitUntil: 'domcontentloaded', timeout: 45000 });
await page.waitForSelector('[data-report-ready]', { state: 'visible', timeout: 30000 });
const rows = await page.locator('table tbody tr').evaluateAll(trs => trs.map(tr =>
[...tr.querySelectorAll('th,td')].map(cell => cell.textContent.trim())
));
console.log(JSON.stringify(rows));
} finally {
await browser.close();
}
For pages with lazy loading, scroll deliberately and verify that the expected item count has stabilized. For authenticated pages, provide credentials through a secure session mechanism and never print cookies or tokens.
9. Validate the extracted output
- Check that required fields exist and have the expected type.
- Compare a sample of values with the visible page.
- Detect duplicate records and unexpected empty strings.
- Check dates, currencies, decimal separators, and Unicode normalization.
- Record the source URL, retrieval time, parser version, and fixture identifier.
- Run the extractor against representative page variants after DOM changes.
def require(record, fields):
missing = [f for f in fields if record.get(f) in (None, '')]
if missing:
raise ValueError(f'Missing required fields: {missing}')
require({'name': 'Widget', 'price': '$10'}, ['name', 'price'])
There is no universal accuracy threshold. Define acceptance checks for your schema and monitor failures over time.
10. Sanitize output and review access rights
Extracted HTML is untrusted. Sanitize it before displaying or inserting it into another document. Prefer text or a restricted allowlist when formatting is unnecessary. A parser or screenshot service can fetch a page, but that capability does not grant permission to scrape, store, or republish its content. Review the target site’s terms, robots guidance where relevant, authentication requirements, copyright, privacy obligations, and rate limits.
11. Or skip the browser setup
ScreenshotNeo captures a rendered page through one API request, so you can inspect what a browser sees before writing extraction logic. See the ScreenshotNeo documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const body = Buffer.from(await res.arrayBuffer());
Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. The MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. One thousand screenshots per month are free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
12. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Expected text is absent | JavaScript renders it after the response | Use Playwright or another browser, wait for a readiness selector, then extract. |
| Readability returns null or the wrong region | The page is not article-shaped or has noisy markup | Use selectors or structured data and validate against fixtures. |
| Selector returns zero items | Class names changed, content is in an iframe, or markup differs | Inspect the current DOM, prefer semantic attributes, and handle iframe content explicitly. |
| Values are duplicated | Desktop and mobile copies or nested cards are both selected | Select the smallest record container and deduplicate by a stable key. |
| Prices or dates parse incorrectly | Locale formatting or hidden accessibility text | Preserve raw text, identify locale, then normalize with explicit rules. |
| Browser times out | Slow resources, consent flow, bot challenge, or an application error | Set bounded timeouts, wait on a specific condition, capture diagnostics, and classify the failure. |
| Output contains unsafe markup | Untrusted HTML was rendered directly | Sanitize with an allowlist or return text only. |
| Requests are blocked | Access controls, authentication, or rate limits | Use permitted credentials, respect limits, and review the site’s terms before continuing. |
13. Performance, reliability, and cost
- Fetch once and reuse saved fixtures during development.
- Prefer the initial HTML when it contains the fields; browser rendering is slower and more resource intensive.
- Wait for the smallest reliable readiness condition instead of a long fixed sleep.
- Limit concurrency to what the target site and your infrastructure can support.
- Cache immutable or slowly changing pages, with a documented TTL.
- Use retries only for transient failures, with exponential backoff and a maximum attempt count.
- Log status, redirect chain, response size, render duration, selector counts, and validation failures without logging secrets.
- For managed services, compare rendering support, output control, page coverage, reliability evidence, operational scale, and cost. Do not infer a universal winner from vendor claims.
14. A repeatable extraction checklist
- Fields and output schema are defined.
- A representative page is saved locally.
- Initial HTML was checked for the required content.
- Page type and extraction method match.
- Selectors use meaningful, maintainable anchors.
- JavaScript rendering is used only when necessary.
- Missing, duplicate, malformed, and changed fields are validated.
- Untrusted HTML is sanitized.
- Access rights, terms, privacy, and rate limits were reviewed.
- Metrics and failure diagnostics are recorded.
FAQ
Can I use Readability for any web page?
No. It is designed for article-like content. Use selectors or structured data for listings, catalogs, tables, and dashboards.
Why can a browser show data that requests.get() cannot?
The browser executes JavaScript and may make additional API requests after the initial HTML arrives. A plain HTTP parser sees only the original response.
Should I extract from visible text or JSON-LD?
Use both when available. Structured data can provide stable fields, while visible content lets you verify what users actually see.
How do I keep an extractor working when a site changes?
Keep fixtures, validate required fields, monitor selector counts, prefer semantic anchors, and update the parser when representative pages change.
Does rendering a page make collection permissible?
No. Rendering is a technical step. Permission depends on the site’s terms, applicable law, authentication context, and how you use the data.


