How to Build Production-Ready Web Scrapers in 30 Minutes
Build a dependable scraper in 30 minutes with permission checks, timeouts, retries, validation, observability, and a clear path to browser rendering.

Short answer: define a small data contract, verify that crawling is allowed, fetch with a session and explicit timeouts, parse into a validated schema, add bounded retries and rate limits, then run a smoke test with logs and metrics. For server-rendered pages, Python Requests plus an HTML parser is usually the fastest path. Use Scrapy for broad crawls and Playwright when JavaScript or interaction creates the data.
The 30-minute target is a focused first production version, not a throughput benchmark. The goal is a scraper that fails visibly, respects the site, and can be operated safely.
What you can build in 30 minutes
- Minutes 0–3: define fields, URL scope, freshness, output format, and stop conditions.
- Minutes 3–6: check the official API or feed, then review
/robots.txt, terms, and authorization. - Minutes 6–12: create one reusable HTTP session with connect and read timeouts.
- Minutes 12–18: parse stable selectors or documented JSON and normalize values.
- Minutes 18–24: add bounded retries, jitter, caching, and conservative rate limits.
- Minutes 24–30: run a small sample, validate rows, save a raw fixture, and emit metrics.
1. Define the scraper contract first
Write down the exact fields and their types before writing selectors. Specify the allowed hostnames and URL patterns, maximum pages, freshness requirement, and what should stop a run. A contract prevents a parser from silently accepting a redesign.

| Decision | Example |
|---|---|
| Fields | title (string), price (decimal), published_at (ISO-8601) |
| Scope | Only example.com/articles/; no external links |
| Freshness | Data must be less than 24 hours old |
| Stop conditions | 100 pages, 3 consecutive ban pages, or 20% validation failures |
Prefer an official API, feed, or bulk export when available. Scrapy’s optimization guidance explains that these options are often faster and cheaper for the site than crawling pages (Scrapy practices).
2. Check permission and crawl policy
Fetch https://host.example/robots.txt before crawling. Find the user-agent group that applies to your crawler and obey the most-specific matching allow or disallow rule. RFC 9309 describes robots.txt as a crawler access protocol; its rules are requests to crawlers, not an access-control system (RFC 9309). Read the site’s terms and confirm your intended use is authorized.
If robots.txt cannot be reached because of a server or network error, RFC 9309 says crawlers must assume complete disallow. Do not treat a missing or malformed file as permission to crawl. Translate any Crawl-delay or Request-rate directives into your own delay and concurrency settings; Scrapy does not automatically act on every such directive.
3. Start with Requests and an HTML parser
For a small, server-rendered target, use one requests.Session. A session reuses connections and keeps cookies, while Requests supplies decompression, proxies, streaming, and other transport features (Requests advanced usage). Always set an explicit timeout; without one, a call can hang indefinitely (Requests timeouts).
from __future__ import annotations
import json
import random
import time
from dataclasses import dataclass
from datetime import datetime, timezone
from decimal import Decimal
from urllib.parse import urljoin, urlparse
import requests
from bs4 import BeautifulSoup
@dataclass
class Record:
url: str
title: str
price: Decimal | None
fetched_at: str
class Scraper:
def __init__(self) -> None:
self.session = requests.Session()
self.session.headers.update({
'User-Agent': 'ExampleResearchBot/1.0 (+https://example.com/bot-info)'
})
self.timeout = (5, 20) # connect, read seconds
def fetch(self, url: str) -> requests.Response:
response = self.session.get(url, timeout=self.timeout)
response.raise_for_status()
return response
def parse(self, url: str, html: str) -> Record:
soup = BeautifulSoup(html, 'html.parser')
title_node = soup.select_one('h1')
if not title_node:
raise ValueError('required field missing: h1')
title = ' '.join(title_node.get_text(' ', strip=True).split())
price_node = soup.select_one('[data-price]')
price = None
if price_node:
price = Decimal(price_node['data-price'])
return Record(
url=url,
title=title,
price=price,
fetched_at=datetime.now(timezone.utc).isoformat(),
)
def main(urls: list[str]) -> None:
scraper = Scraper()
rows = []
for url in urls:
started = time.monotonic()
try:
response = scraper.fetch(url)
row = scraper.parse(response.url, response.text)
rows.append(row.__dict__)
print(json.dumps({'url': url, 'status': response.status_code,
'bytes': len(response.content),
'seconds': time.monotonic() - started}))
except (requests.RequestException, ValueError) as exc:
print(json.dumps({'url': url, 'error': str(exc)}))
time.sleep(random.uniform(1.0, 2.0))
with open('records.json', 'w', encoding='utf-8') as output:
json.dump(rows, output, indent=2, default=str)
if __name__ == '__main__':
main(['https://example.com/articles/first'])
Replace selectors and the example URL with the target’s documented structure. Keep the raw response for failed samples so selector regressions are diagnosable.
4. Normalize and validate data
Normalize whitespace, dates, currencies, and encodings at the boundary. Validate required fields, types, uniqueness, and freshness before writing output. Quarantine malformed records with the source URL and fetch timestamp instead of emitting partial rows. For JSON endpoints, validate the documented keys and preserve the original payload alongside the normalized record.
5. Add safe retries, rate limits, and caching
Retry only transient failures: connection resets, timeouts, and selected 5xx responses. Do not blindly retry authentication errors, authorization failures, or permanent 4xx responses. Use bounded exponential backoff with jitter, and honor Retry-After on 429 and 503 responses.

import random
import time
import requests
RETRYABLE = {429, 500, 502, 503, 504}
def get_with_backoff(session, url, attempts=4):
for attempt in range(attempts):
try:
response = session.get(url, timeout=(5, 20))
if response.status_code not in RETRYABLE:
response.raise_for_status()
return response
retry_after = response.headers.get('Retry-After')
delay = float(retry_after) if retry_after and retry_after.isdigit() else 2 ** attempt
except (requests.Timeout, requests.ConnectionError):
delay = 2 ** attempt
if attempt == attempts - 1:
raise
time.sleep(min(delay, 60) + random.uniform(0, 0.5))
Start with low concurrency and increase gradually while watching 429/503 rates, ban pages, retries, and latency. Cache responses or fingerprint requests so a rerun does not download identical pages. Scrapy includes caching, duplicate-request filtering, download delays, auto-throttling, and robots middleware for larger crawls (AutoThrottle, downloader middleware).
6. Choose Requests, Scrapy, or Playwright
| Tool | Use it when | Trade-offs |
|---|---|---|
| Requests + parser | One site, server-rendered HTML, modest page count | Smallest operational surface; no browser JavaScript |
| Scrapy | Many pages or domains, pipelines, concurrency, retries, and crawl policy | More configuration; excellent scheduling and observability hooks |
| Playwright | Data appears after JavaScript runs or requires clicks, scrolling, or login | Browser CPU and memory costs; more complex failure modes |
Playwright exposes page request and response events so you can find the API calls that produce data (Playwright network events). Browser actions default to a 30-second timeout unless configured (Playwright timeouts).
7. JavaScript-rendered pages with Playwright
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
page.set_default_timeout(15_000)
page.goto('https://example.com/catalog', wait_until='domcontentloaded', timeout=30_000)
page.locator('[data-product]').first.wait_for()
records = page.locator('[data-product]').evaluate_all("""els => els.map(el => ({
name: el.querySelector('[data-name]')?.textContent?.trim(),
price: el.querySelector('[data-price]')?.getAttribute('data-price')
}))""")
print(records)
browser.close()
Prefer the underlying JSON request when the page exposes one; it is usually simpler and cheaper than rendering every page. If you must use a browser, block unnecessary resources, reuse browser contexts, and close pages promptly.
8. Smoke tests and observability
Run five to ten representative URLs before a full crawl. Assert a minimum row count, required fields, valid types, and expected hostnames. Save one successful raw response as a fixture. Emit structured logs containing a run ID, URL, status, elapsed time, response bytes, retry count, parser result, and validation errors.
Alert on elevated 429/503 responses, ban pages, rising latency, zero-row runs, parser exceptions, and validation failures. Track freshness and duplicate rates as well as HTTP metrics. Keep fixtures from different templates and rerun them after a site redesign.
Common errors and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Request hangs | No timeout or stalled upstream | Set separate connect/read timeouts and cap retries |
| 403 or 429 | Permission issue, excessive rate, or bot defense | Confirm authorization, slow down, honor Retry-After, and stop on persistent blocks |
| Empty fields | JavaScript-rendered content or changed selectors | Inspect raw HTML, use stable attributes, or switch to Playwright/API data |
| Duplicate rows | Pagination loops or repeated URLs | Canonicalize URLs and maintain a request fingerprint set |
| Wrong encoding | Incorrect charset detection | Use the response encoding metadata and preserve raw bytes for diagnosis |
| Parser suddenly fails | Template change | Compare a saved fixture, update selectors, and rerun smoke tests |
Performance, reliability, and cost
Connection pooling and response caching reduce latency and bandwidth. Concurrency is not automatically faster: raising it can trigger throttling, bans, and more retries. Measure pages per minute together with error rate and server response time. Browser rendering costs more CPU and memory than HTTP fetching, so reserve it for pages that need JavaScript.
For reliability, make each page operation idempotent, checkpoint progress, and write outputs atomically. Separate transport, parsing, and validation errors so operators know whether a target is unavailable or your selector is wrong. Keep credentials out of logs and pass them through environment variables or a secret manager.
Or skip the browser setup
If your goal is a clean screenshot or PDF of a page rather than extracting structured fields, ScreenshotNeo provides a single GET request. It accepts cookie and consent banners like a visitor, then removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response reports the result in X-Page-Verdict and X-Billed headers.
See the full option list and parameter names in the ScreenshotNeo API documentation.
cURL
curl -G 'https://api.screenshotneo.com/v1/shot' -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get('https://api.screenshotneo.com/v1/shot', params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'}, timeout=90)
r.raise_for_status()
open('shot.webp', 'wb').write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const file = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', file));
ScreenshotNeo also supports full-page captures with lazy images loaded, CSS-element capture, dark mode, 12 device presets and custom viewports, retina scale, PDF paper and margin settings, custom CSS and JavaScript, clicks, selector or network-idle waits, ad and tracker blocking, custom headers, cookies, user agents and Authorization, timezone and geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture for 100 URLs per call, a usage API, and an OpenAPI specification. An MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.
There are 1,000 screenshots per month free with no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free. Create a free ScreenshotNeo account.
FAQ
Should I parse HTML or call an API?
Use the official API, feed, or export when it contains the fields you need. It is generally less fragile and creates less load than page crawling.
Can robots.txt grant permission?
No. It communicates crawler preferences and is not authorization. You still need to review terms and obtain permission for your use.
How many retries are enough?
Use a small bounded number, such as three or four, with backoff. More retries can turn a temporary fault into a prolonged overload.
When should I move to Scrapy?
Move when you need scheduling, many URLs, concurrency controls, pipelines, duplicate filtering, or built-in crawl middleware.
When is Playwright unavoidable?
Use it when the required content is generated only after JavaScript execution or requires browser interactions that an HTTP client cannot reproduce.