ScreenshotNeo

BlogGuides

Web Scraping: What It Is, How It Works, and Best Practices

Learn how web scraping requests, parses, validates, and stores page data, with runnable code, responsible crawling guidance, and troubleshooting.

By the ScreenshotNeo team1 October 20269 min read

Web scraping is the automated extraction of selected information from web pages or services. A typical scraper sends an HTTP request, receives HTML or structured data, parses it, selects the fields it needs, validates the result, and stores the smallest useful dataset. The technical workflow is straightforward; whether a particular collection is permitted depends on the data, method, terms, purpose, and jurisdictions involved.

1. What is web scraping?

Web scraping uses software to collect specific values from web content instead of copying them manually. A scraper might collect article titles, product prices, event dates, documentation links, or metadata from a known set of pages.

Scraping is different from browsing. A browser displays a page for a person; a scraper applies repeatable rules to many responses. It is also different from an official API. An API is a documented access route with defined request and response formats. If an API supplies the data you need under clear conditions, it is usually the better first choice.

Scraping, crawling, and APIs

Method What it does Typical trade-off
Scraping Parses page content for selected fields Flexible, but page markup can change
Crawling Discovers and visits many linked pages Requires scope, scheduling, rate control, and deduplication
Official API Requests documented resources and fields More stable, but limited to the provider’s access and data model
Undocumented endpoint Uses an internal request discovered in a site May change without notice and can have different access or legal considerations

The Canadian privacy commissioners describe APIs as a controlled route that can support credentials, logging, and monitoring, while noting that APIs still have terms and limits. Compare availability, access conditions, documentation, rate limits, response structure, and maintenance before choosing a method.

2. How web scraping works

  1. Define the fields. Write down the exact values, source pages, retention period, and permitted purpose.
  2. Choose the access route. Check for an official API before parsing HTML.
  3. Check site instructions. Review robots.txt, terms, privacy notices, and any published crawling guidance.
  4. Request a page. Send an HTTP GET with a descriptive user-agent and a timeout.
  5. Inspect the response. Check status code, content type, encoding, redirects, and whether the expected content is present.
  6. Parse the document. Use stable selectors for text, attributes, links, and structured data.
  7. Normalize and validate. Convert dates and numbers, remove irrelevant whitespace, reject malformed records, and timestamp results.
  8. Store minimally. Save only fields needed for the stated purpose, with source URL and retrieval time.
  9. Monitor change. Track parse failures, response codes, schema changes, and request volume.

Static pages can be handled with an HTTP client and an HTML parser. Pages that render data through JavaScript may require a browser or a documented data endpoint, but technical access does not establish permission to collect or reuse the content.

3. A complete Python scraper for static HTML

This example collects article titles and links from a page, validates the response, and writes JSON. Replace the example URL and selectors with a target you are permitted to access.

import json
import time
from urllib.parse import urljoin

import requests
from bs4 import BeautifulSoup

URL = "https://example.com/news"
HEADERS = {
    "User-Agent": "ExampleResearchBot/1.0 (+https://example.com/contact)"
}

response = requests.get(URL, headers=HEADERS, timeout=30)
response.raise_for_status()

content_type = response.headers.get("content-type", "")
if "html" not in content_type:
    raise ValueError(f"Expected HTML, received {content_type!r}")

soup = BeautifulSoup(response.text, "html.parser")
records = []
for node in soup.select("article h2 a"):
    title = node.get_text(" ", strip=True)
    href = node.get("href")
    if not title or not href:
        continue
    records.append({
        "title": title,
        "url": urljoin(URL, href),
        "retrieved_at": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
    })

with open("results.json", "w", encoding="utf-8") as output:
    json.dump(records, output, ensure_ascii=False, indent=2)

print(f"Saved {len(records)} records")

Install dependencies with python -m pip install requests beautifulsoup4. Keep selectors narrow and test what happens when an element is missing. Do not assume every response is a successful page.

4. Equivalent requests in cURL and Node.js

cURL: inspect a response before parsing

curl --fail-with-body --location \
  --max-time 30 \
  --user-agent 'ExampleResearchBot/1.0 (+https://example.com/contact)' \
  'https://example.com/news'

Use -D headers.txt to save headers and -o page.html to save the body. Add a delay between requests in a crawl.

Node.js: fetch and parse with Cheerio

import * as cheerio from "cheerio";

const url = "https://example.com/news";
const response = await fetch(url, {
  headers: {
    "User-Agent": "ExampleResearchBot/1.0 (+https://example.com/contact)"
  },
  signal: AbortSignal.timeout(30_000)
});

if (!response.ok) {
  throw new Error(`HTTP ${response.status} for ${url}`);
}
const type = response.headers.get("content-type") || "";
if (!type.includes("text/html")) {
  throw new Error(`Expected HTML, received ${type}`);
}

const html = await response.text();
const $ = cheerio.load(html);
const records = $("article h2 a").map((_, element) => ({
  title: $(element).text().replace(/\\s+/g, " ").trim(),
  url: new URL($(element).attr("href"), url).href
})).get().filter(record => record.title && record.url);

console.log(JSON.stringify(records, null, 2));

Install Cheerio with npm install cheerio. For a browser-rendered page, identify whether a documented endpoint provides the same data before introducing browser automation.

5. Crawling multiple pages safely

A crawler adds URL discovery, a queue, deduplication, limits, retries, and pacing to the single-page workflow.

  • Start with an explicit allowlist of hosts and URL patterns.
  • Normalize URLs and keep a visited set to prevent loops.
  • Set maximum pages, depth, bytes, and runtime before starting.
  • Use one descriptive user-agent and a contact address where appropriate.
  • Cache responses so a retry does not fetch the same page unnecessarily.
  • Use exponential backoff for transient failures and honor Retry-After when supplied.
  • Pause on HTTP 429. Treat repeated 403 responses as a signal to stop and review access conditions.
  • Stop if the site owner asks you to stop.

AWS gives context-dependent examples of one request every 10–15 seconds for small or medium sites and 1–2 requests per second for larger sites or sites where you have explicit permission. These are examples, not universal limits. Follow site-specific instructions and choose the lowest rate that meets your need.

6. JavaScript-rendered pages

View-source HTML may not contain data that appears after scripts run. Before using a browser, inspect the page’s documented API or network requests and confirm that the route is permitted. If browser rendering is necessary, use a bounded wait, a specific readiness selector, and a maximum page count. Avoid collecting session data, hidden fields, or content unrelated to your purpose.

For visual capture rather than field extraction, a screenshot service can remove browser setup. ScreenshotNeo is a website screenshot API and MCP server. Its capture options include full-page shots with lazy images loaded, CSS element capture, custom JavaScript and CSS, waits, request blocking, headers, cookies, user agents, timezone, geolocation, and PDF output.

7. Or skip the browser setup

For a rendered page image or PDF, call ScreenshotNeo’s API. See the ScreenshotNeo API documentation for all parameters.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers identify the page verdict and whether the shot was billed. ScreenshotNeo also provides an MCP server with take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

8. Robots.txt, terms, and permission

robots.txt communicates crawl preferences. Google’s documentation says it tells search engine crawlers which URLs they may access; it does not enforce behavior, and a disallowed URL may still appear in search results. It is not an access-control mechanism for confidential material. Do not treat a permissive file as permission to reuse data.

Review the site’s terms, privacy policy, API documentation, and published contact or crawling guidance. CNIL says data scraping is not prohibited per se and requires case-by-case analysis. Other rules may apply, including copyright, database rights, contractual terms, computer-access laws, and data-protection requirements.

9. Personal data and privacy

Public availability does not automatically remove privacy obligations. The EDPB states that GDPR can apply when scraping involves processing personal data, including collection, storage, organisation, or retrieval. Depending on the project, consider purpose limitation, transparency, data minimisation, accuracy, retention, lawful basis, and safeguards for special-category data. The EDPB material cited here focuses on AI-development data; apply its details to the relevant context rather than assuming every recommendation is universal.

Define a collection purpose before crawling. Exclude unnecessary fields, avoid sensitive categories unless you have a documented basis and safeguards, timestamp records, validate accuracy, restrict access, and delete data when it is no longer needed. Get jurisdiction-specific legal advice for high-risk projects.

10. Reliability, performance, and cost

Reliability checklist

  • Set connect and read timeouts.
  • Record URL, status, content type, retrieval time, parser version, and error reason.
  • Retry only transient failures, with a cap and backoff.
  • Do not retry 401, 403, or repeated 429 responses blindly.
  • Validate required fields and send malformed records to a review queue.
  • Keep fixtures from known responses so parser changes can be detected.

Performance checklist

  • Request only the pages and fields you need.
  • Reuse HTTP connections where your client supports it.
  • Cache immutable or recently fetched pages.
  • Bound concurrency per host and honor rate limits.
  • Prefer an API when it eliminates HTML parsing or browser rendering.
  • Compress stored output and avoid retaining raw pages when they are unnecessary.

Costs include API quotas, bandwidth, storage, proxy or browser infrastructure, and engineering time spent repairing selectors. Browser rendering generally consumes more resources than a plain HTTP request. Measure response size, latency, error rate, and records per request for your own workload instead of assuming a universal benchmark.

11. Troubleshooting common errors

Symptom Likely cause Fix
403 Forbidden Access policy, blocked user-agent, or disallowed route Review terms and robots guidance, identify yourself, reduce scope, and stop if access remains denied.
429 Too Many Requests Rate limit exceeded Pause, honor Retry-After, lower concurrency, and resume only when permitted.
200 response but no records Wrong selector, consent page, login page, or JavaScript rendering Save the response, inspect its title and structure, verify selectors, and find a permitted API or rendering method.
Unexpected encoding Missing or incorrect charset metadata Inspect response headers and parser encoding; preserve Unicode and test accented text.
Duplicate rows Repeated links, pagination overlap, or retries Normalize URLs and deduplicate with a stable key.
Parser breaks after a redesign Markup or class names changed Use semantic selectors, fixture tests, schema checks, and alerts for sudden drops.
Timeouts Slow origin, large response, or blocked connection Set separate connect/read limits, cap retries, and investigate before increasing concurrency.

12. Best-practice checklist

  • ☐ State the purpose and fields before collecting.
  • ☐ Prefer an official API when it provides the required data.
  • ☐ Review robots.txt, terms, privacy notices, and contact guidance.
  • ☐ Identify your crawler and provide contact information where appropriate.
  • ☐ Use conservative pacing, bounded concurrency, and backoff.
  • ☐ Handle 403, 429, redirects, logins, and non-HTML responses deliberately.
  • ☐ Minimise personal data and define retention and deletion rules.
  • ☐ Validate fields, timestamp records, and monitor parser health.
  • ☐ Stop when the owner asks or access conditions are unclear.

13. Frequently asked questions

There is no universal answer. The result depends on the data, access method, terms, purpose, and applicable jurisdictions. Treat legal review as a project requirement when personal data, copyrighted material, restricted systems, or large-scale collection are involved.

Does robots.txt give permission to scrape?

No. It communicates crawler preferences and does not replace authentication, contractual terms, or other access controls.

Should I use an API or scrape HTML?

Use the official API when it supplies the needed data under workable conditions. Scrape HTML when the required information is not available through a suitable documented route and the collection is permitted.

How do I scrape a page that needs JavaScript?

First look for a permitted documented endpoint. If rendering is required, use a controlled browser workflow with explicit waits, resource limits, and a narrow data scope.

How often should a crawler make requests?

Follow the site’s instructions and use the lowest practical rate. AWS publishes different example rates for small and large sites; they are operational examples, not a universal standard.

Sources and further reading