ScreenshotNeo

BlogGuides

What Is Web Scraping? A Beginner’s Guide

Web scraping extracts selected data from web pages and turns it into structured records. Learn the workflow, tools, Python basics, and responsible practices.

By the ScreenshotNeo team4 October 20269 min read

Web scraping is the automated process of requesting web pages, extracting selected information from their HTML or rendered content, and organizing it into usable data. A scraper might collect article titles and dates from a permitted page, then save them as JSON or CSV. It does not necessarily download an entire site.

Crawling and scraping are related but distinct: a crawler discovers pages and follows links; a scraper extracts chosen fields. One program can do both. For a small static page, an HTTP request and an HTML parser are often enough. For many pages, pagination, or browser-rendered content, you may need a crawler framework or browser automation.

1. What web scraping does

A scraper turns a page’s structure into data your program can use. Common fields include product names, public event dates, article headings, or documentation links, subject to the site’s terms, applicable law, privacy obligations, and intended use.

The basic pipeline is:

  1. Define the question and the minimum fields needed.
  2. Check the site’s terms, robots.txt, access controls, and relevant legal and privacy requirements.
  3. Request a permitted page or use an authorized data feed or API.
  4. Parse the HTML or rendered page and select fields.
  5. Normalize and validate values, then store them in JSON, CSV, or a database.
  6. Monitor results and adjust when the page structure changes.

Downloading page source is only one step. Useful scraping means extracting a defined set of fields and checking that the output is accurate.

2. A small Python scraper for a static page

When the needed content is already present in the initial HTML response, Python’s requests and Beautiful Soup are a straightforward starting point. Install the dependencies:

python -m pip install requests beautifulsoup4

Save this as scrape.py. It requests a page, checks the HTTP response, extracts headings and links, validates the result, and writes JSON. Replace the example URL with a page you are permitted to access.

import json
from urllib.parse import urljoin

import requests
from bs4 import BeautifulSoup

URL = "https://example.com/"

response = requests.get(
    URL,
    headers={"User-Agent": "LearningScraper/1.0 (contact: you@example.com)"},
    timeout=(5, 20),
)
response.raise_for_status()

soup = BeautifulSoup(response.text, "html.parser")
records = []

for heading in soup.select("h2"):
    link = heading.find("a", href=True)
    records.append({
        "title": heading.get_text(" ", strip=True),
        "url": urljoin(response.url, link["href"]) if link else None,
    })

if not records:
    raise RuntimeError("No records found; check the page and selectors")

with open("results.json", "w", encoding="utf-8") as file:
    json.dump(records, file, ensure_ascii=False, indent=2)

print(f"Saved {len(records)} records to results.json")

The example uses CSS selectors through soup.select. If it produces no records, inspect the response HTML and adjust the selector to match the page. Selectors are assumptions about a site’s structure, not a stable contract.

Run it and inspect the output

python scrape.py
cat results.json

For CSV output, use Python’s csv.DictWriter and choose explicit field names. For a database, validate and normalize each record before inserting it; make writes idempotent where possible so reruns do not create accidental duplicates.

3. Extract data with cURL or Node.js

For diagnosis, cURL can fetch a page and save its HTML. This does not parse the content by itself:

curl --fail --show-error --location --max-time 25 \
  --user-agent "LearningScraper/1.0 (contact: you@example.com)" \
  "https://example.com/" -o page.html

Node.js can fetch and parse HTML with a library such as Cheerio. Install it with npm install cheerio, then save this as scrape.mjs and run node scrape.mjs:

import * as cheerio from "cheerio";

const url = "https://example.com/";
const response = await fetch(url, {
  headers: { "User-Agent": "LearningScraper/1.0 (contact: you@example.com)" },
  signal: AbortSignal.timeout(20000),
});
if (!response.ok) throw new Error(`HTTP ${response.status} for ${url}`);

const html = await response.text();
const $ = cheerio.load(html);
const records = [];

$("h2").each((_, element) => {
  const title = $(element).text().trim();
  const href = $(element).find("a[href]").first().attr("href");
  records.push({ title, url: href ? new URL(href, response.url).href : null });
});

if (records.length === 0) throw new Error("No records found; inspect HTML and selectors");
console.log(JSON.stringify(records, null, 2));

Keep requests limited and identify your client clearly where appropriate. Do not use these examples to bypass access controls or collect data you are not permitted to use.

4. Choose a tool for the page and scale

Situation Starting point Why
One or a few static pages HTTP client plus HTML parser Small setup; content is already in the response HTML.
Many pages, pagination, or link following Scrapy Provides crawling, request scheduling, pipelines, and export workflows.
Content appears only after browser JavaScript runs Check for an authorized API or feed first; then browser automation if needed A simple HTTP response may not contain the displayed data.
One-off visual record of a rendered page A browser capture workflow A screenshot records appearance rather than extracting structured fields.

Scrapy’s official example selects fields with CSS or XPath, follows a pagination link, and exports JSON Lines. Its scheduler supports asynchronous requests, and settings such as download delay and per-domain concurrency help control request behavior. Start with its official overview when building a repeatable multi-page crawl.

For beginner tutorials on Python parsing and JavaScript-rendered pages, see Real Python’s web scraping tutorials and The Carpentries’ web scraping lesson.

5. Handle pagination, dynamic pages, and changing markup

Pagination and crawling

A crawler starts from one or more allowed URLs, extracts links that match a defined scope, and schedules further requests. Bound the crawl: define allowed domains and URL patterns, set a maximum page count or depth, and detect repeated URLs. For a site with a supported API or feed, prefer that documented interface when it meets the need.

Do not follow every link indiscriminately. Pagination URLs, filters, calendars, and session parameters can create very large or effectively unbounded URL spaces.

JavaScript-rendered content

If a value is missing from the downloaded HTML but visible in a browser, determine how it is supplied. An authorized JSON endpoint or public feed may be simpler and more stable than controlling a browser. If browser rendering is genuinely required and permitted, tools such as Playwright or Selenium can run page scripts before extraction. Browser automation costs more time and resources than parsing an HTTP response, so use it only for the pages that need it.

Markup changes and data quality

Use selectors tied to meaningful structure where possible, and validate critical fields. Check that required values exist, numbers parse, dates use an expected format, and records are not unexpectedly empty or duplicated. Log counts and a small sample of output. A selector can continue running after a site redesign while returning incorrect or incomplete data, so validate results instead of treating a successful process exit as proof.

6. Responsible scraping and site access

Before collecting data, review the site’s terms and robots.txt and consider copyright, privacy, the intended use, and the relevant jurisdiction. Avoid collecting personal or sensitive data unless there is a clear lawful basis and appropriate safeguards. Collect the minimum fields needed and avoid unnecessary load with low concurrency, delays, and caching where appropriate.

Google Search Central describes robots.txt as a file that tells search engine crawlers which URLs they can access. It is a crawler instruction mechanism, not an access-control system or legal permission. Google’s robots.txt documentation also says it should not be used to keep a page secure. A site’s robots rules do not replace checking terms, law, privacy obligations, or technical access controls.

Legality depends on factors such as the data collected, how it is accessed, how it is used, and applicable local law. The reviewed research paper discusses U.S.-based social science research and should not be generalized as a universal legal rule. For consequential commercial or research collection, obtain advice specific to the jurisdiction and use case.

7. Reliability, performance, and cost

  • Request count: Request only necessary pages. Cache responses when permitted and useful, and avoid refetching unchanged material without a reason.
  • Concurrency: More parallel requests may finish sooner but increase load and the risk of errors. Set conservative per-domain concurrency and delays; increase only when appropriate.
  • Timeouts and retries: Set connection and read timeouts. Retry transient failures sparingly with backoff, and do not retry permanent client errors indefinitely.
  • Validation: Track response status, record counts, missing fields, and parse failures. Keep logs sufficient to diagnose changes without storing unnecessary personal data.
  • Browser work: Rendering pages uses more resources than parsing the original response. Restrict it to pages where scripts are necessary.
  • Maintenance: The main ongoing cost is often updating selectors and checks when page markup or behavior changes. Keep representative examples and recheck output after changes.

There is no universal safe request rate or performance figure: site capacity, permission, page size, network conditions, and crawl scope differ. Tune for minimal impact and reliable data rather than maximizing throughput.

8. Troubleshooting common problems

Symptom Likely cause What to do
HTTP 403 or 429 The server denies or limits the request. Stop rapid retries. Review terms and access guidance, reduce request frequency, and use an authorized feed or API if available. Do not attempt to evade a block.
Request times out Slow server, network issue, or unsuitable timeout. Use explicit connect/read timeouts, check connectivity, and retry transient failures sparingly with backoff.
Parser finds no elements Selector mismatch, changed markup, wrong page, or content rendered by JavaScript. Save and inspect the response HTML, verify the final URL and selector, and check whether the data is present before browser execution.
Browser shows data but script does not The initial response does not include browser-generated content. Look for an authorized API or feed; otherwise use permitted browser rendering and wait for a specific element rather than an arbitrary long delay.
Output has duplicates Repeated links, pagination overlap, or no stable record key. Normalize URLs, track visited pages, and deduplicate records using an appropriate key.
Records are malformed or incomplete Unexpected markup, locale-specific formats, or missing values. Validate required fields and formats, preserve raw values for diagnosis when appropriate, and handle missing data explicitly.
Too many pages are discovered Filters or query parameters create URL loops or unbounded combinations. Restrict allowed paths and parameters, track visited URLs, and set crawl limits.

9. Or skip the browser setup

If you need a visual screenshot of a rendered page rather than structured fields, ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. One GET request returns a PNG, JPEG, WebP, or PDF. It can accept cookie and consent banners as a visitor and remove more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and responses report page verdict and billing headers. AI agents can use its MCP server tools take_screenshot, get_page_info, and capture_pdf.

See the ScreenshotNeo API documentation for options. This cURL example saves a WebP screenshot:

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com/ \
  -o shot.webp

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com/"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write("shot.webp", res);

These capture page appearance; they do not replace a parser when you need structured records such as titles, prices, or dates. ScreenshotNeo’s plans include 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, with no card.

10. Frequently asked questions

Is scraping the same as crawling?

No. Crawling discovers and visits pages; scraping extracts selected information. A program can do both.

Can a beginner scrape a page without a framework?

Yes. For a small permitted task where the data is in the initial HTML, an HTTP client and parser are enough to learn the workflow.

Does robots.txt give permission to scrape?

No. It communicates crawler preferences. Review the site’s terms and applicable obligations separately.

When should I move to Scrapy?

Consider it when you need repeatable multi-page crawling, pagination, request scheduling, pipelines, and structured exports.

Should I use a screenshot to extract data?

Usually not for structured fields. Screenshots preserve visual appearance; an HTML parser or authorized data interface is better suited to records.