ScreenshotNeo

BlogHow-to

How to Scrape a Website and Extract Web Content

Learn how to extract structured data from websites with Python, cURL, and Node.js, choose between HTML, data sources, and browser rendering, and troubleshoot common failures.

By the ScreenshotNeo team4 October 202610 min read

To scrape a website, request a page, parse the response, and extract only the fields you need. First check whether those fields are already in the returned HTML. If not, look for a data source the page requests; use a headless browser only when the content remains unavailable through those approaches and is present after browser rendering.

This guide shows a one-page extraction in Python, cURL, and Node.js, then explains how to extend it to repeated pages and dynamic content. Scraping means extracting selected information from pages; crawling means following links to discover or visit additional pages.

1. Define the fields and check access rules

Before writing code, record the page or page pattern you need and the exact fields to collect. For example: title, price, availability, and source URL. Begin with one representative page and confirm that you are authorized to access and use it.

Check the site’s terms and its root-level robots.txt. The Robots Exclusion Protocol provides crawler instructions; it does not grant access authorization. RFC 9309 states, “These rules are not a form of access authorization.” Robots directives, terms, applicable law, and any privacy obligations are separate considerations. See the IETF Robots Exclusion Protocol and MDN’s robots.txt glossary.

Use an authorized API or dataset when available. Keep requests proportionate, collect only necessary information, and reduce or stop activity if the site signals a problem. The reviewed sources do not establish one universally safe request rate.

2. Inspect the response before choosing a scraper

Fetch a page and inspect its response body. A browser may display content that a basic HTTP request does not contain, because JavaScript can fetch data and update the page after the initial response. MDN describes this pattern in its guide to making network requests with JavaScript.

  1. Try the ordinary response HTML first.
  2. If the required fields are missing, inspect the page’s network requests and look for the data source that supplies them. Use it only where permitted.
  3. If the data source is not usable but the information appears in the browser DOM, render the page with a headless browser.

This data-source-first, browser-as-fallback sequence is also recommended in Scrapy’s documentation on selecting dynamically loaded content. It is a decision process, not a guarantee that one approach is always faster or more reliable.

3. Extract fields from ordinary HTML

The examples below fetch one public page, parse its HTML, and print the document title and visible text from paragraph elements. They preserve the source URL so you can trace and validate each result. Selectors are examples: inspect the target page and replace them with selectors for the fields you actually need.

Python with Requests and Beautiful Soup

Install dependencies with python -m pip install requests beautifulsoup4. Save as scrape.py and run python scrape.py.

import json
from urllib.parse import urlparse

import requests
from bs4 import BeautifulSoup

URL = "https://example.com/"

response = requests.get(
    URL,
    headers={"User-Agent": "ExampleResearchBot/1.0 (contact: you@example.com)"},
    timeout=(5, 20),
)
response.raise_for_status()

content_type = response.headers.get("Content-Type", "")
if "html" not in content_type.lower():
    raise ValueError(f"Expected HTML, received {content_type!r}")

soup = BeautifulSoup(response.text, "html.parser")
title = soup.title.get_text(" ", strip=True) if soup.title else None
paragraphs = [p.get_text(" ", strip=True) for p in soup.select("p")]

record = {
    "source_url": response.url,
    "host": urlparse(response.url).hostname,
    "title": title,
    "paragraphs": paragraphs,
}
print(json.dumps(record, ensure_ascii=False, indent=2))

For a real target, replace https://example.com/ and the selectors. For structured fields, prefer selectors tied to meaningful page structure such as an article container or labeled data element over broad selectors such as every paragraph. Missing fields should be represented deliberately, such as null, instead of silently dropping a record.

cURL: retrieve and inspect the response

cURL fetches the response; it does not parse HTML into structured fields by itself. Save the body and inspect it or pass it to a parser.

curl --fail --location --show-error --silent \
  --user-agent 'ExampleResearchBot/1.0 (contact: you@example.com)' \
  --max-time 25 \
  --output page.html \
  --write-out 'HTTP %{http_code}; content type %{content_type}; final URL %{url_effective}\n' \
  'https://example.com/'

Check page.html for the fields you need. If it contains only a shell or placeholder and the browser shows more, move to data-source inspection or browser rendering. Avoid treating a successful HTTP status as proof that the desired data was returned.

Node.js: fetch and extract basic fields

This runnable example uses the built-in fetch API and a small title extractor for the example page. For robust selector-based HTML parsing, install a parser such as Cheerio (npm install cheerio) and use its selectors against the response body.

const url = 'https://example.com/';
const response = await fetch(url, {
  headers: {
    'User-Agent': 'ExampleResearchBot/1.0 (contact: you@example.com)',
  },
  signal: AbortSignal.timeout(20000),
});

if (!response.ok) {
  throw new Error(`HTTP ${response.status} ${response.statusText}`);
}
const contentType = response.headers.get('content-type') || '';
if (!contentType.toLowerCase().includes('html')) {
  throw new Error(`Expected HTML, received ${contentType}`);
}

const html = await response.text();
const titleMatch = html.match(/<title(?:\\s[^>]*)?>([\\s\\S]*?)<\\/title>/i);
const title = titleMatch
  ? titleMatch[1].replace(/<[^>]*>/g, '').trim()
  : null;

console.log(JSON.stringify({ source_url: response.url, title }, null, 2));

For production extraction, use an HTML parser rather than regular expressions: HTML can contain entities, comments, unusual whitespace, and nested markup. With Cheerio, the core pattern is const $ = cheerio.load(html); const title = $('title').text().trim();, then use inspected selectors for the desired fields.

4. Extract JavaScript-loaded content

If the fields are absent from the initial HTML, open the page’s developer tools and inspect its network activity while the page loads. Look for a request whose response contains the data, such as JSON. If an authorized endpoint provides the needed fields, requesting and parsing that response can avoid running a full browser.

When no usable data source is available, but the content appears after the page runs JavaScript, use a headless browser. It loads and renders the page, after which you can inspect the DOM. This adds a browser runtime and more operational complexity, so use it only for pages that require it.

Python browser-rendering example with Playwright

Install Playwright and its Chromium browser with python -m pip install playwright and python -m playwright install chromium.

import asyncio
import json
from playwright.async_api import async_playwright

async def main():
    url = "https://example.com/"
    async with async_playwright() as p:
        browser = await p.chromium.launch()
        page = await browser.new_page()
        await page.goto(url, wait_until="domcontentloaded", timeout=30000)
        await page.locator("body").wait_for(state="visible", timeout=10000)
        result = await page.locator("body").inner_text()
        title = await page.title()
        print(json.dumps({"source_url": page.url, "title": title, "text": result}, ensure_ascii=False, indent=2))
        await browser.close()

asyncio.run(main())

Replace the example URL and wait condition with the target page and a selector that indicates the required content is ready. Waiting for domcontentloaded alone does not mean that asynchronous page data has arrived.

5. Crawl multiple pages and validate records

Once one page works, add only the links or pagination needed for the task. Keep an explicit scope, avoid revisiting URLs, and store each source URL with its extracted values. For each page, handle errors independently so one unavailable page does not erase successful records.

  1. Normalize and validate URLs before requesting them; restrict the crawler to the intended host and paths.
  2. Track visited URLs and use the page’s actual pagination links or documented API parameters.
  3. Set connection and read timeouts, limit concurrency, and use a proportionate request pace. Do not invent a universal request interval; follow the site’s guidance and observed responses.
  4. Write records incrementally in a structured format such as JSON Lines or CSV, including the source URL and a collection timestamp when useful.
  5. Validate samples against their source pages, including missing fields, alternate layouts, and the final page of pagination.

Page markup can change, fields can be absent, and pagination can repeat or skip results. Track expected fields and report parse failures rather than quietly emitting misleading records. Keep only information you are authorized to collect and need for the stated purpose.

6. Choose the right approach

Page situation Approach Trade-off
Fields are in the HTTP response HTML HTTP client plus HTML parser Lightweight and direct; selectors may need maintenance when page structure changes.
Many repeated pages need link following and response handling A crawling framework such as Scrapy Provides crawl-oriented request and response handling; requires setup and maintenance.
Fields are missing in HTML but appear in an authorized data response Inspect and request the data source Avoids rendering a full browser; the source format or endpoint can change.
Fields appear only after browser-side rendering Headless browser Can expose rendered DOM; adds runtime, resource, and operational complexity.

7. Troubleshooting

Symptom Likely cause What to do
Expected text is missing The response contains a JavaScript shell, or the selector does not match the page. Inspect the response and selector. Find the data source; if unavailable and the content renders in the DOM, use a headless browser.
HTTP 403 or 429 The site denied the request or indicated that the request rate is excessive. Stop or reduce requests, review site rules and terms, and use an authorized API or access path. Do not try to evade access restrictions.
HTTP 404 The URL is stale, mistyped, or the page was removed. Check the source link and pagination logic; record the failed URL and continue only where appropriate.
Timeout or connection error Slow response, network issue, or an unsuitable timeout. Use bounded connect/read timeouts, retry transient failures sparingly with backoff, and avoid retrying permanent client errors.
Parser reports no title or fields Markup differs from assumptions, content type is not HTML, or the selector is too narrow. Check status, final URL, content type, and a saved response sample; update selectors and handle absent values explicitly.
Duplicate or missing records Pagination links repeat, URL variants represent the same page, or crawl state is incomplete. Normalize URLs, track visited pages, inspect next-page links, and validate record identifiers across page boundaries.
Playwright cannot launch Browser binaries are not installed or the runtime lacks required dependencies. Run the Playwright browser installation command for the selected browser and check its installation guidance for the environment.
Rendered page is still incomplete The script waited for initial navigation rather than the target data. Wait for a specific content selector or a relevant data response, with a finite timeout; verify the page state before extracting.

8. Performance, reliability, and cost

Direct HTTP requests generally avoid the extra browser process and page rendering, while a headless browser must load and execute page resources. The appropriate choice depends on whether the required data is present in HTML, available from a permitted data source, or only exposed in the rendered DOM. Compare methods on the actual fields, scale, resource needs, and maintenance burden rather than assuming one always wins.

Reliability comes from bounded timeouts, careful error handling, incremental output, explicit crawl scope, and validation against source pages. Retries should be limited to transient failures; repeated attempts can add load and are not a fix for denied access or broken selectors. Parsing should tolerate missing optional fields but flag missing required ones.

For a small one-off extraction, a local script may be sufficient. At larger scale, account for request volume, browser runtime, storage, and maintenance. This guide does not assign a universal cost or safe rate: both depend on the target, infrastructure, and authorized access method.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. It returns a PNG, JPEG, WebP, or PDF from one GET request. A screenshot can help when you need a visual record of a rendered page; it is not a replacement for structured extraction when you need fields as data. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses identify the page verdict and billing status in headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000; all features are on every plan, and yearly billing gives two months free. Sign up for 1,000 free screenshots a month with no card.

FAQ

Is scraping the same as crawling?

No. Scraping extracts selected information from a page; crawling follows links to discover or visit pages. A project may do both.

Does robots.txt give permission to scrape a page?

No. It communicates crawler instructions and is not access authorization. Check the site’s terms, applicable law, and authorized access options as well.

Should I always use a headless browser?

No. Start with response HTML, then look for a permitted data source. Use browser rendering when the required content is only available after client-side rendering.

Can a screenshot API return structured records?

A screenshot is an image or document capture. Use an HTML parser or an authorized data source when your output needs structured fields.