ScreenshotNeo

BlogGuides

Web Scraping Guide: Tools, Techniques, and Best Practices

Choose the right scraping method, build a responsible Python workflow, and troubleshoot common failures across HTTP clients, crawlers, and browsers.

By the ScreenshotNeo team4 October 20268 min read

For a straightforward scraper, fetch a page with an HTTP client and parse its HTML with an HTML parser. Use a crawler framework when you need crawl scheduling and request management, and browser automation when the page depends on JavaScript rendering or interaction. Before collecting data, check the site’s rules, keep requests bounded, and treat every response as untrusted input.

Prefer an official API, export, or feed when it provides the data you need. A robots.txt file gives crawler instructions; it does not grant access or settle whether a project is lawful. The right method depends on the page, request volume, data, and purpose.

How do I scrape a website?

  1. Define the fields and scope. Identify the target pages and the specific data needed. Collect only those fields.
  2. Check access rules. Review the site’s terms and technical restrictions, applicable privacy and copyright obligations, and its robots.txt file. Prefer a documented access method when available.
  3. Fetch a representative page. Use an HTTP client if the response already contains the required content. Inspect the response status, content type, and size.
  4. Parse and validate. Extract only the required elements, normalize values, and handle missing or changed markup explicitly.
  5. Expand carefully. Add pagination or a crawl queue only after the single-page workflow is reliable. Bound concurrency and request frequency, identify your crawler clearly, and stop or reassess if access is blocked or the site signals distress.
  6. Keep provenance when useful. Record retrieval time and source URL when the use case calls for auditability or refreshes.

Do not execute scripts from retrieved pages or unsafely deserialize their contents. Validate scraped values before using them in database queries, URLs, or filesystem paths. Large responses and full-document parse trees can consume substantial memory.

Which web scraping tool should I use?

Need Good starting point Tradeoff
A few pages; data is in the returned HTML Python Requests plus Beautiful Soup Simple and direct, but pagination, retries, and maintenance are yours to manage.
Recurring or larger crawl with request coordination Scrapy Provides a crawler framework; you still need to configure scope, limits, and data handling.
Browser-rendered content or interactions Playwright Automates browser behavior, with additional browser setup and runtime overhead.
Python robots rules checks urllib.robotparser Provides robots.txt rule checking; confirm its behavior fits your project.

Choose based on rendering needs, page count and frequency, pagination, how often the site changes, data sensitivity, and operational complexity. There is no single library that fits every scraping task.

Python example: Requests and Beautiful Soup

This example fetches one page, checks the HTTP response, and extracts links. Install the dependencies with python -m pip install requests beautifulsoup4.

from urllib.parse import urljoin

import requests
from bs4 import BeautifulSoup

url = "https://example.com/"
headers = {"User-Agent": "ExampleResearchBot/1.0 (contact: data@example.org)"}

response = requests.get(url, headers=headers, timeout=(5, 20))
response.raise_for_status()

content_type = response.headers.get("Content-Type", "")
if "html" not in content_type.lower():
    raise ValueError(f"Expected HTML, got {content_type!r}")

soup = BeautifulSoup(response.text, "html.parser")
for anchor in soup.select("a[href]"):
    title = anchor.get_text(" ", strip=True)
    link = urljoin(response.url, anchor["href"])
    print({"title": title, "url": link})

Replace the example domain and selectors with a target you are permitted to access. The example makes one request; it does not implement a crawl queue, retry policy, or robots.txt check.

Check robots.txt in Python

Python’s urllib.robotparser can check whether a user-agent may fetch a URL according to a parsed robots file:

from urllib.robotparser import RobotFileParser

site_root = "https://example.com"
robots = RobotFileParser(f"{site_root}/robots.txt")
robots.read()

user_agent = "ExampleResearchBot"
target = "https://example.com/catalog/item"
if robots.can_fetch(user_agent, target):
    print("Allowed by the parsed robots rules")
else:
    print("Disallowed by the parsed robots rules")

This is a basic rule check, not a complete crawl policy or a legal determination. For a production crawler, handle robots retrieval failures deliberately, use an explicit crawler identity, and apply the rules to the user-agent you actually send.

Do I need a browser automation tool?

Use a browser when the task truly depends on browser behavior: for example, content appears only after client-side rendering, or reaching the target requires a user-like interaction. Playwright automates browsers for these workflows. If the needed data is already present in the HTTP response, browser automation adds setup and runtime overhead without helping the extraction.

For a screenshot rather than structured extraction, a screenshot API can avoid managing a browser capture stack. [ScreenshotNeo](https://screenshotneo.com) is a website screenshot API and MCP server. Its one-request API returns PNG, JPEG, WebP, or PDF; its capture options include full-page and selector screenshots, device presets, waits, custom CSS and JavaScript, and request controls. See the [ScreenshotNeo API documentation](https://screenshotneo.com/docs/).

Or skip the browser setup

If your goal is to capture a page image or PDF rather than extract structured fields, call ScreenshotNeo’s API directly. This Python example saves the response body as an image:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
with open("shot.webp", "wb") as image:
    image.write(r.content)

Use your API key and target URL. The [API documentation](https://screenshotneo.com/docs/) describes the request options. Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents use screenshot tools. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for free and get 1,000 screenshots a month with no card.

How should a crawler handle robots.txt?

RFC 9309 standardizes the Robots Exclusion Protocol. Rules are grouped by user-agent, and the most specific matching path rule applies; when Allow and Disallow rules are equally specific, Allow takes precedence. The standard says: “These rules are not a form of access authorization.” A site’s terms, access controls, and applicable law remain separate considerations.

Robots response RFC 9309 treatment Practical handling
Successful retrieval Parse and follow parseable rules. Apply the rules for your crawler’s user-agent.
4xx response File is unavailable; a crawler may access resources. Do not assume this grants permission under other site rules or laws.
5xx response or network failure File is unreachable; assume complete disallow while the condition applies. Pause crawling and retry cautiously; do not treat an outage as approval.

RFC 9309 says robots.txt should not ordinarily be served from cache for more than 24 hours, unless it is unreachable. If an implementation imposes a parsing limit, the standard requires it to support at least 500 KiB. These are protocol details, not a site-specific request rate.

How do I make a scraper reliable and efficient?

  • Start small. Validate selectors and output on a few pages before expanding scope.
  • Bound work. Limit concurrency, request frequency, page count, and response size. Follow site-specific expectations.
  • Handle failures explicitly. Check status codes and content types, set timeouts, and distinguish transient network problems from persistent blocks or changed content.
  • Parse only what you need. Building a full in-memory tree for a large response can use substantial memory. Avoid downloading unnecessary assets.
  • Expect change. Selectors and page structures can change. Monitor missing fields and extraction errors, and reassess when access is blocked or the site signals distress.
  • Keep data safe. Treat page text and attributes as untrusted. Validate before storage or use, and avoid executing returned content.

For a few pages, Requests plus Beautiful Soup keeps the runtime and operational setup small. A framework can help coordinate a larger recurring crawl, but it does not remove the need to manage limits, failures, site rules, and data safety. Browser automation generally has more setup and runtime overhead, so reserve it for browser-dependent pages.

There is no universal rule that publicly accessible content is always lawful to scrape. The relevant analysis depends on the jurisdiction, site’s terms and technical restrictions, data type, purpose, and downstream use. Personal data can trigger privacy obligations; copyright and contract questions can also matter.

The cited EU court material addresses GDPR processing in a specific factual context, and U.S. Department of Justice material references particular CFAA litigation. Neither resolves every scraping project or jurisdiction. Assess the facts of your project and seek qualified legal advice when the stakes warrant it. A robots.txt rule is a crawler instruction, not access authorization.

Troubleshooting common scraping errors

Symptom Likely cause What to do
Connection hangs or times out Slow origin, network issue, or page that does not respond. Set connect and read timeouts. Reduce request pressure and investigate before retrying.
HTTP 403 or 429 Access denied or request rate rejected. Stop or slow down, review the site’s access rules, and do not try to evade a block.
HTTP 5xx Server-side failure or robots.txt is unreachable. Back off and retry sparingly. For robots.txt, RFC 9309 says to assume complete disallow while it is unreachable.
Expected elements are missing Content is rendered in the browser, selector changed, or response is not the expected page. Inspect the returned HTML and status. Update selectors if markup changed; consider browser automation only if rendering is required.
Unexpected redirects or non-HTML response Login page, redirect, anti-bot response, or wrong URL. Check the final response URL, status, and content type. Respect access restrictions rather than attempting to bypass them.
Memory usage grows sharply Large responses, many retained documents, or large parse trees. Limit response size and concurrency; parse incrementally where appropriate and release objects you no longer need.
Robots check disagrees with expectations Different user-agent group, more specific path rule, stale copy, or retrieval failure. Check the actual crawler user-agent, matching path specificity, and robots retrieval status; apply RFC 9309 failure handling.

Frequently asked questions

Should I use an official API instead of scraping?

Yes, when an official API, export, feed, or documented access method provides the fields and frequency you need. It can reduce parsing and maintenance work.

Can I scrape a page just because it loads in my browser?

Browser accessibility alone does not settle permission. Check the site’s rules, technical restrictions, applicable law, and the data’s intended use.

When should I move from a script to Scrapy?

Consider a crawler framework when recurring or larger work needs coordinated requests and crawl management. The threshold depends on your project’s operational needs.

No. It communicates crawler rules. It is not access authorization and does not answer separate legal or contractual questions.

Primary references