Web Scraping Tools Compared: How to Choose the Right One
Compare Scrapy, Playwright, and hosted scraping APIs by page behavior, crawl needs, maintenance, and cost. Use a practical checklist to choose and validate a tool.
Short answer: Start with Scrapy for recurring crawls of mostly static or request-accessible pages and Python data pipelines. Choose Playwright when the data appears only after JavaScript runs or requires browser interactions. Consider a hosted scraping API when you prefer a provider to operate the execution infrastructure. There is no universal winner: test finalists against the same representative URLs, fields, failure cases, and expected volume.
Quick decision table
| Your situation | Good starting point | What to plan for |
|---|---|---|
| Mostly static pages; many URLs; Python-owned extraction and exports | Scrapy | You own the crawler, deployment, target changes, and operational checks. |
| JavaScript-rendered content, scrolling, clicking, or browser state | Playwright | You write navigation, interaction, pagination, and extraction logic; browser execution uses more resources than simple HTTP requests. |
| You want a hosted service to run jobs and return results | Evaluate a hosted scraping API | Check target support, output and job model, limits, regions, data handling, and total billing for your workload. |
| You only need a visual record of a page, rather than structured data | ScreenshotNeo screenshot API | A screenshot captures pixels or a PDF; it does not extract records into a dataset. |
This is a decision aid, not a scored benchmark. The research for this guide did not establish a neutral, shared-workload speed or reliability winner.
What the three scraping tool categories do
Scrapy: a crawler and extraction framework
Scrapy is a Python application framework for crawling websites and extracting structured data. Its documented capabilities include CSS and XPath selectors, feed exports such as JSON, CSV, and XML, and extension points including middleware and pipelines. It also provides crawl-oriented features such as session handling, caching, robots.txt handling, and crawl-depth restriction. These features make it a natural first candidate for a repeatable crawl and data pipeline. See the Scrapy overview and feed export documentation.
Scrapy does not itself make every target accessible or remove the need to design selectors, handle changed markup, authorize access, and operate the job. For pages whose useful content is assembled in a browser, first check whether the underlying page exposes the needed data through ordinary requests; otherwise you may need browser rendering or a different authorized data source.
Playwright: browser automation with page rendering
Playwright controls browser engines and can run page JavaScript and perform actions such as navigation, clicks, and scrolling. That makes it useful when the desired information only appears after a page executes or responds to interaction. It is browser automation, so a multi-page crawl still needs application code for URL discovery, pagination, retries, extraction, and output management. The official introduction describes its browser automation model and language support.
A browser can render a page; it does not automatically turn a collection of pages into a reliable crawler. Build explicit limits and checkpoints so an unexpected navigation loop or a changed page does not produce an unbounded job.
Hosted scraping APIs: provider-operated execution
A hosted scraping API runs jobs on a provider’s service and returns results or datasets through an API. For example, Scrapy.io’s documentation describes discovering tools, running synchronous or asynchronous jobs, retrieving datasets, and scheduling recurring scrapes without hosting the browsers or proxies yourself. That describes a service model, not a guarantee that a given provider supports your target or workflow. Confirm its current capabilities, limits, output format, job lifecycle, data handling, and pricing before committing.
Hosted execution can reduce the infrastructure your team operates, but introduces vendor dependency and service-specific constraints. Compare total cost at your likely volume, including failed or retried work where the provider’s billing rules make that relevant.
Compare the tools against your actual requirements
| Decision axis | Scrapy | Playwright | Hosted API |
|---|---|---|---|
| Core role | Python crawling and structured extraction framework | Browser automation that renders and interacts with pages | Provider-hosted execution returning results or datasets |
| Best initial fit | Static or request-accessible pages and recurring crawl workflows | JavaScript-rendered pages and interaction-heavy flows | Teams that want a hosted service instead of operating execution infrastructure |
| Crawl orchestration | Built-in crawl concepts and export features | Usually requires custom navigation, pagination, and extraction flow | Depends on provider: inspect jobs, schedules, dataset retrieval, and target support |
| Maintenance owner | Your team maintains spiders, deployment, and target-specific logic | Your team maintains browser flows and target-specific logic | Provider operates its service; your team still owns configuration, data validation, and integration |
| Main trade-off | Development and operation remain yours | Browser work and custom crawl logic add resource and maintenance demands | Vendor dependency, service limits, data handling, and billing need evaluation |
The Scrapy and Playwright distinction is about their documented roles and workflow features; it is not evidence that one is faster across sites. Compare using your own pages and required fields rather than vendor claims or a result from an unrelated workload.
How to choose: a practical evaluation
- Write down the output. List the fields, formats, refresh schedule, completeness requirements, and downstream system that will consume the data. If the deliverable is a visual snapshot, use a screenshot or PDF tool instead of a structured-data scraper.
- Inspect representative pages. Include a simple page, a page with JavaScript-rendered content, a paginated or long page, and a known edge case. Determine whether the required content arrives in the initial HTML or only after rendering or interaction.
- Pick two or three candidates. Try Scrapy for request-accessible crawl workflows, Playwright where browser behavior is essential, and a hosted API if operating execution infrastructure is a priority.
- Run the same small trial. Use identical authorized URLs, fields, and expected records. Record completeness, failed pages, retries, maintenance effort, output quality, and cost at a realistic projected volume. Do not infer a general speed or reliability ranking from a small trial.
- Test change and failure handling. Check what happens when a selector disappears, a page is slow, a response is empty, or pagination stops behaving as expected. Decide how the job reports partial output and how a rerun avoids duplicating records.
- Review ownership and data handling. For self-hosted code, estimate development, deployment, monitoring, and maintenance. For a hosted service, review supported regions, retention and processing terms, limits, and billing. Use the approach that meets your privacy and operational requirements.
Runnable starting points
These examples show a small authorized extraction from a page containing an element with the CSS class product-name. Replace the example domain, selector, and output fields with a site you are permitted to access. They are learning starting points, not a complete production crawler.
Scrapy in Python
Install Scrapy with python -m pip install scrapy. Save as product_spider.py, then run scrapy runspider product_spider.py -O products.json. The command exports collected items as JSON.
import scrapy
class ProductSpider(scrapy.Spider):
name = "products"
allowed_domains = ["example.com"]
start_urls = ["https://example.com/catalog"]
def parse(self, response):
for product in response.css(".product"):
yield {
"name": product.css(".product-name::text").get(default="").strip(),
"url": response.urljoin(product.css("a::attr(href)").get(default="")),
}
next_page = response.css("a.next::attr(href)").get()
if next_page:
yield response.follow(next_page, callback=self.parse)
The spider follows a next-page link only when one exists. Inspect the site’s actual markup and add explicit page and item limits when exploring unfamiliar pagination. Scrapy’s feed exports support multiple output formats.
Playwright in Python
Install with python -m pip install playwright and install a browser with python -m playwright install chromium. Save as capture_page.py and run python capture_page.py. This example waits for a selector and reads rendered text.
import asyncio
from playwright.async_api import async_playwright
async def main():
async with async_playwright() as p:
browser = await p.chromium.launch()
page = await browser.new_page()
await page.goto("https://example.com/catalog", wait_until="domcontentloaded")
await page.locator(".product-name").first.wait_for()
names = await page.locator(".product-name").all_text_contents()
print(names)
await browser.close()
asyncio.run(main())
For a page that loads content after an action, perform the authorized interaction explicitly and wait for a meaningful selector or state. Avoid arbitrary long sleeps where a selector or other specific readiness condition can be used. Add pagination and bounded URL discovery in your own crawler logic.
Hosted API request shape
Hosted providers use different endpoints and parameters, so there is no single runnable request that works across them. Follow the chosen provider’s current documentation for authentication, target URL, extraction instructions, job mode, and result retrieval. Before sending production data, confirm its request and response schemas and test with a small authorized job.
When the job is a screenshot rather than data extraction
A scraper returns data fields; a screenshot API returns a visual image or PDF of a page. If you need a shareable or archived rendering rather than records, ScreenshotNeo is the first alternative to try: it removes cookie and consent banners, newsletter popups, and chat widgets before capture, bills only clean shots, and has a $5 paid plan for 3,000 shots.
Or skip the browser setup
If your deliverable is a screenshot or PDF, ScreenshotNeo takes a page URL in one request. It is a screenshot API and MCP server from Yorker Media. See the ScreenshotNeo API documentation for options and current request details.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));
- Cookie banners, popups, and chat widgets are removed before the shot; each cleaning step can be turned off.
- Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing; response headers identify the page verdict and billing status.
- An MCP server offers
take_screenshot,get_page_info, andcapture_pdfto Claude, Cursor, and other MCP clients. - The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000. Every feature is on every plan.
Sign up for ScreenshotNeo’s free plan: 1,000 screenshots a month, no card required.
Configuration and edge cases to plan for
- Rendering readiness: A successful navigation does not always mean client-rendered content is ready. Wait for a selector or state tied to the data you need.
- Pagination: Confirm the next-page condition and stop condition. Set maximum pages and items for exploratory runs; guard against repeated URLs and loops.
- Dynamic or lazy content: Check whether scrolling or interaction is required. Browser automation can perform those actions, but your code must decide what to do and when the page is ready.
- Sessions and cookies: Use only authorized session state. Avoid logging secrets or placing credentials in source control. Scrapy documents cookies and session support; browser tools also expose browser-context state.
- Changed markup: Validate required fields and flag missing values. A selector returning an empty string should not silently count as a successful complete record.
- Large crawls: Bound concurrency, retries, depth, and total URLs. Save progress in a restartable format and make downstream writes idempotent.
- Output consistency: Normalize whitespace and URLs, define how missing values are represented, and deduplicate according to a stable key.
- Service limits: For a hosted API, check request size, job duration, concurrency, dataset retention, and rate or volume limits in the current provider documentation.
Permission, robots.txt, and responsible access
A tool choice does not grant permission to access or reuse a site’s data. Review applicable law, the site’s terms, privacy obligations, access authorization, and published crawler rules for your use case. RFC 9309, the IETF Robots Exclusion Protocol standard, explains that robots.txt rules “are not a form of access authorization.” A robots.txt allow rule is not legal permission, and a disallow rule by itself does not determine every legal question. Read RFC 9309 and use official APIs or licensed data where appropriate. Do not treat CAPTCHAs, blocks, or rate limits as an invitation to bypass a site’s controls.
Performance, reliability, and cost
Performance depends on target behavior, page weight, network conditions, crawl concurrency, browser startup and rendering, retries, and output work. This guide has no independent benchmark across a shared set of URLs, so do not assume a universal speed winner. A request-based crawl avoids the cost of rendering a browser when the needed content is already available in responses; browser execution is justified when rendering or interaction is required. Measure end-to-end completion and data quality on your workload.
Reliability comes from explicit readiness checks, bounded retries, timeouts, page and item limits, duplicate handling, validation, and resumable output. For hosted APIs, include job failures, retries, storage or retrieval limits, and service-specific billing in your estimate. For self-managed tools, include engineering time, compute, monitoring, and ongoing repairs when a target changes. Compare total cost per valid record or completed job, not only the visible request price.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Scrapy returns no target text | The selector does not match the returned HTML, or content appears only after browser JavaScript. | Inspect the response and selector against actual markup. If the content is browser-rendered, evaluate Playwright or an authorized data endpoint. |
| Playwright times out waiting for content | The selector is wrong, the page failed to load the state, or the wait condition does not match the site. | Inspect the page and console/network errors, use a selector tied to the needed content, and set a bounded timeout appropriate to the workflow. |
| Only the first page is collected | Pagination discovery or the stop condition is missing or incorrect. | Inspect the next-page link or cursor, follow it explicitly, track visited URLs, and set a maximum page count. |
| Some records have empty fields | Markup differs between pages, content has not loaded, or a selector is too broad or narrow. | Validate required fields, handle known page variants, and report incomplete records rather than silently accepting them. |
| A crawl repeats pages or grows unexpectedly | Pagination loops, URL variants, or unbounded link discovery. | Normalize URLs, track visited pages, constrain allowed domains and crawl depth, and cap page and item counts. |
| Hosted job is rejected or produces an unexpected result | Unsupported target or option, invalid request shape, expired credentials, or provider limits. | Check the provider’s current schema, authentication, supported targets, limits, and job status; reproduce with one small authorized URL. |
| Access is blocked or a CAPTCHA appears | The site is restricting automated access or the request pattern is disallowed. | Stop and review authorization and site rules; reduce load or use an official API or licensed source where suitable. Do not bypass controls. |
Frequently asked questions
Which web scraping tool should I use?
Use the page behavior and operational model to decide: Scrapy for crawl-oriented Python extraction, Playwright for browser-dependent rendering and interaction, or a hosted API if provider-operated execution fits your requirements.
Do I need a browser to scrape a JavaScript website?
Only if the data you need is unavailable through an authorized request or data endpoint and appears after browser execution or interaction. Inspect the response and page behavior before adding a browser.
Is Playwright a replacement for Scrapy?
They have different roles. Playwright automates browsers; Scrapy supplies crawler and extraction workflow features. A project may use one or both, depending on target behavior and architecture.
Should I use a scraping API or build my own scraper?
Choose based on the infrastructure you want to operate, target support, output needs, data handling, limits, and total cost at realistic volume. Validate a provider on representative pages before relying on it.
Where can I learn more about Python scraping?
Web Scraping with Python, 3rd Edition by Ryan Mitchell is an optional intermediate-to-advanced O’Reilly book published in February 2024; its catalog listing describes 352 pages and includes chapters on scraper construction and legal and ethical questions. See the publisher catalog.
