ScreenshotNeo

BlogHow-to

How to Use Headless Browsers with Scrapy

Build a Scrapy crawler that renders JavaScript with Playwright only when needed, while keeping Scrapy scheduling, parsing, and resource controls.

By the ScreenshotNeo team29 September 20268 min read

How to Use Headless Browsers with Scrapy

Use scrapy-playwright when a Scrapy request needs JavaScript execution, browser events, or interaction. It lets selected requests run in Playwright while preserving Scrapy’s scheduler, duplicate filtering, middleware, callbacks, and item pipeline. Keep ordinary requests as normal Scrapy downloads, and opt only JavaScript-dependent URLs into a browser.

Scrapy’s own guidance says reproducing the underlying data request is preferred when practical because it transfers less data and avoids browser overhead. When that is difficult, or when the result exists only after browser execution, Scrapy recommends scrapy-playwright for integration. Read Scrapy’s dynamic-content guidance before choosing a browser.

What a headless browser adds to Scrapy

A headless browser is a browser controlled by an automation API without a visible window. Playwright is the automation library; scrapy-playwright is the Scrapy download handler that connects Playwright to Scrapy’s normal request and response workflow.

scrapy-playwright renders selected requests while Scrapy keeps its normal parsing pipeline.
scrapy-playwright renders selected requests while Scrapy keeps its normal parsing pipeline.

A regular Scrapy downloader receives the server’s initial HTML. A browser can execute JavaScript, wait for network activity, click controls, fill forms, and expose the resulting DOM. The response delivered to your callback can then be parsed with the same CSS and XPath selectors you already use.

Requirement Best first choice Reason
Data is in initial HTML Normal Scrapy request Lowest CPU, memory, and transfer overhead
Data comes from a reproducible JSON, GraphQL, or API request Call that endpoint directly Structured data is easier to parse and usually faster
Content appears after JavaScript or browser events scrapy-playwright Executes the page in a real browser while retaining Scrapy’s workflow
You need a screenshot or PDF Playwright or a screenshot API The browser artifact is the output

Install compatible versions

The current scrapy-playwright documentation lists these minimum versions:

  • Python 3.10 or newer
  • Scrapy 2.7 or newer
  • Playwright 1.40 or newer

Create an isolated environment, install the integration, and download the browser engines:

python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
pip install scrapy-playwright
playwright install

You can install a subset of engines when you do not need all of them:

playwright install chromium firefox

Playwright can also drive an existing branded Google Chrome or Microsoft Edge installation. Those branded browsers are not installed by the normal playwright install command.

Configure Scrapy to use Playwright

Playwright is asyncio-based, so select Scrapy’s asyncio reactor and register the Playwright download handler. In settings.py:

TWISTED_REACTOR = "twisted.internet.asyncioreactor.AsyncioSelectorReactor"

DOWNLOAD_HANDLERS = {
    "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
    "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
}

# Browser engine used by new contexts.
PLAYWRIGHT_BROWSER_TYPE = "chromium"

# Keep this below the memory available to your worker.
PLAYWRIGHT_MAX_PAGES_PER_CONTEXT = 8

# Optional defaults for every browser context.
PLAYWRIGHT_CONTEXTS = {
    "default": {
        "viewport": {"width": 1440, "height": 900},
        "java_script_enabled": True,
    },
}

# Normal Scrapy concurrency still applies to non-browser requests.
CONCURRENT_REQUESTS = 16

The handler is global, but requests are not rendered automatically. Add meta={"playwright": True} only to requests that need a browser.

Build a complete JavaScript-rendered spider

The following spider requests a catalog in Chromium, waits for product cards, and parses the rendered HTML with ordinary Scrapy selectors. Scrapy 2.13 introduced async def start(); older projects can use start_requests() instead.

import scrapy


class ProductSpider(scrapy.Spider):
    name = "products"
    allowed_domains = ["example.com"]

    async def start(self):
        yield scrapy.Request(
            "https://example.com/catalog",
            meta={
                "playwright": True,
                "playwright_page_methods": [
                    {"method": "wait_for_selector", "args": ["article.product"]},
                ],
            },
            errback=self.errback,
        )

    def parse(self, response):
        for card in response.css("article.product"):
            yield {
                "name": card.css("h2::text").get(),
                "url": response.urljoin(card.css("a::attr(href)").get()),
                "price": card.css(".price::text").get(),
            }

    async def errback(self, failure):
        request = failure.request
        self.logger.error("Request failed: %s", request.url)
        # Do not retain a Page object here. The integration can close it.
        yield {"url": request.url, "error": repr(failure.value)}

For projects that need direct page operations, pass a page callback and close the page in every path. For simple waits and clicks, use Playwright page methods through request metadata so the integration manages the lifecycle.

Wait for the right condition

Waiting for a fixed delay is easy but fragile. Prefer a condition that represents the content you will parse:

  • Selector: wait for the card, table, or status element that proves rendering finished.
  • Network idle: useful for applications that finish loading through a burst of requests, but risky for pages with analytics or long polling.
  • URL or navigation: use after a click that changes routes.
  • Short delay: a last resort for animation-driven content with no reliable selector.
meta = {
    "playwright": True,
    "playwright_page_methods": [
        {"method": "wait_for_selector", "args": ["div.results[data-ready='true']"]},
        {"method": "click", "args": ["button.load-more"]},
        {"method": "wait_for_selector", "args": ["button.load-more[disabled]"]},
    ],
}

Make selectors specific enough to describe completion, not merely page structure. A selector for an empty container can return before data arrives.

Browser contexts, profiles, and concurrency

A browser context is an isolated session with its own cookies, storage, permissions, and pages. Use a named context when a group of requests must share login state or locale. Use separate contexts when sessions must not contaminate one another.

yield scrapy.Request(
    url,
    meta={
        "playwright": True,
        "playwright_context": "logged_in",
    },
)

Context configuration belongs in PLAYWRIGHT_CONTEXTS. Persistent contexts can retain a browser profile between runs, but they also retain cookies and local storage, so treat the profile directory as stateful data. Set a browser type of chromium, firefox, or webkit according to the site you must support.

PLAYWRIGHT_MAX_PAGES_PER_CONTEXT is a hard resource boundary. Every open page counts, including pages left open after an exception. If the limit is reached, requests can wait indefinitely and the crawl can appear frozen. Keep pages short-lived, add an errback when you retain page objects, and close pages deterministically.

Use the browser only where it is necessary

A practical architecture has two request paths:

  1. Fetch listing pages and detail pages with ordinary Scrapy requests whenever their HTML or API response contains the required fields.
  2. Mark only JavaScript-dependent URLs with playwright=True.
  3. Use browser interaction to reveal the data, then return to normal Scrapy parsing and item processing.

Inspect your browser’s network panel before adding Playwright. If a page fetches a JSON endpoint containing the complete product data, reproduce that request with Scrapy and parse JSON. This usually reduces startup time, memory use, and transferred resources. Browser rendering is justified when the endpoint is difficult to reproduce, protected by browser state, or the required result depends on DOM execution and interaction.

Remote Chromium and connection settings

The integration can connect to a remote Chromium instance with PLAYWRIGHT_CDP_URL. In CDP mode, the browser type must remain Chromium and local launch options are ignored. CDP cannot be combined with PLAYWRIGHT_CONNECT_URL. Use a remote browser when browser processes belong on a separate worker, but keep connection failures and browser capacity in your crawler’s retry and monitoring design.

Common errors and fixes

Symptom Likely cause Fix
ModuleNotFoundError: scrapy_playwright The package is missing from the active environment. Activate the intended virtual environment and run pip install scrapy-playwright.
Browser executable is missing Playwright’s browser binaries were not downloaded. Run playwright install, or install only the engine selected in settings.
Reactor mismatch error Scrapy started with a non-asyncio reactor. Set TWISTED_REACTOR before the crawler starts and remove code that installs another reactor.
Callback sees an empty shell The page was not opted into Playwright, or parsing ran before rendering. Add meta={"playwright": True} and wait for a content selector.
Timeout while waiting The selector never appears, the page is slow, or the selector is wrong. Verify the selector in a browser, wait for a stable condition, and set a suitable timeout.
Crawl freezes after failures Pages remain open and consume the per-context limit. Add an errback, avoid retaining page objects, and close pages in finally blocks.
Login state leaks between users Requests share a browser context. Assign separate named contexts or clear storage between sessions.
Remote connection options do nothing CDP mode ignores local launch options. Configure the remote Chromium service itself and keep the browser type Chromium.
A hosted screenshot service can handle browser cleanup before producing the final capture.
A hosted screenshot service can handle browser cleanup before producing the final capture.

Performance, reliability, and cost planning

Launching and driving a browser costs more CPU and memory than an HTTP request. Set concurrency from observed memory headroom, then increase it gradually. A lower page limit with stable throughput is better than opening pages until the host swaps or the crawl stalls.

  • Reuse browser contexts when the same session can safely share cookies.
  • Keep pages focused: navigate, wait for the required state, extract, and close.
  • Block unnecessary resources only when doing so does not remove the data you need.
  • Use Scrapy’s retry and timeout settings for transient navigation failures.
  • Record the URL, wait condition, browser engine, and failure type so a retry is diagnosable.
  • Test selectors against layout changes; a successful HTTP status does not prove that the desired data rendered.

There is no universal browser throughput number. Capacity depends on the target site, page weight, JavaScript workload, browser engine, context count, and machine resources. Treat the compatibility versions above as requirements, not performance guarantees.

Or skip the browser setup

If your goal is a screenshot or PDF rather than extracted items, ScreenshotNeo provides a hosted website screenshot API. One GET request returns PNG, JPEG, WebP, or PDF. It accepts 63 options, including full-page capture with lazy images, CSS-element capture, device presets, dark mode, custom CSS and JavaScript, clicks, selector or network-idle waits, request blocking, headers, cookies, user agents, timezone, geolocation, resizing, caching, signed links, asynchronous jobs, webhooks, bulk capture, and PDF page settings.

See the ScreenshotNeo API documentation for the complete parameter list.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const image = Buffer.from(await res.arrayBuffer());

ScreenshotNeo accepts and removes cookie or consent banners, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000.

Create a free ScreenshotNeo account and start with 1,000 screenshots per month at no charge.

FAQ

Can I use Playwright directly inside a spider?

Yes, but direct playwright-python calls bypass much of Scrapy’s scheduling, duplicate filtering, and middleware. Use scrapy-playwright when you want browser rendering inside the normal Scrapy workflow.

Does every Scrapy request open a browser?

No. Only requests with meta={"playwright": True} are handled by Playwright when the integration is configured.

Which browser should I install?

Chromium is the usual starting point. Install Firefox or WebKit when the target site’s behavior or your compatibility requirements call for it.

Why does a browser request return status 200 but no data?

Status 200 only confirms that navigation succeeded. Wait for the selector or browser state that proves the application finished rendering, then parse the response.

How do I prevent a long crawl from exhausting memory?

Limit pages per context, avoid retaining page objects, close pages after extra operations, and render only URLs that truly require JavaScript.