ScreenshotNeo

BlogHow-to

How to Capture Webpage Screenshots with Scrapy

Capture rendered webpages in Scrapy with scrapy-playwright, including full-page, lazy-loaded and infinite-scroll screenshots.

By the ScreenshotNeo team1 October 20268 min read

Use Scrapy with scrapy-playwright when you need a screenshot of the rendered page. Scrapy fetches HTML efficiently, while Playwright runs a real browser so JavaScript, layout and client-rendered content appear in the image. Scrapy’s dynamic-content guide recommends this integration because driving Playwright separately can bypass Scrapy components such as middleware and duplicate filtering.

The two useful patterns are:

  • Schedule PageMethod("screenshot", ...) in request metadata and read the resulting bytes from PageMethod.result.
  • Expose the Playwright Page object to the callback with playwright_include_page=True, then call page.screenshot() yourself.

What you need

  1. Python 3.8 or newer.
  2. A Scrapy project.
  3. The scrapy-playwright integration and its browser binaries.
python -m pip install scrapy scrapy-playwright
playwright install

Enable the Playwright download handler in your Scrapy settings. The exact settings can be placed in settings.py:

DOWNLOAD_HANDLERS = {
    "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
    "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
}

TWISTED_REACTOR = "twisted.internet.asyncioreactor.AsyncioSelectorReactor"

The integration’s official README documents these settings and the request metadata used below.

Method 1: schedule a screenshot with PageMethod

Use this method when the screenshot is a fixed page action that can run before Scrapy invokes your callback. The screenshot bytes are available from the PageMethod.result value.

import scrapy
from scrapy_playwright.page import PageMethod


class ScreenshotSpider(scrapy.Spider):
    name = "screenshots"

    async def start(self):
        yield scrapy.Request(
            "https://example.org",
            meta={
                "playwright": True,
                "playwright_page_methods": [
                    PageMethod(
                        "screenshot",
                        path="example.png",
                        full_page=True,
                    ),
                ],
            },
        )

    def parse(self, response):
        screenshot_method = response.meta["playwright_page_methods"][0]
        screenshot_bytes = screenshot_method.result

        # The path above already saved the file. This also shows how to
        # write the returned bytes yourself when you omit path=.
        with open("example-from-bytes.png", "wb") as output:
            output.write(screenshot_bytes)

        yield {
            "url": response.url,
            "bytes": len(screenshot_bytes),
        }

full_page=True captures the complete document height. Without it, Playwright captures the current viewport. The scrapy-playwright documentation shows that method arguments and keyword arguments are passed to the Playwright page method.

Capture a viewport instead of the full document

PageMethod(
    "screenshot",
    path="viewport.png",
    full_page=False,
)

Viewport dimensions come from the browser context. Set them per request when you need a desktop or mobile layout:

yield scrapy.Request(
    "https://example.org",
    meta={
        "playwright": True,
        "playwright_context_kwargs": {
            "viewport": {"width": 1440, "height": 900},
        },
        "playwright_page_methods": [
            PageMethod("screenshot", path="desktop.png"),
        ],
    },
)

Method 2: take the screenshot in the callback

Use callback access when the capture depends on response data, a selector, scrolling, conditional waits or several screenshots. Set playwright_include_page=True, retrieve the page from response.meta, and close it when finished.

import scrapy


class CallbackScreenshotSpider(scrapy.Spider):
    name = "callback-screenshots"

    async def start(self):
        yield scrapy.Request(
            "https://example.org",
            meta={
                "playwright": True,
                "playwright_include_page": True,
            },
        )

    async def parse(self, response):
        page = response.meta["playwright_page"]
        try:
            image_bytes = await page.screenshot(
                path="callback.png",
                full_page=True,
            )
            yield {
                "url": response.url,
                "screenshot_bytes": image_bytes,
            }
        finally:
            await page.close()

Included pages stay open until you close them. Forgetting this can exhaust browser contexts or file descriptors during a crawl. When you do not include the page, the integration closes it after processing.

Wait for the page to become capture-ready

A screenshot taken immediately after navigation can miss content rendered by JavaScript. Prefer a page-specific condition over a large fixed sleep.

Wait for a selector

from scrapy_playwright.page import PageMethod


meta = {
    "playwright": True,
    "playwright_page_methods": [
        PageMethod("wait_for_selector", "main article"),
        PageMethod("screenshot", path="article.png", full_page=True),
    ],
}

Wait for a short delay

PageMethod("wait_for_timeout", 1500),
PageMethod("screenshot", path="delayed.png", full_page=True),

A delay is useful for a known animation or deferred widget, but it is less reliable across network conditions than waiting for a selector.

Wait for network idle in a callback

async def parse(self, response):
    page = response.meta["playwright_page"]
    try:
        await page.wait_for_load_state("networkidle")
        image_bytes = await page.screenshot(path="idle.png", full_page=True)
        yield {"url": response.url, "bytes": len(image_bytes)}
    finally:
        await page.close()

Network idle can remain pending on pages with analytics, polling or streaming connections. Use a selector or a bounded timeout when the page never becomes idle.

Full-page screenshots with lazy-loaded content

full_page=True expands the image to the document’s current height. It does not guarantee that images or sections which load only after scrolling have appeared. Scroll, wait for a page-specific marker, then capture.

import scrapy


class LazyPageSpider(scrapy.Spider):
    name = "lazy-pages"

    async def start(self):
        yield scrapy.Request(
            "https://example.org/catalog",
            meta={
                "playwright": True,
                "playwright_include_page": True,
            },
        )

    async def parse(self, response):
        page = response.meta["playwright_page"]
        try:
            await page.wait_for_selector(".product-card")

            # Scroll in steps so intersection observers and lazy images run.
            previous_height = 0
            for _ in range(20):
                height = await page.evaluate("document.body.scrollHeight")
                if height == previous_height:
                    break
                previous_height = height
                await page.evaluate("window.scrollTo(0, document.body.scrollHeight)")
                await page.wait_for_timeout(400)

            await page.wait_for_selector(".catalog-end")
            image_bytes = await page.screenshot(
                path="catalog-full.png",
                full_page=True,
            )
            yield {"url": response.url, "bytes": len(image_bytes)}
        finally:
            await page.close()

Adapt the selector, number of scroll steps and wait condition to the target site. An infinite-scroll page may never have a stable final height, so define an explicit stopping rule such as a final marker, a known item count or a maximum number of scrolls.

Capture one element

Playwright can capture an element instead of the entire viewport. This is useful for a chart, product card or article body.

async def parse(self, response):
    page = response.meta["playwright_page"]
    try:
        await page.locator("main article").screenshot(path="article-only.png")
        yield {"url": response.url}
    finally:
        await page.close()

Wait for the element and any data it contains before calling locator.screenshot(). If the element is outside the current viewport, Playwright scrolls it into view.

Control browser context and output

Need Setting or method Why it matters
Desktop or mobile layout playwright_context_kwargs.viewport Controls responsive breakpoints.
Full document full_page=True Captures beyond the viewport.
Element only locator(selector).screenshot() Limits the image to one component.
Stable timing wait_for_selector() Waits for a meaningful page condition.
Raw image bytes Return value of screenshot() Lets pipelines upload or hash the image without rereading a file.

Playwright’s Python screenshot guide covers the path and full-page options: playwright.dev/python/docs/screenshots.

Managing many URLs

For a crawl, yield one request per URL and keep browser work bounded. Each included page consumes resources until it closes. Use a maximum concurrency appropriate for the target site and your machine, and always close pages in a finally block.

CONCURRENT_REQUESTS = 4
PLAYWRIGHT_MAX_CONTEXTS = 4
PLAYWRIGHT_MAX_PAGES_PER_CONTEXT = 4

These limits are operational safeguards rather than universal performance values. Raise them only after observing memory, CPU, navigation time and target-site responses.

Common errors and fixes

Error or symptom Likely cause Fix
ModuleNotFoundError: scrapy_playwright The integration is not installed in the active environment. Run python -m pip install scrapy-playwright in the same environment used to start Scrapy.
Browser executable is missing Playwright browser binaries were not installed. Run playwright install; install only the browser required by your deployment if your environment supports that choice.
Request returns HTML but no screenshot meta["playwright"] is missing or download handlers are not configured. Enable both HTTP and HTTPS Playwright handlers and set "playwright": True on the request.
Screenshot is blank or incomplete JavaScript content has not rendered, or the page is still loading lazy content. Wait for a selector, scroll to trigger lazy loading, then capture.
Callback hangs on networkidle Analytics, polling or streaming requests never stop. Replace it with a selector wait or use a bounded timeout.
Memory grows during a crawl Included pages were not closed. Close every page in finally; reduce concurrency if needed.
Mobile layout is not captured The browser context still uses a desktop viewport. Set playwright_context_kwargs with the desired viewport before navigation.
Full-page image stops too early Content appears only after scrolling or a later API response. Scroll in steps, wait for a final marker or item count, and then call screenshot(full_page=True).

Reliability, performance and cost considerations

  • Reliability: Use deterministic selectors and explicit stopping conditions. Fixed sleeps alone are sensitive to network and server timing.
  • Performance: Browser rendering costs substantially more CPU and memory than ordinary HTTP crawling. Reuse contexts where appropriate, limit concurrency, and avoid opening a page when HTML extraction is enough.
  • Output size: Full-page images can become very tall and large. Capture an element or viewport when a complete document is unnecessary.
  • Retries: Retry navigation failures selectively. Repeating a page with a broken selector will not fix the selector and can overload the target.
  • Compliance: Respect the target site’s access rules, authentication requirements and rate limits.
  • Cost: Self-hosting has infrastructure and maintenance costs. A hosted screenshot API can be simpler when you need screenshots without operating browsers.

Or skip the browser setup

ScreenshotNeo provides a hosted screenshot endpoint when you want one request instead of managing Scrapy, Playwright, browser binaries and page cleanup. It removes cookie and consent banners, newsletter popups and chat widgets before capture. Bot checks, blank pages, failed loads, timeouts and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server lets Claude, Cursor and other MCP clients call take_screenshot, get_page_info and capture_pdf.

See the ScreenshotNeo API documentation for request options. The basic call returns a WebP image:

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

ScreenshotNeo supports full-page capture, CSS element capture, dark mode, device presets, custom viewports, retina scale, custom CSS and JavaScript, selector waits, delays, network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, caching, signed links, asynchronous jobs, webhooks, bulk capture and a usage API. The free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000.

Create a free ScreenshotNeo account and start with 1,000 screenshots per month at no charge.

FAQ

Can Scrapy take a screenshot without Playwright?

Scrapy alone downloads responses and parses them. A rendered screenshot requires a browser integration such as scrapy-playwright.

Should I use PageMethod or include the page?

Use PageMethod for a fixed sequence of actions. Include the page when callback logic must decide what to wait for, scroll, capture or return.

Why is my full-page screenshot missing images?

Those images may be lazy-loaded. Scroll the page, wait for the relevant selector or final marker, and capture after the content is present.

Do I need to close a page after every screenshot?

Only pages exposed with playwright_include_page=True require your code to close them. Close included pages even when capture raises an exception.

When is a hosted API a better fit?

Use one when you want screenshots without installing browsers, tuning crawl concurrency or maintaining page lifecycle code. ScreenshotNeo is the hosted option described above.