ScreenshotNeo

BlogHow-to

Take screenshots of lazy-loaded pages with Playwright in Python

Scroll to trigger deferred content, wait for the right signals, and capture a complete page with Playwright in Python.

By the ScreenshotNeo team4 October 20269 min read

To screenshot a lazy-loaded page with Playwright in Python, first scroll through the page to trigger content that loads near the viewport, wait for page-specific signals that the required content has appeared, then capture with page.screenshot(path="page.png", full_page=True). The full_page option sets the output to the full scrollable page; it does not itself trigger scrolling or guarantee that deferred images and content have loaded. The scroll-and-verify workflow below is a practical approach, and its wait conditions must match the site you are capturing. See Playwright’s screenshot documentation and locator documentation.

1. Install Playwright

Install the Python package and its browser binaries. This example uses Chromium:

python -m pip install playwright
python -m playwright install chromium

Save the complete script below as capture_lazy_page.py and run it with python capture_lazy_page.py. Change the URL and the page-specific selectors to match your target.

2. Scroll, wait for content, and capture

import asyncio
from pathlib import Path
from playwright.async_api import async_playwright

URL = "https://example.com"
OUTPUT = Path("page.png")

async def main():
    async with async_playwright() as p:
        browser = await p.chromium.launch()
        page = await browser.new_page(viewport={"width": 1440, "height": 900}, device_scale_factor=1)
        response = await page.goto(URL, wait_until="domcontentloaded", timeout=60_000)
        if response is not None and not response.ok:
            raise RuntimeError(f"Navigation returned HTTP {response.status}: {URL}")

        # Replace this with a selector that appears when the page's main content is ready.
        await page.locator("main").wait_for(state="visible", timeout=15_000)

        # Visit successive viewport regions to trigger viewport-dependent loading.
        # The iteration cap prevents an unbounded page from scrolling forever.
        previous_height = 0
        stable_passes = 0
        max_passes = 100
        for _ in range(max_passes):
            height = await page.evaluate("document.documentElement.scrollHeight")
            step = max(500, await page.evaluate("window.innerHeight"))
            for y in range(0, height, step):
                await page.evaluate("y => window.scrollTo(0, y)", y)
                # Replace this fallback with a known per-section signal where possible.
                await page.wait_for_timeout(200)

            await page.wait_for_timeout(300)
            new_height = await page.evaluate("document.documentElement.scrollHeight")
            if new_height == previous_height:
                stable_passes += 1
                if stable_passes >= 2:
                    break
            else:
                stable_passes = 0
            previous_height = new_height

        # Optional: wait for images that are already in the DOM to finish loading.
        # This cannot force a site to create images that have not yet been requested.
        await page.evaluate("""async () => {
            const images = Array.from(document.images);
            await Promise.all(images.map(img => {
                if (img.complete) return Promise.resolve();
                return new Promise(resolve => {
                    img.addEventListener('load', resolve, { once: true });
                    img.addEventListener('error', resolve, { once: true });
                });
            }));
        }""")

        await page.evaluate("window.scrollTo(0, 0)")
        await page.screenshot(path=str(OUTPUT), full_page=True, animations="disabled")
        await browser.close()
        print(f"Saved {OUTPUT}")

asyncio.run(main())

The 200 ms pause is only a fallback to give scroll-triggered work a chance to start; it is not a universal readiness guarantee. Prefer an explicit locator or image condition when you know what indicates completion. The script is an illustrative pattern, not a claim that it has been run against the example URL or will work unchanged on every site.

Use a meaningful readiness signal

For a page that appends cards while scrolling, wait for an expected card count, a “loaded” marker, or a known final item. For images, inspect the specific image after it is expected to exist:

image = page.locator("img.product-photo").last
await image.wait_for(state="attached", timeout=10_000)
await image.evaluate("img => img.decode()")

decode() waits for decoding of that image element, but it does not guarantee that every image elsewhere on the page has loaded. A broken image can also reject decoding; decide whether that should fail your capture or be reported as a missing asset.

Playwright locators auto-wait for relevant actions and assertions, but locator.all() returns immediately and can give unpredictable results while a dynamic list is changing. Wait for a page-specific completion condition before enumerating a growing list.

3. Pick the right screenshot method

Need Python Behavior
Visible viewport await page.screenshot(path="viewport.png") Captures the current viewport.
Whole document await page.screenshot(path="full.png", full_page=True) Captures the full scrollable page bounds; deferred content still needs to be triggered and checked.
One element await page.locator(".report").screenshot(path="report.png") Scrolls the element into view and captures it. A scrollable element shows only its currently visible scrolled content.
In-memory bytes data = await page.screenshot(full_page=True) Returns image bytes for processing or upload instead of requiring a path.

The screenshot API also supports a clip rectangle, type (png or jpeg), quality for JPEG, omit_background, and animations. Quality applies to JPEG, not PNG. Use a clip when you need a fixed region; use an element locator screenshot when you need a particular component. The full-page option is appropriate for the whole document, not for exposing all content inside a nested scrolling panel.

4. Handle common lazy-loading patterns

Images loaded by scrolling

Scroll each relevant region into view, then wait for the expected image elements to be present and decoded. Some sites use data-src or a background image and only set the final URL after intersection with the viewport; checking for a non-empty src before triggering that intersection can wait forever. Trigger the load first, then inspect the site’s actual image state.

Infinite scrolling

An infinite feed has no natural end. Define a stopping rule before capture: a maximum number of scrolls, a target item count, a known end marker, or a maximum elapsed time. After each scroll, wait for the next item or for the item count to increase. Stop and report partial capture if the limit is reached without the desired completion signal. A loop based only on page height may miss content appended inside a fixed-height nested container.

Nested scroll containers

If the page uses a scrollable panel, scroll that element rather than window. For example:

panel = page.locator(".results-panel")
await panel.evaluate("el => el.scrollTo(0, el.scrollHeight)")

Repeat while checking the panel’s scrollHeight and content state. A locator screenshot of the panel captures its current visible portion, not every off-screen row in that panel. If all rows are needed, load them, adjust the capture strategy, or capture sections separately.

Network activity and delayed requests

wait_until="networkidle" can be useful for a page that becomes quiet after loading, but analytics, polling, and long-lived connections can prevent network idle. Conversely, network idle does not prove that a lazy section you have not scrolled to has loaded. Use a selector or application-specific state as the decisive condition when available.

Sticky headers and animations

Full-page captures can show sticky elements in ways that differ from a normal single viewport, and animated content may vary between runs. Disable finite animations with animations="disabled" for more repeatable screenshots. For sticky headers that obscure content or repeat awkwardly, use a capture-specific stylesheet or hide the header, while noting that the result no longer represents the unmodified page.

5. Synchronous Python, cURL, and JavaScript alternatives

The async API is useful in async applications. The synchronous API is simpler for a standalone script:

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch()
    page = browser.new_page(viewport={"width": 1440, "height": 900})
    page.goto("https://example.com", wait_until="domcontentloaded")
    page.locator("main").wait_for(state="visible")
    page.evaluate("""async () => {
        const step = Math.max(window.innerHeight, 500);
        for (let y = 0; y < document.documentElement.scrollHeight; y += step) {
            window.scrollTo(0, y);
            await new Promise(resolve => setTimeout(resolve, 200));
        }
        window.scrollTo(0, 0);
    }""")
    page.screenshot(path="page.png", full_page=True)
    browser.close()

Playwright’s official screenshot examples also document async and sync Python forms. A cURL request cannot run Playwright’s browser-side scrolling procedure by itself. If your task is simply to request a screenshot from a screenshot API, ScreenshotNeo offers a single GET request; it is a different capture workflow from running your own Playwright browser:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

Equivalent Python request:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Equivalent Node.js request:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

See the ScreenshotNeo API documentation for request options. ScreenshotNeo supports full-page capture with lazy images loaded, plus element capture and other screenshot controls.

6. Troubleshooting

Symptom Likely cause What to change
Images are blank or placeholders The capture happened before scrolling triggered the image request, or before decode completed. Scroll the image into view, wait for its real source or decode, then capture.
The bottom of the page is missing Content expanded after the loop measured height, or the page is infinite. Recheck height after each pass, add a bounded repeat, and use an explicit end condition.
The script waits forever A selector never appears, an image failed, or the page keeps network activity open. Use finite timeouts, log the failed condition, and wait for a page-specific signal instead of unconditional network idle.
Some cards are absent A dynamic list was enumerated before it stabilized. Wait for a target count, next-item signal, or end marker before reading the list.
Only part of a panel is captured It is an independently scrollable element; page full-page bounds do not expand the panel’s internal scroll area. Scroll/load the panel and capture it in sections or change the page layout for capture.
Browser launch fails The Playwright package is installed but the matching browser binary is missing, or the runtime lacks required system dependencies. Run python -m playwright install chromium; install required OS dependencies in the environment as directed by Playwright’s installation documentation.
Capture is inconsistent between runs Animations, ads, personalized content, or timing-dependent layout changes. Disable animations, control viewport and locale if relevant, wait on stable content, and avoid relying on arbitrary sleeps alone.

7. Performance, reliability, and cost

Each scroll step, wait, and full-page raster capture adds time and memory use. Keep the viewport and device scale factor appropriate to the required output, set navigation and locator timeouts, and cap loops for feeds that can grow without limit. Large pages may produce large image buffers; write directly to a file when you do not need the bytes in memory. For repeated captures, control inputs such as viewport, cookies, and timing conditions to make outputs more comparable.

Self-hosted Playwright has no per-screenshot API fee, but you operate the browser process and its runtime. Runtime cost depends on your compute, concurrency, and how much of the page you load; no universal benchmark applies. ScreenshotNeo’s plans are Free: 1,000 shots/month with no card; Starter: $5 for 3,000; Growth: $15 for 15,000; Pro: $39 for 60,000; Scale: $99 for 250,000; Business: $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan. Its response includes page-verdict and billed headers; the stated billing policy charges only clean shots, with bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits costing nothing. See ScreenshotNeo for product details.

Or skip the browser setup

Make one request to ScreenshotNeo to receive a screenshot. See the API documentation for options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

Cookie banners are accepted like a visitor and removed before the shot, along with 60+ known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, and failed loads are never billed. An MCP server lets Claude, Cursor, and other MCP clients take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000. Sign up free for ScreenshotNeo.

FAQ

Does full_page=True trigger lazy loading?

No. It captures the full scrollable bounds. Scroll the relevant areas and verify the content before taking the screenshot.

Can I capture a page that loads forever?

Yes, but only after defining a finite boundary such as a target item count, known end marker, scroll limit, or time limit.

How do I get screenshot bytes instead of a file?

Call data = await page.screenshot(full_page=True) and pass the returned bytes to your image-processing or upload code.

Can a locator screenshot capture every row in a scrollable panel?

Not automatically. It captures the element’s visible scrolled content; load and capture sections separately if the entire panel is required.