ScreenshotNeo

BlogHow-to

How to Capture Bulk Screenshots of Pages with Lazy-Loaded Images

Scroll through each page to trigger lazy-loaded content, wait for images to render, then capture and check full-page screenshots in a repeatable batch.

By the ScreenshotNeo team4 October 202613 min read

To capture bulk screenshots that include lazy-loaded images, automate a browser to visit each URL, scroll through the page in viewport-sized steps, wait for newly revealed content to load, and then take a full-page screenshot. A full-page option captures the scrollable document as one tall image, but it does not guarantee that a page has requested content that only loads after scrolling. Playwright supports full-page screenshots in JavaScript and Python; the scroll-and-wait workflow is what triggers many lazy-loading patterns. Playwright’s screenshot guide and Page API reference document navigation and capture options.

This guide uses Playwright and includes runnable batch scripts in Node.js and Python. Both scripts save one screenshot per URL, keep a CSV-style log of successes and failures, and scroll until the document stops growing or reaches a configured limit. Adjust the readiness checks for the sites you capture: image loading, animations, authentication, overlays, and infinite scroll are site-specific.

1. Why scrolling comes before a full-page screenshot

Pages commonly defer image downloads until an image approaches the viewport. A full-page screenshot tells the browser to capture the whole document; it does not necessarily simulate a visitor scrolling through it first. As a result, an image farther down the page may still be blank, show a placeholder, or be absent when capture begins. A third-party guide in the research dossier describes this scroll-triggered failure mode; treat scrolling as a practical preparation step, not a guarantee for every site.

The reliable sequence is: navigate, scroll down in manageable steps, allow new content to appear, check whether images have loaded, scroll again if the page grew, then capture. Playwright’s full-page option is fullPage: true in JavaScript and full_page=True in Python. Its docs describe a full-page screenshot as an image of the full scrollable page. See the screenshot guide.

2. Prepare a repeatable batch

  1. Make a URL list. Store one absolute URL per line in a text file. Remove blank lines and comments before processing.
  2. Choose a consistent viewport. Use the same width, height, browser, and device scale for pages you intend to compare. Responsive layouts change with viewport size.
  3. Set output and limits. Choose an output directory, a maximum scroll depth, a per-step pause, navigation timeout, and image format. A limit prevents infinite-scroll pages from growing without bound.
  4. Keep a failure log. Record each URL, output filename, and error. One inaccessible page should not prevent the rest of the batch from running.
  5. Review the files. Inspect for empty image areas, placeholders, clipped content, login pages, or a layout that was still shifting.

Use a dedicated folder for the batch. The example scripts derive filenames from each URL and add a numeric prefix, so repeated paths on different hosts do not overwrite each other.

3. Node.js: runnable Playwright batch script

Install Playwright and its Chromium browser in a new project:

npm init -y
npm install playwright
npx playwright install chromium

Save the URLs below as urls.txt, one URL per line. Save this script as capture.mjs and run node capture.mjs.

import { chromium } from 'playwright';
import { mkdir, readFile, writeFile } from 'node:fs/promises';
import path from 'node:path';

const INPUT = 'urls.txt';
const OUTPUT_DIR = 'screenshots';
const VIEWPORT = { width: 1440, height: 900 };
const NAVIGATION_TIMEOUT_MS = 45_000;
const STEP_PAUSE_MS = 700;
const MAX_SCROLL_STEPS = 80;
const MAX_STABLE_ROUNDS = 3;

function safeName(url, index) {
  const parsed = new URL(url);
  const base = `${parsed.hostname}${parsed.pathname}`
    .replace(/[^a-z0-9.-]+/gi, '-')
    .replace(/^-+|-+$/g, '')
    .slice(0, 100) || 'page';
  return `${String(index + 1).padStart(3, '0')}-${base}.png`;
}

async function prepareLazyContent(page) {
  let previousHeight = 0;
  let stableRounds = 0;

  for (let step = 0; step < MAX_SCROLL_STEPS; step++) {
    const height = await page.evaluate(() => document.documentElement.scrollHeight);
    const viewport = page.viewportSize()?.height ?? 900;
    const y = Math.min((step + 1) * viewport, height);
    await page.evaluate(scrollY => window.scrollTo(0, scrollY), y);
    await page.waitForTimeout(STEP_PAUSE_MS);

    const newHeight = await page.evaluate(() => document.documentElement.scrollHeight);
    if (newHeight === previousHeight && newHeight === height) {
      stableRounds++;
    } else {
      stableRounds = 0;
    }
    previousHeight = newHeight;

    if (stableRounds >= MAX_STABLE_ROUNDS && y >= newHeight - viewport) break;
  }

  // Return to the top before capture. This can trigger scroll-linked layout changes.
  await page.evaluate(() => window.scrollTo(0, 0));
  await page.waitForTimeout(STEP_PAUSE_MS);

  // Wait for image elements currently in the DOM to finish or fail. This does not
  // prove that every CSS background or site-specific image widget has loaded.
  await page.evaluate(async () => {
    const images = Array.from(document.images);
    await Promise.all(images.map(img => {
      if (img.complete) return Promise.resolve();
      return new Promise(resolve => {
        img.addEventListener('load', resolve, { once: true });
        img.addEventListener('error', resolve, { once: true });
      });
    }));
  });
}

const urls = (await readFile(INPUT, 'utf8'))
  .split(/\r?\n/)
  .map(line => line.trim())
  .filter(line => line && !line.startsWith('#'));

await mkdir(OUTPUT_DIR, { recursive: true });
const browser = await chromium.launch({ headless: true });
const context = await browser.newContext({ viewport: VIEWPORT, deviceScaleFactor: 1 });
const page = await context.newPage();
page.setDefaultNavigationTimeout(NAVIGATION_TIMEOUT_MS);
const results = [];

try {
  for (let index = 0; index < urls.length; index++) {
    const url = urls[index];
    const file = safeName(url, index);
    try {
      const response = await page.goto(url, { waitUntil: 'domcontentloaded' });
      await prepareLazyContent(page);
      await page.screenshot({
        path: path.join(OUTPUT_DIR, file),
        fullPage: true,
        animations: 'disabled'
      });
      results.push({ url, file, status: response?.status() ?? 'no-response', error: '' });
      console.log(`Saved ${file} (${url})`);
    } catch (error) {
      const message = String(error?.message ?? error).replaceAll(',', ';');
      results.push({ url, file: '', status: 'error', error: message });
      console.error(`Failed ${url}: ${message}`);
    }
  }
} finally {
  await browser.close();
  const csv = [
    'url,file,status,error',
    ...results.map(row => [row.url, row.file, row.status, row.error]
      .map(value => `"${String(value).replaceAll('"', '""')}"`).join(','))
  ].join('\n');
  await writeFile(path.join(OUTPUT_DIR, 'results.csv'), csv);
}

The loop scrolls by one viewport and pauses to give newly revealed content time to request and render. It checks document height and requires several stable rounds at the bottom before stopping. MAX_SCROLL_STEPS is a safety cap, especially useful for infinite-scroll pages. The image wait handles ordinary <img> elements present in the DOM; it cannot identify whether a failed image is a deliberate site error, nor does it cover every CSS background or custom image component.

4. Python: runnable Playwright batch script

Install the Python package and Chromium browser:

python -m pip install playwright
python -m playwright install chromium

Use the same urls.txt format. Save the following as capture.py and run python capture.py.

import asyncio
import csv
import re
from pathlib import Path
from urllib.parse import urlparse
from playwright.async_api import async_playwright

INPUT = Path("urls.txt")
OUTPUT_DIR = Path("screenshots")
VIEWPORT = {"width": 1440, "height": 900}
NAVIGATION_TIMEOUT_MS = 45_000
STEP_PAUSE_MS = 700
MAX_SCROLL_STEPS = 80
MAX_STABLE_ROUNDS = 3


def safe_name(url: str, index: int) -> str:
    parsed = urlparse(url)
    base = re.sub(r"[^a-zA-Z0-9.-]+", "-", parsed.netloc + parsed.path)
    base = base.strip("-")[:100] or "page"
    return f"{index + 1:03d}-{base}.png"


async def prepare_lazy_content(page):
    previous_height = 0
    stable_rounds = 0

    for step in range(MAX_SCROLL_STEPS):
        height = await page.evaluate("() => document.documentElement.scrollHeight")
        viewport = page.viewport_size["height"]
        y = min((step + 1) * viewport, height)
        await page.evaluate("scrollY => window.scrollTo(0, scrollY)", y)
        await page.wait_for_timeout(STEP_PAUSE_MS)

        new_height = await page.evaluate("() => document.documentElement.scrollHeight")
        if new_height == previous_height and new_height == height:
            stable_rounds += 1
        else:
            stable_rounds = 0
        previous_height = new_height

        if stable_rounds >= MAX_STABLE_ROUNDS and y >= new_height - viewport:
            break

    await page.evaluate("() => window.scrollTo(0, 0)")
    await page.wait_for_timeout(STEP_PAUSE_MS)

    # Wait for DOM image elements to load or fail. CSS backgrounds and custom
    # image widgets can need a site-specific readiness condition.
    await page.evaluate("""async () => {
      const images = Array.from(document.images);
      await Promise.all(images.map(img => {
        if (img.complete) return Promise.resolve();
        return new Promise(resolve => {
          img.addEventListener('load', resolve, { once: true });
          img.addEventListener('error', resolve, { once: true });
        });
      }));
    }""")


async def main():
    urls = [line.strip() for line in INPUT.read_text(encoding="utf-8").splitlines()
            if line.strip() and not line.lstrip().startswith("#")]
    OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
    results = []

    async with async_playwright() as p:
        browser = await p.chromium.launch()
        context = await browser.new_context(viewport=VIEWPORT, device_scale_factor=1)
        page = await context.new_page()
        page.set_default_navigation_timeout(NAVIGATION_TIMEOUT_MS)
        try:
            for index, url in enumerate(urls):
                filename = safe_name(url, index)
                try:
                    response = await page.goto(url, wait_until="domcontentloaded")
                    await prepare_lazy_content(page)
                    await page.screenshot(
                        path=str(OUTPUT_DIR / filename),
                        full_page=True,
                        animations="disabled",
                    )
                    status = response.status if response else "no-response"
                    results.append({"url": url, "file": filename,
                                    "status": status, "error": ""})
                    print(f"Saved {filename} ({url})")
                except Exception as exc:
                    results.append({"url": url, "file": "",
                                    "status": "error", "error": str(exc)})
                    print(f"Failed {url}: {exc}")
        finally:
            await browser.close()

    with (OUTPUT_DIR / "results.csv").open("w", newline="", encoding="utf-8") as handle:
        writer = csv.DictWriter(handle, fieldnames=["url", "file", "status", "error"])
        writer.writeheader()
        writer.writerows(results)


if __name__ == "__main__":
    asyncio.run(main())

Both examples use the same sequence and settings so output remains comparable across a URL list. For controlled environments, pin your Playwright dependency version and use the same installed browser revision on each worker.

5. Readiness: how long should you wait?

There is no single wait condition that proves every visual element is ready. A fixed pause is easy to use but can be too short on slow pages and unnecessarily long on fast ones. Prefer a concrete condition when you know the page: wait for a particular image to be visible, for a loading indicator to disappear, or for the expected number of cards to appear. For lazy images, you can check that their elements are complete and have a nonzero natural width; for custom image components, use the site’s own selectors or state.

Do not treat networkidle as a universal visual-readiness signal. Playwright discourages relying on it alone as a test readiness criterion: some sites keep requests open, while other pages finish network activity before their layout and image work is done. Use domcontentloaded or another navigation milestone to begin the workflow, then apply content-specific waits. See Playwright’s Page API documentation.

Useful capture settings

  • fullPage: true / full_page=True: capture the full scrollable document as one image.
  • path: save directly to a file; omit it if you want image bytes for post-processing.
  • type: choose a supported image output such as PNG or JPEG; check the installed Playwright API for exact options and quality behavior.
  • animations: 'disabled': reduce animation differences during capture; dynamic content can still change for other reasons.
  • viewport and deviceScaleFactor: establish the browser’s CSS viewport and pixel density consistently.
  • timeout: bound navigation and waits. A timeout should become a recorded failure with enough context to retry or investigate.

Full-page capture creates a single tall image. If the purpose is visual review, a series of viewport-sized captures can make sections easier to compare and can avoid extremely tall output. That is a workflow choice, not a requirement.

6. Infinite scroll and other page-specific behavior

An infinite-scroll feed can keep growing whenever the browser reaches the bottom. Define a stopping rule instead of trying to scroll forever. The sample scripts use both a maximum number of steps and repeated checks for unchanged document height. Other useful limits include a known target item count or maximum page height.

  • Nested scroll containers: scrolling the window may not move a feed inside a panel. Identify and scroll the element that owns the overflow.
  • Fixed headers: sticky bars may cover content or repeat in a full-page image depending on site layout and browser behavior. Inspect output; consider segment captures when the page’s sticky behavior makes one tall image misleading.
  • Cookie notices and popups: they can cover content. Dismiss them through the page’s normal controls when appropriate, or record that the page was obstructed. Do not assume every overlay can be removed generically.
  • Authentication: provide the needed session state only for pages you are authorized to access. A login wall is not evidence that the target content loaded.
  • Animations and carousels: disable or wait for them when a stable state matters. A screenshot captures one moment, not every animated state.
  • Anti-automation checks: a site may challenge or block browser automation. The scripts record navigation errors, but they cannot guarantee access or bypass a site’s controls.
  • Height-changing content: lazy images, ads, and asynchronous widgets can move the page after it appears to settle. Use a meaningful site-specific condition or take a second height check before capture.

7. Check screenshots and make the batch reliable

A successful file write does not prove that every image rendered. Review the resulting screenshots, especially after the first run against a new site. Look for blank image boxes, low-resolution placeholders, sections that end abruptly, and pages that show a consent screen or challenge instead of the intended content.

  • Keep the URL-to-file mapping in the results log.
  • Capture the final document height and image status in your own logs when diagnosing a recurring site issue.
  • Retry only transient failures, with a bounded retry count and delay. Avoid retrying malformed URLs or predictable authorization failures indefinitely.
  • Use a fresh context per batch or per URL if cookies and local storage from one page could affect another; reuse a context when shared login state is intentional.
  • Limit parallel pages to what the machine and target sites can handle. More concurrency consumes more memory and can add load to sites; start conservatively.
  • For very long documents, consider segment screenshots or a maximum height to limit memory use and output size.

8. Troubleshooting

Symptom Likely cause What to change
Images are blank or placeholders remain The capture began before scrolling triggered the request, or the image did not finish rendering. Scroll through the document first. Increase the per-step wait, then wait for a known image or content selector. Check failed image requests and the site’s loading state.
The page is cut off Full-page capture was not enabled, or the page uses a nested scroll container or virtualized content. Set the full-page option. Scroll the actual container; virtualized lists may remove off-screen items, making a single full-page image insufficient.
The script stops too early The page height stayed constant briefly even though content arrives later. Increase the pause and stable-round threshold, or wait for a page-specific content condition. Keep a maximum scroll-step cap.
The script never reaches a stopping point Infinite scroll continues adding items or height changes continuously. Use a maximum depth, expected item count, or maximum number of scroll steps. Log when the cap is reached.
Navigation times out The server is slow, a request remains active, or the URL is inaccessible from the runner. Use a bounded timeout and a navigation milestone such as domcontentloaded; inspect the URL and record the error. Do not assume a longer timeout repairs an inaccessible page.
Screenshot shows a consent dialog, login form, or challenge The visible page state is an overlay, authentication boundary, or anti-automation response. Handle the site’s legitimate consent or authentication flow when authorized; otherwise record the page as blocked or unavailable.
Output files overwrite one another Filenames are based only on path or are not unique. Include hostname and a stable index or ID in each filename, as the examples do.
Python cannot import Playwright or Chromium fails to launch The package or browser binary is missing from that Python environment. Install the package in the active environment and run python -m playwright install chromium.
Node reports that Playwright or its browser is missing The dependency or browser was not installed for the project. Run npm install playwright and npx playwright install chromium from the project directory.

9. Performance, reliability, and cost

Each URL requires navigation, multiple scroll-and-wait steps, image readiness checks, and a capture. Batch runtime therefore depends on page count, page size, delays, site response, and your concurrency. Increasing parallelism can reduce wall-clock time but also raises browser memory use and request load. Measure your own workload and choose a bounded concurrency level.

Full-page screenshots can be large, especially for long pages or high device scale factors. Use a lower pixel density or JPEG when smaller output is more important than lossless detail; keep PNG where crisp text or pixel-level review matters. These are format tradeoffs, not fixed file-size guarantees.

With self-managed Playwright, account for the compute used to run browsers, storage for output, and maintenance of browser dependencies. Add limits, logs, and retry policies so one exceptional page does not hold the whole batch open. If a hosted API better fits the workflow, compare its actual batch, scrolling, readiness, output, and failure-reporting capabilities rather than assuming that a screenshot endpoint scrolls lazy content automatically.

10. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. Its API can capture a URL in one GET request and return an image or PDF. For this page, the one-call option is:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request options. Cookie banners, newsletter popups, and chat widgets are removed before the shot; those cleanup steps can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response includes page-verdict and billing headers. An MCP server lets AI agents, including Claude, Cursor, and other MCP clients, take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.

Create a free ScreenshotNeo account and get 1,000 screenshots a month with no card.

11. Frequently asked questions

Does full-page capture automatically load every lazy image?

No. It captures the full scrollable document, but content may only be requested after scrolling. Scroll first, then check that the page-specific content is ready.

Should I always wait for network idle?

No. Network activity alone is not a universal signal that visual content is ready. Use a concrete selector, image state, loading indicator, or other condition that matches the page.

Can one script handle every site?

No. Nested scrolling, virtualized lists, authentication, overlays, anti-automation checks, and custom image components require site-specific handling.

When should I use separate viewport captures?

Use them when a tall single image is unwieldy, when you need section-by-section review, or when sticky and virtualized layouts do not represent well in one full-page image.

Can I capture an authenticated page?

Yes, if you are authorized and configure the browser with the necessary session state. Keep credentials and session data out of source control and logs.