ScreenshotNeo

BlogHow-to

How to Bulk Screenshot Websites That Use Client-Side Rendering

Bulk screenshot JavaScript-rendered websites by waiting for application-specific readiness, then capturing URLs through a bounded queue with per-page error handling.

By the ScreenshotNeo team4 October 202610 min read

To bulk screenshot websites that use client-side rendering, render each page in a real browser, wait for a signal that the specific content you need is ready, then capture it through a bounded queue. A successful navigation event only marks a browser navigation milestone; it does not prove that client-rendered components have finished loading. For reliable batches, define the capture scope, use page-specific readiness checks, save a result or error for every URL, and retry only likely transient failures.

1. Define the capture job

Before launching a browser, make a canonical list of URLs and decide what each output should contain. These choices affect both readiness checks and capture time.

  • Scope: the current viewport, the full scrollable page, or one element selected by CSS.
  • Format: PNG for lossless detail, JPEG for smaller photographic images, or another format your workflow supports.
  • Viewport: choose a consistent width and height, and decide whether device scale factor should be greater than one.
  • Readiness: identify a selector or application condition that means the content you care about is visible and stable.
  • Output: set a predictable filename scheme and decide how results, failures, and retries will be recorded.

For a single site, a selector such as [data-page-ready] can be a useful signal if the application actually sets it after rendering the target content. Across unrelated sites, there is no universal selector: provide per-site rules or capture a clearly defined generic milestone and accept that it may not mean the same thing everywhere.

2. Build a bounded Playwright batch

Playwright provides browser-based screenshots and supports viewport, full-page, and element captures. The example below uses a bounded worker pool, waits for a configurable readiness selector per URL, records errors by URL, and writes a JSON manifest. It uses domcontentloaded for navigation and then waits for the application-specific signal. Replace the sample URLs and selectors with real targets.

import { chromium } from 'playwright';
import { mkdir, writeFile } from 'node:fs/promises';

const jobs = [
  { url: 'https://example.com/app', ready: '[data-page-ready]' },
  { url: 'https://example.org/dashboard', ready: '#dashboard-content' },
];
const outputDir = 'screenshots';
const concurrency = 3;
const timeoutMs = 30_000;

await mkdir(outputDir, { recursive: true });
const browser = await chromium.launch();
const results = new Array(jobs.length);
let next = 0;

async function worker() {
  while (true) {
    const index = next++;
    if (index >= jobs.length) return;
    const job = jobs[index];
    const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
    try {
      page.setDefaultTimeout(timeoutMs);
      await page.goto(job.url, { waitUntil: 'domcontentloaded', timeout: timeoutMs });
      await page.locator(job.ready).waitFor({ state: 'visible', timeout: timeoutMs });
      const file = `${outputDir}/${String(index + 1).padStart(4, '0')}.png`;
      await page.screenshot({ path: file, fullPage: true, type: 'png' });
      results[index] = { url: job.url, status: 'ok', file };
    } catch (error) {
      results[index] = { url: job.url, status: 'error', error: String(error) };
    } finally {
      await page.close();
    }
  }
}

try {
  await Promise.all(Array.from({ length: Math.min(concurrency, jobs.length) }, worker));
} finally {
  await browser.close();
  await writeFile(`${outputDir}/manifest.json`, JSON.stringify(results, null, 2));
}

Install Playwright and its browser in the project environment before running the script. For example, with npm:

npm install playwright
npx playwright install chromium
node bulk-screenshots.mjs

Save the script as bulk-screenshots.mjs. For a production batch, load jobs from a file or database, validate URLs and selectors, and persist results as each job completes rather than keeping all results only in memory. The example assigns filenames by input index, so reordering the input changes filenames; use a stable ID or sanitized URL hash if you need repeatable names across runs.

Bound concurrency and isolate failures

Start with a small number of workers and increase it only if the machine and target sites can handle the load. A bounded queue limits browser tabs and memory use; launching one page for every URL at once can exhaust resources or overwhelm sites. Keep a separate browser context per job when cookies, authentication, or storage must be isolated. Reuse a context only when shared session state is intended.

The sample catches errors per URL so one broken page does not discard the rest of the batch. In a durable worker, store the original error and attempt count, then retry only errors that appear transient, such as a temporary network failure. A missing readiness selector may be a bad job configuration or a page change; retrying it repeatedly usually does not help.

3. Choose the right readiness signal

Navigation states such as domcontentloaded and load describe browser lifecycle milestones. They are not guarantees that a JavaScript application has rendered the specific component you want. Use a web assertion or locator wait tied to the page’s relevant content. Playwright discourages using networkidle as a general readiness test; pages may keep requests open, or become network-quiet before the content of interest is ready. See the Playwright Page API.

  • Use a visible-content selector when the target element appears after client rendering.
  • Wait for a state change when an existing loading indicator disappears or an application attribute changes.
  • Use a bounded delay only as a fallback when there is no observable application signal; fixed sleeps can be too short on slow runs and waste time on fast ones.
  • Handle lazy content deliberately. Full-page capture does not guarantee every site’s below-the-fold images or data have loaded. Validate on representative pages; scrolling or site-specific interaction may be needed.

For a set of unrelated sites, keep readiness rules in the job data, such as a per-host selector and timeout. If a page requires login, a click, or a consent decision to reach the content, encode that step explicitly and only for sites where it is appropriate.

4. Choose viewport, full-page, or element capture

Playwright documents full-page screenshots and screenshots of page elements. Puppeteer also documents viewport, full-page, and element screenshots. Choose based on the result consumers need, not just convenience. Full-page images can be very tall and memory intensive; element captures are easier to compare when the surrounding page changes.

Capture scope Use it when Watch for
Viewport You need the visible fold, a consistent visual comparison, or a quick preview. Content outside the viewport is omitted.
Full page You need the complete scrollable document in one image. Very long pages can use substantial memory; lazy content may need explicit loading first.
Element You need a chart, card, report panel, or other specific component. The selector must resolve to the intended visible element; check for duplicates and clipping.

Playwright’s documented screenshot options include fullPage, type, quality for supported lossy formats, omitBackground, animations, scale, and clip, alongside element locator screenshots. Consult the Playwright screenshot documentation for the current API details. Puppeteer’s screenshot behavior and options are documented in its screenshot guide.

5. Other ways to run a batch

Approach Best fit Check before choosing
Playwright Custom automation, page-specific readiness, and control over capture scope. Browser installation, runtime, queueing, and maintenance.
Puppeteer Browser automation workflows built around its navigation and screenshot APIs. Readiness strategy, dimensions, output handling, and runtime setup.
shot-scraper A command-line batch driven by a YAML URL list. Whether its documented interactions and readiness behavior fit the sites. See its documentation.
Hosted APIs Teams that want to outsource browser operations or use a documented batch workflow. Request limits, concurrency, async completion, rendering options, storage, price, and data handling.

For hosted batch options, see the official documentation for ScreenshotOne bulk screenshots, Browshot multiple and batch endpoints, and Urlbox requests and asynchronous processing. Product behavior and limits can change, so confirm the current documentation before scheduling a large batch. No neutral benchmark across these tools is established by the cited documentation.

6. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server for developers. Send one GET request for a URL to receive an image or PDF. Its cookie and consent handling accepts the banner like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

For bulk work, ScreenshotNeo supports up to 100 URLs per call, async jobs with signed webhooks, caching with a chosen TTL, and a usage API. This managed API does not replace a site-specific readiness check for every custom application; use a browser workflow when your capture depends on bespoke interaction or application state. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

The cURL request is directly runnable after replacing the API key. The Python snippet needs requests. The Node.js example uses Bun’s file writer; with Node.js, write the returned bytes using Buffer.from(await res.arrayBuffer()) and node:fs/promises. The one-URL call illustrates the API request; use ScreenshotNeo’s documented bulk or async workflow when submitting a batch. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. All features are on every plan. Start with 1,000 free screenshots a month, no card required.

7. Reliability, performance, and cost

Reliability

  • Persist one result record per input, including URL, timestamp, status, output path, and error details.
  • Set explicit navigation and readiness timeouts; a hung page should not block the entire queue.
  • Use bounded retries with backoff for transient network or browser errors, and avoid retrying deterministic selector failures without changing the job.
  • Keep a dead-letter list for URLs that exhausted retries so they can be inspected separately.
  • On restart, resume unfinished jobs from persisted state rather than repeating successful captures.

Performance

Batch throughput depends on the target pages, browser configuration, capture scope, and machine resources. This research contains no comparative benchmark. Increase concurrency gradually while monitoring memory, CPU, browser crashes, and target-site responses. Reusing a browser process avoids relaunching it for every URL; closing pages and contexts promptly keeps the batch bounded. Full-page captures and high device scale factors produce larger images and can increase processing and storage demands.

Cost and storage

Self-hosted captures consume your compute, storage, and engineering time for browser installation, updates, queues, retries, and monitoring. Hosted services exchange some of that operational work for provider-specific request limits, pricing, data handling, and output rules. Check current rates and limits directly; the research sources do not establish a neutral price comparison. Store outputs with a retention policy and avoid collecting authenticated or personal page content unless the workflow requires it.

8. Troubleshooting

Symptom Likely cause Fix
Screenshot shows a spinner or empty shell Capture began after navigation but before the client app rendered target content. Wait for a page-specific visible selector or state change; inspect the page when the wait times out.
networkidle never arrives Long polling, analytics, or other persistent requests keep activity alive. Use a relevant content assertion instead of treating network quiet as universal readiness.
Readiness selector timeout The selector is wrong, hidden, absent for this route, or appears only after interaction. Verify the selector for that URL, distinguish hidden and visible states, and add required site-specific steps.
Images below the fold are missing The site loads images lazily when they approach the viewport. Scroll through the relevant page or use the site’s own loading signal; validate the result on representative pages.
Browser crashes or the machine runs out of memory Too many pages are open, or full-page captures are very large. Reduce concurrency, close each page/context promptly, and capture only the required scope.
Some URLs fail while others succeed Per-site network, access, content, or configuration differences. Record failures individually, inspect the failing URL and error, then retry only transient cases.
Output files overwrite each other Filename generation is not unique. Use a stable job ID or collision-resistant filename derived from the URL and batch ID.
Hosted batch is delayed or rejected Concurrency, request bucket, payload, or async completion limits. Check current vendor limits, usage endpoints, job status, and webhook handling before increasing submission rate.

9. Frequently asked questions

Does a full-page screenshot trigger lazy loading?

Not reliably for every site. Full-page capture defines the image scope, while the application controls when content and images load. Test representative pages and add scrolling or another site-specific trigger when needed.

Can one readiness selector work for many websites?

Only if the sites share a reliable convention. For unrelated applications, store readiness rules per host or URL and treat missing signals as explicit job failures.

Should I use Playwright or Puppeteer?

Both document browser screenshots, including full-page and element capture. Choose based on your existing runtime and workflow; the sources cited here do not establish a performance winner.

How many browser pages should run at once?

There is no universal safe concurrency. Start small and adjust based on memory, CPU, page behavior, and any constraints that apply to the target sites or hosted API.

Sources