ScreenshotNeo

BlogHow-to

How to Take Bulk Screenshots of URLs with Different Wait Conditions

Capture many URLs in one run while giving each page its own readiness rule, timeout, and output file. Use Playwright to handle asynchronous pages and partial failures.

By the ScreenshotNeo team4 October 202610 min read

Give each capture its own URL, wait rule, timeout, and output path. A Playwright script can navigate to each URL, wait for a navigation event or a page-specific selector, and save the screenshot. Handle errors per URL so one timeout does not discard the rest of the batch.

This guide uses Node.js and Playwright. It covers navigation waits, selector waits, fixed delays, full-page capture, unique filenames, retryable failures, and repeatable output. For a documented configuration-driven alternative, shot-scraper supports URL entries and selector-based waits.

1. Choose a readiness signal for each URL

Navigation completion and application readiness are different. A navigation event tells you that a stage of document loading occurred; it does not guarantee that a client-rendered chart, image, or widget is ready to appear in the screenshot.

Wait rule What it means Use it when
commit The response was received and document loading started. You need the earliest document state and will wait for a specific element afterward.
domcontentloaded The DOMContentLoaded event fired. The page is mostly server-rendered and required markup is available early.
load The load event fired. Resources that participate in the browser load event should be ready.
networkidle No network connections for at least 500 ms. Only when that condition is meaningful for your page. Playwright discourages using it as a general readiness check.
Selector A chosen element reaches the requested state, such as visible. A client-rendered component or page-specific marker indicates readiness.
Fixed delay The script pauses for a set duration. No observable signal is available; treat it as a heuristic.

Use the least restrictive signal that reliably means the content you need is ready. Prefer an application-specific selector over waiting for all network activity to stop. Analytics, polling, and persistent connections can make network-idle waits unreliable, while a fixed delay can be too short on a slow run and waste time on a fast one.

See Playwright’s Page API for navigation wait options and its screenshot API for capture options. The documentation defines network idle as no network connections for at least 500 ms and recommends assessing readiness with web assertions instead.

2. Install Playwright and prepare a capture manifest

Install Playwright in a Node.js project, then install its Chromium browser. The manifest below associates each URL with its own navigation event, optional selector, timeout, delay, and output name. Keep output names unique: repeated names overwrite earlier files.

npm init -y
npm install playwright
npx playwright install chromium

Create captures.json:

[
  {
    "url": "https://example.com/",
    "waitUntil": "domcontentloaded",
    "timeoutMs": 30000,
    "output": "example-home.png",
    "fullPage": true
  },
  {
    "url": "https://example.org/app",
    "waitUntil": "domcontentloaded",
    "waitForSelector": "[data-ready=\"true\"]",
    "timeoutMs": 45000,
    "output": "example-app.png",
    "fullPage": true
  },
  {
    "url": "https://example.net/report",
    "waitUntil": "load",
    "waitForSelector": "main h1",
    "selectorState": "visible",
    "delayMs": 500,
    "output": "example-report.png",
    "fullPage": false
  }
]

Each record uses only the fields it needs. For a page whose readiness is defined by a selector, the navigation event is just the initial navigation boundary; the selector is the meaningful readiness condition. The small delay in the third record is an optional settling heuristic, not a guarantee.

3. Run the bulk screenshot script

Save this as capture.js. It launches Chromium once, processes records sequentially, creates a fresh page for each URL, and writes a JSON report. A failure is recorded against its URL and the loop continues. The script checks that required fields and filenames are valid and creates the output directory.

const fs = require('node:fs/promises');
const path = require('node:path');
const { chromium } = require('playwright');

const inputPath = process.argv[2] || 'captures.json';
const outputDir = process.argv[3] || 'screenshots';
const allowedWaits = new Set(['commit', 'domcontentloaded', 'load', 'networkidle']);
const allowedSelectorStates = new Set(['attached', 'detached', 'visible', 'hidden']);

function validate(item, index) {
  if (!item || typeof item.url !== 'string' || !/^https?:\/\//i.test(item.url)) {
    throw new Error(`Capture ${index}: url must be an http or https URL`);
  }
  if (!item.output || typeof item.output !== 'string' || path.basename(item.output) !== item.output) {
    throw new Error(`Capture ${index}: output must be a filename, not a path`);
  }
  if (item.waitUntil !== undefined && !allowedWaits.has(item.waitUntil)) {
    throw new Error(`Capture ${index}: unsupported waitUntil value`);
  }
  if (item.waitForSelector !== undefined && typeof item.waitForSelector !== 'string') {
    throw new Error(`Capture ${index}: waitForSelector must be a string`);
  }
  if (item.selectorState !== undefined && !allowedSelectorStates.has(item.selectorState)) {
    throw new Error(`Capture ${index}: unsupported selectorState value`);
  }
  for (const key of ['timeoutMs', 'delayMs']) {
    if (item[key] !== undefined && (!Number.isFinite(item[key]) || item[key] < 0)) {
      throw new Error(`Capture ${index}: ${key} must be a non-negative number`);
    }
  }
}

async function main() {
  const captures = JSON.parse(await fs.readFile(inputPath, 'utf8'));
  if (!Array.isArray(captures)) throw new Error('Input JSON must be an array');
  const seen = new Set();
  captures.forEach((item, index) => {
    validate(item, index);
    if (seen.has(item.output)) throw new Error(`Duplicate output filename: ${item.output}`);
    seen.add(item.output);
  });

  await fs.mkdir(outputDir, { recursive: true });
  const browser = await chromium.launch({ headless: true });
  const results = [];
  try {
    for (const [index, item] of captures.entries()) {
      const page = await browser.newPage();
      const timeoutMs = item.timeoutMs ?? 30000;
      try {
        page.setDefaultTimeout(timeoutMs);
        const response = await page.goto(item.url, {
          waitUntil: item.waitUntil ?? 'load',
          timeout: timeoutMs,
        });
        if (response && response.status() >= 400) {
          throw new Error(`Navigation returned HTTP ${response.status()}`);
        }
        if (item.waitForSelector) {
          await page.locator(item.waitForSelector).waitFor({
            state: item.selectorState ?? 'visible',
            timeout: timeoutMs,
          });
        }
        if (item.delayMs) {
          await page.waitForTimeout(item.delayMs);
        }
        const output = path.join(outputDir, item.output);
        await page.screenshot({
          path: output,
          fullPage: item.fullPage ?? true,
          animations: 'disabled',
        });
        results.push({ index, url: item.url, status: 'ok', output });
      } catch (error) {
        results.push({ index, url: item.url, status: 'error', error: String(error.message || error) });
      } finally {
        await page.close();
      }
    }
  } finally {
    await browser.close();
  }

  const reportPath = path.join(outputDir, 'capture-report.json');
  await fs.writeFile(reportPath, JSON.stringify(results, null, 2));
  const failed = results.filter(result => result.status === 'error').length;
  console.log(`Captured ${results.length - failed}/${results.length}; report: ${reportPath}`);
  if (failed) process.exitCode = 1;
}

main().catch(error => {
  console.error(error);
  process.exitCode = 1;
});

Run it with:

node capture.js captures.json screenshots

The script uses sequential capture to keep browser load predictable and the per-URL report straightforward. A nonzero exit code signals that at least one capture failed, while successful image files remain available for review or selective retry.

4. Adapt the workflow to CSV or parallel batches

JSON is convenient when each URL has several optional fields. For a CSV workflow, use columns such as url,waitUntil,waitForSelector,timeoutMs,output, parse the rows with a CSV parser, then map each row to the same record shape before calling the capture loop. Parse quoted CSV fields with a CSV library rather than splitting lines on commas; selectors and URLs can contain punctuation.

For larger lists, parallel pages can improve throughput, but each page consumes browser resources and concurrent navigation can overload your machine or the target sites. Start with a small concurrency limit, preserve unique output paths, and keep per-URL timeouts and error records. Do not blindly retry every failure: a deterministic bad URL or invalid selector will fail again. Retry transient network errors and timeouts selectively, with a bounded attempt count and a backoff delay.

5. Configure consistent captures

  • Viewport: pass viewport: { width, height } to browser.newPage() or create a browser context with a fixed viewport. Use the same dimensions for every page you intend to compare.
  • Full page: set fullPage: true to capture beyond the viewport. Long or infinite-scrolling pages may be slow or may not expose all content until scrolling triggers lazy loading.
  • Element screenshot: use page.locator('selector').screenshot({ path }) when only one element is needed; first wait for that locator to be visible.
  • Animations: the sample disables finite animations during screenshot capture. This can improve visual repeatability, but does not establish that the application has finished rendering.
  • Volatile content: mask timestamps, rotating banners, or other changing regions in visual comparison workflows, or hide them with page styling where appropriate.
  • Authentication and state: use a browser context with the needed cookies or storage state for pages that require a login. Keep credentials out of the manifest and source control.
  • Browser consistency: browser version, operating system, fonts, device scale, headless mode, and hardware can alter rendering. Keep the capture environment consistent when comparing results.

Playwright’s visual comparison guidance notes that browser, platform, fonts, settings, hardware, power state, and headless mode can affect rendering. Screenshot stability controls help with repeatability but do not replace a meaningful readiness condition.

6. cURL, Python, and Node.js options

cURL itself does not render web pages or take browser screenshots. It can call a screenshot API that performs the browser capture. For a do-it-yourself browser workflow, Node.js with Playwright is shown above; Playwright also has language bindings, but the loop and per-URL orchestration must be implemented in the chosen language.

Call ScreenshotNeo with cURL

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

Call ScreenshotNeo with Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Call ScreenshotNeo with Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));

7. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. Its API takes a URL in one GET request and returns an image or PDF. For a batch with per-URL waits, send one request per URL and set the appropriate wait option for each call; see the API documentation for the supported parameter names and configuration.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

Cookie banners are accepted and removed before the shot, along with known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up for 1,000 free screenshots a month, with no card required.

8. Troubleshooting

Symptom Likely cause Fix
Navigation times out The chosen event takes too long, the site is slow, or navigation is blocked. Set a realistic finite timeout. If the needed content appears earlier, use a less restrictive navigation event and wait for its selector. Record and retry only transient failures.
Selector wait times out The selector is wrong, appears only after interaction, or is inside a frame or shadow root. Inspect the live page, correct the selector, and account for frame or shadow DOM boundaries. Confirm that the selected state (visible, attached, hidden, or detached) matches the goal.
Screenshot is blank or incomplete The page navigated but the application content was not ready, or an HTTP error page was captured. Wait for a meaningful page marker, inspect the navigation response status, and save the error details. Check whether authentication or a client-side error prevents rendering.
Some images are missing Images are lazy-loaded or load after the selected readiness signal. Wait for image-specific readiness or scroll the relevant page regions before capture. A load event alone does not guarantee lazy content has been requested.
Some files replace others Two records use the same output filename. Use unique output names. The sample rejects duplicates before starting the browser.
Run stops after one bad URL Error handling surrounds the whole batch instead of each item. Catch errors per record, write a result for every URL, and set a failing exit code after the report is saved.
Network-idle wait never completes The page maintains polling, analytics, or other ongoing connections. Use a selector or application-specific condition. Playwright discourages network idle as a general readiness strategy.
Images differ between runs Fonts, browser environment, animations, or dynamic content changed. Keep the browser and operating environment consistent, disable animations when suitable, and mask or hide volatile regions.
Browser launch fails The Playwright browser binary is not installed or system dependencies are missing. Run npx playwright install chromium in the project environment and follow Playwright’s installation guidance for the operating system.

9. Performance, reliability, and cost

Launching one browser for the batch avoids repeating browser startup for every URL. Sequential processing limits resource use and makes failure reporting simple, but total run time grows with the sum of navigation and wait times. Bounded concurrency can reduce elapsed time at the cost of more memory, CPU, network connections, and pressure on target sites.

Choose selectors that reflect the desired content, finite per-page timeouts, and unique filenames. Keep a report with status and error per URL so a later run can target failed items. Use bounded retries with backoff only for failures likely to be transient. For repeatable visual checks, pin the browser environment and manage animation and volatile content; a screenshot that is stable is not necessarily semantically correct.

With local Playwright, there is no screenshot API charge, but you provide and maintain the machine, browser binaries, storage, and operational retry logic. If using a hosted API, check the provider’s current plan limits and billing rules; those details vary. ScreenshotNeo states that only clean shots are billed, and its plans range from a free 1,000 per month to paid tiers starting at $5 for 3,000. Verify current plan terms in its documentation before building usage assumptions.

10. FAQ

Can each URL use a different wait condition?

Yes. Store the navigation event, selector, or delay on each record and apply only the fields that record needs.

Should I use network idle for every URL?

No. Persistent connections and background requests can prevent it from completing, and it does not prove that the specific content you need is ready. Prefer a meaningful selector or application condition.

Does full-page capture trigger every lazy-loaded image?

Not necessarily. A full-page screenshot captures beyond the viewport, but lazy content may require scrolling or additional readiness checks before it is requested.

Can I keep successful screenshots when a URL fails?

Yes. Catch failures per URL and write the report after processing the batch, as in the sample script.

Can I use this for visual regression?

Yes, if you keep the rendering environment consistent and control animations and changing page regions. Readiness waits and stable rendering are separate concerns.