ScreenshotNeo

BlogHow-to

How to Bulk Screenshot Pages and Wait for a Specific Element to Appear

Capture batches of pages only after the content you need appears. Compare Puppeteer and Playwright workflows, handle timeouts, and save clear per-URL results.

By the ScreenshotNeo team4 October 202610 min read

To bulk screenshot pages after a specific element appears, loop over your URLs, navigate to each page, wait for a site-specific selector to become visible, then save either the full page or just that element. Give each page its own timeout and error result so one missing element does not stop the batch or produce a misleading image.

The examples below use Puppeteer and Playwright. Start sequentially for straightforward debugging; add measured concurrency only when runtime requirements justify it. For Puppeteer, waitForSelector defaults to a 30-second timeout, and its selector wait works across navigations. Puppeteer waitForSelector API

1. Choose a reliable selector and wait condition

Use a selector tied to the content you need, such as a stable ID or a data-* attribute. Avoid selectors that depend on generated class names or visual styling that may change.

Need Wait for Example
Element exists, even if hidden Attached/present in DOM [data-ready="true"]
Element is available to see Visible #report-summary
Specific text or app state Locator assertion or a site-specific readiness condition Playwright toBeVisible()

Presence and visibility are different. A hidden element may already be in the DOM before its content is ready. Conversely, an element can be visible while data inside it is still loading. Wait for the condition that matches the screenshot you need; when necessary, wait for a meaningful attribute, text, or application state as well.

Prefer a stable ID or data attribute. Keep the selector appropriate to the site and the intended capture. Puppeteer offers higher-level locators as well as its lower-level selector wait API; Playwright guidance favors locator waits and web-first assertions for many interactions. Puppeteer page interactions · Playwright Page API

2. Bulk capture with Puppeteer

This runnable ES module reads URLs from a JSON file, reuses one browser page, waits up to 30 seconds for a visible target on each page, and records success or failure per URL. It writes a JSON report so the output file and result stay associated with their source URL.

npm install puppeteer

Save as bulk-shot.mjs:

import puppeteer from 'puppeteer';
import { readFile, mkdir, writeFile } from 'node:fs/promises';

const urls = JSON.parse(await readFile('urls.json', 'utf8'));
const selector = process.env.SELECTOR ?? '[data-ready="true"]';
const timeout = Number(process.env.WAIT_TIMEOUT_MS ?? 30_000);
const outputDir = 'screenshots';

if (!Array.isArray(urls) || urls.some((url) => typeof url !== 'string')) {
  throw new Error('urls.json must be an array of URL strings');
}
if (!Number.isFinite(timeout) || timeout <= 0) {
  throw new Error('WAIT_TIMEOUT_MS must be a positive number');
}

await mkdir(outputDir, { recursive: true });
const browser = await puppeteer.launch();
const results = [];

try {
  const page = await browser.newPage();
  page.setDefaultNavigationTimeout(timeout);

  for (const [index, url] of urls.entries()) {
    const path = `${outputDir}/page-${String(index + 1).padStart(4, '0')}.png`;
    try {
      const response = await page.goto(url, { waitUntil: 'domcontentloaded', timeout });
      if (response && response.status() >= 400) {
        throw new Error(`HTTP ${response.status()}`);
      }
      await page.waitForSelector(selector, { visible: true, timeout });
      await page.screenshot({ path, fullPage: true });
      results.push({ url, ok: true, path, status: response?.status() ?? null });
    } catch (error) {
      results.push({ url, ok: false, error: String(error) });
      console.error(`Screenshot failed for ${url}:`, error);
    }
  }
} finally {
  await browser.close();
}

await writeFile('screenshots-report.json', JSON.stringify(results, null, 2));
console.log(`Finished ${results.filter((r) => r.ok).length}/${results.length} screenshots`);

Create urls.json:

[
  "https://example.com/a",
  "https://example.com/b"
]

Run the batch:

SELECTOR='[data-ready="true"]' node bulk-shot.mjs

The loop is sequential by design. A navigation, wait, or screenshot error is attached to that URL and the next URL is still attempted. The filenames use batch order, while the JSON report preserves the URL-to-file mapping. If input URLs may be duplicated or reordered between runs, use a sanitized stable identifier or a hash of each URL in the filename and retain the URL in the report.

Full page or target element?

page.screenshot({ fullPage: true }) captures the document. To capture only the matching element, use its handle after the wait:

const target = await page.waitForSelector(selector, { visible: true, timeout });
if (!target) throw new Error(`No element matched ${selector}`);
await target.screenshot({ path, type: 'png' });

Element capture is useful for a report card or chart. Full-page capture is better when surrounding context matters. If the target is inside a scrollable container, an element screenshot captures the element’s rendered bounds, not all off-screen content in that container. Puppeteer documents page and element screenshots in its screenshots guide.

3. Bulk capture with Playwright

Use Playwright when it is already part of your project or its browser support fits your needs. Its current guidance recommends locator waits and web-first assertions for many interactions. This example uses a locator wait, captures a full-page image, and records per-URL outcomes.

npm init -y
npm install playwright

Save as bulk-shot.mjs in a project configured for ES modules, or use the equivalent CommonJS imports:

import { chromium } from 'playwright';
import { readFile, mkdir, writeFile } from 'node:fs/promises';

const urls = JSON.parse(await readFile('urls.json', 'utf8'));
const selector = process.env.SELECTOR ?? '[data-ready="true"]';
const timeout = Number(process.env.WAIT_TIMEOUT_MS ?? 30_000);
const outputDir = 'screenshots';

if (!Array.isArray(urls) || urls.some((url) => typeof url !== 'string')) {
  throw new Error('urls.json must be an array of URL strings');
}
await mkdir(outputDir, { recursive: true });
const browser = await chromium.launch();
const results = [];

try {
  const page = await browser.newPage();
  for (const [index, url] of urls.entries()) {
    const path = `${outputDir}/page-${String(index + 1).padStart(4, '0')}.png`;
    try {
      const response = await page.goto(url, { waitUntil: 'domcontentloaded', timeout });
      if (response && response.status() >= 400) {
        throw new Error(`HTTP ${response.status()}`);
      }
      const target = page.locator(selector).first();
      await target.waitFor({ state: 'visible', timeout });
      await page.screenshot({ path, fullPage: true });
      results.push({ url, ok: true, path, status: response?.status() ?? null });
    } catch (error) {
      results.push({ url, ok: false, error: String(error) });
      console.error(`Screenshot failed for ${url}:`, error);
    }
  }
} finally {
  await browser.close();
}

await writeFile('screenshots-report.json', JSON.stringify(results, null, 2));
console.log(`Finished ${results.filter((r) => r.ok).length}/${results.length} screenshots`);

To save only the element, replace the page screenshot with await target.screenshot({ path });. Locator screenshots scroll the element into view and clip to its bounds. For a scrollable element, only the currently scrolled content is captured. See the Playwright screenshots guide and Locator API.

4. Waits, navigation, and page readiness

Choose a navigation condition and selector condition separately. For example, domcontentloaded waits for the initial document parse, then the selector wait covers the particular content needed for the image. A page’s network can remain busy because of analytics, polling, or long-lived requests; a selector tied to the required content is often a more useful readiness signal than waiting for all network traffic to stop.

  • Attached/present: use when existence in the DOM is sufficient.
  • Visible: use when the element must be rendered and visible to the user.
  • Text or app state: use a locator assertion or explicit condition when visibility alone does not mean the data is ready.
  • Navigation: set a bounded navigation timeout and handle navigation errors separately from selector timeouts if the report needs that detail.

Puppeteer’s default waitForSelector timeout is 30 seconds; it throws when the matching condition is not reached in time. You can configure the timeout per call or through page defaults. Avoid disabling timeouts in a batch unless another explicit deadline bounds each page and the overall job. Puppeteer Page.waitForSelector API

5. Batch design: output, retries, and concurrency

Keep results attributable

Store at least the input URL, output path, success flag, and error message for every item. For production jobs, also record timestamps, HTTP status when available, and attempt number. Do not report a timeout as a successful capture. If a page fails, preserve the error in the batch report so the failed subset can be retried.

Retry transient failures carefully

A retry can help with temporary navigation failures, but it will not fix a selector that never exists or a selector that changed after a site update. Retry only a small, bounded number of times, and consider retrying navigation errors separately from selector timeouts. Use a fresh page or reload when a failed attempt may have left the page in an uncertain state. Avoid retrying indefinitely.

Add concurrency only after measuring

Sequential processing is easier to diagnose and limits simultaneous browser work. If throughput is insufficient, use a small worker pool and tune it against the pages, memory, CPU, and browser capacity available in your environment. The official guidance cited here does not establish a universal concurrency limit or winner. Preserve a separate page/context per concurrent job, cap retries, and keep the same per-URL reporting. More concurrent pages can increase memory use and make resource contention or rate limits more likely.

6. cURL, Python, and Node.js alternatives

For self-hosted browser automation, the Puppeteer and Playwright examples above are the browser-based implementations. cURL by itself cannot render a page or wait for a browser element; use it for HTTP retrieval or to call a screenshot API that performs rendering. Python can drive a browser through a browser automation library, but no Python browser workflow is required for this title. The examples below show the API option in cURL, Python, and Node.js.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));

These API examples capture a URL with ScreenshotNeo. To make the capture depend on a selector, configure a selector wait using the supported options in the ScreenshotNeo API documentation. The call above illustrates a one-request capture; do not assume a selector wait unless you configure it.

7. Troubleshooting

Symptom Likely cause Fix
Selector wait times out Wrong selector, content is not loaded, or element is hidden Inspect the page DOM, use a stable selector, confirm whether presence or visibility is required, and set a realistic bounded timeout.
Wait succeeds but screenshot is blank or incomplete Element appeared before its data or images finished rendering Wait for a meaningful ready state or content condition; use full-page capture only when the whole document is needed.
Navigation timeout Page is slow or waits for a load condition that never completes Use a suitable navigation milestone such as domcontentloaded, then wait for the required element with its own deadline.
Element screenshot omits content Target is inside a scrollable container, or only part is rendered Capture the page or scroll the container deliberately and stitch only if the use case requires it.
Batch stops on one bad URL Error handling surrounds the whole loop rather than each URL Catch errors inside the per-URL iteration and write a result for every input.
Output belongs to the wrong URL Filename is not stable or results were reordered during concurrent processing Associate URL and output path in a report; use a stable per-item key rather than completion order.
Browser process remains open after failure Browser close is not in a finally block Use finally to close the browser even if navigation or capture throws.
Playwright locator strictness issue Selector matches multiple elements Make the selector specific, or deliberately choose the intended match, such as first().

8. Performance, reliability, and cost

In a self-hosted workflow, each page consumes browser time and local CPU and memory; full-page images can be larger and slower to produce than a crop. Reuse a browser process for a sequential batch, bound both navigation and element waits, and avoid capturing more pixels than you need. Concurrency can reduce wall-clock time when resources allow, but measure it with representative pages because the sources do not establish a universal safe setting.

Reliability depends on site-specific readiness signals, selector stability, bounded timeouts, per-URL failure records, and cleanup. Dynamic sites can change markup, delay content, or block automated traffic. Keep a retry policy bounded and distinguish a selector timeout from a navigation or HTTP failure.

Browser automation has no per-call screenshot API price in these examples, but it uses your own compute and requires browser setup and maintenance. A hosted screenshot API trades that setup for service pricing and request limits; check current product documentation and plan details before sizing a recurring batch.

9. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. It takes a URL and returns an image or PDF. Its API supports selector waits and other capture options; see the ScreenshotNeo documentation for parameter names and configuration. For a direct capture:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Cookie banners are accepted and removed before the shot, along with known newsletter popups and chat widgets; each of those steps can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for AI agents and MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up free for 1,000 screenshots a month, with no card required.

10. FAQ

Should I use Puppeteer or Playwright?

Use the framework already established in your project when it meets the browser needs. Both support waiting for elements and capturing pages or elements.

Does a visible element mean its data is ready?

No. Visibility confirms a rendered visible element, not that every asynchronous value inside it has finished loading. Wait for a site-specific ready condition when needed.

Can I capture several URLs at once?

Yes. Start sequentially, then introduce a measured worker pool if the batch runtime requires it. There is no universal concurrency value in the cited guidance.

Can I use cURL alone to wait for a page element?

No. cURL does not run a browser rendering engine. Use browser automation or a screenshot API with rendering and wait options.

Sources