ScreenshotNeo

BlogHow-to

How to Screenshot Thousands of URLs in Batches with a Rate Limit

Build a reliable screenshot pipeline for thousands of URLs with paced concurrency, durable progress tracking, retries, and clear cost controls.

By the ScreenshotNeo team4 October 202611 min read

To screenshot thousands of URLs without overwhelming your browser workers or the sites you visit, put URLs in a durable queue, limit concurrency and request pace, and track each URL as its own job. Save each result with metadata, retry only transient failures with bounded backoff, and ramp up gradually while watching errors and target-site behavior.

A batch is a way to submit work; it does not guarantee every URL starts at once or that limits are counted per batch. Provider limits and target-site limits are separate. There is no universal safe requests-per-minute rate or concurrency setting: it depends on the targets, capture mode, browser resources, service plan, and site behavior.

1. Choose an approach

Approach You operate Best fit
Playwright workers Browsers, queue, pacing, retries, storage, monitoring You need control over browser version, network, capture settings, or data locality.
Hosted screenshot API Usually the input queue, pacing policy, result collection, and monitoring You want to reduce browser infrastructure you manage. Check each provider’s batch semantics, limits, retention, retries, and pricing.

Playwright provides navigation and screenshot primitives, including viewport and full-page capture, but leaves batch orchestration to your application. Its screenshot guide and Page screenshot API document the capture options.

For a documented hosted example, ScreenshotRun says its batch endpoint accepts multiple URLs, counts quota and rate limits per URL, may queue work when a batch exceeds current per-minute capacity, and supports polling or per-screenshot webhooks. Those details apply to that provider; verify current terms before choosing any service.

2. Design the batch pipeline

  1. Normalize and deduplicate. Parse URLs, reject malformed or unsupported schemes, and decide whether query parameters and fragments distinguish pages for your use case.
  2. Persist before processing. Store each URL as a job in a durable queue or database so a process restart does not lose submitted work.
  3. Set two pacing controls. Cap simultaneous browser jobs globally and, where appropriate, per domain. Also cap job starts per interval to avoid bursts. Begin conservatively and adjust from observed results.
  4. Record each outcome. Track queued, running, succeeded, failed, and retryable states, plus attempt count, timestamps, final redirected URL, navigation response status or error, capture settings, and output path.
  5. Retry selectively. Respect Retry-After when present. Use bounded exponential backoff with jitter for transient network errors and timeouts. Do not endlessly retry invalid URLs, access-denied pages, or persistent bot checks.
  6. Use stable output identifiers. Key filenames by a stable job ID or URL hash, not raw URL text. Keep a manifest or sidecar record so results can be audited and regenerated.
  7. Ramp and monitor. Start with a small representative set, then increase concurrency and pace only while failure rates and target responses remain acceptable.

Check target-site terms and access policies before capturing at scale. A provider’s API allowance does not override restrictions or rate limits imposed by the destination site.

3. Runnable Playwright example in JavaScript

This minimal Node.js worker processes an input file with a global concurrency limit and a start interval. It records one JSON line per URL and uses bounded exponential backoff for navigation errors and retryable HTTP responses. It is a local demonstration, not a durable queue: for a long-running production batch, persist jobs and outcomes in a database or queue so restarts resume safely.

Install Playwright and its browser, then create urls.txt with one URL per line:

npm init -y
npm install playwright
npx playwright install chromium

Save as batch.mjs:

import { chromium } from 'playwright';
import { createHash } from 'node:crypto';
import { appendFile, mkdir, readFile } from 'node:fs/promises';

const INPUT = process.env.INPUT ?? 'urls.txt';
const OUT = process.env.OUT ?? 'shots';
const CONCURRENCY = Math.max(1, Number(process.env.CONCURRENCY ?? 2));
const START_INTERVAL_MS = Math.max(0, Number(process.env.START_INTERVAL_MS ?? 1500));
const MAX_ATTEMPTS = Math.max(1, Number(process.env.MAX_ATTEMPTS ?? 3));
const NAV_TIMEOUT_MS = Math.max(1000, Number(process.env.NAV_TIMEOUT_MS ?? 30000));
const FULL_PAGE = process.env.FULL_PAGE === '1';
const IMAGE_TYPE = process.env.IMAGE_TYPE ?? 'png';
if (!['png', 'jpeg'].includes(IMAGE_TYPE)) throw new Error('IMAGE_TYPE must be png or jpeg');

const raw = (await readFile(INPUT, 'utf8')).split(/\r?\n/).map(s => s.trim()).filter(Boolean);
const seen = new Set();
const urls = [];
for (const value of raw) {
  let u;
  try { u = new URL(value); } catch { console.error(`Skipping invalid URL: ${value}`); continue; }
  if (!['http:', 'https:'].includes(u.protocol)) { console.error(`Skipping unsupported scheme: ${value}`); continue; }
  u.hash = '';
  const normalized = u.toString();
  if (!seen.has(normalized)) { seen.add(normalized); urls.push(normalized); }
}
await mkdir(OUT, { recursive: true });
const browser = await chromium.launch({ headless: true });
let nextIndex = 0;
let nextStartAt = 0;

async function waitForStartSlot() {
  const now = Date.now();
  const startAt = Math.max(now, nextStartAt);
  nextStartAt = startAt + START_INTERVAL_MS;
  if (startAt > now) await new Promise(resolve => setTimeout(resolve, startAt - now));
}

function retryableStatus(status) {
  return status === 408 || status === 425 || status === 429 || status >= 500;
}

async function capture(url) {
  const id = createHash('sha256').update(url).digest('hex').slice(0, 20);
  const imagePath = `${OUT}/${id}.${IMAGE_TYPE === 'jpeg' ? 'jpg' : 'png'}`;
  let lastError = null;
  for (let attempt = 1; attempt <= MAX_ATTEMPTS; attempt++) {
    await waitForStartSlot();
    const page = await browser.newPage({ viewport: { width: 1365, height: 900 } });
    const startedAt = new Date().toISOString();
    try {
      const response = await page.goto(url, { waitUntil: 'domcontentloaded', timeout: NAV_TIMEOUT_MS });
      const status = response?.status() ?? null;
      if (status !== null && retryableStatus(status) && attempt < MAX_ATTEMPTS) {
        const retryAfter = Number(response.headers()['retry-after']);
        const backoff = Number.isFinite(retryAfter) && retryAfter > 0
          ? retryAfter * 1000
          : Math.min(30000, 1000 * (2 ** (attempt - 1))) * (0.5 + Math.random());
        await new Promise(resolve => setTimeout(resolve, backoff));
        lastError = `HTTP ${status}`;
        continue;
      }
      await page.screenshot({ path: imagePath, fullPage: FULL_PAGE, type: IMAGE_TYPE });
      const record = { url, finalUrl: page.url(), status, attempt, startedAt, finishedAt: new Date().toISOString(), imagePath, fullPage: FULL_PAGE, imageType: IMAGE_TYPE, error: null };
      await appendFile(`${OUT}/manifest.jsonl`, `${JSON.stringify(record)}\n`);
      return;
    } catch (error) {
      lastError = String(error?.message ?? error);
      if (attempt < MAX_ATTEMPTS) {
        const backoff = Math.min(30000, 1000 * (2 ** (attempt - 1))) * (0.5 + Math.random());
        await new Promise(resolve => setTimeout(resolve, backoff));
      } else {
        await appendFile(`${OUT}/manifest.jsonl`, `${JSON.stringify({ url, attempt, finishedAt: new Date().toISOString(), imagePath: null, error: lastError })}\n`);
      }
    } finally {
      await page.close();
    }
  }
}

async function worker() {
  while (true) {
    const index = nextIndex++;
    if (index >= urls.length) return;
    await capture(urls[index]);
  }
}
try {
  await Promise.all(Array.from({ length: Math.min(CONCURRENCY, urls.length) }, worker));
} finally {
  await browser.close();
}
console.log(`Processed ${urls.length} unique URLs. Results and manifest: ${OUT}/`);

Run with the defaults, or tune the environment variables after observing behavior:

node batch.mjs
CONCURRENCY=1 START_INTERVAL_MS=3000 MAX_ATTEMPTS=2 FULL_PAGE=1 node batch.mjs

The example’s limiter controls starts across all workers; it does not impose separate per-domain queues. For mixed domains, add a per-domain scheduler and a durable job store. Also note that Retry-After is handled only when it is a numeric number of seconds; a production implementation should parse both permitted header date and delay forms. The response status recorded here is the main navigation response, not a guarantee that every subresource loaded successfully.

Relevant Playwright capture options

Need Option or approach Trade-off
Visible viewport only Default screenshot Captures the initial viewport; usually smaller than a full-page image.
Entire scrollable document fullPage: true May take longer and produce larger images; validate pages with very long or dynamically growing content.
JPEG output type: 'jpeg', quality: 80 Smaller lossy output; quality applies to JPEG.
High-density pixels scale: 'device' Can increase pixel dimensions and storage. css uses one pixel per CSS pixel.
Specific region clip Useful for consistent comparisons; ensure the clip lies within the page.
Rendered content readiness Wait for a selector or a page-specific condition More reliable than a fixed delay for known pages, but site-specific selectors can fail.

See Playwright’s screenshot API for the current option names and behavior. Avoid choosing networkidle as a universal readiness signal: analytics, streaming, and long polling can keep network activity open. Use a selector or explicit condition when the content requirement is known.

4. Use a hosted API for managed capture

A hosted service can remove browser installation and operation from your workload, but you still need to understand its quota accounting, submission limits, result delivery, and retry behavior. Compare providers on maximum batch size versus actual processing capacity, per-URL or per-request charges, webhooks or polling, error detail, retention, capture settings, and current pricing. Verify these details directly because they differ by provider.

For ScreenshotRun specifically, its batch documentation describes per-URL quota and rate accounting, possible queuing when current capacity is exceeded, and polling or webhooks. Treat those as provider-specific behaviors and confirm current details in its documentation before relying on them.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. It accepts a URL in one GET request and returns an image or PDF. Its API supports 63 options, including full-page capture, format and viewport choices, custom CSS or JavaScript, waits, request blocking, caching, bulk capture of up to 100 URLs per call, async jobs with signed webhooks, and signed links. The API accepts parameter names used by other screenshot APIs to make switching easier. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie and consent banners are accepted and removed before the shot, along with known newsletter popups and chat widgets; each cleanup step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 screenshots.

Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.

5. Manage rate limits and failures

Separate the two kinds of limits

  • Provider/API limits: your plan may constrain request rate, concurrency, monthly quota, batch size, or asynchronous job volume. Read the provider’s current documentation and response headers; do not infer that submitting one batch consumes one unit.
  • Destination-site limits: sites may throttle, block, or disallow automated access. Watch status codes, challenge pages, and error patterns by domain. Slow or stop requests when responses indicate overload or access restrictions.

Classify before retrying

Outcome Typical action
Timeout, connection reset, temporary DNS failure Retry a small bounded number of times with backoff and jitter; inspect whether the target or your workers are the source.
429 or explicit retry guidance Honor Retry-After, reduce pace for that provider or domain, and avoid synchronized retries.
5xx response Retry cautiously; pause or reduce concurrency if failures cluster.
401/403, CAPTCHA, bot check, invalid URL Do not loop retries. Check credentials or access policy, or mark the URL for review.
Navigation succeeds but screenshot is blank or incomplete Check readiness conditions, redirects, script errors, consent overlays, and whether the page requires interaction or authentication.

Make jobs idempotent so a retry does not create confusing duplicate artifacts. Preserve each attempt in the manifest or job history, and publish the final status separately from attempt logs.

6. Performance, reliability, and cost

  • Concurrency and pace are different. Concurrency limits simultaneous work; a start interval limits how quickly jobs begin. Use both if you need to avoid bursts as well as cap browser resource use.
  • Full-page work costs more resources. Tall pages can take longer and create larger files than viewport captures. Lazy-loaded images may require scrolling or page-specific readiness logic. Measure representative URLs before estimating storage or completion time.
  • Browser memory is workload dependent. Close each page after capture, cap worker count, and monitor memory and browser crashes. Recycle browser processes if long runs show resource growth.
  • Storage grows with output volume. Choose PNG for lossless fidelity or JPEG when smaller lossy images are acceptable. Define retention, compression, and cleanup policies before a recurring crawl.
  • Cost depends on what is counted. For a hosted provider, establish whether usage is per submitted URL, successful screenshot, or API call, and account for retries and failed pages. Verify current pricing and quotas directly. For self-managed runs, include compute, storage, engineering time, and operational monitoring.
  • Do not promise a finish time from URL count alone. Page complexity, waits, full-page mode, rate constraints, target behavior, and retry volume all change throughput. Run a representative sample and extrapolate cautiously, then monitor the real batch.

7. Troubleshooting

Symptom Likely cause Fix
Many navigation timeouts Slow targets, overly short timeout, overloaded workers, or blocked traffic. Check domain-level patterns, reduce pace, increase timeout only for known slow pages, and stop retrying persistent blocks.
429 responses appear in waves Workers start in bursts or the provider/domain has a lower limit. Honor retry guidance, add per-domain pacing, lower concurrency, and avoid synchronized retry schedules.
Process restarts lose progress Work state existed only in memory. Persist queued, running, and terminal states before launch; recover stale running jobs on startup.
Overwritten or unsafe filenames Names were derived directly from URLs or collisions occurred. Use a stable hash or job ID and keep the original URL in metadata.
Screenshot misses content Capture began before client rendering or lazy loading completed. Wait for a meaningful selector or explicit condition; use full-page capture when needed and validate representative pages.
Images are unexpectedly huge Full-page mode or device scale generated many pixels. Use viewport capture, CSS scale, JPEG where acceptable, or a resize step; record settings with every output.
Retries make the overload worse Unbounded or synchronized retry loops. Bound attempts, add exponential backoff with jitter, respect retry headers, and pause the affected queue when errors rise.

8. FAQ

Should I put thousands of URLs into one API request?

Only if that provider documents a batch endpoint and its batch size, quota accounting, queuing behavior, and completion mechanism. Otherwise submit manageable jobs through your own queue.

Is there a universally safe requests-per-minute setting?

No. Target sites and providers have different limits, and workload resource needs vary. Start low, observe responses, and tune per domain and service.

Should I capture full pages by default?

No. Use viewport captures when the first screen is the comparison target. Use full-page mode when below-the-fold content matters, then account for greater time and storage.

Can I safely retry every failure?

No. Retry transient errors within a bound. Invalid URLs, access denials, and persistent bot checks need correction or manual review rather than repeated requests.

What should each saved screenshot include?

Keep a record with the source and final URL, capture time, response outcome, attempt count, settings, and image path. This makes failures explainable and reruns reproducible.