ScreenshotNeo

BlogHow-to

How to Bulk Screenshot URLs with a Proxy in Puppeteer

Build a Node.js batch that screenshots a URL list through a Puppeteer proxy, with bounded concurrency, unique filenames, and per-URL results.

By the ScreenshotNeo team4 October 202611 min read

To bulk screenshot URLs through a proxy in Puppeteer, launch Chromium with a proxy argument, then process your URL list with a bounded worker queue. Each worker navigates to one URL, saves a screenshot under a collision-resistant filename, records the result, and closes its page. Puppeteer documents the individual capture operation as Page.screenshot(); the batch loop is application code you build around it.

The example below uses Node.js and Puppeteer, reads one URL per line from urls.txt, routes browser traffic through an HTTP proxy, and writes screenshots and a JSON manifest to an output directory. See the Puppeteer screenshot guide and launch options reference.

1. Install Puppeteer and prepare the URL list

Use a project-local Puppeteer install so the package and browser setup can be managed with the project. Pin the version selected for your application in its lockfile; no particular version is required by this pattern. Puppeteer normally downloads a compatible browser during installation. In a deployment environment, confirm that installation has run and that the runtime can launch the browser.

mkdir screenshot-batch
cd screenshot-batch
npm init -y
npm install puppeteer

Create urls.txt with one absolute URL per line. Blank lines and lines beginning with # are ignored.

https://example.com/
https://www.iana.org/
# Add another URL on its own line

The script accepts only HTTP and HTTPS URLs. It rejects invalid lines before starting Chromium, so typos do not consume browser work. Only run captures for sites you are authorized to access, and follow the site’s terms and your proxy provider’s rules.

2. Run the bulk screenshot script

Save this as capture.mjs. Set PROXY_SERVER to the proxy host and port, such as http://proxy.example:8080. If the proxy requires HTTP Basic credentials, set PROXY_USERNAME and PROXY_PASSWORD in the environment. Do not put credentials in the source file or print them in logs.

import puppeteer from 'puppeteer';
import { readFile, mkdir, writeFile } from 'node:fs/promises';
import { createHash } from 'node:crypto';
import path from 'node:path';

const INPUT_FILE = process.env.URLS_FILE ?? 'urls.txt';
const OUTPUT_DIR = process.env.OUTPUT_DIR ?? 'screenshots';
const CONCURRENCY = Math.max(1, Number.parseInt(process.env.CONCURRENCY ?? '2', 10) || 2);
const NAVIGATION_TIMEOUT_MS = Math.max(1000, Number.parseInt(process.env.NAVIGATION_TIMEOUT_MS ?? '45000', 10) || 45000);
const PROXY_SERVER = process.env.PROXY_SERVER;
const PROXY_USERNAME = process.env.PROXY_USERNAME;
const PROXY_PASSWORD = process.env.PROXY_PASSWORD;

function parseUrls(text) {
  const urls = text.split(/\r?\n/)
    .map((line) => line.trim())
    .filter((line) => line && !line.startsWith('#'));

  const invalid = [];
  const valid = [];
  for (const raw of urls) {
    try {
      const parsed = new URL(raw);
      if (parsed.protocol !== 'http:' && parsed.protocol !== 'https:') {
        throw new Error('Only http and https URLs are supported');
      }
      valid.push({ requestedUrl: raw, normalizedUrl: parsed.href });
    } catch (error) {
      invalid.push({ requestedUrl: raw, status: 'invalid', error: error.message });
    }
  }
  return { valid, invalid };
}

function outputName(url) {
  const parsed = new URL(url);
  const host = parsed.hostname.replace(/[^a-z0-9.-]/gi, '_').slice(0, 80) || 'page';
  const slug = parsed.pathname.split('/').filter(Boolean).join('-')
    .replace(/[^a-z0-9_-]/gi, '_').slice(0, 50) || 'home';
  const hash = createHash('sha256').update(url).digest('hex').slice(0, 12);
  return `${host}-${slug}-${hash}.png`;
}

async function main() {
  const input = await readFile(INPUT_FILE, 'utf8');
  const { valid, invalid } = parseUrls(input);
  if (valid.length === 0) {
    throw new Error(`No valid HTTP(S) URLs found in ${INPUT_FILE}`);
  }
  await mkdir(OUTPUT_DIR, { recursive: true });

  const args = PROXY_SERVER ? [`--proxy-server=${PROXY_SERVER}`] : [];
  const browser = await puppeteer.launch({ headless: true, args });
  const results = [...invalid];
  let next = 0;

  async function worker() {
    while (true) {
      const index = next++;
      if (index >= valid.length) return;
      const item = valid[index];
      const file = outputName(item.normalizedUrl);
      const outputPath = path.join(OUTPUT_DIR, file);
      let page;
      try {
        page = await browser.newPage();
        if (PROXY_USERNAME || PROXY_PASSWORD) {
          if (!PROXY_USERNAME || !PROXY_PASSWORD) {
            throw new Error('Set both PROXY_USERNAME and PROXY_PASSWORD for proxy authentication');
          }
          await page.authenticate({ username: PROXY_USERNAME, password: PROXY_PASSWORD });
        }
        await page.setViewport({ width: 1365, height: 900, deviceScaleFactor: 1 });
        page.setDefaultNavigationTimeout(NAVIGATION_TIMEOUT_MS);
        const response = await page.goto(item.normalizedUrl, {
          waitUntil: 'domcontentloaded',
          timeout: NAVIGATION_TIMEOUT_MS,
        });
        await page.screenshot({ path: outputPath, type: 'png', fullPage: true });
        results.push({
          requestedUrl: item.requestedUrl,
          finalUrl: page.url(),
          status: 'captured',
          httpStatus: response?.status() ?? null,
          outputPath,
        });
      } catch (error) {
        results.push({
          requestedUrl: item.requestedUrl,
          status: 'failed',
          outputPath: null,
          error: error instanceof Error ? error.message : String(error),
        });
      } finally {
        if (page) await page.close().catch(() => {});
      }
    }
  }

  try {
    await Promise.all(Array.from({ length: Math.min(CONCURRENCY, valid.length) }, () => worker()));
  } finally {
    await browser.close();
  }

  const manifestPath = path.join(OUTPUT_DIR, 'manifest.json');
  await writeFile(manifestPath, `${JSON.stringify(results, null, 2)}\n`);
  const successes = results.filter((result) => result.status === 'captured').length;
  console.log(`Captured ${successes}/${results.length} entries. Manifest: ${manifestPath}`);
  if (successes !== results.length) process.exitCode = 1;
}

main().catch((error) => {
  console.error(error instanceof Error ? error.message : error);
  process.exitCode = 1;
});

Set the proxy and run the job. These examples use shell environment variables so secrets stay outside the script; use your platform’s secret manager for scheduled or deployed jobs.

# macOS/Linux
PROXY_SERVER='http://proxy.example:8080' \
PROXY_USERNAME='your-user' PROXY_PASSWORD='your-password' \
CONCURRENCY=2 node capture.mjs

# Windows PowerShell
$env:PROXY_SERVER = 'http://proxy.example:8080'
$env:PROXY_USERNAME = 'your-user'
$env:PROXY_PASSWORD = 'your-password'
$env:CONCURRENCY = '2'
node capture.mjs

Without a proxy, omit PROXY_SERVER. For an unauthenticated proxy, set only PROXY_SERVER. For other proxy authentication schemes or provider-specific session controls, check the provider’s instructions and validate them with the exact proxy and browser setup. Proxy authentication behavior can vary. A third-party Puppeteer proxy guide shows the launch argument approach and discusses authentication; treat its package-specific workarounds as optional and verify compatibility before using them.

3. Understand proxy routing and authentication

The --proxy-server argument is passed to Chromium at browser launch. It routes browser requests for all pages in that browser through the selected proxy. This is a browser-wide setting for this launch. If URLs need different proxy routes, run separate browser instances with different launch arguments, or adopt a proxy architecture that supports the required routing and verify its behavior. Do not assume a Puppeteer installation/download proxy setting routes page traffic.

Puppeteer’s configuration guide describes environment proxy settings in the context of Puppeteer setup and browser download. The browser’s route to target sites is configured separately here through its launch argument. In particular, puppeteer-core does not use Puppeteer configuration; its launch also needs a browser executable path or channel, as described in the launch method reference.

Test the route with a URL you are permitted to use that reports the request’s egress address, then inspect the target page’s behavior. A successful navigation alone does not prove that the intended proxy or location was used.

4. Choose capture, wait, and isolation settings

Setting Example choice When to change it
waitUntil domcontentloaded Use a different navigation lifecycle event when your page requires more loading before capture. A page may continue loading images or client-rendered content after this event; add a deliberate selector wait or delay when needed.
Navigation timeout 45 seconds in the sample Raise or lower it for your network, targets, and job deadline. Timeout does not mean the page is safe to retry without limit.
fullPage true Set to false for viewport-only shots. Full-page images can be much taller and larger, and some pages load content as you scroll.
Viewport 1365 × 900, scale 1 Set consistent dimensions for comparisons. Change device scale factor if a higher-density capture is needed, while accounting for larger output.
Format PNG Puppeteer supports screenshot type selection such as PNG, JPEG, and WebP subject to the API/browser support. PNG is lossless; JPEG/WebP can reduce size when lossy output is acceptable. Quality applies to lossy formats, not PNG.
Output names Host, path slug, URL hash The hash prevents collisions for repeated paths, query variants, and distinct URLs that normalize to similar names. Do not use raw URLs as filesystem paths.
Session isolation New page per URL Each page here is fresh, but pages created in the same browser share its default browser context state. If cookies or storage must not cross URLs, use isolated browser contexts and close each context after its capture.

The documented screenshot options include path, type, quality, clip, fullPage, and background handling. fullPage defaults to false; quality is not applicable to PNG. See the full ScreenshotOptions reference. For an element-only capture, wait for the element and use ElementHandle.screenshot(), as shown in the screenshot guide.

5. Tune reliability and batch behavior

  • Bound concurrency. The sample uses two workers by default. There is no universal safe concurrency number: tune against page size, browser memory, proxy capacity, target response behavior, and your own measurements. Increasing it can reduce elapsed time while increasing resource pressure.
  • Keep failures visible. The manifest records invalid input and per-URL navigation or capture errors. A non-zero process exit code signals that one or more entries did not produce a capture.
  • Retry selectively. The sample does not retry. If you add retries, limit the count, record each attempt, and retry only failures you consider transient. Avoid retry loops for authentication failures, invalid URLs, or access-denied responses.
  • Choose what an HTTP error means. page.goto() can return a response with an HTTP error status. The sample records that status and still captures the resulting page. Change that policy if error pages should be marked failed in your workflow.
  • Close resources. Each page is closed in a finally block, and the browser closes even if workers reject. Keep the manifest and completed images when a batch has partial failures so you can resume deliberately.
  • Respect target limits. Use low concurrency and per-host pacing where appropriate. Proxy use does not grant permission to bypass access controls or rate limits.

For large URL lists, use a durable queue and persist each result as it completes instead of holding all results until process exit. Keep input URL, final redirected URL, status, output path, and error in that record. That makes interrupted jobs easier to audit and resume without recapturing completed entries.

6. Troubleshoot common failures

Symptom Likely cause Fix
Browser fails to launch Browser download did not run, required runtime dependencies are missing, or the configured executable is unavailable. Run the Puppeteer browser installation step for your environment and check that the bundled browser can launch there. With puppeteer-core, provide executablePath or channel.
Proxy connection error or navigation timeout Wrong host/port or scheme, unavailable proxy, blocked egress, or slow target. Confirm the proxy endpoint and scheme with its provider, test a permitted destination, and adjust the navigation timeout only when longer waits are expected.
407 Proxy Authentication Required Credentials were not supplied, are wrong, or the proxy authentication method is not handled by the setup. Check both environment variables and provider requirements. Verify support with your exact proxy configuration; do not assume Basic authentication handling applies to every scheme.
Target returns 403, challenge, or CAPTCHA The target declined the request or requires a permitted interaction or identity. Respect the target’s access rules. Confirm the proxy and account are authorized for the workflow; do not use proxy rotation to evade an access restriction.
Screenshot is blank or missing late content Capture happened before the page’s client rendering, images, or lazy content completed. Wait for a stable selector or a site-specific readiness condition. For lazy-loaded full pages, consider a controlled scroll-and-wait routine, then capture; test it against the page because pages implement loading differently.
Duplicate or overwritten files Filenames based only on hostname or path collide for query variants or repeated paths. Keep a digest of the full normalized URL in the filename, as in the sample, and avoid concurrent runs sharing the same output directory unless names or run directories are isolated.
Memory grows or the job slows at higher concurrency Several large documents and full-page images are resident at once. Reduce worker count, choose viewport captures where sufficient, and measure memory and duration on representative pages before raising concurrency.
Some URLs work while others fail Per-site redirects, network behavior, TLS, content restrictions, or provider routing differ. Use the manifest to identify the failed URL and error, test it individually through the same proxy, and adjust only the relevant timeout, wait, or authorized access configuration.

7. Or skip the browser setup

If your goal is simply to get screenshots from a list of URLs, ScreenshotNeo is a website screenshot API and MCP server. A GET request returns a PNG, JPEG, WebP, or PDF. The same parameter names used by other screenshot APIs work, which can make switching straightforward. See the ScreenshotNeo API documentation for parameters and response details.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
  • Cookie banners are accepted and 60+ known consent platforms are removed before the shot; newsletter popups and chat widgets are removed too. Each step can be turned off.
  • Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are never billed. Responses identify the page verdict and billing status in headers.
  • An MCP server gives AI agents tools named take_screenshot, get_page_info, and capture_pdf.
  • The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots.

Sign up free for 1,000 screenshots a month, with no card required.

FAQ

Does Puppeteer have a built-in bulk screenshot command?

No. Use application code to read the URL list and call Page.screenshot() for each page.

Can the proxy change for every URL?

The launch argument configures the browser route for that browser instance. Use separate launches for separate proxy routes, or validate a routing architecture that supports your use case.

Should every capture use a new browser?

Usually a single browser with controlled pages is simpler and avoids repeatedly starting Chromium. Use separate browser instances when you need separate launch-level proxy routes or stronger process isolation.

Will networkidle always produce a complete screenshot?

No. Some pages keep network connections open or render content after network activity settles. Wait for a page-specific selector or readiness condition when completeness matters.

How many pages can I capture at once?

There is no universal documented concurrency value. Start low and tune against the proxy, target behavior, memory, and output size for your workload.