ScreenshotNeo

BlogHow-to

How to Batch Capture Website Screenshots with Puppeteer on a VPS in India

Build a reliable Node.js screenshot batch on an India-based Linux VPS, with bounded concurrency, retries, stable filenames, and deployment guidance.

By the ScreenshotNeo team4 October 202611 min read

To batch capture website screenshots with Puppeteer on a VPS in India, run a Node.js script on a supported Linux image, launch Puppeteer with its compatible browser, and process a URL list with a small concurrency limit. Give each URL a stable output filename, set a viewport, navigate with a suitable wait condition, save the screenshot, and record failures so one bad page does not stop the batch.

The VPS location determines where the browser runs and the network location sites see. The research for this guide did not verify any India-region VPS provider, price, regional routing behavior, or hosting performance. Choose an image and architecture supported by your Puppeteer and Chrome versions, and check the provider’s current region inventory directly.

1. Choose Puppeteer and prepare the VPS

Puppeteer is a JavaScript library for controlling Chrome or Firefox, and it runs headless by default. Page.screenshot() captures a page; ElementHandle.screenshot() captures a selected element. The puppeteer package normally downloads a compatible Chrome for Testing browser during installation. puppeteer-core does not manage the browser binary, so use it when you want to provide or connect to a browser yourself. See the Puppeteer documentation and installation guide.

Choice Browser handling Use it when
puppeteer Downloads a compatible browser by default. You want the straightforward install and do not already manage a browser binary.
puppeteer-core You provide or manage the browser executable and launch configuration. You need explicit browser management or connect to an externally managed browser.

The current Puppeteer system requirements documentation lists Node.js 22.12 or later and Chrome for Testing support on Debian/Ubuntu Linux x64 and arm64, and openSUSE/Fedora Linux x64 and arm64. Check the requirements and linked Chrome package dependencies against your exact VPS image and browser version; a minimal image may lack required system libraries. If package installation scripts are suppressed, use the documented manual browser installation steps. See Puppeteer system requirements.

  1. Provision a Linux VPS in the India region you need, then confirm the provider offers that region and the selected CPU architecture.
  2. Install a supported Node.js version and the OS libraries required by the browser.
  3. Create a working directory and install the package. For the browser-managed route, run npm init -y and npm install puppeteer.
  4. Create a URL file with one absolute URL per line. Start with a few representative pages before running a large batch.
  5. Ensure the account running the job can write to the output directory and has enough disk space for accumulated screenshots.

2. Create the URL list and batch script

The following runnable example reads newline-delimited URLs, validates that they are HTTP or HTTPS URLs, uses a fixed viewport, limits concurrent pages, retries navigation failures a bounded number of times, and writes a JSON Lines outcome manifest. It creates one browser for the batch, closes each page in a finally block, and closes the browser when workers finish.

// package.json should include: { "type": "module" }
// Save as capture.mjs. Install with: npm install puppeteer
import fs from 'node:fs/promises';
import path from 'node:path';
import crypto from 'node:crypto';
import puppeteer from 'puppeteer';

const inputPath = process.argv[2] ?? 'urls.txt';
const outputDir = process.argv[3] ?? 'screenshots';
const manifestPath = path.join(outputDir, 'results.jsonl');
const concurrency = Number(process.env.CONCURRENCY ?? 2); // Tune for your VPS.
const maxAttempts = 2; // Initial attempt plus at most one retry.
const timeoutMs = 60_000;

if (!Number.isInteger(concurrency) || concurrency < 1) {
  throw new Error('CONCURRENCY must be a positive integer');
}

const lines = (await fs.readFile(inputPath, 'utf8'))
  .split(/\r?\n/)
  .map((line) => line.trim())
  .filter((line) => line && !line.startsWith('#'));

const urls = lines.map((value, index) => {
  let parsed;
  try {
    parsed = new URL(value);
  } catch {
    throw new Error(`Invalid URL on input line ${index + 1}: ${value}`);
  }
  if (!['http:', 'https:'].includes(parsed.protocol)) {
    throw new Error(`Unsupported protocol on input line ${index + 1}: ${value}`);
  }
  return parsed.href;
});

await fs.mkdir(outputDir, { recursive: true });
await fs.writeFile(manifestPath, ''); // Start a fresh manifest for this run.
const browser = await puppeteer.launch();
let nextIndex = 0;

function outputName(url, index) {
  const digest = crypto.createHash('sha256').update(url).digest('hex').slice(0, 12);
  return `${String(index + 1).padStart(5, '0')}-${digest}.png`;
}

async function captureOne(url, index) {
  const filename = outputName(url, index);
  const target = path.join(outputDir, filename);
  let lastError;

  for (let attempt = 1; attempt <= maxAttempts; attempt += 1) {
    const page = await browser.newPage();
    try {
      await page.setViewport({ width: 1365, height: 900, deviceScaleFactor: 1 });
      const response = await page.goto(url, {
        waitUntil: 'domcontentloaded',
        timeout: timeoutMs,
      });
      if (response && response.status() >= 400) {
        throw new Error(`HTTP ${response.status()}`);
      }
      await page.screenshot({
        path: target,
        type: 'png',
        fullPage: true,
      });
      return { index: index + 1, url, status: 'success', file: filename, attempts: attempt };
    } catch (error) {
      lastError = error;
      await fs.rm(target, { force: true });
      if (attempt < maxAttempts) {
        await new Promise((resolve) => setTimeout(resolve, 500 * attempt));
      }
    } finally {
      await page.close().catch(() => {});
    }
  }
  return {
    index: index + 1,
    url,
    status: 'failed',
    attempts: maxAttempts,
    error: String(lastError?.message ?? lastError),
  };
}

async function worker() {
  while (true) {
    const index = nextIndex++;
    if (index >= urls.length) return;
    const result = await captureOne(urls[index], index);
    await fs.appendFile(manifestPath, `${JSON.stringify(result)}\n`);
    process.stdout.write(`${result.status}: ${result.url}\n`);
  }
}

try {
  await Promise.all(Array.from({ length: Math.min(concurrency, urls.length) }, worker));
} finally {
  await browser.close();
}

Run it with node capture.mjs urls.txt screenshots. Set a different worker count with CONCURRENCY=1 node capture.mjs. Each manifest line reports the input URL, outcome, output filename or error, and number of attempts. The filename combines the input order and a URL hash: reruns are easy to compare, and duplicate URLs at different positions do not overwrite one another.

Choose the page-ready condition deliberately

This example waits for domcontentloaded, which avoids requiring every network connection to become idle. It does not guarantee that every image, font, animation, or client-rendered component is ready. Use Puppeteer’s supported navigation conditions according to the page behavior:

  • load: wait for the page load event; useful when the page’s load handler marks the desired state.
  • domcontentloaded: proceed once the initial document is parsed; often useful for dynamic pages where background requests continue.
  • networkidle0 or networkidle2: wait for low network activity; can be unsuitable when pages keep requests or polling open.

For a known application, a selector is often a clearer readiness signal: after navigation, call await page.waitForSelector('main .report', { timeout: 15000 }). If you need lazy-loaded images, scroll through the page in controlled steps before capture and allow content to settle; full-page capture alone does not guarantee that every site lazy-loads every asset. These are workload-specific choices, not guarantees about third-party pages.

3. Pick capture size, format, and scope

Puppeteer’s screenshot API supports viewport screenshots, full-page screenshots, clipping, and file output. The relevant options include path to write a file, fullPage to capture the full document, type for formats supported by the installed browser, and clip for a specific rectangle. Consult the ScreenshotOptions API for the version you install.

// Viewport image (omit fullPage)
await page.screenshot({ path: 'viewport.png', type: 'png' });

// Whole document
await page.screenshot({ path: 'full-page.png', type: 'png', fullPage: true });

// Region in CSS pixels
await page.screenshot({
  path: 'region.png',
  type: 'jpeg',
  quality: 85,
  clip: { x: 0, y: 0, width: 900, height: 600 },
});

// A specific element
const card = await page.$('.pricing-card');
if (!card) throw new Error('Pricing card was not found');
await card.screenshot({ path: 'pricing-card.png', type: 'png' });

Use PNG when you need lossless output, JPEG with a quality setting when a smaller photographic image is acceptable, or another format only if supported by your browser and chosen Puppeteer version. Keep output extensions consistent with the actual format. For deterministic comparisons, hold the viewport, device scale factor, browser version, wait rule, and any page-specific setup steady. Dynamic ads, timestamps, rotating content, and font loading can still cause visual differences.

4. Tune concurrency and make runs reliable

There is no universal concurrency value. Each open page consumes browser processes, memory, CPU, network bandwidth, and disk write capacity, while some sites take longer or continue loading indefinitely. Begin with a low limit, observe the VPS under a representative batch, then raise it only if memory remains available and failure rates stay acceptable. Higher concurrency may shorten completion time but can increase resource pressure and load on target sites.

  • Keep a hard worker limit; do not create one page per URL all at once.
  • Use bounded retries only for transient errors. Repeatedly retrying a persistent 404, access denial, or CAPTCHA wastes resources.
  • Preserve per-URL outcomes so a failed URL can be rerun without losing successful captures.
  • Set navigation and selector timeouts. A timeout bounds how long one troublesome page can occupy a worker.
  • Close pages and the browser in cleanup paths. Restart the process between scheduled batches if you observe browser state or resource accumulation.
  • Monitor free disk space and rotate or archive old screenshots and logs.
  • Respect target site policies and avoid excessive request rates.

For scheduled operation, run the script through the VPS’s service manager or scheduler, send stdout and stderr to persistent logs, and alert on nonzero failures or low disk. Store credentials outside the script if pages require authentication, and restrict access to screenshot output if it contains private data.

5. Deploy safely on Linux

Run Puppeteer under a dedicated, least-privileged account where practical. Browser sandboxing is a security boundary; do not treat --no-sandbox as a routine deployment fix. If Chrome reports a sandbox error, first determine whether the environment supports the required sandbox and configure it correctly. Puppeteer’s Linux troubleshooting documentation discusses sandbox-related errors and the discouraged nature of disabling the sandbox; see Puppeteer troubleshooting.

Before scaling up, capture a small set containing a short page, a long page, a JavaScript-heavy page, and a page with lazy-loaded images. Check the output dimensions and content, then adjust the wait condition or scrolling logic for the actual sites. The code is a practical template; it is not a benchmark or confirmation that any site permits automated capture.

6. Troubleshooting

Symptom Likely cause Fix
Browser executable missing at launch Browser download did not run, or puppeteer-core was installed without a browser path. Use puppeteer with its install flow, follow the official manual browser installation guide if install scripts were blocked, or configure the managed executable when using puppeteer-core.
Chrome fails with missing shared library The VPS image lacks an OS dependency required by Chrome. Check Puppeteer’s system requirements and the Chrome dependency list for the distribution and architecture; install the missing package and relaunch.
Sandbox error or browser exits immediately The process environment cannot initialize the browser sandbox. Configure a supported sandbox environment and permissions. Avoid disabling the sandbox as a default workaround.
Navigation timeout The page is slow, a resource never settles, the timeout is too short, or the selected wait condition is too strict. Use a bounded timeout and a more appropriate condition such as domcontentloaded; wait for a meaningful selector if the application has one. Investigate connectivity separately.
Screenshot is blank or incomplete Capture ran before client rendering or a required element appeared; lazy content may not have loaded. Wait for the relevant selector, scroll to trigger lazy loading, or use a suitable delay after the page is ready. Check the recorded navigation status.
Worker exits with out-of-memory or browser crashes Concurrency is too high for the VPS, or pages are unusually resource intensive. Lower CONCURRENCY, close pages promptly, and inspect memory use before increasing the limit again.
Output files overwrite or are missing Names collide, output directory is not writable, or the process stopped before writing. Use unique names such as the indexed hash in this example, check permissions and disk space, and inspect the JSONL manifest and process logs.
Some sites show CAPTCHA or deny access The site’s access controls rejected the request or require an allowed authenticated session. Follow the site’s access rules; do not attempt to bypass bot protections. Exclude the URL or use an authorized workflow.

7. Cost and performance considerations

Self-hosting shifts the work to your VPS: account for the server, storage, bandwidth, operational time, and maintenance of Node.js, Puppeteer, browser binaries, and OS dependencies. The research does not establish provider prices, benchmark throughput, or a recommended machine size. Measure a representative batch on the exact VPS image and sites you intend to use.

Keep image dimensions and format aligned with downstream needs. Full-page images can be much larger than viewport captures, and retaining every run indefinitely increases disk use. A lower worker count may take longer but reduce memory contention and avoid sending bursts of requests. An India-based server can be useful when the browser should run from India, but actual routing and site behavior depend on the VPS network and target site.

8. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. One GET request returns an image or PDF. Cookie banners are accepted like a visitor and 60+ known consent platforms, newsletter popups, and chat widgets are removed before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server gives Claude, Cursor, and other MCP clients the take_screenshot, get_page_info, and capture_pdf tools.

For a managed capture, use the API key from your account. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo supports bulk capture of up to 100 URLs per call, full-page capture, element selection, device presets and custom viewports, image formats, caching, and more. Every feature is on every plan. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Sign up for 1,000 free screenshots a month, with no card.

Frequently asked questions

Does running the browser on an India VPS guarantee an India-local result?

It places the browser process on that server, but the research does not verify a provider’s regional routing or guarantee how a target site maps the request. Confirm region and network details with the provider and site behavior with an authorized capture.

How many screenshots can Puppeteer run at once?

There is no universal safe number. Start with a small concurrency limit, watch memory, CPU, disk, and failure rate on your workload, then tune gradually.

Can I capture only one element instead of a whole page?

Yes. Find the element with a selector and call its screenshot() method, as shown above. Use a page screenshot with a clip when you need a fixed rectangular region.

Should I use networkidle2 for every URL?

No single wait condition fits every site. Pages with polling or long-lived requests may never become idle, while a dynamic page may need a selector wait or a deliberate delay after navigation.