ScreenshotNeo

BlogHow-to

How to Make Bulk Website Screenshots with Puppeteer on a Linux VPS in India

Build a reliable bulk screenshot job with Puppeteer on an India-region Linux VPS, from Chrome setup and runnable code to retries and troubleshooting.

By the ScreenshotNeo team4 October 202612 min read

To capture many websites with Puppeteer on a Linux VPS, install a compatible Node.js version and Chrome for Testing, read URLs from a file, and process them through a bounded queue. For each URL, open a page, navigate with a readiness condition that suits the site, save a uniquely named screenshot, record any failure, and close the page. Start with low concurrency and increase it only after measuring the pages on your server.

This guide uses Ubuntu or Debian as the Linux example and an India-region server when that location suits your users or workload. Puppeteer’s current system requirements list Node.js 22.12 or newer and Chrome for Testing on Debian/Ubuntu and openSUSE/Fedora, for x64 and arm64. Requirements change, so check them for your operating system and architecture before deploying. Puppeteer system requirements

1. Choose what each screenshot should contain

Decide whether you need the visible viewport, the whole document, or one element. Puppeteer’s Page.screenshot() captures a page; ElementHandle.screenshot() captures a particular element. The fullPage option defaults to false. Set it to true only when you want the full document; a fixed viewport is usually more predictable for monitoring or visual comparisons. The API also supports a clip rectangle, output path, image type, and encoding. Puppeteer screenshot guide · Screenshot options

  • Viewport: captures the visible page at the configured viewport dimensions.
  • Full page: captures the full document, which can create very tall files and take longer to render.
  • Element: useful for a chart, card, or other specific component; make sure the selector exists and the element is visible.
  • Clip: captures a rectangle when you need a specific region.

For repeatable results, set the viewport before navigation. A screenshot can still differ across runs if the page content, fonts, animations, consent state, or external resources change.

2. Create the VPS and install browser dependencies

Select a Linux VPS based on the pages you will capture, expected concurrency, storage for images, and outbound network use. If location in India matters, AWS lists Mumbai (ap-south-1) and Hyderabad (ap-south-2) as India regions. This establishes that those regions are available; it does not establish a cheapest provider, a suitable instance size, or a performance advantage. Compare current provider pricing and measure your own workload.

On an Ubuntu or Debian server, install Node.js 22.12 or newer using a current supported installation method for your OS, then install the system libraries Chrome needs. Puppeteer’s troubleshooting documentation lists dependencies for Debian and Ubuntu; the package names below cover the common library and font requirements and should be checked against that guide when you change the base image.

sudo apt-get update
sudo apt-get install -y ca-certificates fonts-liberation libasound2t64 libatk-bridge2.0-0 libatk1.0-0 libc6 libcairo2 libcups2 libdbus-1-3 libexpat1 libfontconfig1 libgbm1 libgcc-s1 libglib2.0-0 libgtk-3-0 libnspr4 libnss3 libpango-1.0-0 libpangocairo-1.0-0 libstdc++6 libx11-6 libx11-xcb1 libxcb1 libxcomposite1 libxdamage1 libxext6 libxfixes3 libxrandr2 xdg-utils
node --version
npm --version

The exact package names can vary across OS releases. Consult Puppeteer troubleshooting and system requirements if a package is unavailable or Chrome reports a missing shared library.

Puppeteer normally downloads a compatible Chrome for Testing when you install the puppeteer package. Some package managers or deployment environments block install scripts, which can skip that download. Verify browser installation before sending a large batch.

mkdir -p ~/bulk-shots && cd ~/bulk-shots
npm init -y
npm install puppeteer
npx puppeteer browsers install chrome
node -e "const p=require('puppeteer'); console.log(p.executablePath())"

If your package manager disabled lifecycle scripts, explicitly run the browser install command shown above. Keep Puppeteer and its managed browser compatible by using the package’s install process, or manage the browser yourself and follow Puppeteer’s compatibility guidance.

Keep Chrome’s sandbox configured. Puppeteer strongly discourages running Chrome with --no-sandbox. Since this job opens arbitrary web content, do not treat disabling the sandbox as a routine VPS fix. Use an account and host configuration that allow Chrome’s sandbox to work, and investigate the actual permission or kernel issue if launch fails.

3. Prepare a URL list

Put one absolute HTTP or HTTPS URL per line in urls.txt. Use URLs you are permitted to access. The script below skips blank lines and comment lines beginning with #.

https://example.com
https://www.gov.in
# comments are ignored
https://www.wikipedia.org/

Use filenames derived from a stable URL hash rather than the hostname alone. A URL can contain a path or query that distinguishes it from another URL on the same host; hashes also avoid unsafe filename characters and collisions from simple slugging.

4. Run a bounded bulk capture script

Save the following as capture.mjs. It launches one browser, creates a fresh page per URL, applies a per-navigation timeout, writes each screenshot to a unique path, records success and failure in a JSON Lines manifest, and closes pages even when navigation or capture fails. The default concurrency is two workers; treat that as a cautious starting point, not a universal safe limit. Set CONCURRENCY after observing CPU, memory, and failures on your own VPS.

import fs from 'node:fs/promises';
import path from 'node:path';
import { createHash } from 'node:crypto';
import puppeteer from 'puppeteer';

const inputFile = process.argv[2] ?? 'urls.txt';
const outputDir = process.env.OUTPUT_DIR ?? 'screenshots';
const concurrency = Math.max(1, Number.parseInt(process.env.CONCURRENCY ?? '2', 10));
const navigationTimeout = Math.max(1000, Number.parseInt(process.env.NAV_TIMEOUT_MS ?? '45000', 10));
const fullPage = process.env.FULL_PAGE === '1';
const waitUntil = process.env.WAIT_UNTIL ?? 'domcontentloaded';
const allowedWaits = new Set(['load', 'domcontentloaded', 'networkidle0', 'networkidle2']);
if (!allowedWaits.has(waitUntil)) throw new Error(`Invalid WAIT_UNTIL: ${waitUntil}`);
if (!Number.isFinite(concurrency) || !Number.isFinite(navigationTimeout)) {
  throw new Error('CONCURRENCY and NAV_TIMEOUT_MS must be numbers');
}

const raw = await fs.readFile(inputFile, 'utf8');
const urls = raw.split(/\r?\n/).map((line) => line.trim())
  .filter((line) => line && !line.startsWith('#'));
for (const value of urls) {
  const parsed = new URL(value);
  if (!['http:', 'https:'].includes(parsed.protocol)) {
    throw new Error(`Only HTTP and HTTPS URLs are supported: ${value}`);
  }
}
await fs.mkdir(outputDir, { recursive: true });
const manifestPath = path.join(outputDir, 'results.jsonl');
const manifest = await fs.open(manifestPath, 'a');
const browser = await puppeteer.launch({ headless: true });
let nextIndex = 0;

async function worker() {
  while (true) {
    const index = nextIndex++;
    if (index >= urls.length) return;
    const url = urls[index];
    const id = createHash('sha256').update(url).digest('hex').slice(0, 20);
    const file = path.join(outputDir, `${String(index + 1).padStart(5, '0')}-${id}.png`);
    const startedAt = new Date().toISOString();
    let page;
    const record = { index, url, file, startedAt };
    try {
      page = await browser.newPage();
      await page.setViewport({ width: 1365, height: 900, deviceScaleFactor: 1 });
      page.setDefaultNavigationTimeout(navigationTimeout);
      const response = await page.goto(url, { waitUntil, timeout: navigationTimeout });
      record.httpStatus = response?.status() ?? null;
      record.ok = true;
      await page.screenshot({ path: file, type: 'png', fullPage });
    } catch (error) {
      record.ok = false;
      record.error = error instanceof Error ? error.message : String(error);
    } finally {
      if (page) await page.close().catch(() => {});
      record.finishedAt = new Date().toISOString();
      await manifest.appendFile(`${JSON.stringify(record)}\n`);
    }
  }
}

try {
  await Promise.all(Array.from({ length: Math.min(concurrency, urls.length) }, () => worker()));
} finally {
  await browser.close();
  await manifest.close();
}
console.log(`Processed ${urls.length} URL(s). Results: ${manifestPath}`);

The httpStatus field records the main document’s response status when available. A response such as 404 may still produce a screenshot; decide whether your pipeline should treat particular status codes as failures. The script captures only the initial page state. For pages that need a specific element or application state, add a site-appropriate wait before the screenshot.

Run it from the project directory:

node capture.mjs urls.txt
CONCURRENCY=1 FULL_PAGE=1 node capture.mjs urls.txt

Useful settings:

  • CONCURRENCY: number of simultaneous pages. Raise gradually while watching memory, CPU, navigation timeouts, and target-site responses.
  • NAV_TIMEOUT_MS: navigation timeout in milliseconds; default is 45 seconds.
  • WAIT_UNTIL: one of load, domcontentloaded, networkidle0, or networkidle2. The example defaults to domcontentloaded because some sites maintain network activity. Choose based on the page and add a more specific wait when needed.
  • FULL_PAGE=1: captures the full document; otherwise capture the viewport.
  • OUTPUT_DIR: output directory; defaults to screenshots.

For pages that lazy-load images, scroll in steps and wait for content before capture, or wait for a known selector. There is no universal wait that guarantees every third-party image and font has finished loading. For one element, locate it and use its screenshot method:

const element = await page.waitForSelector('.report-card', { timeout: 10000 });
if (!element) throw new Error('Report card was not found');
await element.screenshot({ path: file });

5. Retry failures and operate the job reliably

The sample records failures but does not retry them automatically. That keeps retries from silently multiplying a long batch and allows you to inspect failure patterns first. For transient DNS, connection, or server errors, retry only a small number of times with a delay that increases between attempts. Avoid retrying invalid URLs, blocked requests, or selector errors without changing the cause. Preserve attempt count and the final error in your manifest.

  • Bound the queue: do not create a page for every URL at once. The worker loop bounds active pages.
  • Set a job deadline: navigation timeouts constrain page loads; an external supervisor or job runner can also stop an entire stuck batch.
  • Clean up: close each page in finally and close the browser and manifest even after worker errors.
  • Make outputs traceable: keep the input URL, output filename, timestamps, status, and error together. Store the input list with the job for later audit.
  • Make reruns deliberate: the manifest opens in append mode. Use a new output directory per run or remove old outputs and manifest when you explicitly want a clean rerun.
  • Protect the host: run as a non-root service account where practical, retain Chrome’s sandbox, and restrict access to the output directory and credentials.

For a long-running production task, add process supervision, log rotation, disk-space monitoring, and a retention policy for screenshots. Treat remote pages as untrusted inputs and avoid passing sensitive credentials to URLs.

6. Size and locate the Linux VPS

There is no sourced universal minimum VPS size or safe concurrency number for bulk screenshots. Pages vary widely in script use, image size, fonts, and network behavior, while full-page capture can consume more memory than a viewport capture. Measure with a representative URL sample on the target server:

  1. Run a small batch at concurrency one and note completion time, peak memory, CPU use, output size, and failure rate.
  2. Repeat with the same URLs and capture settings at a modestly higher concurrency.
  3. Choose a level that leaves headroom for slow or unusually heavy pages, and rerun after changing the browser, OS image, or page mix.
  4. Estimate disk needs from measured average output size and retention duration; include room for temporary files and failed-job logs.

Choose an India region if it fits the location needs of the workload. AWS lists Mumbai and Hyderabad as India regions; compare those or other providers using geography, current costs, available CPU and memory, storage, outbound bandwidth, and how you will maintain Chrome dependencies and sandboxing. This guide does not establish current provider prices or recommend a provider or instance size.

7. Performance and cost considerations

Raising concurrency may improve total batch time when the job waits on network responses, but it also raises simultaneous browser memory use, CPU demand, outbound requests, and the chance of timeouts or rate limits. Full-page screenshots and pages with heavy client-side rendering can cost more time and memory than a small viewport capture. Measure with the same URLs and settings you intend to run.

For total operating cost, account for the VPS, storage, outbound bandwidth, engineering time to maintain Node.js and browser dependencies, and time spent investigating flaky pages. Cloud pricing and traffic costs change, so check the chosen provider’s current pricing rather than relying on an old estimate. If only occasional captures are needed, compare the cost of operating the browser stack with a screenshot API.

8. Troubleshooting common failures

Symptom Likely cause What to do
Could not find Chrome or browser executable missing The install script or browser download was skipped, or the install did not finish. Run npx puppeteer browsers install chrome, inspect the executable path with node -e "const p=require('puppeteer'); console.log(p.executablePath())", and ensure the deployed package and browser are compatible.
Chrome exits immediately or reports a shared library error A minimal OS image lacks a Chrome dependency. Install the dependencies listed for your OS in the Puppeteer troubleshooting guide. Inspect unresolved libraries with ldd /path/to/chrome | grep not.
Chrome refuses to start because of sandbox or permissions The host configuration prevents the browser sandbox from working. Configure the host and account so Chrome can use its sandbox. Puppeteer strongly discourages --no-sandbox; do not make that the default workaround.
Navigation timeout The site is slow, never reaches the chosen readiness condition, or has persistent network activity. Choose a readiness condition that matches the page, increase the timeout only when justified, and wait for a required selector or asset instead of assuming network idle is universal.
Screenshot is blank, incomplete, or missing images The page had not rendered the required content, images are lazy-loaded, or the site returned an interstitial or error page. Inspect the captured output and response status, wait for a page-specific selector, scroll to trigger lazy content when appropriate, and record unexpected page states for review.
Out of memory or browser crashes in a batch Too many active pages, unusually heavy sites, full-page captures, or insufficient host memory. Lower concurrency, use viewport capture when suitable, process smaller batches, and measure memory on representative pages before increasing load.
Duplicate files or overwritten results Filename generation is based on a non-unique hostname or multiple runs share an output path. Use a URL-derived unique identifier and a separate output directory per run. The sample includes the input index and a URL hash.
Different images across repeat runs The source page changed or rendered dynamic content, animations, personalized content, or time-sensitive data. Use a stable viewport and state, wait for the content you need, and note that screenshots reflect the page at capture time.

9. Or skip the browser setup

If you need a managed screenshot API instead of maintaining Chrome, [ScreenshotNeo](https://screenshotneo.com) returns an image or PDF from one GET request. Its cookie and consent handling accepts banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, with the page verdict and billing status returned in response headers. An MCP server provides screenshot tools for AI agents, including Claude and Cursor. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Replace YOUR_API_KEY with your key. The Node.js example requires a runtime with built-in fetch; inspect the response and save its body in your application as needed. Sign up for ScreenshotNeo’s free plan: 1,000 screenshots a month, no card required.

10. Frequently asked questions

Can I capture URLs that require login?

Yes, if you have authorized access. You would need to establish the required browser session or credentials before navigation and protect those secrets. Do not put sensitive credentials in the URL list or logs.

Should I use networkidle2 for every page?

No. The Puppeteer screenshot guide demonstrates it, but sites with persistent requests or delayed application content may never reach that condition or may still need a selector-specific wait. Pick readiness logic for each page type.

Can one job create PDFs instead of images?

Puppeteer supports page PDF generation through its PDF API. Use that when the required output is a document; this script is configured to write PNG screenshots.

Will a screenshot always be pixel-identical on rerun?

No. The captured page reflects its content and rendering at that moment. Dynamic data, personalization, fonts, browser versions, and external assets can change the result.