How to Capture Bulk Website Screenshots on a Low-Budget VPS in India
Build a repeatable bulk screenshot pipeline with Playwright on a low-budget India VPS, with bounded concurrency, retries, and realistic capacity planning.
Use a headless browser such as Playwright on a Linux VPS, feed it URLs from a file, and capture each page to a unique filename. Keep concurrency low, set navigation and capture timeouts, retry transient failures, and log every result. There is no universal page-count capacity: page weight, JavaScript, capture mode, browser version, memory, CPU, disk, transfer limits, and network conditions all affect throughput.
This guide builds a runnable Node.js batch script for Chromium, explains how to install it on a VPS, and shows how to estimate capacity with a representative pilot. The AWS Lightsail $5/month Linux bundle is one published price reference, not a capacity recommendation: its listed configuration is 0.5 GB memory, 2 vCPUs, 20 GB SSD, and 1 TB transfer, with the Mumbai region receiving half the displayed transfer allowance. [AWS Lightsail pricing]
1. Choose the workflow and VPS carefully
Playwright is a reasonable default for scripted Chromium screenshots. It supports viewport screenshots, full-page screenshots, element screenshots, and buffers; its browser support includes Chromium, Firefox, and WebKit on Linux, macOS, and Windows. [Playwright screenshots]
For a batch job, the basic pipeline is:
- Read one URL per line from a text file.
- Validate and normalize each URL.
- Open a page with a defined navigation and readiness policy.
- Capture to a predictable, collision-resistant filename.
- Record the outcome and error for each URL.
- Retry only transient failures, with a limit and delay.
- Review memory, CPU, disk, transfer use, and target-site responses before raising concurrency.
When comparing VPS plans, compare monthly cost and billing rules, RAM and CPU, disk space for retained images, bandwidth and regional adjustments, region and IP requirements, and throughput measured on your own page mix. AWS says Lightsail bundles are charged at a fixed hourly rate up to the maximum monthly plan cost. [AWS Lightsail billing]
A low price alone does not establish that an instance can run a particular batch. In particular, 0.5 GB RAM is constrained for a browser workload; use a small instance as a cautious pilot and be prepared to use a larger plan or a managed service. Do not convert the USD price to INR without checking a current exchange rate, taxes, and billing currency.
2. Install headless Playwright on Linux
Use a supported Node.js release available for your Linux distribution, then install the project and Chromium dependencies. Playwright documents that Linux system dependencies must be installed; headless execution is the default, while headed Linux execution requires Xvfb. [Playwright in CI]
mkdir bulk-shots
cd bulk-shots
npm init -y
npm install playwright
npx playwright install --with-deps chromium
Keep the Playwright package and browser binaries installed together: when you update Playwright, run its browser installation command again so the expected browser version is available. On a minimal VPS, installation may need an account with permission to install OS packages. If your distribution is not supported by the installer, follow Playwright’s Linux dependency documentation for the specific distribution.
Create urls.txt with one authorized target per line:
https://example.com/
https://example.org/
https://www.example.net/catalog
3. Runnable Node.js batch script
Save this as capture.mjs. It uses a fixed worker limit, per-navigation timeout, explicit readiness condition, bounded retries, unique output names, and a JSON Lines log. The default is viewport capture; pass --full-page for full-page captures.
import { chromium } from 'playwright';
import { createHash } from 'node:crypto';
import { mkdir, readFile, appendFile } from 'node:fs/promises';
import path from 'node:path';
const args = new Set(process.argv.slice(2));
const fullPage = args.has('--full-page');
const concurrency = Math.max(1, Number(process.env.CONCURRENCY ?? 1));
const timeoutMs = Math.max(1000, Number(process.env.TIMEOUT_MS ?? 30000));
const maxAttempts = Math.max(1, Number(process.env.ATTEMPTS ?? 3));
const outputDir = process.env.OUTPUT_DIR ?? 'screenshots';
const logFile = process.env.LOG_FILE ?? 'results.jsonl';
if (!Number.isInteger(concurrency) || concurrency > 8) {
throw new Error('Set CONCURRENCY to an integer from 1 to 8; raise it only after a pilot.');
}
const raw = await readFile('urls.txt', 'utf8');
const urls = raw.split(/\r?\n/).map(s => s.trim()).filter(s => s && !s.startsWith('#'));
const jobs = urls.map((value, index) => {
let u;
try { u = new URL(value); } catch { throw new Error(`Invalid URL on input line ${index + 1}: ${value}`); }
if (!['http:', 'https:'].includes(u.protocol)) throw new Error(`Unsupported protocol on line ${index + 1}`);
return u.href;
});
await mkdir(outputDir, { recursive: true });
await mkdir(path.dirname(logFile), { recursive: true }).catch(() => {});
const browser = await chromium.launch({ headless: true });
let next = 0;
async function capture(url) {
const key = createHash('sha256').update(url).digest('hex').slice(0, 16);
const filename = `${String(jobs.indexOf(url) + 1).padStart(5, '0')}-${key}.png`;
const target = path.join(outputDir, filename);
let lastError = '';
for (let attempt = 1; attempt <= maxAttempts; attempt++) {
const context = await browser.newContext({ viewport: { width: 1365, height: 900 }, deviceScaleFactor: 1 });
const page = await context.newPage();
page.setDefaultNavigationTimeout(timeoutMs);
page.setDefaultTimeout(timeoutMs);
try {
const response = await page.goto(url, { waitUntil: 'domcontentloaded', timeout: timeoutMs });
if (!response) throw new Error('Navigation returned no main-resource response');
if (response.status() >= 400) throw new Error(`HTTP ${response.status()}`);
await page.screenshot({ path: target, fullPage });
await appendFile(logFile, JSON.stringify({ url, status: 'ok', httpStatus: response.status(), file: target, attempt }) + '\n');
return;
} catch (error) {
lastError = String(error?.message ?? error);
if (attempt < maxAttempts) await new Promise(resolve => setTimeout(resolve, Math.min(1000 * 2 ** (attempt - 1), 8000)));
} finally {
await context.close();
}
}
await appendFile(logFile, JSON.stringify({ url, status: 'failed', error: lastError, attempts: maxAttempts }) + '\n');
}
async function worker() {
while (true) {
const index = next++;
if (index >= jobs.length) return;
await capture(jobs[index]);
}
}
try {
await Promise.all(Array.from({ length: Math.min(concurrency, jobs.length) }, worker));
} finally {
await browser.close();
}
Run it with one worker first:
node capture.mjs
CONCURRENCY=2 TIMEOUT_MS=45000 ATTEMPTS=3 node capture.mjs --full-page
Set CONCURRENCY to a conservative integer and increase only after measuring. The script intentionally reports HTTP status codes of 400 and above as failures. If a site uses an expected status or redirect behavior that differs, adjust that policy deliberately rather than silently treating all navigations as successful.
What the capture settings mean
| Setting | Effect | Trade-off |
|---|---|---|
waitUntil: 'domcontentloaded' |
Starts capture after initial HTML has been parsed. | Fast, but late client-rendered content may be missing. |
load |
Waits for the load event. | Can take longer or time out on pages with slow resources. |
networkidle |
Waits for network activity to settle. | Some sites keep connections or polling active, so it may never be suitable. |
fullPage: true |
Captures the full document height. | Tall pages can use more memory and produce large images. |
| Viewport capture | Captures the visible viewport. | Predictable dimensions and generally less work than a very tall page. |
| Element capture | Captures a selected element via locator screenshot. | Requires the selector to exist and the element to be visible. |
To wait for a known rendered element, insert await page.locator('main').waitFor({ state: 'visible', timeout: timeoutMs }); after navigation and before the screenshot. For element capture, use await page.locator('.product-card').first().screenshot({ path: target });. Prefer a page-specific readiness selector when you know one; there is no single wait policy that fits every site.
4. Python and cURL alternatives
If your existing batch tooling is Python, Playwright provides a Python API. Install the package and Chromium dependencies using the official Python installation flow, then use this small synchronous worker pattern. It is intentionally sequential; add a bounded process or async worker pool only after a pilot.
python -m pip install playwright
python -m playwright install --with-deps chromium
from pathlib import Path
from hashlib import sha256
import json
import time
from playwright.sync_api import sync_playwright
urls = [line.strip() for line in Path('urls.txt').read_text().splitlines()
if line.strip() and not line.lstrip().startswith('#')]
out = Path('screenshots')
out.mkdir(parents=True, exist_ok=True)
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
for i, url in enumerate(urls, 1):
name = f'{i:05d}-{sha256(url.encode()).hexdigest()[:16]}.png'
error = None
for attempt in range(1, 4):
context = browser.new_context(viewport={"width": 1365, "height": 900})
page = context.new_page()
try:
response = page.goto(url, wait_until='domcontentloaded', timeout=30000)
if response is None or response.status >= 400:
raise RuntimeError(f'HTTP status: {None if response is None else response.status}')
page.screenshot(path=str(out / name), full_page=False)
with open('results.jsonl', 'a') as f:
f.write(json.dumps({"url": url, "status": "ok", "file": str(out / name)}) + '\n')
error = None
break
except Exception as exc:
error = str(exc)
if attempt < 3:
time.sleep(min(2 ** attempt, 8))
finally:
context.close()
if error:
with open('results.jsonl', 'a') as f:
f.write(json.dumps({"url": url, "status": "failed", "error": error}) + '\n')
browser.close()
cURL can fetch a URL directly, but it does not execute page JavaScript or render a browser screenshot. Use it to check network reachability or download an already-existing image, not as a substitute for browser capture.
curl --fail --location --max-time 30 https://example.com/ -o page.html
5. Filenames, retries, and safe batch operation
- Stable names: combine the input order with a hash of the full URL. This avoids unsafe URL characters and keeps repeated runs easy to compare. If you need history rather than overwrite behavior, add a run identifier or date directory.
- Retry policy: retry timeouts and temporary network failures with exponential backoff and a finite attempt count. Do not endlessly retry 404s, access denials, or CAPTCHA pages; record and inspect them.
- Idempotency: write each capture to a temporary filename and rename on success if downstream processes may read the directory during a run. Keep a manifest so reruns can skip verified outputs if that fits your workflow.
- Politeness: capture only pages you are authorized to access, use a modest request rate, and respect site access controls. Browser automation can create significant traffic.
- Secrets: do not put credentials or private URLs in a world-readable URL file or logs. Restrict file permissions and avoid logging authorization headers or cookies.
- Disk: estimate output size from a representative sample and leave space for logs, temporary files, and OS updates. Prune old outputs or move them to appropriate object storage if retention requires it.
6. Pilot the workload and estimate capacity
- Select a representative sample: ordinary pages, the heaviest pages, long pages if using full-page mode, and pages with known dynamic content.
- Run sequentially on the actual VPS and browser version. Record elapsed time, failures, output bytes, peak memory, and CPU.
- Repeat with a small increase in worker count. Watch for rising timeout rates, memory pressure, swapping, CPU saturation, and target rate limits.
- Choose the smallest concurrency that meets your schedule reliably, then leave headroom for page variation and other processes.
- Recalculate transfer and storage for the complete batch. India-region plan allowances can differ from headline amounts; check the exact region terms.
Estimate a run from observed pilot data rather than a generic pages-per-hour claim: divide the number of representative pages by the measured completion rate at the chosen concurrency, and add time for retries and slow outliers. That estimate is only useful if the sample resembles the real batch. No source in the research establishes a universal throughput number.
7. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Browser launch fails with missing libraries | Linux browser dependencies are absent. | Run npx playwright install --with-deps chromium with package-install permissions, or install the documented dependencies for your distribution. |
| Executable or browser revision not found | Playwright package and installed browser binaries do not match. | Run npx playwright install chromium after installing or upgrading the package. |
| Navigation timeout | Slow server, long page load, stalled resource, or an unsuitable wait condition. | Use an appropriate explicit readiness condition, increase the timeout modestly for known slow pages, and log failures. Do not set an unlimited timeout. |
| Screenshot is blank or missing dynamic content | Capture occurred before client-side rendering completed. | Wait for a page-specific visible selector or a short, justified delay after DOM readiness. |
| HTTP 403, 429, or CAPTCHA | Target access control or rate limiting blocked the request. | Reduce request rate, verify authorization, and stop retrying blocked pages automatically. |
| Process is killed or the VPS becomes unresponsive | Memory exhaustion, often from too many workers or tall pages. | Lower concurrency, use viewport capture, close contexts promptly, and move to a plan with more memory if required. |
| Output directory fills | Images or retained runs exceed available disk. | Measure average file size, set retention, and provision more storage before the next batch. |
| Intermittent failures repeat in every run | Retries may be masking a stable site problem or a local network issue. | Keep per-URL logs, inspect failed URLs separately, and distinguish temporary failures from permanent status codes. |
8. Performance, reliability, and cost
Concurrency is the main practical throughput control, but each browser page also consumes memory and CPU. Full-page screenshots, heavy scripts, large images, and long waits raise resource use or duration. Start at one worker, test a representative batch, then increase gradually while tracking peak memory, CPU, disk growth, transfer use, and failure rate. Reuse a browser process for a batch, but close each context after a page so page state and resources do not accumulate.
For reliability, keep finite timeouts and retries, persist one outcome per URL, and make failed records easy to rerun. A VPS job can be interrupted by process or machine failure; for long batches, split work into checkpoints and preserve the input manifest and logs. Costs include the instance, any storage or transfer beyond plan terms, and operator time spent maintaining browser dependencies. The Lightsail price is a listed AWS plan price checked in the dossier on 2026-10-03; verify current pricing and regional terms before choosing a plan. [AWS pricing] [billing rules]
9. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. One GET request returns a PNG, JPEG, WebP, or PDF. It accepts cookie and consent banners like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.
For a bulk workflow, the API supports up to 100 URLs per call, and also offers async jobs with signed webhooks, caching with a chosen TTL, a usage API, and an OpenAPI spec. See the ScreenshotNeo API documentation for the full parameter list. One call for an individual page looks like this:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed; an MCP server lets AI agents take screenshots. The Free plan includes 1,000 screenshots a month with no card. Paid plans start at $5 for 3,000 screenshots; all features are on every plan. Sign up free for 1,000 screenshots a month, with no card required.
10. FAQ
Can I run this after closing my SSH session?
Yes. Use a process supervisor or terminal multiplexer appropriate to your server, and keep the logs and input manifest so you can inspect and resume the run.
Should I use full-page capture for every URL?
Only when the entire document is needed. Viewport capture is a better default for predictable output size and lower work on very tall pages.
Can I capture pages behind a login?
Only where you have authorization. Add credentials through a protected browser context or approved authentication flow, and keep secrets out of source files and logs.
How many URLs can a low-cost VPS handle?
There is no reliable universal number. Measure a representative batch on the exact plan, browser version, and capture settings you intend to use.


