How to Take Bulk Website Screenshots with Playwright on an Indian AWS EC2 Instance
Run a bounded Playwright screenshot queue on an EC2 instance in Mumbai, with full-page capture, per-URL results, and practical guidance for reliability and cost.
Direct answer: Launch a Linux EC2 instance in AWS Asia Pacific (Mumbai), region ap-south-1. Install Node.js and the matching Playwright package, then install its Chromium browser and Linux dependencies. Feed URLs to a bounded worker queue; for each URL, navigate with a timeout and save page.screenshot({ fullPage: true }). Keep browser contexts isolated when jobs should not share cookies or storage, record success and failure per URL, and close each context and the browser when finished.
This guide builds a small runnable batch capture script and explains how to deploy and operate it. Instance availability, pricing, and performance depend on current AWS offerings and your workload, so check the current Mumbai region options and measure your own URL set.
1. Choose the Mumbai region and an instance
In the EC2 launch flow, choose AWS Asia Pacific (Mumbai), region code ap-south-1. AWS instance families and their regional availability change, and no single size is right for every batch. Start with a modest instance and low concurrency, run representative pages, and observe memory, CPU, timeouts, and completion rate before changing the machine or worker count. See AWS regional instance type availability.
Use a dedicated instance for the capture job. Restrict SSH ingress to trusted source IP addresses and use a least-privilege instance role if the job needs AWS services. AWS warns against allowing SSH access from anywhere in its security group guidance. Treat input URLs as untrusted: a page can cause the browser to request other network destinations. Do not expose the browser or instance to arbitrary network access, and do not use this workflow to bypass access controls.
2. Install Node.js, Playwright, and Chromium
Connect to the Linux instance using your chosen AMI’s documented setup instructions. Install a supported Node.js runtime, then create a project and install Playwright:
mkdir bulk-screenshots
cd bulk-screenshots
npm init -y
npm install playwright
Install the browser binary and Linux system dependencies with Playwright’s installer:
npx playwright install --with-deps chromium
Playwright’s browser binaries are tied to the installed Playwright version. Pin the project dependency and install the matching browser on the EC2 machine; do not copy a browser cache from a different operating system. Browser downloads take hundreds of megabytes depending on the browser. If the job only needs the Chromium headless shell, Playwright documents an --only-shell installation option; choose it only when that browser mode fits your task. Refer to the official Playwright browser installation documentation.
3. Prepare and validate the URL input
Put one URL per line in urls.txt. The script below accepts only HTTP and HTTPS URLs and uses an index plus a short SHA-256 digest for each output filename. This avoids collisions between URLs with the same hostname and prevents URL text from becoming a filesystem path.
https://example.com/
https://www.wikipedia.org/
https://stripe.com/
Use URLs you are authorized to access. Sites may require login, block automation, rate-limit requests, or return location-dependent content. Do not attempt to evade access controls. For authenticated pages, supply credentials through a secure mechanism and deliberately configure browser context state; do not commit secrets to the source file or manifest.
4. Run a bounded screenshot queue
Save the following as capture.mjs. It launches Chromium once, runs a fixed number of workers, creates a fresh browser context for each URL, captures the whole scrollable page, and writes a JSON-lines manifest. Each job is isolated from the others, and failures are recorded while the rest of the queue continues.
import { chromium } from 'playwright';
import { createHash } from 'node:crypto';
import { mkdir, readFile, writeFile, appendFile } from 'node:fs/promises';
import path from 'node:path';
const inputFile = process.argv[2] ?? 'urls.txt';
const outputDir = process.argv[3] ?? 'screenshots';
const concurrency = Number(process.env.CONCURRENCY ?? 2);
const timeoutMs = Number(process.env.NAVIGATION_TIMEOUT_MS ?? 45_000);
if (!Number.isInteger(concurrency) || concurrency < 1) {
throw new Error('CONCURRENCY must be a positive integer');
}
if (!Number.isInteger(timeoutMs) || timeoutMs < 1) {
throw new Error('NAVIGATION_TIMEOUT_MS must be a positive integer');
}
function parseHttpUrl(value, lineNumber) {
let parsed;
try {
parsed = new URL(value);
} catch {
throw new Error(`Invalid URL on line ${lineNumber}: ${value}`);
}
if (parsed.protocol !== 'http:' && parsed.protocol !== 'https:') {
throw new Error(`Only HTTP and HTTPS are allowed (line ${lineNumber})`);
}
return parsed.href;
}
const raw = await readFile(inputFile, 'utf8');
const urls = raw.split(/\r?\n/)
.map((line) => line.trim())
.filter((line) => line.length > 0 && !line.startsWith('#'))
.map((line, index) => parseHttpUrl(line, index + 1));
await mkdir(outputDir, { recursive: true });
const manifestPath = path.join(outputDir, 'manifest.jsonl');
await writeFile(manifestPath, '');
const browser = await chromium.launch({ headless: true });
let next = 0;
async function record(entry) {
await appendFile(manifestPath, `${JSON.stringify(entry)}\n`);
}
async function worker() {
while (true) {
const index = next++;
if (index >= urls.length) return;
const url = urls[index];
const digest = createHash('sha256').update(url).digest('hex').slice(0, 12);
const screenshotPath = path.join(outputDir, `${String(index + 1).padStart(5, '0')}-${digest}.png`);
const startedAt = new Date().toISOString();
const context = await browser.newContext({
viewport: { width: 1440, height: 900 },
deviceScaleFactor: 1,
});
try {
const page = await context.newPage();
const response = await page.goto(url, {
waitUntil: 'domcontentloaded',
timeout: timeoutMs,
});
await page.screenshot({ path: screenshotPath, fullPage: true, type: 'png' });
await record({
index,
url,
status: 'ok',
httpStatus: response?.status() ?? null,
screenshotPath,
startedAt,
finishedAt: new Date().toISOString(),
});
console.log(`OK ${url} -> ${screenshotPath}`);
} catch (error) {
await record({
index,
url,
status: 'error',
error: error instanceof Error ? error.message : String(error),
startedAt,
finishedAt: new Date().toISOString(),
});
console.error(`FAILED ${url}: ${error instanceof Error ? error.message : error}`);
} finally {
await context.close().catch(() => {});
}
}
}
try {
await Promise.all(Array.from({ length: Math.min(concurrency, urls.length) }, () => worker()));
} finally {
await browser.close();
}
Run it with the default settings (two concurrent jobs and a 45-second navigation timeout):
node capture.mjs urls.txt screenshots
Change the concurrency and timeout for a measured workload:
CONCURRENCY=1 NAVIGATION_TIMEOUT_MS=60000 node capture.mjs urls.txt run-2026-10-04
The values in this example are starting parameters, not performance guarantees. The script’s domcontentloaded condition means the initial document has been parsed; it does not guarantee that every image, animation, or client-rendered element is ready. If a target needs additional rendering time, add a site-appropriate readiness condition or a deliberate delay before the screenshot. Network idle is not a universal signal: analytics, polling, and long-lived connections may keep requests active.
5. Pick capture extent, isolation, and readiness deliberately
| Decision | Use this when | Playwright setting or approach |
|---|---|---|
| Viewport or full page | You need only the visible screen, or the entire scrollable document. | Omit fullPage for viewport capture; set fullPage: true for the full page. |
| One context per job | URLs must not share cookies, local storage, or session state. | Create and close a browser.newContext() for each job, as in the script. |
| Shared context | A sequence of pages intentionally needs the same authenticated session or browser state. | Create one context for that identity and use multiple pages within it; manage state and cleanup explicitly. |
| Headless browser mode | The batch runs without a desktop display. | Launch Chromium with headless: true; install the matching Playwright browser build and dependencies. |
| Page readiness | A page needs more than initial document parsing before capture. | Wait for a known selector or app-specific condition, or use a bounded delay. Select the condition per site. |
Playwright documents the screenshot API, including full-page capture, in its Page screenshot reference. Browser contexts provide isolated browser sessions; see Browser contexts and the browser.newContext() API.
For consistent comparisons, use the same viewport, device scale factor, browser version, operating system image, and headless mode for every run. Rendering can differ across host operating systems, browser versions, hardware, settings, and headless modes. Full-page screenshots can also become large or slow for very long documents; if you only need a first-screen preview, capture the viewport instead.
6. Store outputs and check the batch
The script writes PNGs and a manifest.jsonl file under the output directory. Each manifest line records the input URL, status, HTTP status when available, timestamps, output path, or error. Review the manifest after a run rather than assuming every input produced a valid image. For durable retention, choose an approved object-storage destination based on retention, access, and budget needs; the appropriate storage tier and cost depend on your requirements.
Keep each run in its own output directory if you need to compare results or avoid overwriting earlier captures. Ensure the attached filesystem has room for the expected images and manifest. If the batch is interrupted, the manifest indicates completed jobs; add resume logic keyed by input identity if rerunning the whole queue is undesirable.
7. Control concurrency, reliability, and cost
Concurrency
Do not launch an unbounded number of pages. Each browser page consumes system resources, while target sites may rate-limit or block concentrated traffic. The script uses a small bounded queue; increase its concurrency gradually while watching CPU, memory, navigation failures, and the target sites’ limits. Playwright Test’s --workers option controls its own test worker processes, each of which starts a browser; it does not control concurrency in a custom script. A custom script needs its own queue or semaphore.
Failures and retries
The example records individual navigation and screenshot errors and continues with other URLs. For production batches, consider a capped retry policy for transient failures, with a delay between attempts. Avoid retrying every failure blindly: an invalid URL, access denial, or persistent block is unlikely to improve with repeated requests. Detect browser process failures and restart the browser in a supervised job if necessary. Keep input order and URL identity in the manifest so results remain auditable.
Spend
EC2 compute charges depend on the selected instance, region, runtime, and current AWS pricing; storage and data transfer may add costs. The dossier does not establish a price for this workload. Check current AWS pricing for the machine and storage you choose, set a budget appropriate to the job, and stop or terminate billed resources when the batch is complete. Do not estimate cost from a generic pages-per-minute figure: the URL mix, browser version, viewport, full-page length, and instance all affect runtime.
8. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Playwright says the browser executable is missing. | The browser binary was not installed, or it does not match the installed package. | Run npx playwright install chromium for the installed package version. Reinstall after changing the Playwright version. |
| Chromium exits or reports missing shared libraries. | Linux browser dependencies are absent from the AMI. | Run npx playwright install --with-deps chromium using the AMI’s supported package setup. |
| Navigation times out. | The site is slow, unreachable, blocking automation, or waiting for the selected event longer than the configured limit. | Check network reachability and the target’s response. Choose a suitable readiness condition and a justified timeout; do not increase it without limit. |
| Screenshot misses content below the fold or lazy images. | The capture was viewport-only, or content had not loaded before capture. | Set fullPage: true. For lazy-loaded content, use a site-appropriate readiness condition or controlled scrolling and wait before capture. |
| Some pages are blank or incomplete. | The site may rely on scripts, authentication, consent choices, or a later rendering event. | Inspect the page and response, configure the required authorized session, and wait for the specific content to appear before capturing. |
| Memory pressure or browser crashes increase with batch size. | Too many pages are active at once, or contexts are not being closed. | Lower CONCURRENCY, keep the finally cleanup, and measure the workload before scaling up. |
| Files overwrite or have confusing names. | Output naming is based only on a hostname or an unstable path. | Use a stable per-input identifier such as the index and URL digest used in the example, and separate runs into different directories. |
| SSH is reachable from unexpected networks. | The instance security group allows overly broad ingress. | Restrict SSH to trusted source addresses and review the security group rules. |
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. One GET request returns an image or PDF, so a batch job does not need to install and maintain a browser on EC2. Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. AI agents can take screenshots through its MCP server. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000. Every feature is on every plan. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
Replace YOUR_API_KEY and the target URL. For multiple URLs, make one request per URL with a bounded client-side queue; keep credentials out of source control. Sign up for 1,000 free screenshots a month with no card.
FAQ
Does this require a desktop session on EC2?
No. Playwright launches Chromium in headless mode in the example.
Does full-page capture make one image per screen?
No. fullPage: true asks Playwright to capture the full scrollable page as a screenshot.
Can the screenshots be pixel-identical across runs?
Not necessarily. Host OS, browser version, hardware, settings, headless mode, and changing site content can affect rendering. Keep the capture environment and page state consistent when comparing images.
Can I use the same context for every URL?
Only if sharing cookies and storage is intentional. A fresh context per job avoids accidental state sharing; a shared context can be useful for an intentionally shared authenticated session.


