How to Take Batch Screenshots of Indian Government Website Pages with Playwright
Capture Indian government website pages in batches with Playwright. Learn how to handle readiness, naming, capture modes, failures, and access rules.
To take batch screenshots with Playwright, put the target page URLs in a list, reuse one browser and context, then visit each URL and save a uniquely named image. Choose a readiness condition suited to each site, record failures per URL, and check the actual site’s access rules before running a batch. This guide shows a sequential Node.js script first, then how to adapt it for different capture scopes and larger batches.
Access caveat: No particular Indian government domain was inspected for this guide. It cannot establish permission, request limits, terms of use, or robot rules for a target site. Check each site’s own policies before automated access; start with a small batch and keep request volume controlled.
1. Set up Playwright and prepare the URL list
Install Playwright and its Chromium browser in a Node.js project:
npm install playwright
npx playwright install chromium
Save the following as batch-screenshots.mjs. Replace the example URLs with pages you are allowed to capture. It creates one browser, one context with a consistent viewport and locale, and one page per URL in sequence. Each URL gets its own output name and result record, and an individual failure does not stop the remaining URLs.
import { chromium } from 'playwright';
import { mkdir, writeFile } from 'node:fs/promises';
import path from 'node:path';
const urls = [
'https://example.gov.in/',
'https://example.gov.in/about',
];
const outputDir = path.resolve('screenshots');
const timeoutMs = 30_000;
await mkdir(outputDir, { recursive: true });
const browser = await chromium.launch();
const context = await browser.newContext({
viewport: { width: 1440, height: 1000 },
locale: 'en-IN',
colorScheme: 'light',
});
const results = [];
try {
for (const [index, url] of urls.entries()) {
const filename = `${String(index + 1).padStart(3, '0')}.png`;
const outputPath = path.join(outputDir, filename);
const page = await context.newPage();
try {
const response = await page.goto(url, {
waitUntil: 'domcontentloaded',
timeout: timeoutMs,
});
// Add a site-specific readiness check here if the page renders content later.
await page.screenshot({ path: outputPath, fullPage: true });
const result = {
url,
finalUrl: page.url(),
status: response?.status() ?? null,
outcome: 'saved',
outputPath,
};
results.push(result);
console.log(result);
} catch (error) {
const result = {
url,
finalUrl: page.url(),
outcome: 'failed',
error: String(error),
outputPath,
};
results.push(result);
console.error(result);
} finally {
await page.close();
}
}
} finally {
await context.close();
await browser.close();
await writeFile(
path.join(outputDir, 'index.json'),
JSON.stringify({ capturedAt: new Date().toISOString(), results }, null, 2),
);
}
Run it with node batch-screenshots.mjs. The output directory contains numbered PNGs and an index.json associating each source URL with its final URL, HTTP status when available, outcome, path, and failure detail. Numbered names avoid collisions when different URLs normalize to the same filename.
2. Choose when a page is ready to capture
page.goto() supports navigation conditions such as commit, domcontentloaded, load, and networkidle. The example uses domcontentloaded as a starting point, not a guarantee that every page’s visible content is ready. A page may render content after navigation through client-side scripts, or continue loading nonessential requests.
commit: Continue once the response has started and the document begins loading. Use only when you have another explicit readiness check.domcontentloaded: A practical starting point for many pages; then wait for a meaningful page-specific element if the content is dynamic.load: Wait for the load event, which may take longer when a page has many assets.networkidle: Avoid treating it as a universal visual-readiness signal. Playwright discourages using it for tests; pages with ongoing network activity may not reach it reliably.
For a known page element, add a locator wait after navigation:
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30_000 });
await page.locator('main h1').waitFor({ state: 'visible', timeout: 10_000 });
await page.screenshot({ path: outputPath, fullPage: true });
Use a selector that is meaningful for that actual page. A generic sleep can be useful for a known short animation or delayed widget, but it is less reliable than waiting for the content the capture is meant to show.
3. Select the capture area and image settings
Match the capture mode to the record you need. Playwright’s screenshot APIs provide viewport, full-page, locator, and clipped-region captures. The official references are the Playwright screenshots guide and Page screenshot API.
| Purpose | Setting | Trade-off |
|---|---|---|
| Visible browser area | page.screenshot({ path }) |
Fast and bounded, but content below the viewport is omitted. |
| Entire scrollable page | page.screenshot({ path, fullPage: true }) |
Includes below-fold content; a very long page can produce a very tall image. |
| One component | page.locator('main article').screenshot({ path }) |
Useful for a specific panel or article; the selector must match an element. |
| Fixed rectangle | page.screenshot({ path, clip: { x: 0, y: 0, width: 900, height: 700 } }) |
Captures a defined region relative to the page viewport. |
For format and output control, Playwright supports PNG, JPEG, and WebP. JPEG and WebP support a quality setting; scale: 'css' produces one image pixel per CSS pixel, while device scaling preserves device-pixel sizing. For example:
await page.screenshot({
path: outputPath.replace(/\.png$/, '.jpeg'),
type: 'jpeg',
quality: 85,
fullPage: true,
scale: 'css',
});
Additional screenshot options include animation handling, caret behavior, masks, and injected styles. Use them only when they fit the purpose of the record. For example, masking a changing timestamp can help comparison images, but masking or hiding meaningful page content can make the evidence misleading. Consult the Page screenshot API for the current option definitions.
4. Keep filenames and results traceable
For a fixed list, index-based names such as 001.png and a JSON index are simple and collision-resistant. If names should be recognizable, derive them from the URL but sanitize the result and include a stable suffix or index because distinct URLs can produce the same readable slug.
Keep, at minimum, the requested URL, final URL after redirects, capture timestamp, output path, status/outcome, and capture settings. For comparison work, also record viewport width and height, device scale, browser version, locale, and color scheme. Use separate browser contexts if cookies or storage must not carry between groups of pages; contexts provide isolated sessions, while pages in one context share its setup.
5. Scale carefully for larger batches
The sequential loop is a good default for a modest list because it limits simultaneous navigation and keeps failures easy to associate with inputs. For larger lists, bounded concurrency can reduce total elapsed time, but it also increases resource use and sends more simultaneous requests. There is no universal safe concurrency or request rate for Indian government sites: determine limits from the actual site’s published rules and your operational needs before increasing it.
If introducing concurrency, use a fixed worker pool or semaphore rather than creating a page for every URL at once. Keep the same per-URL timeout, failure logging, and unique output naming. Close each page in a finally block, and close the context and browser after the batch even if the worker reports errors. If one target site’s cookies or session state should not affect another, isolate those captures in separate contexts.
For repeatable visual comparisons, pin the Playwright/browser runtime where practical and keep viewport, device scale, locale, color scheme, and headless settings stable. Rendering may still vary with operating system, browser version, hardware, power source, settings, or headless mode. Dates, rotating content, banners, animations, and personalization can also change between captures. Record the environment and capture options so a screenshot is understood as the output of that setup at that time, not as a canonical rendering for every visitor.
6. Respect the limits of screenshots as evidence
The official Guidelines for Indian Government Websites (GIGW 3.0) portal includes accessibility and mobile-friendliness resources. Its accessibility guidance describes meaningful text equivalents for non-text content and evaluation using manual and accessibility-tool checks. A screenshot documents rendered appearance at a point in time; it does not by itself demonstrate accessibility or GIGW conformance.
For responsive review, deliberately capture multiple viewport sizes and record the dimensions alongside each image. Do not infer mobile behavior from a desktop capture or treat visual appearance alone as proof that assistive technologies can use a page.
7. Troubleshooting common problems
| Problem | Likely cause | Fix |
|---|---|---|
| Navigation times out | The site is slow, a request remains active, or the chosen wait condition is too broad. | Use an explicit timeout, choose a suitable navigation condition, and wait for a page-specific element when needed. Log the failed URL and continue with the rest. |
| Screenshot is blank or missing content | The page content renders after the navigation event, or a selector wait is absent. | Wait for a meaningful visible element or a known site-specific condition before capturing. Confirm the final URL and response status in the index. |
| Some pages fail but the script stops | An error escaped the per-URL handler or cleanup was not placed in finally. |
Catch errors inside the loop per URL, close each page in finally, and close the context/browser in an outer finally. |
| Output files overwrite each other | Names were derived from URLs that sanitize to the same value. | Use an index or append a stable unique suffix; never assume a readable URL slug is unique. |
| Images are too large or unusually tall | Full-page capture includes a long scrollable document or device-pixel scaling increases dimensions. | Use viewport or element capture if that matches the purpose, and choose CSS scale when one output pixel per CSS pixel is desired. |
| Repeated captures differ | Environment, viewport, color scheme, dynamic content, or page personalization changed. | Record and stabilize browser/runtime and context settings; account for changing page content with a documented mask or style only when appropriate. |
| Navigation is blocked or challenged | The target may restrict automated access or require a different access method. | Stop and review the target site’s published policy and access requirements. Do not attempt to evade an access control. |
8. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. One GET request can return a screenshot or PDF, so a batch workflow can call the API for each allowed URL and save each response. See the ScreenshotNeo API documentation for request options and integration details.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.gov.in/ -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.gov.in/"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://example.gov.in/',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', bytes));
ScreenshotNeo accepts cookie and consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the capture; each of those steps can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up for ScreenshotNeo’s free 1,000 screenshots a month, with no card.
9. Frequently asked questions
Can I reuse one Playwright browser for every URL?
Yes. Reusing a browser and a configured context avoids launching a new browser for each page. Create and close pages per URL, or reuse a page if your workflow carefully resets any state that matters.
Should each government website get its own browser context?
Use separate contexts when session cookies or storage must be isolated. A shared context is convenient when the same capture configuration should apply and shared session state is acceptable.
Does a screenshot prove that a page is accessible?
No. It records visual rendering. Accessibility review also requires checks beyond appearance, including text alternatives and manual or tool-assisted evaluation.
What should I do if a page redirects?
Record both the requested URL and page.url() after navigation. Review the destination and response status to understand what the saved image represents.


