ScreenshotNeo

BlogHow-to

How to Batch Screenshot Indian Government Website Pages for Records

Capture a reviewed list of public pages with Playwright, then pair every image with a manifest of URLs, timestamps, and settings.

By the ScreenshotNeo team4 October 20269 min read

Use a reviewed list of public URLs, capture each page with Playwright, and save a manifest beside the resulting images. Record the requested URL, final URL after redirects, capture time and timezone, viewport, browser version, and whether each image shows the viewport or the full page. A screenshot records a rendered appearance at capture time; by itself, it does not establish what a page showed earlier, prove compliance, or determine legal admissibility.

This guide uses Node.js with Playwright. It also includes cURL, Python, and Node.js examples for ScreenshotNeo, a website screenshot API and MCP server for developers from ScreenshotNeo. For a reproducible records workflow, keep your URL list and capture settings under review, preserve failed attempts in the manifest, and inspect the images after the batch completes.

1. Define the record set and capture scope

Write down the public URLs to capture before running the batch. Decide what each image needs to show:

  • Viewport capture: the content visible in the browser viewport at the time of capture.
  • Full-page capture: the full scrollable page, useful for long notices or pages whose relevant content is below the fold.
  • Element capture: one selected page element, when a specific notice or result is the subject of the record.

Use the same viewport, browser, and scale when images need to be comparable. Playwright documents full-page capture, element targeting, output path, image type, and scale options in its Page API. CSS-pixel scale generally produces a smaller image; device-pixel scale preserves a denser output. Choose based on how the image will be reviewed, and record the choice.

2. Create a Node.js batch capture

Install Playwright and its browser, then create a file named capture.mjs. The script below reads one URL per line from urls.txt, captures each page, and writes a JSON Lines manifest. Each row includes success or failure details so a failed navigation does not disappear from the record.

npm init -y
npm install playwright
npx playwright install chromium

Put the URLs in urls.txt, one public URL per line. Lines beginning with # and blank lines are ignored.

https://www.india.gov.in/
https://www.meity.gov.in/
https://www.digitalindia.gov.in/
import { chromium } from 'playwright';
import { readFile, appendFile } from 'node:fs/promises';
import { basename, resolve } from 'node:path';

const inputPath = process.argv[2] ?? 'urls.txt';
const outputDir = resolve(process.argv[3] ?? 'captures');
const fullPage = process.env.FULL_PAGE === '1';
const width = Number(process.env.VIEWPORT_WIDTH ?? 1440);
const height = Number(process.env.VIEWPORT_HEIGHT ?? 1000);
const timeoutMs = Number(process.env.TIMEOUT_MS ?? 30000);

if (!Number.isInteger(width) || width < 1 || !Number.isInteger(height) || height < 1) {
  throw new Error('Viewport width and height must be positive integers.');
}

const urls = (await readFile(inputPath, 'utf8'))
  .split(/\r?\n/)
  .map(line => line.trim())
  .filter(line => line && !line.startsWith('#'));

if (urls.length === 0) throw new Error(`No URLs found in ${inputPath}`);

await import('node:fs/promises').then(({ mkdir }) => mkdir(outputDir, { recursive: true }));
const manifestPath = resolve(outputDir, 'manifest.jsonl');
const browser = await chromium.launch({ headless: true });
const context = await browser.newContext({ viewport: { width, height } });
const browserVersion = browser.version();

try {
  for (let index = 0; index < urls.length; index++) {
    const requestedUrl = urls[index];
    const capturedAt = new Date().toISOString();
    const page = await context.newPage();
    const filename = `${String(index + 1).padStart(4, '0')}.png`;
    const imagePath = resolve(outputDir, filename);
    const record = {
      requestedUrl,
      finalUrl: null,
      capturedAt,
      output: filename,
      viewport: { width, height },
      browser: 'Chromium',
      browserVersion,
      fullPage,
      status: 'failed',
      error: null
    };

    try {
      const response = await page.goto(requestedUrl, {
        waitUntil: 'load',
        timeout: timeoutMs
      });
      record.finalUrl = page.url();
      record.httpStatus = response?.status() ?? null;
      await page.screenshot({ path: imagePath, fullPage, type: 'png' });
      record.status = 'captured';
    } catch (error) {
      record.finalUrl = page.url();
      record.error = error instanceof Error ? error.message : String(error);
    } finally {
      await appendFile(manifestPath, `${JSON.stringify(record)}\n`);
      await page.close();
    }
  }
} finally {
  await context.close();
  await browser.close();
}

console.log(`Finished ${urls.length} URL(s). Manifest: ${manifestPath}`);

Run it with the defaults for viewport-only captures:

node capture.mjs urls.txt captures

Set FULL_PAGE=1 to capture the full scrollable page. The viewport and navigation timeout can also be set per run:

FULL_PAGE=1 VIEWPORT_WIDTH=1365 VIEWPORT_HEIGHT=900 TIMEOUT_MS=45000 node capture.mjs urls.txt captures-full

The script uses Playwright’s documented sequence of launching a browser, navigating to a URL, saving a screenshot to a path, and closing the browser. It waits for the page load event. Some public sites continue loading analytics or other requests afterward; switching to networkidle can help in some cases, but pages with continuous traffic may never reach network idle. If the page has important late-rendering content, use a targeted wait for a known selector or a short delay, then inspect the image.

3. Choose settings that fit the record

Choice Use it when What to record
Viewport or full page Use viewport for the initial visible state; full page for content across the scrollable document. The selected mode; for viewport captures, viewport width and height.
Whole page or element Capture the whole page for general context; target a selector when one region is the relevant subject. The selector and any element-specific capture details.
CSS-pixel or device-pixel scale CSS pixels can reduce file size; device pixels can retain higher-density detail. The selected scale and device scale factor if configured.
PNG, JPEG, or WebP Choose an image type supported by the capture API and suitable for the intended review workflow. File type and output path.

For a selected element, create a locator and call its screenshot method, for example await page.locator('main').screenshot({ path: imagePath }). The selector must match the page actually rendered; record it because it changes what the output contains. Playwright’s screenshot options, including output type, path, full-page mode, scale, and element screenshots, are documented in the official API reference.

4. Preserve and review the manifest

The manifest is workflow guidance for making a batch interpretable, not a prescribed standard. Keep it with the images. In addition to the fields in the sample script, consider recording:

  • The URL list version or a checksum of the input file.
  • Any selected element, extra wait, or browser context settings.
  • The page’s HTTP status when available and whether a redirect changed the final URL.
  • Whether an image was expected but not created, and the error reported by the browser.

Review the batch before relying on it: confirm every requested URL has a manifest row; every successful row points to an image; final URLs do not unexpectedly lead to a login, error, or unrelated page; and full-page images include the expected content. Dynamic pages can change while loading or during capture. A screenshot records a rendered state at its capture time, not all page data, server behavior, or the page’s earlier appearance.

5. Or skip the browser setup

ScreenshotNeo can capture a URL with one GET request. See the ScreenshotNeo API documentation for parameters and configuration.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://www.india.gov.in/ -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://www.india.gov.in/"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://www.india.gov.in/'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

With Node.js on versions that provide global fetch, save the response using Node’s filesystem API:

import { writeFile } from 'node:fs/promises';
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://www.india.gov.in/' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

ScreenshotNeo accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. You still need to keep your own URL list and capture records if your workflow requires a manifest.

Sign up for 1,000 free screenshots a month, with no card required.

6. Troubleshooting batch captures

Symptom Likely cause Fix
Navigation times out The site is slow, unreachable, or still waiting on the selected navigation event. Check the URL and connectivity, increase TIMEOUT_MS, and try a less strict wait event. Preserve the failure and error in the manifest.
Image shows an error, login, or unexpected destination The site redirected, required authentication, or returned an error page. Inspect finalUrl and the HTTP status, then decide whether the URL belongs in the reviewed public set. Do not silently treat it as the intended page.
Important content is missing The content rendered after the screenshot, was below the fold in viewport mode, or is loaded only after interaction. Try full-page capture, wait for a relevant selector, or use a short delay where appropriate. Review the resulting image.
Full-page image is unexpectedly long or incomplete The page has unusual dynamic layout or content changes while scrolling/capture occurs. Inspect the rendered page, retry under consistent settings, and consider capturing a relevant element if that better represents the record.
Browser launch fails The Playwright package or Chromium browser is not installed for the current environment. Run npm install playwright and npx playwright install chromium in the project environment.
Manifest exists but an image does not The page navigation or screenshot step failed after the row was prepared. Read the row’s status and error; rerun only after reviewing the cause, and retain the original failure record.

7. Performance, reliability, and cost

The local Playwright script reuses one browser and browser context, then processes URLs sequentially. This limits simultaneous load and makes the order easy to audit, but a large list takes longer than parallel capture. If you introduce concurrency, keep it bounded, check the target sites’ applicable access rules, and record the concurrency and retry policy so results remain interpretable. Retries can capture a different page state, so preserve the original attempt and timestamp rather than overwriting it.

Each local capture uses browser and network resources; full-page images and device-pixel output can require more memory and produce larger files than viewport captures. There is no per-screenshot service charge in this local Playwright workflow, but the machine, storage, and maintenance still have costs. A hosted screenshot API trades browser setup for an API plan and provider behavior. ScreenshotNeo’s listed plans are Free with 1,000 shots/month, Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000; yearly billing gives two months free, and every feature is on every plan. Check the plan details and API behavior in its documentation before choosing it for a records process.

For reliability, retain the input URL list, manifest, image files, and the exact settings used together. A successful browser call does not by itself show that the intended page was captured; verify the final URL, status, and image. For administrative or legal use, determine applicable records and evidentiary requirements from relevant primary sources or qualified guidance. The workflow here makes no claim about a particular GIGW version, retention period, or legal admissibility.

Frequently asked questions

Does a screenshot prove a government page’s earlier content?

No. It records what the browser rendered at capture time. Keep context such as the requested and final URLs, timestamp, and settings, and do not treat the image alone as proof of an earlier state.

Does this workflow establish GIGW compliance?

No. Capturing a page does not establish compliance. This guide does not interpret a GIGW version or provide a compliance assessment.

Should I capture the full page or only the viewport?

Choose based on the record’s purpose: viewport for the initially visible content, full-page for the scrollable page, or an element for a specific region. Record that choice with each image.

Can I capture pages that require sign-in?

The sample is intended for public URLs and does not configure authentication. Follow the site’s access rules and your organization’s requirements before capturing authenticated pages.