ScreenshotNeo

BlogHow-to

How to Capture Screenshots of Indian Government Websites with Puppeteer

Capture Indian government webpages with Puppeteer, choosing the right viewport, wait condition, and screenshot scope while respecting each site’s access rules.

By the ScreenshotNeo team4 October 20268 min read

Puppeteer captures an Indian government webpage with Page.screenshot(). Set the viewport before navigation, wait for the page state or specific content you need, then choose a viewport, full-page, clipped region, or element screenshot. The example below saves a full-page PNG. Replace the example URL with a page you are authorized to access, and check that portal’s own policies before automating it.

1. Install Puppeteer and capture a page

The following ES module is a complete starting point. It launches Puppeteer’s browser, sets a reproducible viewport, navigates to the target, waits for a page-specific heading, saves a PNG, and closes the browser even if navigation or capture fails.

import puppeteer from 'puppeteer';

const url = 'https://example.gov.in/'; // Replace with the target page.
const outputPath = 'government-page.png';

const browser = await puppeteer.launch();
try {
  const page = await browser.newPage();
  await page.setViewport({
    width: 1365,
    height: 900,
    deviceScaleFactor: 1,
  });

  await page.goto(url, {
    waitUntil: 'domcontentloaded',
    timeout: 60_000,
  });

  // Prefer a stable selector for the content the screenshot must contain.
  // Replace this example with a selector present on the target page.
  await page.waitForSelector('main', { visible: true, timeout: 20_000 });

  await page.screenshot({
    path: outputPath,
    type: 'png',
    fullPage: true,
  });
  console.log(`Saved ${outputPath}`);
} finally {
  await browser.close();
}

Install Puppeteer in a project with npm install puppeteer. Puppeteer’s screenshot guide demonstrates launching a browser, navigating, calling page.screenshot(), and closing the browser. See the Puppeteer screenshot guide and Page.screenshot() API.

The URL and dimensions above are examples, not government-wide defaults. Set the viewport before navigating because changing it later can cause a reload on some pages and affect responsive layout. Record the viewport and device scale factor when screenshots need to be reproducible. See Page.setViewport().

2. Choose when the page is ready

A navigation event does not guarantee that every visual element is ready. Some pages load content after the initial HTML, update sections through JavaScript, or keep network connections open. Pick a condition that matches the part of the page you need:

Condition Use it when Limit
domcontentloaded The document is parsed and you will wait for a specific element or state afterward. Images, fonts, and later scripts may still be loading.
load You need the page load event, including dependent resources that delay it. It does not mean that delayed or application-rendered content is complete.
networkidle2 A page becomes quiet enough for a capture after navigation. It is not a guarantee that every dynamic section is finished, and pages with ongoing requests can make it unsuitable.
waitForSelector() A known element marks the content you need as present or visible. A selector can appear before its contents are complete; choose a meaningful stable target.

Puppeteer documents networkidle2 in its screenshot example and offers Page.waitForSelector() for selector-based readiness. A selector often expresses the capture requirement more directly than a generic quiet-network condition. See Puppeteer’s guide and waitForSelector().

To use the documented network-idle option, replace the navigation call with:

await page.goto(url, {
  waitUntil: 'networkidle2',
  timeout: 60_000,
});

For a site with a known delayed section, navigate first and then wait for that section. If a fixed delay is necessary because the page exposes no dependable readiness signal, keep it explicit and conservative; a delay cannot prove that content has loaded.

await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 60_000 });
await page.waitForSelector('[data-testid="results"]', {
  visible: true,
  timeout: 20_000,
});
// Capture only after the target content is present.

3. Choose screenshot scope and output

Use the narrowest capture that serves the task. A viewport image records what fits on screen at the chosen dimensions. A full-page image records the scrollable page in one capture. An element or clipped rectangle focuses on a component or region.

Need Puppeteer option Example
Visible viewport Default screenshot behavior await page.screenshot({ path: 'viewport.png' })
Entire scrollable page fullPage: true await page.screenshot({ path: 'full.png', fullPage: true })
Known rectangle clip with x, y, width, and height await page.screenshot({ path: 'region.png', clip: { x: 0, y: 0, width: 900, height: 600 } })
Specific DOM element ElementHandle.screenshot() Query an element, then call its screenshot() method.

For element capture, Puppeteer’s ElementHandle.screenshot() attempts to scroll the element into view if it is hidden. Check the handle before capture so a missing selector produces a useful error.

const target = await page.$('main article');
if (!target) {
  throw new Error('Could not find main article');
}
await target.screenshot({ path: 'article.png' });

A rectangle must be within the rendered page area and use the coordinate system you intend to capture. For an element that may move or vary in size, element capture is usually less brittle than hard-coded clip coordinates.

Use PNG for lossless output and when image clarity matters. Puppeteer also supports JPEG and WebP screenshot types where supported by the installed Puppeteer/browser version. The quality option applies to lossy formats and does not apply to PNG. Consult the current ScreenshotOptions API for supported options in your version.

await page.screenshot({ path: 'page.jpg', type: 'jpeg', quality: 82 });
await page.screenshot({ path: 'page.webp', type: 'webp', quality: 82 });

4. Capture responsibly on Indian government sites

Indian government sites include central, state, district, and local administration services. GIGW describes guidance for government websites and apps, including usability, consistency, security, and accessibility; it is not a Puppeteer manual and does not grant blanket permission to automate every portal. Check the particular website’s policy and access rules, especially for authenticated services, personal information, frequent requests, and reuse of captured content. See the Guidelines for Indian Government Websites.

Do not attempt to defeat a CAPTCHA, login control, rate limit, or other restriction. If a protected service blocks access, use an authorized test environment or ask the site operator for permission. GIGW accessibility guidance discusses CAPTCHA alternatives, but that guidance is not authorization to bypass one. See the GIGW Website Policies.

Policies are site-specific. For example, GIGW’s own policy says its content requires due permission for reproduction and that its pages may not be loaded into frames on another site. That statement applies to the GIGW site; do not generalize it to every government portal. A screenshot is a visual record, not proof that a page meets accessibility or GIGW requirements. GIGW accessibility guidance references WCAG 2.1 and text alternatives.

5. Troubleshooting

Symptom Likely cause Fix
Navigation times out The site is slow, keeps requests open, or the chosen lifecycle event is not reached. Use a suitable navigation condition such as domcontentloaded, set a deliberate timeout, then wait for the required selector. Do not treat a longer timeout as proof of readiness.
Screenshot is blank or missing content The capture happened before the target rendered, the selector is wrong, or access was blocked. Wait for a stable visible content selector. Inspect the page state and response manually. If a CAPTCHA or access control appears, stop and use an authorized route.
Full-page image omits lazy content Some content loads only after scrolling or interaction. Check whether the page requires scrolling or an allowed interaction before capture. Wait for the actual content selector; do not assume a full-page option itself triggers every site-specific load.
Layout differs from the browser you expected Viewport dimensions or device scale factor differ, or responsive breakpoints changed. Set viewport before navigation and record width, height, and deviceScaleFactor.
Element screenshot throws or captures the wrong area The selector matched nothing, matched the wrong element, or the layout shifted. Check the element handle, use a specific selector, wait for visibility, and capture the element rather than relying on stale clip coordinates.
Output file is unexpectedly large A long full-page PNG contains many pixels and is lossless. Capture only the needed region, use a lossy format with an appropriate quality where acceptable, or reduce viewport/device scale factor. Preserve sufficient detail for the intended use.

6. Performance, reliability, and cost

A full-page screenshot can be much larger than a viewport capture because it includes the entire scrollable page. Element and clip captures reduce the captured area when a full-page record is unnecessary. Use a stable selector to avoid waiting longer than needed, and always close the browser in a finally block. For repeated captures, reuse a browser process where appropriate while creating a fresh page for each capture; close pages and the browser cleanly when finished.

Capture time depends on the target page, its resources, and the readiness condition; the dossier provides no benchmark or universal timing. Set a timeout appropriate to your job, log the URL and failure stage, and distinguish a navigation failure from a missing selector or screenshot write error. Puppeteer is software you run, so costs depend on the compute and storage you provide; there is no fixed screenshot price established here.

7. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. Cookie and consent banners, newsletter popups, and chat widgets are removed before capture; each cleaning step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server lets Claude, Cursor, or another MCP client use take_screenshot, get_page_info, and capture_pdf.

Here is a cURL request for a permitted target URL. See the ScreenshotNeo API documentation for the API details.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.gov.in/ \
  -o government-page.webp

Equivalent Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.gov.in/"},
    timeout=90,
)
r.raise_for_status()
open("government-page.webp", "wb").write(r.content)

Equivalent Node.js (Node 18 or later, which includes fetch):

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.gov.in/',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('government-page.webp', bytes));

ScreenshotNeo has 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000 screenshots. All features are available on every plan. Its 63 options include full-page capture with lazy images loaded, element capture by CSS selector, viewport and device presets, retina scale, PDF settings, custom CSS and JavaScript, selector waits, custom headers and cookies, caching with a chosen TTL, signed links, async jobs with signed webhooks, bulk capture for up to 100 URLs per call, a usage API, and an OpenAPI specification. Check the docs and the target site’s policy before capturing.

Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.

8. FAQ

Can Puppeteer prove a government page is accessible?

No. A screenshot records rendered pixels; it does not assess accessibility or prove compliance with GIGW or WCAG.

Should I use a screenshot as evidence that a page was unchanged?

A screenshot can be part of a visual record, but preserve the capture time, target URL, viewport, and relevant capture conditions if you need to interpret or reproduce it. It does not establish why a page changed or who changed it.

Can I capture pages behind a login?

Only where you are authorized and the site’s rules permit it. Do not automate around CAPTCHA or other access restrictions.