How to capture screenshots of Indian government websites with an AI agent
Use Playwright or an AI agent to capture a government website screenshot, verify the page and capture scope, and handle protected flows responsibly.
Use a browser automation library such as Playwright to navigate to the government website and capture the rendered page. Set fullPage: true for the full scrollable document; omit it for the visible viewport, or capture a locator for one component. Verify the final host and route, record the browser and viewport, and stop for human review if the site presents a CAPTCHA or another protected step.
This guide shows a runnable Playwright workflow, an AI-agent prompt that records useful context, and checks for responsible handling. A screenshot documents one rendered state; it does not prove how the page appears for every visitor or establish permission to republish its contents.
1. Choose the right capture scope
| Need | Capture | Playwright option |
|---|---|---|
| What is visible in the browser window | Viewport screenshot | page.screenshot({ path: 'page.png' }) |
| The entire scrollable document | Full-page screenshot | page.screenshot({ path: 'page.png', fullPage: true }) |
| A particular panel or component | Element screenshot | locator.screenshot({ path: 'element.png' }) |
| Image bytes for a later processing step | Screenshot buffer | page.screenshot() |
Use the smallest scope that answers the question. A viewport image is easier to compare across fixed screen sizes. A full-page image is useful for documenting a long page, but may be very tall. An element image isolates a region while leaving its surrounding context out of the image. The official Playwright screenshot guide documents page, full-page, buffer, and element captures.
2. Set up Playwright
These commands create a small Node.js project and install Playwright with its Chromium browser. Run them from a new directory:
npm init -y
npm install playwright
npx playwright install chromium
Save the following as capture.mjs. Pass the target URL as the first argument. The script checks that the input is an HTTP or HTTPS URL, records the final URL after redirects, uses a fixed viewport, waits for the page’s load event, and writes either a viewport or full-page PNG. It deliberately does not solve or bypass human checks.
import { chromium } from 'playwright';
const input = process.argv[2];
const scope = process.argv[3] ?? 'full';
if (!input) {
console.error('Usage: node capture.mjs <https-url> [full|viewport]');
process.exit(2);
}
let target;
try {
target = new URL(input);
} catch {
console.error('The target must be a valid absolute URL.');
process.exit(2);
}
if (!['http:', 'https:'].includes(target.protocol)) {
console.error('Only HTTP and HTTPS URLs are supported.');
process.exit(2);
}
if (!['full', 'viewport'].includes(scope)) {
console.error('Scope must be full or viewport.');
process.exit(2);
}
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage({ viewport: { width: 1440, height: 1000 } });
const response = await page.goto(target.href, {
waitUntil: 'load',
timeout: 60000,
});
// A response may be absent for certain navigation failures. Keep that visible
// in the run notes rather than treating it as a successful HTTP response.
const status = response?.status() ?? 'no response';
const finalUrl = page.url();
const output = scope === 'full' ? 'page-full.png' : 'page-viewport.png';
await page.screenshot({ path: output, fullPage: scope === 'full' });
console.log(JSON.stringify({
requestedUrl: target.href,
finalUrl,
capturedAt: new Date().toISOString(),
viewport: { width: 1440, height: 1000 },
scope,
httpStatus: status,
output,
title: await page.title(),
}, null, 2));
} finally {
await browser.close();
}
Run it with a verified URL:
node capture.mjs https://www.india.gov.in/ full
node capture.mjs https://www.india.gov.in/ viewport
The example uses India.gov.in as an input example only; it does not claim that this page or any other specific government site was tested. Check the actual site’s current URL, route, policies, and access behavior before running a capture.
3. Give an AI agent a bounded task
An agent can run the script or operate a Playwright browser, but the task should specify the target and stopping conditions. Ask it to report what it observed rather than infer that a capture is valid simply because a file exists.
Capture a screenshot of this page for documentation: <verified URL>.
1. Confirm that the final host and path after navigation are the expected ones.
2. Use a 1440 by 1000 viewport and capture the full scrollable page as PNG.
3. Wait for the page load event; if the page has a clear, page-specific readiness
signal, wait for that signal as well.
4. If a CAPTCHA, login, access challenge, or other human verification appears,
stop. Do not solve it, evade it, or try alternate routes around it. Report the
visible state and ask an authorized person to review or complete the step.
5. Save the screenshot and report the requested URL, final URL, capture time,
browser and version, viewport, capture scope, visible errors or redirects,
and whether the page appeared complete.
6. Do not claim that the image represents other browsers, screen sizes, dates,
or users. Do not publish or redistribute it unless the applicable policy and
rights permit that use.
For reproducible work, include the Playwright version and browser version in your run record. Keep the screenshot with its metadata so later reviewers can tell what it represents.
4. Verify the site and handle government content carefully
The Guidelines for Indian Government Websites and Apps (GIGW) apply to government websites and applications at central, state, and local levels. The guidelines discuss user-centricity, accessibility, security, responsive presentation, and testing across browsers, operating systems, connection speeds, and screen resolutions. A screenshot captures one browser and viewport state, so it cannot demonstrate that the page works or looks the same in all those contexts. See the official GIGW guidance.
Check the exact host, page ownership, and route before capture. GIGW describes a website URL as a strong indicator of authenticity and status; gov.in and nic.in domains are useful signals, with specified exceptions for some eligible educational and research institutions. A domain signal alone does not verify the ownership or status of every linked page. Follow links and redirects carefully and record the final URL.
Do not automate through CAPTCHA or other human verification. GIGW’s accessibility guidance includes alternatives for CAPTCHA. If a protected flow blocks the capture, stop and hand it to an authorized person rather than attempting to evade the check. Preserve relevant accessibility context in your notes: an image alone may omit text alternatives, accessible document versions, or other information needed to understand the page.
Before publishing or redistributing the image, read the target site’s terms, copyright policy, and privacy or security notices. The IGOD terms explain that external linked sites may have separate policies and that NIC cannot authorize third-party copyrighted material hosted there. Those terms do not create blanket permission for every government website or reuse. Cite the source page and capture date in reporting, and obtain permission from the relevant rights holder where required.
5. Make captures repeatable and useful
Record the capture context
- Requested URL and final URL after redirects.
- Capture timestamp and timezone.
- Browser name and version, Playwright version, and operating system.
- Viewport width and height, device scale factor if configured, and whether the capture was viewport, full-page, or element-only.
- Navigation status, visible interstitials or errors, and any readiness condition used.
Choose waits that match the page
waitUntil: 'load' waits for the load event, but it does not prove that every dynamic widget, delayed image, or third-party component has finished. When the page has a known stable element, wait for that locator to become visible before capturing. Avoid using an arbitrary long delay as a substitute for a readiness condition. Some sites continue background requests indefinitely, so waiting for all network activity to stop can be unreliable.
Compare responsive layouts fairly
For responsive review, capture the same route and comparable page state at each intended viewport. Keep browser and viewport details with every image. GIGW’s cross-browser and screen-resolution guidance is a reminder that one desktop capture cannot stand in for testing at other sizes or environments.
6. Use ScreenshotNeo to skip browser setup
ScreenshotNeo is a website screenshot API and MCP server. Its API takes a URL and returns an image or PDF; the MCP server exposes screenshot tools for Claude, Cursor, and other MCP clients. The service can remove known consent banners, newsletter popups, and chat widgets before capture, with each step configurable. Its response headers identify the page verdict and billing status.
For an API request, create an API key and use the endpoint and parameter names in the ScreenshotNeo documentation. This cURL example saves a WebP response:
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://www.india.gov.in/ \
-o india-gov-full.webp
The API accepts full-page capture options; consult the docs for the current parameter names and options. Python and Node.js examples using the same endpoint follow.
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://www.india.gov.in/"},
timeout=90,
)
r.raise_for_status()
with open("india-gov.webp", "wb") as f:
f.write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://www.india.gov.in/',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await (await import('node:fs/promises')).writeFile('india-gov.webp', bytes);
Check the final URL and the returned X-Page-Verdict and X-Billed headers when interpreting an API response. A CAPTCHA or protected flow still requires human review; a capture tool is not permission to circumvent access controls. Review the target’s terms before reusing an image.
7. Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| Navigation times out | Slow server, stalled resource, or a page that does not finish loading | Check the URL and network access. Use a site-specific readiness signal where possible. Do not treat a timeout as a complete capture; record it and retry only when appropriate. |
| Screenshot contains an access challenge or CAPTCHA | The site requires human verification or blocks the automated browser | Stop automation and ask an authorized person to review or complete the step. Do not bypass the challenge. |
| Final URL is unexpected | Redirect, canonical route, login flow, or interstitial | Inspect and record the final host and path. Do not assume the requested page was captured. |
| Some content is missing | Client-rendered content, lazy loading, delayed images, or a readiness check that was too early | Wait for a known page-specific element or appropriate load state, then capture again. Record the readiness condition used. |
| Full-page image is unusually tall or awkward | The document contains long feeds, repeated content, or fixed-position elements | Decide whether a viewport or element capture better fits the task. Keep full-page scope only when the whole document is needed. |
| Image differs between runs | Page content, time-dependent banners, viewport, browser, or network conditions changed | Compare metadata and page state. Capture again at the same viewport and browser context, and include the capture time in any report. |
| Screenshot is present but unusable | The browser captured an error page, blank state, or partial render | Inspect the image and navigation status before treating the file as evidence. Report the visible failure state. |
| Playwright cannot launch Chromium | Browser binaries were not installed for the local Playwright package | Run npx playwright install chromium in the project and retry. |
8. Performance, reliability, and cost
Local Playwright avoids an external screenshot-service charge, but requires maintaining Node.js, Playwright, browser binaries, and a runtime environment with network access. Capture duration depends on the target page, its resources, and the chosen wait condition; this guide makes no timing or success-rate claim. Full-page output can consume more time, memory, and storage than a viewport or element image, especially for very long pages. Use the narrowest useful capture scope and avoid unnecessary retries against the site.
Reliability depends on the target site’s availability and behavior as well as the browser environment. Record failures rather than silently treating them as successful evidence. For repeated captures, store metadata alongside each file and compare equivalent browser and viewport conditions. Check the site’s access rules before scheduling automation.
ScreenshotNeo’s plans are Free: 1,000 shots per month with no card; Starter: $5 for 3,000; Growth: $15 for 15,000; Pro: $39 for 60,000; Scale: $99 for 250,000; and Business: $249 for 1,000,000. Yearly billing gives two months free. Every feature is on every plan. The service bills clean shots only; bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, with page verdict and billing information in response headers. See the docs for configuration and response details.
9. Frequently asked questions
Can an AI agent screenshot a gov.in page?
It can use a browser automation tool such as Playwright when the page is publicly accessible and the site’s rules permit the capture. Verify the exact host and stop at human verification.
Does a full-page screenshot prove the website is accessible?
No. It records a visual state in one browser and viewport. Accessibility review requires more than an image, including attention to alternatives for non-text content and accessible documents.
Should I publish a screenshot of a government page?
Check the specific page’s terms, copyright notices, privacy and security policies, and any third-party rights before publishing or redistributing it. A government domain alone does not settle reuse rights.
Can I automate around a CAPTCHA if I only need a screenshot?
No. Stop and ask an authorized person to handle the protected step. Do not instruct an agent to solve or evade human verification.


