Puppeteer Screenshot of an Indian Government Website with a CAPTCHA
Capture the CAPTCHA page Puppeteer displays, choose the right screenshot scope, and understand what the image can—and cannot—prove.
Puppeteer can save a screenshot of the page as it currently appears, including a CAPTCHA challenge. It captures the rendered browser state; it does not solve the CAPTCHA, prove verification succeeded, establish that automation is permitted, or show whether the site meets accessibility requirements. This guide shows how to capture and inspect the visible state without bypassing the challenge.
No target URL was supplied, so the examples use a placeholder. Confirm the exact URL and site owner before describing a page as an Indian government website. The Guidelines for Indian Government Websites and apps (GIGW) describe gov.in and nic.in domains as appropriate for government sites and the URL as a strong authenticity indicator, not proof by itself: GIGW guidelines.
Capture the current page with Puppeteer
Install Puppeteer, navigate to the page, wait for a useful page state, save the screenshot, then close the browser. The official Puppeteer guide documents this flow and the screenshot options: Puppeteer screenshots and Page.screenshot() API.
npm install puppeteer
// screenshot.js
const puppeteer = require('puppeteer');
(async () => {
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
await page.setViewport({ width: 1365, height: 900, deviceScaleFactor: 1 });
const response = await page.goto('https://example.gov.in/', {
waitUntil: 'domcontentloaded',
timeout: 60000,
});
console.log('HTTP status:', response?.status() ?? 'no main response');
// Give client-rendered content a brief chance to settle. This does not
// solve or interact with a CAPTCHA.
await new Promise(resolve => setTimeout(resolve, 1500));
await page.screenshot({ path: 'government-page.png' });
} finally {
await browser.close();
}
})().catch(error => {
console.error(error);
process.exitCode = 1;
});
Run it with node screenshot.js. Replace the sample URL with the URL you are authorized to access. If the page requires a human verification step, the screenshot records the state Puppeteer can see; use the site’s own supported process and terms for any further access.
Choose the capture scope
| Method | Use it when | Trade-off |
|---|---|---|
| Viewport screenshot | You need the visible browser area and the CAPTCHA at its on-screen size. | Content below the viewport is omitted. |
fullPage: true |
You need a tall image of the entire document. | The CAPTCHA can become small in the overall image, making details harder to inspect. |
clip |
You need a specific rectangular region. | Coordinates depend on page layout and viewport state. |
| Element screenshot | You have identified a specific element, such as a challenge container. | The element must exist and remain attached; layout changes can alter what is captured. |
// Full document
await page.screenshot({ path: 'full-page.png', fullPage: true });
// A rectangle in page coordinates
await page.screenshot({
path: 'captcha-region.png',
clip: { x: 300, y: 220, width: 520, height: 300 },
});
// A selected element; replace the selector after inspecting the page
const challenge = await page.$('.captcha-container');
if (challenge) {
await challenge.screenshot({ path: 'captcha-element.png' });
}
An element screenshot scrolls the element into view when needed. If the selector is unknown or the challenge is inside a cross-origin frame, inspect the page structure and frames first; do not assume a particular vendor or selector. For screenshots that must be comparable, hold the viewport, device scale factor, page state, and capture scope constant.
Return screenshot bytes instead of writing a file
Page.screenshot() can return image bytes, which is useful when sending the capture to storage or another service. The API also documents base64 output. Avoid logging or exposing a screenshot if it contains personal or session data.
const imageBytes = await page.screenshot({ type: 'png' });
// imageBytes is a Uint8Array; pass it to your storage or response layer.
Wait for the state you intend to document
A screenshot represents one moment. Navigation completion does not guarantee that every script, image, or challenge widget has finished rendering. Puppeteer supports navigation wait conditions such as domcontentloaded, load, and networkidle0/networkidle2; choose based on the page, and use a bounded timeout. Pages with polling or long-lived requests may never become network-idle, so waiting for a known selector or a short explicit delay can be more reliable.
// Wait for a known page landmark, then capture
await page.goto('https://example.gov.in/', {
waitUntil: 'domcontentloaded',
timeout: 60000,
});
await page.waitForSelector('main', { timeout: 15000 });
await page.screenshot({ path: 'page.png' });
If the page displays a CAPTCHA, capture that state as evidence of what appeared. Do not automate challenge solving or infer that a screenshot means the check passed. Puppeteer documents page interactions through mouse, touch, and keyboard input, and recommends locators for selecting and interacting with elements, but the cited guidance does not establish CAPTCHA-specific instructions or permission for any particular government site: Puppeteer page interactions. Check the target site’s terms and contact its support channel if access is blocked.
What a screenshot can establish
- It can document the rendered content visible in a particular browser state, viewport, and time.
- It can show that a CAPTCHA was displayed in that capture, if the challenge is legible in the image.
- It cannot establish that verification succeeded, that the site permits automated access, or that the page is accessible to people with disabilities.
- It cannot establish official ownership from appearance alone. Record the exact URL and, when known, the owning organization.
GIGW adopts WCAG 2.1 success criteria and includes an explicit CAPTCHA accessibility checkpoint. It calls for text alternatives that explain the purpose and alternative CAPTCHA forms using different sensory output modes. The guideline also describes manual evaluation; a screenshot alone does not certify a site: GIGW accessibility guidance.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request returns an image or PDF. Its capture flow accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each of those steps can be turned off. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. An MCP server gives AI agents tools including take_screenshot, get_page_info, and capture_pdf. Review the ScreenshotNeo API docs and use a page you are authorized to capture.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.gov.in/ \
-o shot.webp
Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.gov.in/"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://example.gov.in/',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', new Uint8Array(await res.arrayBuffer()));
For a Node.js environment without Bun, write the response bytes with Node’s fs/promises:
const { writeFile } = require('node:fs/promises');
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.gov.in/' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
ScreenshotNeo includes full-page capture, element selection, viewport and device presets, retina scale, custom waits, headers, cookies, user agent, caching, and more. Its free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000 shots. See ScreenshotNeo or sign up for 1,000 free screenshots a month, no card required.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Navigation times out | The site is slow, has long-lived requests, or blocks the request. | Use a bounded timeout and a suitable wait condition such as domcontentloaded; then wait for a specific page landmark. Check the URL and site terms. |
| Screenshot is blank or incomplete | Capture occurred before client rendering, or content is lazy-loaded. | Wait for a relevant selector or content state. For long documents, use fullPage and verify the resulting image dimensions. |
| CAPTCHA is too small to read | A full-page image scales a tall page into a small preview. | Capture the viewport or clip around the challenge, and use a stable viewport. |
| Element screenshot fails | The selector matched nothing, the element detached during rendering, or it is in a frame. | Wait for a verified selector, re-query immediately before capture, and inspect frames when appropriate. |
| Image differs between runs | Viewport, page timing, dynamic content, or device scale factor changed. | Set viewport and device scale factor explicitly and synchronize on a page landmark. Record capture time and URL. |
| Access stops at a challenge | The site is requiring verification or otherwise limiting access. | Document the displayed state and use the site’s supported human process or support channel. Do not treat the screenshot as successful verification. |
Performance, reliability, and cost
For Puppeteer, browser startup and page loading are part of each capture’s work. Reusing a browser process for a controlled batch can avoid repeated startup overhead, but isolate pages and always close them; cap concurrency to the resources available. Choose the smallest capture scope that answers the question: a viewport or element image is usually easier to inspect and store than a very tall full-page image. No benchmark is implied here; the target site’s response time and rendered content determine the actual capture time.
Reliability improves when navigation and selector waits have explicit timeouts, browser cleanup is in a finally block, and the capture records URL, timestamp, viewport, and relevant response status. Treat a missing main-document response or non-success status as a signal to inspect, not as proof of a particular failure cause. Screenshots can contain sensitive data, so store and share them according to the site’s and your organization’s requirements.
Puppeteer itself is open-source software, but this method still consumes the compute and storage of the machine running the browser. ScreenshotNeo’s stated plans are Free: 1,000 shots/month; Starter: $5 for 3,000; Growth: $15 for 15,000; Pro: $39 for 60,000; Scale: $99 for 250,000; Business: $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan. Only clean shots are billed; consult its docs for request options and response headers.
FAQ
Does Puppeteer solve the CAPTCHA?
No. This method captures the page state; it does not solve the challenge or confirm a successful verification.
Does seeing a CAPTCHA prove the website is official?
No. Check the exact domain and owning organization. GIGW identifies gov.in and nic.in as appropriate government domains and the URL as a strong indicator, but an unspecified screenshot is not proof.
Can a screenshot prove CAPTCHA accessibility?
No. GIGW’s CAPTCHA checkpoint calls for explanatory text alternatives and alternative sensory modes, and describes manual evaluation. An image of one state cannot establish conformance.
Which capture is best for evidence?
Use the viewport when legibility of the displayed challenge matters. Use full-page capture when the surrounding document context matters, and include a focused capture if the challenge becomes too small to inspect.


