How to Archive Indian Government Web Pages as Screenshots with Browserless
Capture a government page with Browserless, choose viewport or full-page output, wait for dynamic content, and keep useful context with the image.
To save a visual snapshot of an Indian government web page with Browserless, send a POST request to its /screenshot REST endpoint with the page URL and screenshot options. Set fullPage to true for the full document or omit it for a viewport capture. Wait for important content and images, save the image response, and record the URL and capture settings beside it. A screenshot is a visual rendition of a page, not a complete web archive or proof of authenticity.
This guide shows the Browserless REST API workflow, including cURL, Python, and Node.js examples. It also covers waiting, output formats, incomplete captures, and practical recordkeeping. Browserless documents the Screenshot API and its options.
1. Prepare the page and capture context
Before capturing, confirm the exact page URL and identify the organization responsible for the content. GIGW identifies gov.in and nic.in domains as indicators of official status and says ownership information is part of the minimum content expected on subsequent pages. These are identity checks, not proof that a screenshot is authentic. See the Guidelines for Indian Government Websites and Apps (GIGW).
Decide what the screenshot needs to show:
- Viewport capture: the page as rendered in a particular browser window. Use this when the visible screen and responsive layout are the point.
- Full-page capture: the document from top to bottom. Use
fullPage: true. A long page can produce a very tall image.
Choose and record a viewport deliberately when responsive layout matters. Different widths can change navigation, line wrapping, and which content appears. The example below uses a 1365 × 900 viewport as a chosen setting, not as a prescribed standard.
2. Get a Browserless token and capture the page
Browserless requires an API token. The current REST endpoint uses POST with JSON, a target url, and optional Puppeteer-style screenshot options. Replace YOUR_API_TOKEN and the example URL with your token and target page. The examples save the raw image response to a file.
cURL
curl --fail-with-body --silent --show-error \
-X POST 'https://production-sfo.browserless.io/screenshot?token=YOUR_API_TOKEN' \
-H 'Cache-Control: no-cache' \
-H 'Content-Type: application/json' \
--data '{
"url": "https://www.india.gov.in/",
"options": {
"fullPage": true,
"type": "png",
"viewport": { "width": 1365, "height": 900 }
},
"scrollPage": true,
"waitForImages": true
}' \
--output government-page.png
Browserless documents scrollPage and waitForImages as top-level request settings; fullPage, output type, and viewport are screenshot options. Scrolling can trigger lazy-loaded content before the full-page image is made. If you only need the initially visible viewport, set fullPage to false and omit scrollPage.
Python
import os
from pathlib import Path
import requests
TOKEN = os.environ["BROWSERLESS_TOKEN"]
endpoint = f"https://production-sfo.browserless.io/screenshot?token={TOKEN}"
payload = {
"url": "https://www.india.gov.in/",
"options": {
"fullPage": True,
"type": "png",
"viewport": {"width": 1365, "height": 900},
},
"scrollPage": True,
"waitForImages": True,
}
response = requests.post(
endpoint,
headers={
"Cache-Control": "no-cache",
"Content-Type": "application/json",
},
json=payload,
timeout=90,
)
response.raise_for_status()
content_type = response.headers.get("Content-Type", "")
if not content_type.startswith("image/"):
raise RuntimeError(f"Expected an image response, got {content_type!r}: {response.text[:500]}")
Path("government-page.png").write_bytes(response.content)
print(f"Saved government-page.png ({content_type}, {len(response.content)} bytes)")
Install the dependency with python -m pip install requests, then set the BROWSERLESS_TOKEN environment variable before running the script. The response check prevents an error payload from being saved under a .png filename.
Node.js
import { writeFile } from 'node:fs/promises';
const token = process.env.BROWSERLESS_TOKEN;
if (!token) throw new Error('Set BROWSERLESS_TOKEN first');
const endpoint = new URL('https://production-sfo.browserless.io/screenshot');
endpoint.searchParams.set('token', token);
const payload = {
url: 'https://www.india.gov.in/',
options: {
fullPage: true,
type: 'png',
viewport: { width: 1365, height: 900 },
},
scrollPage: true,
waitForImages: true,
};
const response = await fetch(endpoint, {
method: 'POST',
headers: {
'Cache-Control': 'no-cache',
'Content-Type': 'application/json',
},
body: JSON.stringify(payload),
signal: AbortSignal.timeout(90_000),
});
if (!response.ok) {
throw new Error(`Browserless returned HTTP ${response.status}: ${(await response.text()).slice(0, 500)}`);
}
const contentType = response.headers.get('content-type') ?? '';
if (!contentType.startsWith('image/')) {
throw new Error(`Expected an image response, got ${contentType}`);
}
await writeFile('government-page.png', Buffer.from(await response.arrayBuffer()));
console.log(`Saved government-page.png (${contentType})`);
Run this with a Node.js version that provides the built-in fetch API, and set BROWSERLESS_TOKEN in the process environment. Do not commit the token or print the complete request URL to logs because the token is in its query string.
3. Choose capture options that fit the page
| Need | Setting or approach | What to keep in mind |
|---|---|---|
| Whole document | options.fullPage: true |
Captures the full document height. Very long pages can result in large images. |
| Only the current screen | options.fullPage: false or omit the option |
Use a deliberate viewport if page layout is important. |
| Images or lazy-loaded sections | Top-level waitForImages: true; use scrollPage: true for lazy content |
Waiting and scrolling may add time. Inspect the output for missing content. |
| Output format | options.type: png, jpeg, or webp |
Match the filename extension to the requested type. Browserless returns an image response. |
| Responsive layout | options.viewport with width and height |
Changing viewport dimensions can change what is visible and how content wraps. |
| Specific page element | Top-level selector containing a CSS selector |
Browserless waits for the element and crops to its bounding box. If the selector is absent, capture cannot represent the intended element. |
| Fixed rectangle | options.clip with x, y, width, and height |
Useful for a known region; coordinates depend on the rendered page. |
| JPEG size and quality | options.quality |
Quality applies to JPEG, not PNG. Choose an output type and quality appropriate to how the image will be used. |
| Content that needs time or a condition | Browserless shared wait settings, such as waiting for a selector, event, function, or timeout | Wait for the specific content that matters rather than assuming navigation completion means all content is ready. |
| Navigation behavior | gotoOptions |
Use when the page needs particular navigation behavior; consult Browserless’s current API reference for accepted fields. |
| Unneeded requests | rejectResourceTypes or rejectRequestPattern |
Blocking requests may speed a capture but can also remove styles, images, or data the page needs. |
| Continue after an async wait fails | bestAttempt |
This may allow a response despite a failed or timed-out asynchronous condition; inspect the resulting image carefully. |
These options affect the rendered image. They do not establish that the website’s underlying records, linked files, or interactive behavior have been preserved.
4. Verify the image and retain its context
Open the saved file and check that it depicts the expected page, not a browser error, access-denied response, CAPTCHA, blank page, or partial load. Browserless calls these signs that automation may be blocked or that the capture may not represent the expected content. Do not label such an image as a successful capture of the target page.
Keep a small sidecar record alongside each image. This is practical recordkeeping advice, not a formal standard:
- Exact URL captured
- Capture time in UTC
- Viewport width and height
- Whether full-page capture and scrolling were enabled
- Output format and any wait settings
- Whether the page displayed an error or appeared incomplete
A screenshot captures a visual state. It does not include all linked documents, the page’s behavior, or the underlying records. For a records program, consult the responsible department’s policy. GIGW says government organizations should define how long content remains online, when it moves to offline archives, and whether or when it may be deleted; the template does not set one universal retention period. A screenshot alone should not be treated as satisfying a department’s retention policy or as establishing legal admissibility.
5. Troubleshooting Browserless screenshots
| Symptom | Likely cause | What to try |
|---|---|---|
| HTTP authentication error | Missing, invalid, expired, or incorrectly encoded API token. | Check the token in the endpoint query parameter and confirm it is the token for the Browserless account you intend to use. Keep it out of source control and logs. |
| HTTP error or an error body saved as an image | The request failed, but the client saved the response without checking it. | Check HTTP status before writing the file. Inspect the response body for diagnostics, and verify the response content type begins with image/. |
| Blank or white image | The page may not have rendered, may have timed out, or may be blocking automation. | Open the target normally to confirm it is available; wait for a relevant element or images, then inspect the result again. A blank image is not a successful capture of page content. |
| CAPTCHA, access denied, or 403 page | The site is presenting a bot check or denying the automated request. | Respect the site’s access controls. Do not describe the challenge or denial screen as the requested page. Capture only where access is permitted. |
| Images or lower-page content missing | Content may load only after scrolling or after an image or page condition completes. | Try scrollPage: true and waitForImages: true, or wait for the specific element needed. Then inspect the full-page output. |
| Layout differs from a regular browser view | The rendered viewport may differ, or the page may render responsive or time-dependent content. | Set and record viewport dimensions, wait for the relevant content, and note any remaining difference. |
| Screenshot is unexpectedly huge | A full-page capture of a very long document can be tall and large. | Use viewport capture if a screen view is sufficient, or capture a relevant element with selector. Keep full-page mode when the complete vertical page is needed. |
| Element capture fails or is empty | The selector may not match, or the target element may not have appeared by capture time. | Confirm the CSS selector on the rendered page and wait for that element before relying on the result. |
| Request times out | The page or a wait condition took longer than the configured time budget. | Wait for a narrower condition, review navigation settings, and allow an appropriate client timeout. A longer client timeout cannot guarantee that the page will load. |
| Output format and file extension disagree | The requested type and output filename do not match. |
Use .png, .jpg, or .webp consistently with the chosen format and confirm the response content type. |
6. Performance, reliability, and cost considerations
- Waiting improves completeness but adds latency. Full-page scrolling and waiting for images can take longer than capturing the initial viewport. Wait for the content that matters rather than adding broad delays by default.
- Large captures need more storage and transfer. Full-page images can be tall; JPEG or WebP may be appropriate when a compact image matters, while PNG can be useful when preserving crisp rendered details matters. Check the actual output.
- Network and page behavior affect repeatability. Government pages can change, load dynamic content, or show access checks. Record the capture time and inspect every result; a successful HTTP response alone does not prove the expected page was captured.
- Check your Browserless plan for its current usage and pricing terms. The supplied capture documentation explains the endpoint and options but does not establish a price or per-capture cost here.
- Plan for failures. Record failed attempts separately, avoid treating error images as captures, and retry only where appropriate and permitted by the site.
7. ScreenshotNeo alternative
Browserless gives you a direct browser screenshot API. If you prefer a single screenshot call with cleanup and explicit page verdicts, ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. Its documented API accepts a URL and returns PNG, JPEG, WebP, or PDF; see the ScreenshotNeo API docs.
Or skip the browser setup
Use the one-call API request below, replacing the example URL with the government page you are permitted to capture:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://www.india.gov.in/ -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://www.india.gov.in/"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://www.india.gov.in/' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo returned HTTP ${res.status}`);
const image = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', image));
ScreenshotNeo accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers state the page verdict and whether it was billed. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up for ScreenshotNeo’s free 1,000 screenshots a month, with no card required.
Frequently asked questions
Does a screenshot preserve the entire government web page?
No. It preserves a visual rendering. It does not preserve linked files, page behavior, or underlying records.
Does full-page capture prove that a page was official or unchanged?
No. A screenshot by itself does not establish authenticity or prove that the content has not changed. Keep its URL and capture context, and follow the requirements for your intended use.
How long should I keep screenshots?
There is no universal retention duration in the cited GIGW archival policy template. Consult the policy of the responsible department and the requirements for your records program.
Can I capture any government page with the API?
Do not assume every site permits automated access. Respect its access controls and applicable policies; an access-denied or CAPTCHA result is not the requested page.


