How to Generate Website Previews for a List of Indian Government Websites
Build a consistent preview set from verified government URLs with browser automation, a metadata manifest, and clear handling for redirects and failures.
To generate previews for a list of Indian government websites, start with URLs verified against the responsible departments, then capture each page at the same viewport using browser automation. Save the image and a manifest row for every URL, including failures, redirects, the UTC capture time, and the viewport. A screenshot is useful for visual comparison and triage; it does not establish accessibility, security, correctness, or compliance.
1. Build and verify the URL inventory
Use an existing authorized inventory or find candidate sites through the Integrated Government Online Directory (IGOD). IGOD describes itself as a directory of government websites, and its policy says users should verify information with the relevant department or another source. Treat each directory entry as a lead, not as permanent proof that a URL is current.
Keep the submitted URL and the eventual destination separately. A useful CSV has columns such as:
organization,level,source_url,verified_on,submitted_url,notes
Ministry example,central,IGOD,2026-10-04,https://example.gov.in/,Verified with department
Replace the illustrative row with real URLs you are authorized to process. Do not guess domains from organization names. Check ownership, current destination, and any posted policies before running a batch.
2. Choose what the preview should show
| Choice | Use it when | Tradeoff |
|---|---|---|
| Initial viewport | You want compact, side-by-side comparisons of the first screen. | Content below the fold is not represented. |
| Full page | You need to inspect long-page structure and lower sections. | Images can be very tall and larger to store; sticky elements may appear differently in a full-page capture. |
| One-off batch | You need a point-in-time inventory or triage set. | It cannot show how a page changed over time. |
| Scheduled recapture | You need change tracking. | It adds traffic, storage, and operational work. Keep each run dated and comparable. |
Choose a fixed viewport width and height, browser/device scale, color scheme, and wait strategy. Record those settings with the results. Use the same settings for every site in a run so that differences in the previews are less likely to come from your capture configuration.
3. Capture a list with Playwright and Python
This runnable example reads a newline-delimited URL file, creates one PNG per input, and writes a JSON Lines manifest. It uses a fixed-size viewport and a bounded navigation timeout. It records redirects and failures instead of silently omitting them. The example uses domcontentloaded; that avoids waiting forever for sites whose analytics or streaming connections keep the network busy.
Install
python -m venv .venv
# macOS or Linux:
. .venv/bin/activate
# Windows PowerShell:
# .venv\Scripts\Activate.ps1
pip install playwright
playwright install chromium
Prepare urls.txt
https://www.india.gov.in/
https://www.example.gov.in/
Replace the example with verified URLs. The script accepts only HTTP and HTTPS URLs, deduplicates exact normalized inputs, and preserves the submitted input in the manifest.
Save as capture.py
import asyncio
import json
import re
from datetime import datetime, timezone
from pathlib import Path
from urllib.parse import urlsplit, urlunsplit
from playwright.async_api import async_playwright
INPUT = Path("urls.txt")
OUT = Path("previews")
MANIFEST = OUT / "manifest.jsonl"
WIDTH, HEIGHT = 1440, 1000
TIMEOUT_MS = 30_000
DELAY_SECONDS = 1.0
FULL_PAGE = False
def normalize(raw):
value = raw.strip()
if not value:
return None
parts = urlsplit(value)
if parts.scheme.lower() not in ("http", "https") or not parts.netloc:
raise ValueError("URL must be absolute HTTP or HTTPS")
# Normalize scheme and host casing, retaining path, query, and fragment.
netloc = parts.netloc.lower()
return urlunsplit((parts.scheme.lower(), netloc, parts.path or "/", parts.query, parts.fragment))
def safe_name(url, index):
host = urlsplit(url).netloc
stem = re.sub(r"[^a-zA-Z0-9.-]+", "-", host).strip("-") or "site"
return f"{index:04d}-{stem}.png"
async def main():
OUT.mkdir(parents=True, exist_ok=True)
raw_urls = INPUT.read_text(encoding="utf-8").splitlines()
urls, seen = [], set()
for raw in raw_urls:
if not raw.strip():
continue
url = normalize(raw)
if url not in seen:
seen.add(url)
urls.append((raw.strip(), url))
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True)
context = await browser.new_context(
viewport={"width": WIDTH, "height": HEIGHT},
device_scale_factor=1,
color_scheme="light",
)
with MANIFEST.open("w", encoding="utf-8") as manifest:
for index, (submitted, url) in enumerate(urls, start=1):
row = {
"submitted_url": submitted,
"requested_url": url,
"final_url": None,
"captured_at_utc": datetime.now(timezone.utc).isoformat(),
"viewport": {"width": WIDTH, "height": HEIGHT},
"full_page": FULL_PAGE,
"file": None,
"title": None,
"http_status": None,
"outcome": "error",
"error": None,
}
page = await context.new_page()
try:
response = await page.goto(
url,
wait_until="domcontentloaded",
timeout=TIMEOUT_MS,
)
row["final_url"] = page.url
row["title"] = await page.title()
row["http_status"] = response.status if response else None
# A received HTTP error page is still recorded as a capture;
# inspect its status before treating the image as usable.
filename = safe_name(url, index)
await page.screenshot(path=str(OUT / filename), full_page=FULL_PAGE)
row["file"] = filename
row["outcome"] = "captured" if (response is None or response.status < 400) else "http_error_captured"
except Exception as exc:
row["final_url"] = page.url
row["error"] = f"{type(exc).__name__}: {exc}"
row["outcome"] = "navigation_or_capture_error"
finally:
row["finished_at_utc"] = datetime.now(timezone.utc).isoformat()
manifest.write(json.dumps(row, ensure_ascii=False) + "\n")
manifest.flush()
await page.close()
await asyncio.sleep(DELAY_SECONDS)
await context.close()
await browser.close()
if __name__ == "__main__":
asyncio.run(main())
Run it with python capture.py. Results appear in previews/; manifest.jsonl has one JSON object per non-duplicate input. Inspect the manifest for http_error_captured and navigation_or_capture_error. The script deliberately does not retry access denials or attempt to bypass authentication, bot checks, or CAPTCHAs.
Common adjustments
- For full-page output, set
FULL_PAGE = True. Be aware that very long documents can consume substantial memory and produce unwieldy files. - For a more settled page, use
wait_until="load"or wait for a known selector after navigation. Keep the overall timeout bounded. Avoid a blanket infinite wait for network idle: long-lived requests can prevent it. - For lazy-loaded content, a project may need a controlled scroll-and-wait routine before the screenshot. Make that choice explicit and apply it consistently; scrolling can trigger more requests and change page state.
- To run concurrently, use a small fixed number of workers and separate pages. Start sequentially, then raise concurrency only if site policies and observed behavior support it. The example is sequential and waits one second after each URL.
4. Keep outputs auditable and useful
Keep the manifest alongside the image files. The example records submitted and normalized URLs, final URL, UTC start and finish times, viewport, title when available, HTTP status, output path, and outcome. Add columns for organization, verification date, run identifier, browser version, or operator if those matter to your project. Use stable filenames that do not expose query parameters, which may contain tokens or personal data.
A contact sheet can make visual triage faster, but keep the original captures and manifest as the source of truth. Include failed URLs in reports so the number of successful previews is not mistaken for total coverage. Restrict access to captures if pages may contain personal or otherwise sensitive material; do not publish the output by default.
5. Operate the batch carefully
- Review site policies and your authorization for automated requests. There is no single rate limit or blanket permission rule established for every Indian government domain in the sources reviewed.
- Begin with a small sample and modest sequential traffic. If a site blocks or objects, stop for that site and review the situation rather than trying to evade the block.
- Use bounded timeouts and retain failure records. A slow or unreachable site should not stall the entire inventory.
- For recurring runs, schedule them at a reasonable cadence, retain dates, and compare like-for-like viewports. More frequent captures add traffic and storage without necessarily improving the decision you need to make.
- Do not send credentials or capture login-only pages unless your authorization and data-handling process explicitly cover them. Never work around access controls.
Local browser automation gives control over browser settings and storage, while a managed capture API can reduce browser installation and job-management work. For a managed service, check its current data handling, retention, regional processing, availability, and terms before sending URLs or page content. Do not assume a vendor has a particular retention or reliability property without verifying its documentation.
6. Treat previews as visual evidence only
The Guidelines for Indian Government Websites and Apps (GIGW) cover government organizations at central, state, district, and local levels and set goals including usability, user-centricity, and universal accessibility. A thumbnail cannot demonstrate that a site meets those goals. It may help reveal a broken layout, expired notice, missing logo, or rendering difference, but it does not prove keyboard access, working links, valid markup, security, or compliance.
Use the GIGW tools and resources for separate HTML, CSS, broken-link, accessibility, mobile-friendliness, and assistive-technology checks when those are part of the project. NIC's S3WaaS is for government entities creating and hosting primarily informational sites; it is not a generic bulk screenshot service. SugamyaWeb is an accessibility evaluation service with eligibility and service terms to check, not a screenshot generator.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. Its one-call endpoint can return a PNG, JPEG, WebP, or PDF; see the ScreenshotNeo API documentation for parameters and configuration.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://www.india.gov.in/ -o preview.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://www.india.gov.in/"}, timeout=90)
open("preview.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://www.india.gov.in/' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('preview.webp', new Uint8Array(await res.arrayBuffer()));
Use one request per verified URL and keep your own manifest of requested URL, final result metadata, capture time, and outcome for the batch. ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for AI agents. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. See current options and sign up for 1,000 free screenshots a month, no card required.
Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| Invalid URL or navigation failure | Missing scheme, malformed host, DNS issue, or the destination is unavailable. | Check the exact URL with the owning department, retain the submitted value, and correct the inventory rather than guessing a replacement. |
| Timeout after navigation starts | Slow page, stalled resources, or a readiness condition that never completes. | Keep the timeout bounded, try a less strict readiness event such as domcontentloaded, and record the timeout. Do not wait indefinitely for every request to finish. |
| HTTP 403, 429, or CAPTCHA | The site denied or limited automated access. | Stop requests to that site and review authorization and site policy. Do not rotate identities or bypass the challenge. |
| Image exists but page looks incomplete | The screenshot happened before client rendering, fonts, or key content appeared. | Wait for a stable, site-specific selector or a short bounded delay after DOM readiness; record the changed wait rule for comparable runs. |
| Blank or mostly empty capture | Navigation may have failed, content may require interaction, or the page may intentionally render later. | Check the recorded HTTP status, final URL, title, and browser error; use an authorized readiness condition. Keep it marked as questionable instead of counting it as a good preview. |
| Missing content near the bottom | Only the initial viewport was captured, or lower content is lazy-loaded. | Enable full-page capture or use a controlled scroll-and-wait routine, then apply that method consistently across the batch. |
| Huge images or browser memory pressure | Very long pages or excessive parallel tabs. | Prefer viewport captures for overview, process sequentially or with low concurrency, and split unusually long pages into separate handling. |
| Many near-duplicate files | URLs differ by tracking parameters or duplicate directory entries. | Deduplicate only after deciding which query parameters affect the page; preserve original inputs and document normalization. |
Frequently asked questions
Can a screenshot prove a government website is GIGW-compliant?
No. It records a visual state at one viewport and time. Use dedicated accessibility, link, markup, and other relevant evaluations for compliance work.
Should I use IGOD as the only source of URLs?
Use it to discover candidates, then verify each destination and ownership with the relevant department or another source, as IGOD's own policy recommends.
Is NIC S3WaaS the right tool for screenshotting existing sites?
No. It supports eligible government entities creating and hosting sites. Its template preview is part of that creation workflow, not a general URL-list capture service.
How often should I recapture the list?
Set the cadence from the decision you need to support. Point-in-time inventories need one run; change tracking needs dated repeat runs. Account for the extra traffic and storage, and follow each site's policies.


