ScreenshotNeo

BlogHow-to

How to Schedule Screenshots of Indian Government Service Websites

Use Playwright and your system scheduler to save dated screenshots of public service pages. Keep captures consistent and check each site’s access guidance first.

By the ScreenshotNeo team4 October 202610 min read

Use a browser automation script to open a public service page, save a timestamped screenshot, and run the script with a scheduler on the computer or server where it will execute. Playwright can capture a viewport or a full page. Keep the browser, operating system, viewport, and capture settings consistent when comparing images over time.

Before scheduling captures, read the target site’s own terms and access guidance. Official guidance reviewed for this article does not establish blanket permission to automate screenshots across every Indian government website. Keep the job passive: do not log in, submit forms, or interact with a service transaction. If the site denies access, stop and contact its owner.

1. Choose the page and define what you need to observe

Start with a specific public informational or service landing page, and write down why you need recurring captures: for example, to keep a dated visual record or review page changes. The National Government Services Portal describes itself as a single-window interface to services from government departments. It can help locate a service, but its linking policy does not grant permission to automate capture on every linked site.

Check the individual department or service domain for its access rules, terms, robots guidance, or contact details. The Government of India’s GIGW guidance covers website quality and accessibility across government levels; it is not a universal authorization for automated screenshots on each site.

  • Capture only public pages that you are permitted to access.
  • Do not include credentials, personal data, form submission, or transaction steps in a passive monitoring job.
  • Choose a modest cadence suited to your purpose and the site’s published guidance. There is no universally appropriate interval.
  • If access is blocked or the site asks automated clients to stop, stop the job and ask the site owner.

2. Set up a repeatable Playwright capture

The following Node.js example uses Playwright to open a public page, wait for its load event, and save a full-page PNG in a directory organized by date. It also writes a JSON sidecar with the URL, timestamp, and capture settings. Replace the example URL with a page you are permitted to capture.

mkdir government-page-captures
cd government-page-captures
npm init -y
npm install playwright
npx playwright install chromium

Save this as capture.mjs:

import { chromium } from 'playwright';
import { mkdir, writeFile } from 'node:fs/promises';
import path from 'node:path';

const targetUrl = process.env.TARGET_URL ?? 'https://services.india.gov.in/';
const outputRoot = process.env.OUTPUT_DIR ?? './captures';
const viewport = { width: 1365, height: 900 };
const capturedAt = new Date();
const date = capturedAt.toISOString().slice(0, 10);
const time = capturedAt.toISOString().replaceAll(':', '-').replace(/\.\d{3}Z$/, 'Z');
const outputDir = path.join(outputRoot, date);
const imagePath = path.join(outputDir, `capture-${time}.png`);

await mkdir(outputDir, { recursive: true });
const browser = await chromium.launch({ headless: true });
try {
  const page = await browser.newPage({ viewport, deviceScaleFactor: 1 });
  const response = await page.goto(targetUrl, {
    waitUntil: 'load',
    timeout: 60_000,
  });

  if (!response) {
    throw new Error('Navigation returned no main-document response.');
  }
  if (!response.ok()) {
    throw new Error(`Page returned HTTP ${response.status()} ${response.statusText()}`);
  }

  await page.screenshot({ path: imagePath, fullPage: true, type: 'png' });
  const metadata = {
    capturedAt: capturedAt.toISOString(),
    url: page.url(),
    status: response.status(),
    viewport,
    deviceScaleFactor: 1,
    fullPage: true,
    browser: 'Playwright Chromium',
  };
  await writeFile(`${imagePath}.json`, JSON.stringify(metadata, null, 2));
  console.log(`Saved ${imagePath}`);
} finally {
  await browser.close();
}

Run it once manually before scheduling:

TARGET_URL='https://services.india.gov.in/' node capture.mjs

For an exact service URL, set TARGET_URL to that URL. Keep the URL and settings in the metadata so later reviewers can tell what each file represents. Avoid putting secrets in command history or metadata; this workflow should normally need no login.

3. Select viewport or full-page capture

A viewport screenshot records what is visible at a fixed browser size. A full-page screenshot extends below the initial viewport and is useful for long landing pages. Playwright documents both page screenshots and full-page capture in its Page API.

Choice Use it when Trade-off
Viewport You want a stable record of the initial view or a fixed region. Content below the fold is omitted.
Full page You need to include content below the fold in one image. Long or dynamic pages can produce very tall images, and lazy content may not appear unless it loads while scrolling.

For lazy-loaded pages, scrolling through the page before capture can trigger more content to load, but also changes page state and can increase runtime. Only add such behavior if it is appropriate for the target page and does not interact with a service workflow. For example, a simple scroll-to-bottom pass can be added before the screenshot:

await page.evaluate(async () => {
  const step = Math.max(300, window.innerHeight);
  for (let y = 0; y < document.body.scrollHeight; y += step) {
    window.scrollTo(0, y);
    await new Promise(resolve => setTimeout(resolve, 150));
  }
  window.scrollTo(0, 0);
});
await page.screenshot({ path: imagePath, fullPage: true, type: 'png' });

This is not a guarantee that every lazy resource will load: some pages require a particular interaction or use changing content. Keep captures passive, and do not click controls that submit, authorize, or change a service request.

4. Schedule the script in its runtime environment

Use the scheduler available on the machine or hosting service that runs the browser. The machine must be available at the scheduled time, and the scheduler needs the same Node.js installation, project directory, permissions, environment variables, and installed Playwright browser binary as your manual run.

For a Unix-like system with cron, a daily job can be expressed as:

17 6 * * * cd /absolute/path/government-page-captures && TARGET_URL='https://services.india.gov.in/' /usr/bin/node /absolute/path/government-page-captures/capture.mjs >> /absolute/path/government-page-captures/capture.log 2>&1

Replace paths and the URL for your environment. Cron implementations differ; confirm the syntax and timezone used by your machine. On other operating systems, use that system’s built-in task scheduler or the scheduler provided by your hosting environment. The sources cited here document screenshot capture, not one scheduler setup that applies to every host.

  1. Run the script manually as the same account that will run the scheduled job.
  2. Use absolute paths for the project, Node.js executable, output directory, and log file.
  3. Set an explicit timezone if the scheduler supports it, and record timestamps in UTC as the script does.
  4. Confirm the output directory has enough space and that the scheduler account can write to it.
  5. Review logs and output after the first scheduled run, then check them periodically.

5. Make comparisons meaningful

Browser rendering can vary across host operating systems, browser versions, settings, hardware, power source, and headless mode. Playwright’s visual comparisons guidance recommends controlling the environment for screenshot comparisons. Use the same host, browser version, viewport, device scale factor, locale, timezone, and capture settings for your baseline and later runs.

Playwright browser binaries are associated with Playwright versions. After installing or updating the Playwright package, install the corresponding browser binaries as described in its browser documentation. Record package versions and update them deliberately so a browser update is not mistaken for a page change.

  • Keep a dated archive and a small metadata sidecar for each capture.
  • Choose PNG for lossless pixel comparison; use another format only when storage or downstream tooling calls for it.
  • Expect differences from rotating banners, time-dependent content, randomized recommendations, ads, and asynchronous widgets.
  • If using visual-diff tooling, compare captures from the same environment and review meaningful regions; raw pixel differences can be caused by rendering noise.
  • A screenshot is a visual record, not proof that links, forms, accessibility, or service transactions work.

6. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. Its GET endpoint can return an image or PDF from a URL; the examples below save an image response. See the ScreenshotNeo API documentation for request options. Use it only for a public page you are permitted to capture, and schedule the request with your chosen scheduler if you need recurring captures.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://services.india.gov.in/ -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://services.india.gov.in/"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://services.india.gov.in/',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', bytes));

ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Check the target site’s access rules before sending any capture request.

Sign up for 1,000 free screenshots a month, with no card required.

7. Troubleshooting

Symptom Likely cause What to do
Browser executable missing The Playwright browser binary was not installed for the package version in use. Run npx playwright install chromium in the project, and repeat after relevant package updates.
Works manually but not on schedule The scheduler uses another working directory, account, Node.js path, environment, or permissions. Use absolute paths, run as the scheduler account, log stdout and stderr, and verify write access.
Navigation timeout The site is slow, unreachable, waiting on resources, or blocking the browser. Check the URL and connectivity. A longer timeout may help with a slow response, but do not repeatedly retry a denied or blocked page; check with the site owner.
HTTP error or unexpected page The page moved, returned an error, or the host served a denial page. Inspect the response status and URL, confirm the current public page, and respect any access denial.
Screenshot is blank or incomplete Navigation ended before the page rendered, content loads later, or a resource failed. Inspect the page in a normal browser and use a suitable wait condition for the public content. Avoid fixed long sleeps unless the page has a known delay.
Images or sections missing in full-page capture Content may be lazy-loaded or conditional on scrolling. Consider a careful scroll pass before capture, then verify it does not trigger controls or transactions.
Images differ even though the page seems unchanged Browser, OS, font, viewport, scale, dynamic content, or headless settings changed. Restore the baseline environment and settings; document browser updates and time-dependent regions.
Output files overwrite each other The filename has insufficient timestamp precision or a fixed name. Use a timestamped filename and a date-based directory as in the example.
Disk usage grows unexpectedly Full-page images accumulate without retention. Estimate the archive size from your own captures and set a retention policy appropriate to your recordkeeping needs.

8. Performance, reliability, and cost

Each local capture starts or uses a browser process, navigates the page, waits for content, and writes an image plus metadata. Full-page screenshots and scroll-triggered lazy loading can take longer and use more memory than viewport captures. Schedule at a cadence the target site permits and that your runner can sustain. Avoid overlapping runs: if a capture can exceed the scheduler interval, add a lock or otherwise ensure only one instance runs at a time.

Reliability depends on the machine being online, browser dependencies being present, the site being reachable, and the scheduler running successfully. Capture logs, check for a newly created image, and alert through your own operational process if a run fails. A saved screenshot alone does not prove the page was fully healthy; retain the HTTP status and timestamp, and inspect failures rather than silently treating them as normal captures.

Local automation has no per-screenshot API charge, but it uses compute, storage, and maintenance time. A persistent server or managed runner may be useful when a personal computer cannot stay available, but the appropriate hosting and its cost depend on your environment. For ScreenshotNeo, the published plans are Free: 1,000 shots/month; Starter: $5 for 3,000; Growth: $15 for 15,000; Pro: $39 for 60,000; Scale: $99 for 250,000; Business: $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan. Only clean shots are billed; refer to the response’s billing headers when accounting for requests.

9. What a screenshot can and cannot tell you

GIGW’s scope describes guidance for Indian government websites and applications across central, state, district, and local levels, and it cites WCAG 2.1 among the standards informing its guidelines. A screenshot can help preserve a visual state, but it cannot establish that interactive controls work, that assistive technologies can use the page, or that time-dependent behavior is accessible. See the official GIGW scope and accessibility guidance when assessing those dimensions.

FAQ

Can I schedule captures of any Indian government service site?

Do not assume so. The sources reviewed do not establish blanket permission. Check the individual site’s access guidance and stop if access is denied.

Should I use screenshots as an accessibility audit?

No. They document appearance at a point in time, but cannot test keyboard operation, screen-reader behavior, or many other accessibility requirements.

How often should the job run?

Choose an interval that matches your monitoring purpose and the target’s published guidance. The reviewed sources specify no universal schedule.

Will a full-page capture always include every image?

No. Lazy-loaded or conditional content may not appear unless the page loads it during capture. Verify the output and use only passive loading behavior appropriate to the page.