ScreenshotNeo

BlogHow-to

How to Automate Scheduled Screenshots of Indian Municipal Service Websites

Build a repeatable screenshot job with Playwright, schedule it with cron or CI, and keep timestamps, capture settings, and failures interpretable.

By the ScreenshotNeo team4 October 202610 min read

Use Playwright to render each public municipal service page and save a screenshot, then use cron, a CI workflow, or a managed service to run the capture script on a schedule. Store each image with its capture time, final URL, and browser settings. Treat visual differences as signals for human review, and log failed captures separately from actual page changes.

This guide uses Python and Playwright for the do-it-yourself workflow. Before automating any particular municipal website, check that site’s current terms and automated-access rules. There is no universal permitted request rate established for these sites.

1. Decide what each capture should record

Start with a small inventory. For each page, decide what evidence you need and how often you need it. A screenshot captures a rendered view under specific conditions; it does not explain whether a difference is important.

Decision What to specify
URLs Exact public page URLs and the service or information each represents.
Capture scope Viewport, a specific element, or full page.
Cadence How often to capture, based on the monitoring need and the site’s access conditions.
Retention How long images, metadata, and failure records should be kept.
Review and alerts Who reviews meaningful differences and who is notified when capture fails.
Runtime conditions Browser version, operating system, viewport, scale, locale, and other settings needed to interpret or compare images.

Check the terms and automated-access restrictions for every site in the inventory. This guide does not establish permission to automate any specific municipal website.

2. Install Playwright and its browser

Create a project and install Playwright for Python. Install the browser binaries in the same environment that will run the scheduled job.

python -m venv .venv
source .venv/bin/activate
python -m pip install playwright
python -m playwright install chromium

On Windows, activate the environment with .venv\\Scripts\\activate. In a hosted CI environment, install the required browser and system dependencies as part of the workflow; a browser installed on a developer’s laptop is not automatically available to the runner.

3. Write a repeatable capture script

Save the following as capture_sites.py. It reads a JSON list of pages, captures a viewport or full page, and writes an image and a JSON sidecar for each attempt. The sidecar records the requested URL, final URL if navigation succeeded, timestamp, settings, and any failure.

import json
import os
import re
from datetime import datetime, timezone
from pathlib import Path
from urllib.parse import urlparse

from playwright.sync_api import sync_playwright

CONFIG_PATH = Path(os.environ.get("SITES_CONFIG", "sites.json"))
OUTPUT_DIR = Path(os.environ.get("CAPTURE_DIR", "captures"))
NAVIGATION_TIMEOUT_MS = int(os.environ.get("NAVIGATION_TIMEOUT_MS", "45000"))


def safe_name(value: str) -> str:
    parsed = urlparse(value)
    stem = parsed.netloc + parsed.path
    stem = re.sub(r"[^A-Za-z0-9._-]+", "_", stem).strip("._-")
    return (stem or "page")[:140]


def capture_all() -> int:
    sites = json.loads(CONFIG_PATH.read_text(encoding="utf-8"))
    OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
    failures = 0

    with sync_playwright() as playwright:
        browser = playwright.chromium.launch(headless=True)
        context = browser.new_context(
            viewport={"width": 1440, "height": 1000},
            device_scale_factor=1,
            locale="en-IN",
            color_scheme="light",
        )
        page = context.new_page()
        page.set_default_navigation_timeout(NAVIGATION_TIMEOUT_MS)

        for site in sites:
            url = site["url"]
            name = site.get("name", url)
            scope = site.get("scope", "full_page")
            selector = site.get("selector")
            captured_at = datetime.now(timezone.utc).isoformat(timespec="seconds")
            stamp = datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%SZ")
            base = OUTPUT_DIR / f"{safe_name(url)}-{stamp}"
            record = {
                "name": name,
                "requested_url": url,
                "captured_at_utc": captured_at,
                "browser": "Chromium via Playwright",
                "viewport": {"width": 1440, "height": 1000},
                "device_scale_factor": 1,
                "locale": "en-IN",
                "color_scheme": "light",
                "scope": scope,
                "status": "failed",
            }

            try:
                response = page.goto(url, wait_until="domcontentloaded")
                # Allow the page to finish its initial layout. This is not a
                # guarantee that every third-party or lazy resource is ready.
                page.wait_for_timeout(1000)
                record["final_url"] = page.url
                record["http_status"] = response.status if response else None

                if scope == "element":
                    if not selector:
                        raise ValueError("scope 'element' requires a selector")
                    page.locator(selector).screenshot(path=str(base.with_suffix(".png")))
                elif scope == "viewport":
                    page.screenshot(path=str(base.with_suffix(".png")))
                elif scope == "full_page":
                    page.screenshot(path=str(base.with_suffix(".png")), full_page=True)
                else:
                    raise ValueError(f"Unsupported scope: {scope}")

                record["status"] = "captured"
            except Exception as exc:
                failures += 1
                record["error"] = f"{type(exc).__name__}: {exc}"
                record["final_url"] = page.url
            finally:
                base.with_suffix(".json").write_text(
                    json.dumps(record, indent=2, ensure_ascii=False),
                    encoding="utf-8",
                )

        context.close()
        browser.close()

    return failures


if __name__ == "__main__":
    failed = capture_all()
    print(f"Capture failures: {failed}")
    raise SystemExit(1 if failed else 0)

The script deliberately records a navigation or capture error as a failed attempt. An access-denied page, browser error, or timeout should not be mistaken for proof that the municipal site’s content changed. A returned HTTP status is recorded as context; decide separately whether a particular status should count as a failed capture for your workflow.

Create sites.json beside the script. Use only URLs you have reviewed for access conditions:

[
  {
    "name": "Example public service page",
    "url": "https://example.gov.in/service",
    "scope": "full_page"
  },
  {
    "name": "Example service status panel",
    "url": "https://example.gov.in/status",
    "scope": "element",
    "selector": "main .service-status"
  }
]

The example hostnames are placeholders. Replace them with your reviewed target URLs. Supported scopes in this script are full_page, viewport, and element.

4. Choose viewport, element, or full-page capture

Scope Use it when Trade-off
Viewport The first visible screen is the record you need. Content below the viewport is not included.
Element A known service panel or page region is the subject. The selector must continue to identify the intended element.
Full page Below-the-fold content matters. Long pages can create large images and may include content that changes independently.

For element captures, use a selector tied to a stable part of the page and handle the case where it is missing. The sample treats a missing selector as a failed attempt rather than silently saving an unrelated screenshot.

5. Run the capture on a schedule

Option A: cron on a machine you manage

First confirm the script runs manually from the project directory. Then edit that machine’s crontab with crontab -e. This example runs daily at 06:00 UTC; choose a cadence and time appropriate to your use case and site policies.

0 6 * * * cd /path/to/project && /path/to/project/.venv/bin/python capture_sites.py >> /path/to/project/capture.log 2>&1

Use absolute paths, ensure the cron account can write to the output directory, and arrange log rotation and image retention. A machine that is powered off, disconnected, or out of disk space cannot produce the scheduled archive.

Option B: GitHub Actions

A CI workflow can install the browser, run the same script, and retain output as a workflow artifact. Save this as .github/workflows/capture.yml:

name: Scheduled website captures

on:
  workflow_dispatch:
  schedule:
    - cron: "0 6 * * *"

jobs:
  capture:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with:
          python-version: "3.12"
      - name: Install Playwright
        run: |
          python -m pip install playwright
          python -m playwright install --with-deps chromium
      - name: Capture pages
        run: python capture_sites.py
      - name: Upload capture archive
        if: always()
        uses: actions/upload-artifact@v4
        with:
          name: municipal-site-captures
          path: captures/
          if-no-files-found: warn
          retention-days: 30

Scheduled CI jobs may run later than the requested time. Retention settings also determine how long uploaded artifacts remain available; choose storage and retention that meet your archive needs. If you need a durable record, deliver images and sidecars to storage intended for that purpose rather than relying only on short-lived job artifacts.

6. Keep visual comparisons interpretable

Keep the browser, viewport, scale, color scheme, locale, and other relevant settings fixed between runs. Record them alongside each image. Differences can come from a real page update or from changes in the execution environment, such as the browser version or operating system.

If you add baseline image comparison, treat the result as a review signal. Have a person inspect meaningful changes and intentionally update the baseline when an accepted site change becomes the new expected appearance. Screenshots help inspect visual layout; they do not describe page structure or establish accessibility or legal compliance. Automated accessibility checks catch only some issues, so they do not replace manual assessment.

7. Reliability, performance, and cost

  • Keep the job small at first. Start with a low request rate and increase only when the use case and target site’s policies justify it. The available research does not establish a universal permitted rate.
  • Separate capture failures from page changes. Save errors, timestamps, and any available final URL and HTTP status. A timeout or blocked page is an operational outcome to investigate, not a visual change baseline.
  • Bound time and storage. Use a navigation timeout, keep a defined retention period, and remove or archive old images. Full-page PNGs can be large, especially for long pages.
  • Control repeatability. A fixed environment and capture settings make comparisons easier to interpret, but cannot guarantee pixel-identical rendering in every run.
  • Plan for maintenance. Browser binaries, CI dependencies, selectors, and target pages can change. Review failures and update the capture configuration when the page structure changes.
  • Account for operating cost. A self-managed job uses compute, storage, and maintenance time. CI artifact retention or a separate storage destination also has limits and costs set by the provider. The research does not establish a current price comparison for managed screenshot services.

8. Troubleshooting

Symptom Likely cause Fix
Executable doesn't exist or browser launch fails Chromium was not installed in the environment running the job. Run python -m playwright install chromium; in CI, install the browser and required system dependencies in the workflow.
Navigation timeout The page did not reach the chosen navigation milestone before the timeout, or the site was slow/unreachable. Check reachability and the failure record. Increase the timeout only if justified, and keep failures distinct from content changes.
Screenshot is blank or shows an error page The page may have returned an error, access-denied content, or incomplete rendering. Inspect the final URL, HTTP status, and page visually. Do not classify this image as evidence of a normal page update.
Element selector not found The page changed, loaded a different layout, or the selector is incorrect. Inspect the current page and update the selector only after verifying it identifies the intended service region. Preserve the failed attempt in the log.
Images or content are missing Lazy-loaded content or delayed resources may not be ready after navigation. Wait for a known selector or a page-specific condition before capture. A fixed delay can help with a known short delay but is not a guarantee that every resource has loaded.
Images differ across runs without an obvious site change Browser, operating system, viewport, scale, fonts, or other rendering conditions changed. Pin and record the runtime and capture settings; review differences before accepting a new baseline.
Cron runs manually but not on schedule Cron uses a different working directory, environment, PATH, or account permissions. Use absolute paths, invoke the virtual environment’s Python directly, and check the scheduled log and output-directory permissions.
CI job passes but no archive appears The capture output path is wrong, the script wrote elsewhere, or no files were created. Check the configured output directory and artifact path. The workflow uses if: always() so available failure records can also be uploaded.

9. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. It can capture a URL in one request and return an image or PDF. The request below saves a WebP screenshot; see the API documentation for request options and scheduling integrations you build around the call.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

Replace the example URL with a municipal page whose access conditions you have checked. Schedule the request with your existing scheduler and retain the resulting image with its timestamp and capture settings. ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are never billed; its MCP server lets AI agents take screenshots; and 1,000 screenshots a month are free with no card, with paid plans starting at $5 for 3,000. Sign up for free and get 1,000 screenshots a month with no card.

10. Frequently asked questions

Should I use screenshots to monitor every municipal page?

Capture pages and regions tied to a clear review need. A large archive without an owner, retention plan, or response process is difficult to interpret.

Does a screenshot prove a service is available to residents?

No. It records what the browser rendered at capture time and under its access conditions. It does not prove that a resident can complete a transaction or that the service is accessible.

Can I compare screenshots to detect a change automatically?

Yes, image comparison can flag visual differences, but the result should be reviewed. Rendering conditions and incidental page changes can also produce differences.

What should I do if a site blocks the capture?

Record it as a failed or blocked capture, review the site’s access rules, and contact the site’s operator if appropriate. Do not treat the block page as the municipal service page’s new baseline.