How to Schedule Selenium Screenshots of Indian Panchayat Websites
Schedule Selenium to capture public Panchayat pages with page-specific waits, timestamped files, reliable cleanup, and a deployment-specific scheduler.
To schedule Selenium screenshots of Indian Panchayat websites, write a script that opens each public URL in a headless browser, waits for the page content you need, saves a timestamped screenshot, and always closes the browser. Then run that script on a machine or job service with a scheduler configured for your required frequency. The example below uses Python and Chrome; scheduling commands depend on the operating system or service hosting it.
Indian Panchayat-related pages do not all share one layout. The Ministry of Panchayati Raj describes eGramSwaraj as an integrated information system with profile, planning, progress reporting, accounting, asset-directory, and user-management modules. Choose the exact public page you want to monitor and a readiness condition that fits that page. [Ministry of Panchayati Raj: eGramSwaraj; eGramSwaraj portal]
1. Choose public pages and capture expectations
Write down each target URL and what a successful screenshot should contain: for example, a page heading, a report table, or a visible status panel. Use public pages you are authorized to monitor. This guide is for viewing and capturing pages; do not automate login, CAPTCHA solving, or privileged administrative actions.
eGramSwaraj is one possible target, alongside state or local Panchayat websites. NIC describes eGramSwaraj as a web-based application for Panchayat digitisation and says it is available 24×7; that description does not guarantee that every local website has the same template or behavior. [NIC eGramSwaraj overview]
2. Install Selenium and configure Chrome
Use a supported Python version and install Selenium in the same environment that will run the scheduled job:
python -m venv .venv
# Linux or macOS
. .venv/bin/activate
# Windows PowerShell: .venv\Scripts\Activate.ps1
python -m pip install --upgrade selenium
Selenium WebDriver controls a real browser session and can drive browsers locally or remotely. Its Chrome documentation lists --headless=new as a commonly used argument. Browser and driver compatibility matter when deploying, so keep the browser installation and Selenium version maintained together. [Selenium WebDriver; Selenium Chrome options]
Runnable Python capture script
Save as capture_panchayat.py. Replace the example URL and the readiness selector with the public pages and stable page element you intend to capture. The script creates an output directory, uses a consistent viewport, records a UTC timestamp in each filename, and quits Chrome even if navigation or capture fails.
from datetime import datetime, timezone
from pathlib import Path
import json
from selenium import webdriver
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.common.by import By
URLS = [
"https://egramswaraj.gov.in/", # Replace with the exact public page to monitor.
]
OUTPUT_DIR = Path("screenshots")
WAIT_SECONDS = 30
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
options = webdriver.ChromeOptions()
options.add_argument("--headless=new")
options.add_argument("--window-size=1440,1000")
# Selenium Manager can manage the driver for a local Chrome installation.
driver = webdriver.Chrome(options=options)
manifest = []
try:
for index, url in enumerate(URLS, start=1):
driver.get(url)
# Baseline navigation check. Prefer a meaningful, page-specific element below.
WebDriverWait(driver, WAIT_SECONDS).until(
lambda d: d.execute_script("return document.readyState") == "complete"
)
# Example if the target page has a stable heading; replace or remove this
# line after inspecting the actual public page:
# WebDriverWait(driver, WAIT_SECONDS).until(
# EC.visibility_of_element_located((By.CSS_SELECTOR, "main h1"))
# )
stamp = datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%SZ")
filename = f"page-{index}-{stamp}.png"
path = OUTPUT_DIR / filename
if not driver.save_screenshot(str(path)):
raise RuntimeError(f"Screenshot could not be saved: {path}")
manifest.append({"url": url, "file": str(path), "captured_at_utc": stamp})
finally:
driver.quit()
(OUTPUT_DIR / "latest-manifest.json").write_text(
json.dumps(manifest, indent=2), encoding="utf-8"
)
print(f"Saved {len(manifest)} screenshot(s) in {OUTPUT_DIR.resolve()}")
Install the expected_conditions and By imports as shown if you enable the sample heading wait. Remove those imports if you keep only the ready-state baseline. The code deliberately does not assume a selector shared across Panchayat sites.
3. Wait for the visible content, not just navigation
Selenium’s navigation wait concerns document readiness. JavaScript can update the page after that state is reached, so a completed navigation is not proof that the content you want has appeared. Use an explicit wait for a page-specific condition such as a known heading becoming visible, a report table appearing, or a loading indicator disappearing. Selenium describes explicit waits as polling for a condition until it succeeds or times out. Avoid mixing implicit and explicit waits because Selenium warns that doing so can produce unpredictable wait times. [Selenium waits; Selenium driver options]
For example, after inspecting a page, replace the sample selector with one that is stable for that specific target:
WebDriverWait(driver, 30).until(
EC.visibility_of_element_located((By.CSS_SELECTOR, "main h1"))
)
If a page has no stable element, use a condition tied to the visual state you need, or use a short, bounded delay after the baseline navigation wait. Avoid an unbounded sleep: it wastes runtime on fast loads and may still be too short on slow ones. A selector that changes frequently can make otherwise healthy captures fail, so review it when the site changes.
4. Schedule the script on its host
There is no universal scheduler command: Linux cron or systemd timers, Windows Task Scheduler, and managed job services have different setup and logging behavior. Pick the host first, then follow that platform’s current official scheduling instructions. The research sources for this article do not establish a particular scheduler’s current setup steps, so this guide does not provide unverified commands.
- Confirm the script works manually from the same account and working directory the scheduled job will use.
- Use the absolute path to the Python interpreter in the virtual environment and to the script. Set the working directory explicitly if your scheduler supports it.
- Provide a writable destination for screenshots and logs. If using a relative output directory, confirm what directory the scheduler starts in.
- Choose a capture frequency that meets the monitoring need without repeatedly loading the site unnecessarily.
- Configure the scheduler to record standard output and errors, and decide how to notify an operator or retry after a failed run.
- Set a retention policy for old images and manifests so repeated runs do not fill the disk.
Test one scheduled run and inspect its log, exit status, output path, and screenshot before relying on a recurring schedule. For a job that must survive reboots or host changes, document the Python and browser versions, target URLs, schedule, and storage location.
5. Capture multiple pages and identify the results
The script loops over URLs in one browser session to reduce repeated startup overhead. Each saved file has a sequence number and UTC time. The manifest records the URL and capture time so reviewers can identify what each image represents. If one failing page should not prevent later URLs from being captured, wrap each URL’s navigation and screenshot in its own error handler, record the failure in the manifest, and choose an appropriate nonzero exit status if any target failed.
Selenium’s Python WebDriver API documents saving the current browser view to a file with save_screenshot. This captures the current window viewport. Keep the window size fixed for comparisons; if you need a full-page image, verify the approach against the target browser and page, since viewport screenshots and full-page capture are different requirements. [Selenium Python Chrome WebDriver API]
6. Alternatives for where the browser runs
| Approach | Useful when | Considerations |
|---|---|---|
| Local browser on the scheduled machine | A small set of pages, direct control over files, and a host you already operate | You maintain the runtime, browser compatibility, disk space, logs, and scheduler configuration. |
| Remote WebDriver | The browser should run on a separate machine or browser host | Network access, remote browser lifecycle, file transfer, and service cost depend on the chosen environment. |
| Managed scheduled job | You want the scheduler and execution environment managed elsewhere | Confirm browser support, storage, logs, retry behavior, and pricing with that provider; this research does not establish a particular provider. |
Selenium documents local and remote WebDriver operation, but the cost and reliability of a deployment depend on the selected host and schedule. No screenshot-frequency benchmark or service-specific cost figure is established by the sources used here. [Selenium WebDriver]
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. One GET request returns an image or PDF. Its capture flow accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://egramswaraj.gov.in/ -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://egramswaraj.gov.in/"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://egramswaraj.gov.in/' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
These examples use the service’s GET endpoint and save the response body. Check the response status and headers when integrating it into a recurring job; use the docs for supported parameters and response details. ScreenshotNeo offers 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000. The same features are on every plan. Learn about ScreenshotNeo.
Sign up for 1,000 free screenshots a month with no card.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Chrome or driver fails to start | Chrome is missing, cannot run in the host environment, or browser and driver versions are incompatible. | Install a supported browser, check Selenium’s current Chrome guidance, and test startup under the same account and environment as the scheduled job. |
| Wait times out | The selector does not exist on this page, the page is still loading, or the site changed. | Inspect the public page, choose a stable page-specific condition, and set a bounded timeout suitable for the host and page. |
| Screenshot is blank or incomplete | Navigation completed before the desired JavaScript-rendered content appeared, or the page failed to load. | Wait for the actual visible content; save diagnostic logs and verify the URL and screenshot manually. |
| Images differ between runs | Viewport, dynamic content, timestamps, animations, or page data changed. | Keep viewport and browser configuration consistent; compare only stable regions or capture at a consistent point in the page’s update cycle. |
| Script works manually but not when scheduled | The job uses another working directory, Python environment, account, or output permission. | Use absolute paths, set the job’s working directory, activate or directly invoke the intended virtual environment, and inspect captured stdout and stderr. |
| Repeated runs exhaust disk space | Images and logs accumulate without retention. | Define a retention policy and monitor available storage. |
| One URL prevents the rest from running | An exception exits the URL loop. | Handle errors per URL, record failures, continue where appropriate, and report the overall job as failed if required captures are missing. |
Reliability, performance, and cost
- Reliability: use explicit waits, a guaranteed
driver.quit(), clear logs, and a defined retry or alert policy. Do not mix Selenium implicit and explicit waits. - Performance: a browser startup has overhead; reusing one session across a modest URL list avoids starting Chrome for every page. Long waits and unnecessary capture frequency increase runtime. Measure the job in its actual host environment rather than assuming a fixed duration.
- Storage: timestamped images are useful for history but grow over time. Retain only the history you need and keep manifests and logs manageable.
- Cost: Selenium itself does not establish hosting cost. Compute, storage, and any remote browser or scheduler charges depend on your deployment and run frequency. ScreenshotNeo has a free allowance of 1,000 shots monthly with no card; paid plans are Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000. Yearly billing gives two months free.
FAQ
Can this run every day?
Yes. Configure the host’s scheduler for the desired frequency and verify a scheduled run’s logs and output. The exact setup depends on the host.
Does the script capture an entire long page?
save_screenshot saves the current browser view. Full-page capture needs a browser-specific or tool-specific approach and should be validated on the pages you monitor.
Can I use the same selector for every Panchayat site?
Only if the pages actually share that stable element. Inspect each target and use page-specific waits where their layouts differ.
What does eGramSwaraj usage data say about screenshot automation?
Nothing directly. A Ministry release dated 17 March 2026 reports that, as of 11 March 2026, 2,54,604 of 2,64,211 Gram Panchayats and equivalents (96.36%) had uploaded FY 2025–26 GPDPs, and 2,42,871 (91.92%) had made ₹38,491 crore in payments through eGramSwaraj–PFMS. Those are portal activity figures, not Selenium coverage or screenshot success statistics. [Ministry of Panchayati Raj release via PIB]


