How to Automate Screenshots of Indian School Websites with Selenium
Use Selenium to capture authorized school webpages as PNGs, wait for dynamic content, handle full-page needs, and run repeatable batches reliably.
Selenium can automate screenshots of Indian school websites by opening an authorized page in a real browser, waiting for the content you need, and saving the browser window as a PNG. In Python, the core call is driver.save_screenshot("screenshots/page.png"). That captures the current browser window; it does not automatically mean the whole document. Selenium’s Python Firefox API also documents a full-page screenshot method, but that capability is browser-specific.
This guide uses Python and Chrome for the ordinary viewport capture. The example URL is a placeholder: choose the exact public page you are allowed to capture, decide whether the viewport or entire document is needed, and adapt the wait condition to the page. For UDISE+ or any authenticated portal, establish authorization first and consider its official API rather than treating account-protected pages as public scraping targets.
1. Decide what to capture and confirm access
Before automating, write down the target URL, expected page state, browser, viewport size, and output path. “A school website” is not one uniform layout or rendering pattern: the right wait condition and screenshot dimensions depend on the actual page.
- Use a viewport screenshot when you need what a visitor sees in the browser window at a particular size. Selenium’s common WebDriver method saves the current window to a PNG.
- Use a full-document screenshot when content below the fold must appear in one image. Selenium documents a dedicated method in its Python Firefox API; do not assume it works identically in every browser.
- Use an official API when your goal is structured school data and an appropriate API is available. UDISE+ has an API portal. Its School Directory Management entry page asks users to continue securely with a UDISE+ account, so do not assume directory data is openly accessible.
- Confirm permission and site rules for independent school websites. The available sources do not establish a universal permission rule or the access terms for every school site.
UDISE+ describes India’s school education information system. Its homepage reports 14.67 lakh schools, 10.05 lakh government schools, 24.72 crore students, and 1.02 crore teachers for academic year 2025–26. These are portal figures about the education system, not counts of school websites or evidence that all school data is public. UDISE+
2. Install Selenium and a browser
Install Selenium in the Python environment that will run your script:
python -m pip install selenium
Install a supported browser such as Chrome or Firefox. Selenium’s browser setup and driver management can vary with the installed browser and environment; consult its current WebDriver documentation if the browser does not start. The script below uses Chrome and Selenium’s Selenium Manager support where available.
3. Capture a page with Python and Selenium
This complete example creates the output directory, opens an example page, waits for the document to reach the browser’s complete state, sets a repeatable viewport, checks the Boolean result from save_screenshot(), and closes the browser even if an error occurs. Replace the example URL and, when possible, replace the generic document-ready wait with a condition for the specific content you need.
from pathlib import Path
from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
TARGET_URL = "https://example.org/"
OUTPUT = Path("screenshots/school-homepage.png")
VIEWPORT = {"width": 1440, "height": 1000}
OUTPUT.parent.mkdir(parents=True, exist_ok=True)
options = webdriver.ChromeOptions()
# Uncomment for a headless run, such as on a server:
# options.add_argument("--headless=new")
# Selenium Manager may locate or manage the browser driver automatically.
driver = webdriver.Chrome(options=options)
try:
driver.set_window_size(VIEWPORT["width"], VIEWPORT["height"])
driver.get(TARGET_URL)
# This waits for the document lifecycle state. It does not guarantee that
# every image, client-side widget, or page-specific data request is ready.
WebDriverWait(driver, 30).until(
lambda browser: browser.execute_script("return document.readyState") == "complete"
)
saved = driver.save_screenshot(str(OUTPUT))
if not saved:
raise OSError(f"Selenium could not write screenshot to {OUTPUT}")
print(f"Saved {OUTPUT}")
finally:
driver.quit()
Selenium documents save_screenshot(filename) as saving the current window to PNG. It recommends a full path ending in .png; the method returns False if an I/O error occurs. The script uses a path relative to the current working directory, so use an absolute path if the process may start from a different directory. Selenium Python WebDriver API
Wait for the content you actually need
document.readyState == "complete" is a useful baseline, but it does not prove that a client-rendered section, delayed image, carousel, or API-backed school detail has appeared. If the page has a stable element that marks readiness, wait for it:
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
WebDriverWait(driver, 30).until(
EC.visibility_of_element_located((By.CSS_SELECTOR, "main .school-profile"))
)
Change the selector to an element that exists on the target page and indicates the state you intend to capture. If the page has no reliable marker, a short fixed delay can be used as a fallback, but it is less reliable because page response times vary:
import time
time.sleep(2)
Avoid treating a successful navigation as proof that every visual element is ready. For a meaningful batch, inspect representative pages and choose a page-specific wait condition.
4. Capture the full page in Firefox when needed
The standard save_screenshot() call captures the current window. Selenium’s Python Firefox WebDriver API documents save_full_page_screenshot(filename) for a full-document image. Use Firefox explicitly for this path and check the result just as you would for a viewport capture:
from pathlib import Path
from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
TARGET_URL = "https://example.org/"
OUTPUT = Path("screenshots/school-full-page.png")
OUTPUT.parent.mkdir(parents=True, exist_ok=True)
driver = webdriver.Firefox()
try:
driver.get(TARGET_URL)
WebDriverWait(driver, 30).until(
lambda browser: browser.execute_script("return document.readyState") == "complete"
)
saved = driver.save_full_page_screenshot(str(OUTPUT))
if not saved:
raise OSError(f"Firefox could not write screenshot to {OUTPUT}")
finally:
driver.quit()
Check the current Firefox API documentation and your Selenium version when using this method. Do not copy it into a Chrome workflow and assume that browser supports the same API. Selenium Python Firefox WebDriver API
5. Run a repeatable batch
For multiple authorized URLs, keep the browser lifecycle outside the loop to avoid starting a fresh browser for every page. Use distinct output filenames, wait for each page’s content, record failures, and always quit the driver. This example continues after a failed page and prints a simple per-URL result:
from pathlib import Path
from urllib.parse import urlparse
import re
from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait
URLS = [
"https://example.org/",
"https://www.example.com/",
]
OUT_DIR = Path("screenshots")
OUT_DIR.mkdir(parents=True, exist_ok=True)
def filename_for(url: str) -> str:
host = urlparse(url).netloc or "page"
safe_host = re.sub(r"[^A-Za-z0-9.-]+", "_", host)
return f"{safe_host}.png"
options = webdriver.ChromeOptions()
# options.add_argument("--headless=new")
driver = webdriver.Chrome(options=options)
try:
driver.set_window_size(1440, 1000)
for url in URLS:
destination = OUT_DIR / filename_for(url)
try:
driver.get(url)
WebDriverWait(driver, 30).until(
lambda browser: browser.execute_script("return document.readyState") == "complete"
)
if not driver.save_screenshot(str(destination)):
raise OSError(f"Could not write {destination}")
print(f"OK {url} -> {destination}")
except Exception as exc:
print(f"FAILED {url}: {type(exc).__name__}: {exc}")
finally:
driver.quit()
For production batches, persist results and errors in a log or structured file, and decide whether a failed URL should be retried. Avoid putting credentials in source code or logs. A single driver reduces startup work, but a browser crash can affect later URLs; for long batches, process URLs in bounded groups and restart the browser between groups if operational experience shows that is necessary.
6. Handle login and protected pages carefully
If an intended page requires authentication, use only credentials and access that you are authorized to use. Keep secrets outside the script, such as in environment variables or a secret manager, and do not publish screenshots that expose student or staff information without appropriate authorization.
UDISE+ is a specific example where the School Directory Management entry page requests secure account continuation. The official portal also describes APIs for school and related data. Determine whether the official API meets your need before automating an account workflow. UDISE+ School Directory Management · UDISE+ API portal
7. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. One GET request can return an image or PDF. For example, this cURL request saves a WebP screenshot:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.org -o shot.webp
Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.org"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.org' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));
See the ScreenshotNeo API documentation for request options. Cookie banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status. Its MCP server provides screenshot tools for AI agents. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.
Sign up for ScreenshotNeo and get 1,000 free screenshots a month with no card.
8. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
save_screenshot() returns False |
The path is invalid, the directory does not exist, or the process cannot write there. | Create the parent directory, use a writable absolute path ending in .png, and check the return value. |
| Browser does not start or driver error appears | The browser is missing, incompatible, or unavailable to Selenium’s driver setup. | Install or update the browser and Selenium; consult Selenium’s current browser setup documentation and check the runtime environment. |
| Screenshot is blank or missing page content | The page had not rendered the target content when capture ran, or navigation did not reach the intended state. | Wait for a page-specific visible element, verify the final URL and page state, and capture again. |
| Content below the fold is absent | The ordinary method captures the current window, not necessarily the entire document. | Use a full-page method supported by the chosen browser, such as the documented Python Firefox method, or capture the page in sections. |
| One URL fails and stops the batch | An exception escaped the loop, such as a timeout or navigation error. | Handle exceptions per URL, record the failure, and decide whether to retry. Keep driver.quit() in a finally block. |
| Images or delayed widgets are missing | Document readiness occurred before those resources or client-side components finished. | Wait for the relevant element or use a site-specific readiness condition; do not rely on one fixed delay for every site. |
| Access denied or sign-in page appears | The destination requires an account, permission, or a supported access route. | Confirm authorization and use the official API where suitable. Do not attempt to bypass access controls. |
9. Performance, reliability, and cost
- Browser startup: Starting the browser once for a batch avoids repeated startup overhead. Reusing a driver also means a driver failure may interrupt several captures, so log each URL and keep batches manageable.
- Waits: A precise element wait can reduce unnecessary idle time while avoiding captures that are too early. A fixed sleep is simple but can be too short on slow pages and wasteful on fast ones.
- Output size: A larger viewport and full-document image produce larger files. Set the viewport intentionally and choose the screenshot scope that answers the task.
- Reliability: Check screenshot return values, use timeouts, close the driver in a
finallyblock, and retain a record of successes and failures. Validate representative pages before a large run. - Cost: Selenium itself is software, but running a browser consumes local or server compute, storage, and bandwidth. The research sources do not provide a universal cost estimate. ScreenshotNeo offers a no-card free plan of 1,000 shots per month; paid plans begin at $5 for 3,000 shots, with yearly billing giving two months free. All listed features are available on every plan.
10. Frequently asked questions
Does Selenium save screenshots as PNG?
Yes. Selenium’s Python save_screenshot() documentation describes a PNG image of the current window.
Does this script capture every school website in India?
No. The examples capture only the URLs you provide and can access. Each site’s layout, access requirements, and rendering behavior may differ.
Can I screenshot UDISE+ school data without signing in?
The cited School Directory Management page asks users to continue securely with a UDISE+ account. Check the official API portal and the access requirements for your use case.
Can I use the Firefox full-page method in Chrome?
Do not assume so. The cited full-document method is documented for Selenium’s Python Firefox WebDriver API.


