Capture Indian Exam Result Webpages in Bulk with Python Selenium
Capture a reviewed list of exam result pages with Selenium, wait for the right content, and record each success or failure for reliable follow-up.
Use Selenium to visit a reviewed list of result URLs, wait for a portal-specific result condition, and save a per-page record that includes success or failure. Navigation finishing does not prove JavaScript-rendered result content is ready. There is no universal selector or permission rule for Indian exam portals: inspect each portal’s current markup and terms, and adapt its configuration individually.
This guide uses Python and Chrome. The example selectors and timeouts are placeholders to adapt after reviewing each authorized target page; they are not verified selectors or recommended values for every portal. The Government of India’s results gateway describes results from multiple examination bodies and links to terms and policies, which is one reason to check each source separately.
1. Define a reviewed batch and its success condition
Start with a bounded list of official or otherwise authorized URLs. For each portal, specify:
- The requested URL and the fields or visual evidence needed.
- A success condition tied to the expected result view, such as a result heading, status message, or known result row.
- Portal-specific access guidance, selector configuration, and any official download or API route.
Do not assume every page has the same markup, that public visibility permits bulk automation, or that a high request volume is acceptable. Use modest sequential pacing consistent with the portal’s stated rules. The sources here establish no universal rate limit.
2. Install Selenium and prepare the browser
Install Selenium in a virtual environment. Selenium Manager can manage browser drivers for supported setups; for locked-down environments, follow the setup instructions for the browser and driver available there.
python -m venv .venv
# Activate the environment for your shell, then:
python -m pip install selenium
For example, activate with source .venv/bin/activate on macOS or Linux, or .venv\Scripts\activate in Windows Command Prompt. The script below uses Chrome. A browser must be installed in the execution environment.
3. Capture pages with explicit waits and an audit record
Save this as capture_results.py. Replace each example URL and CSS selector with values verified for that portal. This runnable structure captures each URL independently, logs navigation or wait failures, and continues through the bounded input list. It stores page text and a browser-window screenshot on success; review whether those artifacts contain personal information before retaining or sharing them.
from datetime import datetime, timezone
import json
import re
from pathlib import Path
from selenium import webdriver
from selenium.common.exceptions import (
TimeoutException,
WebDriverException,
)
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
# Replace these examples with reviewed, authorized URLs and selectors.
PAGES = [
{
"url": "https://example.gov.in/results",
"result_selector": "YOUR_RESULT_SELECTOR",
},
]
OUTPUT = Path("captures")
PAGE_LOAD_TIMEOUT_SECONDS = 30
RESULT_WAIT_SECONDS = 20
def safe_name(value: str) -> str:
value = re.sub(r"[^A-Za-z0-9._-]+", "_", value)
return value[:100] or "page"
def make_driver():
options = webdriver.ChromeOptions()
# Keep the default page-load strategy (normal) unless you have
# a robust explicit readiness condition for the specific portal.
return webdriver.Chrome(options=options)
def capture_one(driver, page: dict) -> dict:
requested_url = page["url"]
selector = page["result_selector"]
record = {
"requested_url": requested_url,
"captured_at_utc": datetime.now(timezone.utc).isoformat(),
"status": "error",
}
name = safe_name(requested_url)
try:
driver.set_page_load_timeout(PAGE_LOAD_TIMEOUT_SECONDS)
driver.get(requested_url)
record["resolved_url"] = driver.current_url
record["title"] = driver.title
result = WebDriverWait(driver, RESULT_WAIT_SECONDS).until(
EC.visibility_of_element_located((By.CSS_SELECTOR, selector))
)
record["status"] = "success"
record["result_text_file"] = f"{name}.txt"
(OUTPUT / record["result_text_file"]).write_text(
result.text, encoding="utf-8"
)
record["screenshot_file"] = f"{name}.png"
if not driver.save_screenshot(str(OUTPUT / record["screenshot_file"])):
record["screenshot_file"] = None
record["screenshot_note"] = "Browser did not save screenshot"
except TimeoutException as exc:
record["status"] = "timeout"
record["error"] = str(exc) or "Navigation or result condition timed out"
record["resolved_url"] = driver.current_url
record["title"] = driver.title
except WebDriverException as exc:
record["status"] = "webdriver_error"
record["error"] = str(exc)
try:
record["resolved_url"] = driver.current_url
record["title"] = driver.title
except WebDriverException:
pass
return record
def main():
OUTPUT.mkdir(parents=True, exist_ok=True)
records = []
driver = make_driver()
try:
# Do not configure an implicit wait when using WebDriverWait.
driver.implicitly_wait(0)
for page in PAGES:
records.append(capture_one(driver, page))
finally:
driver.quit()
(OUTPUT / "audit.json").write_text(
json.dumps(records, ensure_ascii=False, indent=2),
encoding="utf-8",
)
print(f"Recorded {len(records)} page outcomes in {OUTPUT / 'audit.json'}")
if __name__ == "__main__":
main()
Run it with python capture_results.py. Each entry in audit.json records the requested URL, UTC capture timestamp, status, and—when available—the resolved URL, title, evidence filenames, or error. Avoid logging candidate identifiers or full result content in the audit file unless required.
4. Choose a readiness condition that proves the result appeared
Selenium’s navigation readiness and a page’s application-level readiness are different. A document can reach its ready state before client-side code adds or updates the results panel. Use WebDriverWait with a condition appropriate to that portal, such as visibility of a result element, presence of a completion marker, or a custom predicate that checks a known state.
The sample uses visibility_of_element_located. Other useful expected conditions include presence in the DOM, clickability, or invisibility of a loading indicator. Choose the condition that matches what a successful result page actually looks like. Treat timeout as an explicit outcome, not as an empty result.
A fixed sleep can waste time when a page is fast and still fail when it is slow. Selenium warns that mixing implicit and explicit waits can produce unpredictable timeout behavior; keep implicit wait at zero when using explicit waits.
5. Configure page loading and browser behavior
| Choice | Behavior | When to consider it |
|---|---|---|
normal (default) |
Waits for document readiness state complete. |
Good starting point when pages load normally and the result condition is checked afterward. |
eager |
Returns when readiness reaches interactive; some resources may still load. |
Consider only when waiting for nonessential assets is costly and the explicit result condition is reliable. |
none |
Does not block on document readiness. | Requires careful synchronization; otherwise navigation and capture can race. |
These strategies change when navigation returns; none replaces checking for the result state. If using eager or none, test the readiness condition against each portal and keep timeouts bounded. Also set a finite page-load timeout so a stalled navigation does not hold the batch indefinitely.
6. Adapt for different portals and page structures
Use a per-portal configuration
Keep selectors and success conditions beside each portal’s URL rather than assuming one selector works everywhere. If a portal changes its markup, update that portal’s adapter and preserve the failed record so the change is visible.
Check for iframes
If an element is visible in the browser but Selenium cannot locate it, inspect whether it is inside an iframe. Switch into the relevant frame before locating the field, then switch back to the top-level document:
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
wait = WebDriverWait(driver, 15)
wait.until(EC.frame_to_be_available_and_switch_to_it(
(By.CSS_SELECTOR, "YOUR_IFRAME_SELECTOR")
))
try:
result = wait.until(EC.visibility_of_element_located(
(By.CSS_SELECTOR, "YOUR_RESULT_SELECTOR")
))
print(result.text)
finally:
driver.switch_to.default_content()
Frame and result selectors are portal-specific. Nested frames require switching through each parent frame in order.
Choose text extraction or visual evidence
DOM text is useful when you need specific visible text or fields and can verify their structure. A PNG is useful as visual evidence of the browser window at that moment. A screenshot is not proof that the entire page or every lazy-loaded image is visible. Selenium’s save_screenshot captures the current window; use a separate, portal-appropriate approach if full-page evidence is required.
Handle downloads separately
If a portal provides an official downloadable result file, consider using that supported route where appropriate. Remote Selenium Grid environments have separate managed-download behavior; an immediate file listing is not proof that a download has completed. Wait for a known completion condition and record the file outcome.
7. Troubleshoot common failures
| Symptom | Likely cause | Action |
|---|---|---|
| Navigation returned, but result content is absent | JavaScript populated the page after document readiness. | Wait for a result-specific element or state with an explicit wait. |
| Wait times out on a selector | Selector is outdated, the page is a different state, or the element is in a frame. | Inspect the current page and selector; check frames; record timeout rather than treating it as no result. |
| Navigation hangs on assets | Ancillary resources delay the chosen navigation readiness state. | Keep a page-load timeout. Consider eager only with a robust explicit condition, then validate per portal. |
| CAPTCHA, login, denial, or unexpected page appears | The portal presented an access or verification state, or the URL redirected. | Record the resolved URL and state, stop or follow the portal’s official process. Do not attempt to defeat access controls. |
| Screenshot exists but misses content | The window viewport does not show the full page, or content had not loaded. | Wait for the required content and treat the PNG as a viewport image, not complete-page proof. |
| WebDriver cannot start | Browser installation, driver setup, permissions, or runtime environment is incompatible. | Check the installed browser and Selenium setup instructions; verify the environment can launch that browser. |
| Batch stops after one problematic page | An exception escaped the per-page handler or cleanup interrupted the run. | Keep each page isolated, catch expected navigation and wait failures, and always close the driver in finally. |
8. Performance, reliability, and privacy
- Bound the work: use a reviewed list and sequential or modest batches governed by the source’s rules. There is no universal safe rate established for all portals.
- Bound the waits: set page-load and result-condition timeouts. Record timeout distinctly from success and from an empty result.
- Make reruns auditable: persist per-page status and evidence filenames. On rerun, decide explicitly whether to overwrite evidence or create a new timestamped run directory.
- Reduce browser overhead carefully: reuse one driver for a bounded batch, as in the sample, while isolating outcomes per URL. If a browser becomes unhealthy, restart it and record the interruption rather than silently skipping remaining pages.
- Protect result data: screenshots and text may expose personal exam information. Minimize captured fields, restrict artifact access, and set a retention period suitable for the purpose.
The browser workflow has operational costs in browser runtime, storage, and review of exceptions. No throughput or accuracy benchmark is established here; measure on the specific authorized pages and environment. Check current portal terms and technical guidance when the workflow is first run and when it is updated.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. For a public result page where a screenshot is the evidence you need, one GET request can return an image or PDF. It does not replace portal-specific authorization, result extraction, or a browser workflow that must inspect page data.
cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://results.gov.in/ -o result.webp
Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://results.gov.in/"},
timeout=90,
)
r.raise_for_status()
open("result.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://results.gov.in/',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('result.webp', res);
See the ScreenshotNeo API documentation for parameters and response details. Before capture, it accepts cookie or consent banners like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for AI agents. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000.
Sign up for 1,000 free screenshots a month, with no card.
FAQ
Can one Selenium selector work across Indian exam portals?
Do not assume so. Inspect each current portal and maintain a separate selector and success condition where its structure differs.
Does a screenshot prove that the result is correct?
No. It records what appeared in the browser window. Validate the expected result state and keep the requested and resolved URLs with the evidence.
Should I use a fixed sleep to wait for results?
Use a condition-based explicit wait as the primary readiness check. A fixed delay does not confirm that the result state appeared.
Can I automate any publicly accessible result page?
Public accessibility alone does not establish permission for automated collection. Check the current terms and use the portal’s official process.


