How to automate Selenium screenshots of Indian real estate listing pages
Build a careful Selenium workflow to capture real estate listing pages or elements, with stable locators, waits, permission checks, and troubleshooting.
Use Selenium WebDriver to open each listing URL, wait for the content you need, and save either the browser window or a specific page element as a PNG. Start with a small, explicit URL list, choose a locator tied to stable markup, and confirm that the portal permits your planned automation and screenshot use. Selenium can control the browser; that capability does not itself grant permission to automate a site.
1. Confirm permission and define the capture scope
Before writing a capture loop, identify the portal, the exact pages, how often you will capture them, how long you will retain screenshots, and whether you will share or republish the images. Read the portal’s current terms and obtain permission where required. Do not assume that publicly visible pages can be automated or that screenshots can be republished.
For example, eRealtor’s terms state that they are governed by Indian law and in force from 25 July 2026. They prohibit bots, crawlers, or automated tools against the platform except ordinary public-search-engine indexing, and prohibit scraping, copying, or republishing listings and media without written permission. Those restrictions are specific to eRealtor; check the terms for the actual portal you plan to use. eRealtor terms and conditions
Keep the target URLs explicit. Selenium’s documentation lists link spidering among practices it discourages. A screenshot workflow for a known set of pages does not need to discover and follow every link on a portal. Selenium discouraged practices
2. Install Selenium and a browser
The example below uses Python and Chrome. Install the Selenium package and make sure Chrome is available in the execution environment. Selenium’s current driver-management behavior can locate or manage a compatible driver in common setups; restricted or offline environments may need an explicitly installed browser driver configured for your environment. See the Selenium WebDriver documentation.
python -m pip install selenium
Save the following as capture_listings.py. Replace the sample URLs and selectors with pages and markup you are authorized to capture.
3. Capture a window or a page element
from pathlib import Path
from urllib.parse import urlparse
import re
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
URLS = [
"https://example.com/property/listing-1",
"https://example.com/property/listing-2",
]
OUTPUT_DIR = Path("screenshots")
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
# Use a stable selector from the portal's actual markup.
HEADING_SELECTOR = "h1"
ELEMENT_SELECTOR = "main"
CAPTURE_ELEMENT = False
WAIT_SECONDS = 25
def filename_for(url: str) -> str:
"""Create a readable filename; the index prevents same-path collisions."""
parsed = urlparse(url)
path = re.sub(r"[^a-zA-Z0-9]+", "-", parsed.path).strip("-") or "page"
host = re.sub(r"[^a-zA-Z0-9]+", "-", parsed.netloc)
return f"{host}-{path}"[:120]
options = webdriver.ChromeOptions()
# For a server/container without a display, enable headless mode:
# options.add_argument("--headless=new")
# Set a fixed viewport for repeatable window screenshots.
options.add_argument("--window-size=1440,1000")
driver = webdriver.Chrome(options=options)
try:
for index, url in enumerate(URLS, start=1):
driver.get(url)
wait = WebDriverWait(driver, WAIT_SECONDS)
# Wait for meaningful listing content, not an arbitrary sleep.
wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, HEADING_SELECTOR)))
stem = f"{index:03d}-{filename_for(url)}"
if CAPTURE_ELEMENT:
target = wait.until(
EC.visibility_of_element_located((By.CSS_SELECTOR, ELEMENT_SELECTOR))
)
destination = OUTPUT_DIR / f"{stem}-element.png"
if not target.screenshot(str(destination)):
raise RuntimeError(f"Element screenshot failed: {url}")
else:
destination = OUTPUT_DIR / f"{stem}-window.png"
if not driver.save_screenshot(str(destination)):
raise RuntimeError(f"Window screenshot failed: {url}")
print(f"Saved {destination}")
finally:
driver.quit()
Selenium describes WebDriver as its browser-control interface. The basic lifecycle is to create a driver, navigate to a URL, capture what you need, and end the session. Selenium WebDriver
4. Choose the screenshot scope and locator
| Choice | Use it when | Trade-off |
|---|---|---|
| Window screenshot | The surrounding page context, navigation, or multiple visible sections matter. | It captures the current browser window; it is not automatically a full-document capture. |
| Element screenshot | You need one listing card, detail block, or other specific element. | The element must exist and be visible; validate that the selected element encloses the intended content. |
Selenium’s Python API documents save_screenshot(filename) as saving the current window as a PNG and returning whether the operation succeeded. The WebElement screenshot method saves the selected element. Python WebDriver API · Selenium element interactions
For locators, prefer a unique, predictable ID when one exists. Otherwise use a concise CSS selector based on stable attributes or structure. Avoid generated styling classes unless you have confirmed they remain stable. Selenium notes that XPath can be harder to debug and complex DOM traversal can be costly. Selenium locator guidance
5. Make the capture wait for the right condition
Pages can render in stages. A successful navigation does not guarantee that the listing title, price, or image has loaded. Wait for the condition that matters to the screenshot: for example, a heading is present, an element is visible, or a particular image has completed loading. The sample waits for the heading to exist; if the capture depends on an image, use a condition appropriate to that image and the site’s markup.
A fixed sleep can make a workflow slower while still failing on a slower page. Use an explicit wait with a timeout, then handle timeout as a page-level failure so one stalled URL does not silently produce an incomplete artifact. Avoid relying on a universal readiness signal: portals may populate content asynchronously, and this research does not establish loading behavior or selectors for any particular Indian property portal.
6. Adapt the workflow safely
Headless execution
On a machine without a graphical display, uncomment --headless=new. Keep the same explicit viewport argument so window captures have a predictable size. A browser’s headless rendering can differ from a desktop session, so inspect representative outputs before relying on a batch.
Capture one element
Set CAPTURE_ELEMENT = True and choose a selector for the listing region. The code waits for visibility before saving. If a card selector matches several results, make the target more specific or select the intended result explicitly; do not assume the first match is the right listing.
Capture many known URLs
Add only the pages that are in scope to URLS. The index and URL-derived filename reduce accidental overwrites, but the output name is not a record of permission or provenance. For auditability, keep a separate capture manifest containing the source URL, capture time, and intended use, subject to your retention policy.
Use a Grid for a real scaling need
A local browser is a reasonable starting point for one repeatable capture path. Selenium Grid allocates browsers across machines and supports execution across browsers and operating systems. Use Grid when distributed runs or browser/OS coverage are requirements; it adds infrastructure and is unnecessary for a small single-browser job. Selenium Grid documentation
7. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Driver cannot start or browser version mismatch | The browser or driver is missing, incompatible, or unavailable in the environment. | Install a supported browser, check Selenium’s driver setup guidance, and configure the driver explicitly where automatic management cannot work. |
| Timeout waiting for heading | The selector is wrong, the page is still loading, the content is in a different browsing context, or access was denied. | Inspect the authorized page’s live DOM, verify the selector, and wait for the actual content condition. Check whether the page requires a frame switch or shows an access/error page. |
| Screenshot is blank or incomplete | The capture happened before relevant content rendered, the browser window is not the expected size, or a page overlay obscures the content. | Wait for the target content, set the viewport explicitly, and inspect the saved image. Handle overlays only when doing so is permitted by the portal and consistent with your use. |
| Element screenshot fails | The element was detached during a rerender, hidden, or matched incorrectly. | Wait for visibility, reacquire the element immediately before capture, and use a stable selector that identifies the intended content. |
| Files overwrite each other | Names were derived from a non-unique title or path. | Include a sequence number or unique listing identifier and preserve a URL-to-file manifest. |
| Portal blocks the session or shows a CAPTCHA | The portal’s access controls rejected automation or require a permitted access path. | Stop and review the portal’s terms and authorization. Do not attempt to bypass the block; request permission or use an approved method. |
8. Performance, reliability, and cost
For a small URL set, one driver session avoids repeated browser startup and should always be closed with quit(), including when a capture raises an exception. Keep waits bounded, save each result as it completes, and log failures by URL so a later retry does not require recapturing successful pages. Run captures at a rate allowed by the portal and your authorization; no universal safe request interval can be inferred for every site.
Local execution costs depend on the machine and browser resources you already have. Grid or hosted browser infrastructure introduces service or operating costs and is justified by a need for parallel capacity or browser/OS coverage. This research establishes no vendor pricing or performance benchmark, so compare actual requirements and current provider terms before selecting infrastructure.
Screenshot retention and reuse can create costs beyond the browser run: storage, review, and operational handling. Keep only what the permitted purpose needs, protect any listing data in the resulting images, and follow the applicable terms and retention rules.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. One request returns an image or PDF. Cookie and consent banners, newsletter popups, and chat widgets are removed before capture; those steps can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture.
For a permitted listing URL, this cURL call saves a WebP image. See the ScreenshotNeo API documentation for parameters and response details.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/property/listing-1 -o listing.webp
Python equivalent:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com/property/listing-1"},
timeout=90,
)
open("listing.webp", "wb").write(r.content)
Node.js equivalent:
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://example.com/property/listing-1'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('listing.webp', Buffer.from(await res.arrayBuffer())));
ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are never billed; its MCP server lets AI agents take screenshots; and 1,000 screenshots a month are free with no card, with paid plans starting at $5 for 3,000. Create a free ScreenshotNeo account.
FAQ
Does a screenshot prove that I had permission to use a listing?
No. Permission depends on the portal’s terms, authorization, and how you use the image; the capture itself establishes none of those.
Should I use a window or element screenshot for a listing card?
Use an element screenshot when the card itself is the artifact. Use a window screenshot when its page context is part of what you need to preserve.
When should I move from local Selenium to Grid?
Move when you need distributed execution or coverage across browsers and operating systems. A single repeatable browser capture does not require Grid.


