How to Take Batch Screenshots of Indian Real Estate Listing Pages from URLs
Capture a list of real estate listing URLs consistently with Playwright, track failures, and check portal terms before automating.
Use a browser automation script to read listing URLs from a file, open each page, wait for the content you need, and save a screenshot with a stable filename. The Playwright example below captures full pages, records a timestamp and outcome for every URL, and retries failures a limited number of times. Before automating any portal, check its current terms and obtain permission where required.
1. Choose what each screenshot should record
Choose the capture scope before you start so every image serves the same purpose:
- Viewport: the initial visible screen. Use it when you need to preserve what a user sees at the top of a listing.
- Full page: the scrollable page, including below-the-fold details. This can produce tall images and may include content loaded as you scroll.
- Element: a selected listing card, details panel, or other element. Use this when the surrounding page chrome is irrelevant.
Playwright supports viewport, full-page, and element screenshots. Its Python API also exposes full-page capture with full_page=True. See the Playwright screenshot documentation and Playwright for Python screenshots.
2. Check portal terms and define the record
A screenshot is an automated access to page content. Do not assume screenshot-only capture is exempt from a portal’s automation rules. Magicbricks’ current General Terms & Conditions say automated software to extract or download data is prohibited without prior written consent. Check the current terms and obtain consent or use an authorized method if required: Magicbricks terms.
The research for this guide did not establish current automation rules for 99acres or NoBroker. Check their terms directly before running a batch. Housing.com describes itself as an advertising and information platform and says it does not validate listing authenticity; independently verify property and project details. A screenshot records what appeared at a time, not whether the listing is accurate or current. See Housing.com terms.
Keep a manifest with one row per input URL. At minimum, record the original URL, capture timestamp, output filename, and success or failure. If the images may be used as records, also retain the terms or authorization decision, capture environment, and any relevant notes. Use stable filenames such as a zero-padded row number plus a short label, rather than relying on an address that could contain unsafe filename characters.
3. Batch full-page screenshots with Playwright and Python
This script reads one URL per line from urls.txt, captures a full-page PNG, retries each page at most twice after the first attempt, and writes manifest.csv. It uses a document-ready navigation condition and a short settling delay; neither guarantees that a portal’s listing data is ready. For a specific portal, replace the delay with a meaningful selector after checking that the selector is appropriate and authorized.
Install
python -m venv .venv
# macOS or Linux:
source .venv/bin/activate
# Windows PowerShell:
# .venv\Scripts\Activate.ps1
python -m pip install playwright
python -m playwright install chromium
Create urls.txt
https://example.com/listing-one
https://example.com/listing-two
Use only URLs you are authorized to access and capture. Then save this as batch_screenshots.py:
import asyncio
import csv
import os
import re
from datetime import datetime, timezone
from pathlib import Path
from urllib.parse import urlparse
from playwright.async_api import async_playwright
INPUT = Path("urls.txt")
OUTPUT_DIR = Path("screenshots")
MANIFEST = Path("manifest.csv")
VIEWPORT = {"width": 1440, "height": 1000}
MAX_ATTEMPTS = 3 # initial try plus at most two retries
SETTLE_MS = 1200
def safe_label(url: str) -> str:
host = urlparse(url).netloc or "page"
label = re.sub(r"[^a-zA-Z0-9]+", "-", host).strip("-").lower()
return (label or "page")[:48]
async def main():
urls = [line.strip() for line in INPUT.read_text(encoding="utf-8").splitlines()
if line.strip() and not line.lstrip().startswith("#")]
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
rows = []
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True)
context = await browser.new_context(viewport=VIEWPORT, device_scale_factor=1)
page = await context.new_page()
for index, url in enumerate(urls, start=1):
filename = f"{index:04d}-{safe_label(url)}.png"
path = OUTPUT_DIR / filename
status, error = "failed", ""
captured_at = datetime.now(timezone.utc).isoformat()
for attempt in range(1, MAX_ATTEMPTS + 1):
try:
response = await page.goto(url, wait_until="domcontentloaded", timeout=45000)
await page.wait_for_timeout(SETTLE_MS)
await page.screenshot(path=str(path), full_page=True, animations="disabled")
status = "success"
error = ""
break
except Exception as exc:
error = f"attempt {attempt}/{MAX_ATTEMPTS}: {type(exc).__name__}: {exc}"
if attempt < MAX_ATTEMPTS:
await page.wait_for_timeout(800 * attempt)
rows.append({
"input_url": url,
"captured_at_utc": captured_at,
"output_file": str(path) if status == "success" else "",
"status": status,
"error": error,
})
await context.close()
await browser.close()
with MANIFEST.open("w", newline="", encoding="utf-8") as f:
writer = csv.DictWriter(f, fieldnames=["input_url", "captured_at_utc", "output_file", "status", "error"])
writer.writeheader()
writer.writerows(rows)
succeeded = sum(row["status"] == "success" for row in rows)
print(f"Finished: {succeeded}/{len(rows)} successful. See {MANIFEST}.")
if __name__ == "__main__":
asyncio.run(main())
Run it with python batch_screenshots.py. The example uses one page at a time, which keeps resource use and access rate modest. Its host-based filenames are intentionally simple; use a row number as the stable identity, and add a reviewed short label if you need to distinguish multiple properties on one portal.
Wait for a known listing element
When a listing is rendered after initial navigation, a selector wait can be more meaningful than a fixed delay. Identify a stable selector by inspecting the page you are authorized to capture, then replace the settling line with something like:
await page.locator("YOUR_LISTING_SELECTOR").wait_for(state="visible", timeout=20000)
If the selector is optional or varies among pages, handle that explicitly and record that the page was captured without the expected element. A timeout should not silently be treated as a complete, comparable screenshot.
4. Adapt the capture mode
Viewport-only capture
await page.screenshot(path="listing.png", full_page=False)
Capture a single element
await page.locator("YOUR_LISTING_SELECTOR").screenshot(path="listing-card.png")
The locator must identify the intended element. If it matches several cards, select the correct one explicitly; if it matches none, wait for it and record the failure rather than saving an unrelated image.
Stabilize visual comparisons
Use the same browser engine and version, viewport dimensions, device scale factor, headless setting, and wait rule for every URL. Playwright notes that screenshots may vary with operating system, browser version, settings, hardware, power source, and headless mode. Environment consistency reduces avoidable differences but does not freeze changing page content, listing availability, personalized responses, or third-party widgets. See Playwright’s visual comparison guidance.
5. Alternatives for simple captures and other languages
Chrome Headless command line
For a small number of individual viewport captures, Chrome Headless provides a --screenshot option and a window-size setting. For example, on a system where the Chrome executable is named chrome:
chrome --headless --window-size=1440,1000 --screenshot="listing.png" "https://example.com/listing"
Executable names and installation paths differ by operating system. This command is suitable for a simple single capture; a script around it is needed for a URL list, manifest, retries, and custom readiness checks. See Chrome Headless documentation.
Playwright with Node.js
For a JavaScript workflow, save this as batch.mjs. It accepts URLs from urls.txt, uses the same stable browser setup for each, and records outcomes in a JSON manifest.
import { chromium } from 'playwright';
import { readFile, writeFile, mkdir } from 'node:fs/promises';
import { hostname } from 'node:url';
const urls = (await readFile('urls.txt', 'utf8'))
.split(/\r?\n/).map(s => s.trim()).filter(s => s && !s.startsWith('#'));
await mkdir('screenshots', { recursive: true });
const browser = await chromium.launch({ headless: true });
const context = await browser.newContext({ viewport: { width: 1440, height: 1000 }, deviceScaleFactor: 1 });
const page = await context.newPage();
const manifest = [];
for (let i = 0; i < urls.length; i++) {
const url = urls[i];
const label = (new URL(url).hostname || 'page').replace(/[^a-z0-9]+/gi, '-').toLowerCase();
const file = `screenshots/${String(i + 1).padStart(4, '0')}-${label}.png`;
let status = 'failed';
let error = '';
const capturedAt = new Date().toISOString();
for (let attempt = 1; attempt <= 3; attempt++) {
try {
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 45000 });
await page.waitForTimeout(1200);
await page.screenshot({ path: file, fullPage: true, animations: 'disabled' });
status = 'success'; error = ''; break;
} catch (e) {
error = `attempt ${attempt}/3: ${e.name}: ${e.message}`;
if (attempt < 3) await page.waitForTimeout(800 * attempt);
}
}
manifest.push({ input_url: url, captured_at_utc: capturedAt, output_file: status === 'success' ? file : '', status, error });
}
await context.close();
await browser.close();
await writeFile('manifest.json', JSON.stringify(manifest, null, 2));
console.log(`Finished: ${manifest.filter(x => x.status === 'success').length}/${urls.length} successful.`);
Install with npm install playwright and npx playwright install chromium, then run node batch.mjs. As in the Python example, replace the fixed delay with a portal-appropriate selector wait when you know what indicates that the desired listing content is ready.
6. Or skip the browser setup
ScreenshotNeo accepts one GET request with a URL and returns an image or PDF. Its API parameters also accept the names used by other screenshot APIs, which can make switching simpler. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/listing -o listing.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com/listing"},
timeout=90,
)
open("listing.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/listing' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('listing.webp', bytes));
For batches, call the API once per input URL or use its bulk capture option for up to 100 URLs per call; keep your own URL-to-output manifest so each result stays tied to the right listing. Available options include full-page capture with lazy images loaded, element capture by CSS selector, viewport and device presets, custom waits, and image format choices including PNG, JPEG, and WebP. ScreenshotNeo can accept cookie banners before capture and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each of those steps can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and whether the shot was billed. An MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.
ScreenshotNeo is a website screenshot API and MCP server from ScreenshotNeo. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. Sign up for ScreenshotNeo’s free plan.
7. Performance, reliability, and cost
- Keep concurrency modest. The sample is sequential. If you add parallel pages, do so only where permitted, start conservatively, and monitor failures and rate limits; aggressive concurrency can increase resource use and may violate portal terms.
- Set bounded timeouts and retries. Navigation can fail or a page can remain incomplete. A small retry count with a brief backoff avoids endless loops. Do not retry authorization denials or repeated bot challenges as if they were transient network errors.
- Keep output and disk use in mind. Full-page images can be much taller and larger than viewport captures. Use a viewport or element capture when that is the record you need, and choose a format and scale deliberately.
- Expect page changes. Availability, prices, images, and personalized page content can change between captures. Store UTC capture time and environment information when comparisons matter.
- Plan workflow cost. Playwright and Chrome Headless are browser tools; this workflow’s direct costs depend on the machine and storage you provide. The research sources establish no tested speed or total-cost comparison. ScreenshotNeo’s stated plans are Free: 1,000 shots/month; Starter: $5 for 3,000; Growth: $15 for 15,000; Pro: $39 for 60,000; Scale: $99 for 250,000; Business: $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan. Only clean shots are billed under the product’s stated billing rules.
8. Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| Navigation times out | The page is slow, blocked, or waiting on resources that never settle. | Use a bounded navigation timeout and a readiness condition for the content you need. Record failure. Do not treat timeout as a successful capture. |
| Screenshot is blank or missing listing details | The capture ran before client-rendered content appeared, or content requires scrolling or interaction. | Wait for a meaningful visible selector. For lazy content, use full-page capture and confirm the page’s behavior; where needed, scroll in controlled increments before capture. |
| Some URLs fail while others work | Bad or expired URL, redirect, access restriction, transient network issue, or portal automation controls. | Preserve the exact URL and error in the manifest, retry transient failures only a limited number of times, and check portal authorization. Do not try to evade access controls. |
| Images differ between runs | Browser, OS, viewport, scale, headless mode, hardware, page state, or changing content differs. | Pin the capture environment and settings; record them with the manifest. Expect live listings to change. |
| Element capture errors or captures the wrong card | The selector matches zero or multiple elements, or the page structure changed. | Wait for the selector, verify uniqueness, and select the intended match explicitly. Log selector failures. |
| Output image is unexpectedly large | Full-page capture included a long feed or a tall page. | Use viewport or element capture when appropriate, or resize/compress for downstream storage while retaining an original if required. |
| Batch stops before writing the manifest | An error occurred outside the per-page retry block, such as input or filesystem failure. | Check that urls.txt is readable and the output directory is writable. For a durable job, write manifest rows incrementally so an interruption does not lose completed results. |
9. FAQ
Does a screenshot prove a listing is genuine?
No. It preserves a visual state at a particular time. It does not independently verify ownership, availability, price, or any other claim.
Should I capture every listing page in a portal?
Only if the portal’s terms and your authorization permit it. Check current rules and use an authorized access method.
Is a viewport screenshot enough for later comparison?
It is enough only if the information you need is visible in that viewport. Capture a full page or a specific element when the record needs more context.
Can I compare screenshots from different computers?
You can, but rendering differences may come from the capture environments as well as the page. Standardize and record browser and viewport settings for more meaningful comparisons.


