ScreenshotNeo

BlogHow-to

How to Capture Screenshots of Indian Real Estate Listings in Bulk

Capture permitted property listings consistently with Playwright, organized filenames, and a record of each source URL and capture time.

By the ScreenshotNeo team4 October 20269 min read

To capture screenshots of Indian real estate listings in bulk, first confirm that the portal permits your intended automated access and later use. Then give a browser automation script a list of authorized listing URLs, wait for the relevant content, and save each page or listing element using a predictable filename. Keep an index of each source URL and capture time.

Playwright can capture the visible viewport, a selected element, or the full scrollable page. The example below uses Python and Chromium to capture full pages from a CSV file. Use automation only where the portal’s current terms and your permission allow it. A page being publicly visible does not, by itself, grant permission to automate access or copy its contents.

1. Check permission and define the purpose

Before collecting URLs, decide whether the screenshots are for private comparison, a client file, publication, or another purpose. Check the portal’s current terms for automated access, copying, storage, and later sharing. Permission to view a page does not necessarily include permission to collect it in bulk or reuse its photos and text.

Policies differ by portal. The reviewed PropertyHub India terms prohibit bots and similar technologies for access, copying, monitoring, or extraction, and prohibit circumventing CAPTCHA, rate limits, or security measures. The reviewed eRealtor terms prohibit unpermitted scraping, copying, and bots. These are examples of individual platform policies, not a rule about every Indian real estate site. Check the current terms for each portal and get written permission or use an authorized export or API if your intended capture is not permitted.

  • Use only listing URLs you are authorized to access.
  • Confirm that automated browsing is allowed for your purpose.
  • Check whether storing, sharing, or publishing screenshots is permitted separately.
  • Do not bypass a login, CAPTCHA, rate limit, or other technical protection.
  • Stop if the portal blocks automation or asks for a human check.

2. Choose what each screenshot needs to show

Capture type Use it when Trade-off
Viewport You need a consistent view of what appears on screen. Content below the fold is not included.
Listing element You need a focused capture of a listing card or another specific region. Surrounding page context may be omitted.
Full page You need the whole scrollable listing page as displayed. The image can be very tall and include unrelated page content.

Choose the smallest scope that preserves the context you need. If you are documenting a listing’s displayed details, a full-page image may be useful. If you are comparing cards in a results page, an element capture can keep the output focused. These are records of what a page displayed at capture time, not proof that its claims are accurate.

3. Prepare a URL list and output index

Save one authorized URL per row in a CSV file named listings.csv:

url
https://example.com/listing/authorized-listing-1
https://example.com/listing/authorized-listing-2

Replace the example URLs with URLs that you are permitted to capture. Avoid adding contact details or unrelated personal information to the index unless it is needed and authorized. The script below creates an index.csv with the source URL, capture timestamp, output filename, and outcome.

4. Install Playwright and run a bulk capture

Install Playwright for Python and its Chromium browser. Save the following as capture_listings.py in the same directory as listings.csv.

python -m pip install playwright
python -m playwright install chromium
import csv
import re
import time
from datetime import datetime, timezone
from pathlib import Path
from urllib.parse import urlparse

from playwright.sync_api import TimeoutError as PlaywrightTimeoutError
from playwright.sync_api import sync_playwright

INPUT_CSV = Path("listings.csv")
OUTPUT_DIR = Path("screenshots")
INDEX_CSV = Path("index.csv")

# Adjust these for pages you are permitted to capture.
NAVIGATION_TIMEOUT_MS = 45_000
CONTENT_WAIT_MS = 1_000
DELAY_BETWEEN_PAGES_SECONDS = 2
FULL_PAGE = True

# Set this to a selector that identifies the listing content on your target portal,
# for example: "main" or a portal-specific listing container. Leave empty to skip.
CONTENT_SELECTOR = ""


def safe_slug(value: str) -> str:
    value = re.sub(r"[^A-Za-z0-9_-]+", "-", value).strip("-_")
    return value[:80] or "listing"


def filename_for(url: str, index: int) -> str:
    parsed = urlparse(url)
    path_part = safe_slug(parsed.path.rstrip("/").split("/")[-1])
    return f"{index:04d}-{path_part}.png"


def read_urls(path: Path) -> list[str]:
    with path.open(newline="", encoding="utf-8-sig") as file:
        return [row["url"].strip() for row in csv.DictReader(file) if row.get("url", "").strip()]


def main() -> None:
    urls = read_urls(INPUT_CSV)
    OUTPUT_DIR.mkdir(parents=True, exist_ok=True)

    with INDEX_CSV.open("w", newline="", encoding="utf-8") as index_file:
        fields = ["url", "captured_at_utc", "filename", "status", "detail"]
        writer = csv.DictWriter(index_file, fieldnames=fields)
        writer.writeheader()

        with sync_playwright() as playwright:
            browser = playwright.chromium.launch(headless=True)
            page = browser.new_page(viewport={"width": 1440, "height": 1000}, device_scale_factor=1)
            page.set_default_navigation_timeout(NAVIGATION_TIMEOUT_MS)

            for number, url in enumerate(urls, start=1):
                filename = filename_for(url, number)
                output_path = OUTPUT_DIR / filename
                captured_at = datetime.now(timezone.utc).isoformat()
                status = "error"
                detail = ""

                try:
                    response = page.goto(url, wait_until="domcontentloaded")
                    if response is not None and response.status >= 400:
                        raise RuntimeError(f"HTTP status {response.status}")

                    if CONTENT_SELECTOR:
                        page.locator(CONTENT_SELECTOR).first.wait_for(state="visible", timeout=NAVIGATION_TIMEOUT_MS)
                    else:
                        page.locator("body").wait_for(state="visible", timeout=NAVIGATION_TIMEOUT_MS)

                    # A short, fixed pause can allow late page updates to appear.
                    # Use only a delay appropriate to the portal's rules and page behavior.
                    page.wait_for_timeout(CONTENT_WAIT_MS)

                    page.screenshot(path=str(output_path), full_page=FULL_PAGE)
                    status = "captured"
                except (PlaywrightTimeoutError, RuntimeError, Exception) as error:
                    detail = f"{type(error).__name__}: {error}"
                    if output_path.exists():
                        output_path.unlink()
                finally:
                    writer.writerow({
                        "url": url,
                        "captured_at_utc": captured_at,
                        "filename": filename if status == "captured" else "",
                        "status": status,
                        "detail": detail,
                    })
                    index_file.flush()

                if number < len(urls):
                    time.sleep(DELAY_BETWEEN_PAGES_SECONDS)

            browser.close()


if __name__ == "__main__":
    main()

Run the script with:

python capture_listings.py

Successful images go in screenshots/; index.csv records the source and result for each row. The pause is an operational setting, not a guarantee that a site allows a particular request rate. Follow the portal’s rules, and do not reduce the delay to get around a block.

5. Capture a listing element instead of a whole page

If the portal permits automated capture and you only need a listing region, set CONTENT_SELECTOR to a selector that uniquely identifies the target element and replace the screenshot call:

listing = page.locator(CONTENT_SELECTOR).first
listing.wait_for(state="visible", timeout=NAVIGATION_TIMEOUT_MS)
listing.screenshot(path=str(output_path))

Use a selector that matches the intended listing, not a broad container that could include navigation or unrelated listings. If a page contains multiple cards, decide which card corresponds to the input URL or capture each authorized card separately. A selector that stops matching after a portal redesign should fail visibly; do not silently save a screenshot of the wrong region.

6. Make output consistent and auditable

  • Use stable names. Include a sequence number and a listing identifier when available, such as 0001-listing-id.png. Avoid using a full address or personal details in filenames unless necessary.
  • Keep the index. Store the exact source URL and UTC capture time alongside each output.
  • Choose dimensions deliberately. Keep the viewport and device scale factor fixed when comparing screenshots. Increase the viewport only when that better fits the intended record.
  • Preserve the original capture. If you annotate an image for internal use, retain the unaltered screenshot and label the edited copy so it cannot be mistaken for the original.
  • Protect the files. Limit access and retention to what your purpose and permissions require, especially if pages show personal contact information.
  • Review failures. Check the index for timeouts, HTTP errors, and missing content rather than assuming every URL produced a valid image.

7. Verify listing and regulatory claims separately

A screenshot records the page as it appeared; it does not establish that the price, ownership, approvals, amenities, or registration claims are correct. Verify RERA registration on the appropriate official state regulator portal and check other material facts independently. The reviewed Telangana RERA disclaimer cautions that portal content should not be treated as a legal statement. RealtorCentral’s reviewed terms also direct users to check RERA status through the official state portal. For a decision with legal consequences, consult the relevant authority or a qualified professional.

8. Common errors and fixes

Symptom Likely cause What to do
Navigation times out The page is slow, unreachable, or waiting on resources that do not finish. Check whether the URL loads manually. Increase the timeout only if appropriate, use a wait condition for the actual listing content, and retry later. Do not repeatedly retry a site that is blocking automation.
HTTP error recorded The server returned an error response, or the listing has moved or been removed. Check the URL manually and update the authorized URL list if the listing has changed. Do not treat an error-page image as a listing capture.
Screenshot is blank or incomplete Content may render after initial navigation, require scrolling, or depend on scripts. Wait for a visible content selector rather than relying only on page load. If content appears only after scrolling, inspect the page’s behavior and use a permitted, deliberate scroll workflow. Stop if a human check or access control appears.
Element selector times out The selector is incorrect, the page changed, or the element is not visible. Inspect the page manually and update the selector. Test it on one permitted page before processing a batch.
Screenshot contains the wrong listing A broad selector matched the first of several cards, or the URL-to-card mapping is wrong. Use a more specific selector and validate the URL and image together. Keep the source URL in the index.
Access denied, CAPTCHA, or rate-limit response The portal restricts automated access or requests human verification. Stop automation. Do not attempt to evade the restriction. Seek permission or use an allowed export/API.
Browser executable missing Playwright’s Chromium browser was not installed for the current environment. Run python -m playwright install chromium in the environment used for capture.
CSV parsing error The file lacks a url header or contains malformed rows. Ensure the first row is exactly url and each subsequent row contains one URL.

9. Performance, reliability, and cost

Bulk work is mainly constrained by page load time, screenshot size, available memory, and the portal’s permitted request rate. Full-page captures can use more memory and storage than viewport or element images. Keep the browser workload modest, write each result as it finishes, and retain failure details so an interrupted run does not require guessing which URLs succeeded.

For reliability, start with a small permitted batch, review the images and index, and then proceed with the remaining authorized URLs. Keep the browser and capture settings consistent across a comparison set. Add retries only for transient failures, with a limited retry count and a pause; never use retries to circumvent site blocking or rate limits.

Playwright is software for browser automation; this workflow has no per-screenshot API price stated here. Your actual costs may include the machine, storage, and time spent reviewing results. The reviewed sources provide no benchmark for capture speed, so plan based on a small sample from the sites you are allowed to access.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. For sites and uses you are authorized to capture, a single request can return a screenshot. Cookie banners, newsletter popups, and chat widgets are removed before capture; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. See the API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/authorized-listing -o listing.webp

Sign up for 1,000 free screenshots a month, with no card required.

FAQ

Can I capture listings that require a login?

Only if you are authorized to access and capture them, and the portal’s terms allow your method and intended use. Do not bypass authentication or other technical controls.

Should I use a screenshot as evidence of a property claim?

It can document what a page displayed at a particular time, but it does not verify the claim. Check official records and other reliable sources separately.

How should I share screenshots with a client?

Confirm that both the portal’s terms and your permission cover sharing. Include the source URL and capture timestamp so the image’s origin is clear.

Can every Indian property portal be captured with the same script?

No. Page structure, access rules, and technical behavior vary. Review the current policy for each portal and adapt selectors only where automated capture is permitted.