How to Schedule Recurring Screenshots of Indian Real Estate Listing Websites
Build a dated archive of Indian property listings with Playwright and GitHub Actions. Choose a capture cadence, preserve comparable screenshots, and handle failures.
To schedule recurring screenshots of Indian real estate listings, use a browser script to capture each listing and a scheduler to run that script on a repeat cadence. This guide uses Playwright with Python and GitHub Actions. It saves timestamped images, records run results, and explains how to choose a capture scope and interval. A screenshot is a record of how a page rendered in one browser session; it does not verify the listing’s claims.
1. Choose the listing pages and capture cadence
Start with specific listing URLs you are authorized to view. Put each URL in a small configuration file with a stable, human-readable identifier. A listing’s direct page is usually easier to compare over time than a search-results page, where inventory and ordering can change.
Choose an interval that fits the change you want to observe. Daily captures may suit a personal record; a longer interval can reduce storage and avoid unnecessary requests. Do not use an aggressive cadence. Check each portal’s current terms, account rules, and access restrictions before automating. If a site blocks or challenges the browser, stop and use an authorized route rather than trying to bypass the restriction.
| Decision | Practical choice |
|---|---|
| Target | A specific listing URL with a stable identifier |
| Cadence | As infrequently as meets your record-keeping need |
| Timezone | Use Asia/Kolkata for India-local scheduling where supported, or convert the desired time to UTC |
| Capture scope | Viewport for compact consistency, full page for below-the-fold details, or one element for a stable known region |
2. Set up a Playwright screenshot script
Playwright can capture the viewport, a selected element, or the full scrollable page. This example takes full-page PNGs, uses one browser viewport consistently, and includes the listing ID and UTC timestamp in each filename. See the Playwright screenshot documentation for capture options.
Install dependencies
python -m venv .venv
source .venv/bin/activate
python -m pip install playwright
python -m playwright install chromium
On Windows, activate the virtual environment with .venv\\Scripts\\activate. Save this as listings.json:
[
{
"id": "listing-123",
"url": "https://example.com/property/listing-123"
},
{
"id": "listing-456",
"url": "https://example.com/property/listing-456"
}
]
Replace the example URLs with listing pages you are allowed to access. Save the following as capture.py:
import json
import os
import re
import sys
from datetime import datetime, timezone
from pathlib import Path
from urllib.parse import urlparse
from playwright.sync_api import sync_playwright
CONFIG = Path("listings.json")
OUTPUT = Path("screenshots")
OUTPUT.mkdir(exist_ok=True)
def safe_id(value):
return re.sub(r"[^A-Za-z0-9_-]+", "-", value).strip("-") or "listing"
def main():
listings = json.loads(CONFIG.read_text(encoding="utf-8"))
failures = []
timestamp = datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%SZ")
with sync_playwright() as playwright:
browser = playwright.chromium.launch(headless=True)
context = browser.new_context(
viewport={"width": 1440, "height": 1000},
device_scale_factor=1,
locale="en-IN",
timezone_id="Asia/Kolkata",
)
page = context.new_page()
for listing in listings:
listing_id = safe_id(listing["id"])
url = listing["url"]
try:
parsed = urlparse(url)
if parsed.scheme not in ("http", "https") or not parsed.netloc:
raise ValueError("URL must be an absolute http or https URL")
response = page.goto(url, wait_until="domcontentloaded", timeout=60000)
# Give client-rendered listing content a short, bounded settling period.
page.wait_for_timeout(1500)
path = OUTPUT / f"{listing_id}_{timestamp}.png"
page.screenshot(path=str(path), full_page=True, animations="disabled")
status = response.status if response else "no-main-response"
print(f"OK {listing_id} status={status} file={path}")
except Exception as error:
failures.append(listing_id)
print(f"ERROR {listing_id}: {error}", file=sys.stderr)
browser.close()
if failures:
print("Failed listing IDs: " + ", ".join(failures), file=sys.stderr)
return 1
return 0
if __name__ == "__main__":
raise SystemExit(main())
Run it once locally with python capture.py. Confirm that the images show the intended content and that the output directory contains one file per listing. The script continues after an individual listing fails, reports the failures, and exits unsuccessfully if any failed so a scheduler can flag the run.
Choose the screenshot scope
- Viewport: use
page.screenshot(path="shot.png")for a consistent, compact view of the visible area. - Full page: use
page.screenshot(path="shot.png", full_page=True)to include content below the fold. Long pages can produce very tall files, and lazy-loaded content may not appear unless it has loaded. - Element: locate a stable region and call
locator.screenshot(path="shot.png"). This focuses comparison, but a changed selector or page layout can make the capture fail.
Playwright also supports image type, quality, scale, masking, and injected styles. Keep the same settings for every run. Masking or hiding dynamic regions can make comparisons easier, but may conceal a meaningful change to the listing; preserve an unmasked original if those details matter. Consult the API docs before using options such as quality, which apply to particular output formats.
3. Schedule the script with GitHub Actions
GitHub Actions can run a workflow on a cron schedule. Scheduled workflow times default to UTC; the documentation also supports an IANA timezone. GitHub documents a minimum interval of five minutes, but scheduled events can be delayed during high load and queued runs can be dropped. Treat the schedule as a trigger, not a minute-exact delivery guarantee. See GitHub’s schedule event documentation.
Create .github/workflows/capture.yml in the repository’s default branch:
name: Recurring listing screenshots
on:
schedule:
- cron: "30 3 * * *"
timezone: "Asia/Kolkata"
workflow_dispatch:
jobs:
capture:
runs-on: ubuntu-latest
permissions:
contents: read
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.12"
- name: Install Playwright and Chromium
run: |
python -m pip install playwright
python -m playwright install --with-deps chromium
- name: Capture listing pages
run: python capture.py
- name: Upload screenshots
if: always()
uses: actions/upload-artifact@v4
with:
name: listing-screenshots-${{ github.run_id }}
path: screenshots/
if-no-files-found: ignore
retention-days: 30
This example requests a daily run at 03:30 Asia/Kolkata time. Check the timezone support and current syntax in GitHub’s documentation when setting up the workflow. The manual workflow_dispatch trigger lets you run it on demand. The artifact retention setting is an example; choose a period that fits your needs and the repository’s policies. GitHub Actions artifacts are not a permanent archive.
- Commit the script, listing configuration, and workflow to the repository’s default branch.
- Run the script manually and review the output before enabling recurring captures.
- Run the workflow manually from the Actions tab and check its logs and uploaded artifact.
- After the scheduled run, review both the workflow status and the actual files. A successful trigger alone does not prove that each page loaded correctly.
4. Keep the archive useful and comparable
- Keep originals: filenames use listing IDs and UTC capture timestamps, so later runs do not silently replace earlier screenshots.
- Keep settings fixed: viewport, device scale, locale, timezone, image type, and capture scope affect the output. Keep them stable to make visual changes easier to interpret.
- Choose durable storage: download workflow artifacts to a controlled archive or use storage with a retention period that meets your needs. Protect the archive and limit access.
- Record context: keep the workflow run ID, timestamp, listing ID, response status, and any error alongside each capture. A small CSV or JSON manifest can make later review easier.
- Review cautiously: rotating banners, recommendations, timestamps, and layout changes can alter pixels even when the property terms did not change. Compare the relevant listing details in context.
Do not share screenshots that expose phone numbers, names, account details, or other personal information. Use captures for a legitimate purpose and secure stored copies. A screenshot does not establish when a listing first appeared, whether its claims are accurate, or whether an offer is legally binding. Magicbricks’ terms, for example, say listing and RERA information may not be verified by the platform and ask users to verify independently; check the current Magicbricks terms and the terms of every other portal you use. Do not generalize one portal’s terms to all Indian property websites.
5. cURL, Python, and Node.js capture alternatives
The scheduled workflow above uses Python and Playwright. These examples show other ways to start a single capture; they do not include a scheduler. Use your chosen scheduler to run the relevant command or script. cURL fetches a URL but does not render JavaScript like a browser, so it is not a substitute for browser screenshots of dynamic listing pages.
cURL: fetch the page HTML
curl --fail --location --max-time 60 \
"https://example.com/property/listing-123" \
--output listing.html
This saves the server response HTML, not a screenshot. To produce a rendered image with cURL, call a screenshot API. For a clean browser screenshot with ScreenshotNeo, see the optional section below.
Python: one-off Playwright capture
from playwright.sync_api import sync_playwright
url = "https://example.com/property/listing-123"
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page(viewport={"width": 1440, "height": 1000})
page.goto(url, wait_until="domcontentloaded", timeout=60000)
page.wait_for_timeout(1500)
page.screenshot(path="listing.png", full_page=True)
browser.close()
Node.js: one-off Playwright capture
npm install playwright
npx playwright install chromium
const { chromium } = require('playwright');
(async () => {
const browser = await chromium.launch({ headless: true });
const page = await browser.newPage({
viewport: { width: 1440, height: 1000 },
deviceScaleFactor: 1,
});
try {
await page.goto('https://example.com/property/listing-123', {
waitUntil: 'domcontentloaded',
timeout: 60000,
});
await page.waitForTimeout(1500);
await page.screenshot({ path: 'listing.png', fullPage: true });
} finally {
await browser.close();
}
})();
6. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. One GET request can return a PNG, JPEG, WebP, or PDF. Its cookie and consent handling removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and responses identify the page verdict and billing status in headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents using Claude, Cursor, or another MCP client.
Use this cURL example for a one-off shot (replace the URL and API key). To make it recurring, run the request from a scheduler and store each dated result. See the ScreenshotNeo API documentation for parameters and response details.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.com/property/listing-123 \
-o shot.webp
The same request in Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={
"access_key": "YOUR_API_KEY",
"url": "https://example.com/property/listing-123",
},
timeout=90,
)
r.raise_for_status()
with open("shot.webp", "wb") as output:
output.write(r.content)
And in Node.js:
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://example.com/property/listing-123',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', bytes));
Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Every feature is on every plan. Sign up for 1,000 free screenshots a month, with no card required.
7. Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| Browser executable missing | Playwright package installed but Chromium was not installed | Run python -m playwright install chromium; in CI, install browser dependencies with --with-deps. |
| Navigation timeout | Slow page, network issue, or the page keeps background requests open | Use a bounded timeout and a practical wait condition such as domcontentloaded; inspect the run log. Do not retry rapidly or bypass a site challenge. |
| Screenshot is blank or missing listing details | Client rendering or lazy content had not settled when the screenshot was taken | Wait for a known listing selector if one is stable, or add a modest bounded delay. Check the screenshot manually. Full-page capture alone does not guarantee every lazy-loaded image has loaded. |
| Some listings failed but others were saved | One or more URLs returned errors or triggered an exception | Review the per-listing error and response status. The example script continues through the list but returns a failing exit status if any listing failed. |
| No scheduled workflow runs | Workflow file is not on the default branch, cron/timezone is misconfigured, or the schedule has not triggered yet | Check the default branch and workflow syntax; run workflow_dispatch to confirm setup. Scheduled events are not guaranteed to start at an exact minute. |
| Workflow says success but expected images are absent | Capture logic did not write files, configuration was empty, or the artifact step had no files | Inspect step logs, confirm listings.json entries, and verify the output directory before relying on the archive. |
| Visual differences are noisy | Dynamic banners, recommendations, ads, or timestamps changed | Compare relevant listing details, keep capture settings fixed, and cautiously mask only regions that are not important to your record. |
| Portal shows a bot check or access challenge | The site is restricting automated access | Stop automated attempts and use an authorized access method. Do not attempt to evade the restriction. |
8. Performance, reliability, and cost
Each browser launch and page load takes time and compute. For a short list, one browser and a sequential loop are simple to operate. Keep timeouts bounded, avoid overly frequent schedules, and add concurrency only if the target site permits it and you can manage the additional load. Full-page images can be large; viewport shots generally use less storage, while element shots can keep an archive focused.
Reliability depends on more than the scheduler: network access, page behavior, browser installation, and artifact retention all matter. GitHub documents possible schedule delays and dropped queued runs under high load. Preserve workflow logs and image files, inspect failures, and use a second alert or an archive outside temporary workflow storage if missing a capture matters.
The DIY approach has no per-screenshot API charge, but it consumes runner time and requires maintenance and storage. The GitHub example retains artifacts for 30 days, which is only a sample configuration. A hosted screenshot service can remove browser-runtime maintenance, but check its current quotas, schedule options, retention, privacy handling, target compatibility, and price before adopting it. ScreenshotNeo offers 1,000 shots per month free with no card; paid plans are $5 for 3,000, $15 for 15,000, $39 for 60,000, $99 for 250,000, and $249 for 1,000,000. Yearly billing gives two months free. These are product plan details; confirm current terms when choosing a service.
FAQ
Will a screenshot prove that the asking price was accurate?
No. It records what appeared in a browser session. Verify price, ownership, approvals, and other material claims through appropriate authoritative sources and due diligence.
Can I reliably capture a listing at an exact minute?
No scheduler in this guide guarantees that. GitHub says scheduled events may be delayed or dropped during high load, so do not use this setup where minute-exact capture is essential.
Should I capture the listing’s search results too?
Only if changes to search ranking or inventory are part of your goal. Search results can reorder as inventory changes, so they are harder to compare as a record of one property.
How long should I retain the screenshots?
Choose retention based on your purpose, storage capacity, privacy obligations, and any applicable site terms. Keep access restricted and delete captures when they are no longer needed.


