How to Automate Screenshot Capture for Indian Property Portals Using URLs
Turn a list of property listing URLs into repeatable screenshots with Playwright. Learn how to choose capture scope, handle page states, and check portal permissions first.
You can automate screenshots from property listing URLs with a browser script: open each URL, wait for the page content you need, and save a viewport, a specific element, or the full scrollable page. This guide uses Playwright, which can capture those scopes and save PNG, JPEG, or WebP images.
Check permission before automating. Magicbricks’ current General Terms & Conditions state: “Use of automated software to extract or download data from the Site is prohibited without prior written consent.” The terms also prohibit bots or scrapers without permission. A screenshot may raise different questions depending on what you capture, how you automate it, and how you use the resulting image. Do not assume that a technically successful capture is authorized. Review each portal’s current official terms and obtain written permission for your particular use where required. This research does not establish blanket permission, an official API, or current screenshot rules for every Indian portal. [Read the Magicbricks terms](https://property.magicbricks.com/terms/terms.html). Housing.com’s retrieved terms do not settle screenshot permission, and current 99acres permission terms were not established here.
1. Decide what each screenshot needs to show
Before writing the loop, choose the capture scope and output settings. These choices affect the usefulness and size of each file.
| Choice | Use it when | Trade-off |
|---|---|---|
| Viewport | You need the portion visible in the browser window. | Fast and compact, but content below the fold is omitted. |
| One element | You have permission to capture a specific listing card or content section. | Focused output; the selector must match the intended element. |
| Full page | You need the page’s full scrollable content in one image. | Can produce very tall files and may trigger lazy-loaded content as the browser scrolls. |
Use PNG when you need lossless output, JPEG when a smaller photographic image is suitable, and WebP when your downstream tools accept it. Pick a predictable filename scheme, such as a sequential index plus a sanitized listing identifier. Avoid putting private data or URL query tokens in filenames.
Playwright’s screenshot API supports viewport, element, and full-page capture, as well as PNG/JPEG/WebP, output paths, and CSS-pixel or device-pixel scale. See the [Playwright screenshot documentation](https://playwright.dev/docs/screenshots).
2. Set up Playwright
The example below uses Node.js and Playwright. It reads one URL per line from urls.txt, opens each page in a controlled browser session, checks that the document loaded, and saves a full-page PNG. It records failures per URL and continues with the rest of the list.
npm init -y
npm install playwright
npx playwright install chromium
Save this as capture.mjs:
import { chromium } from 'playwright';
import { mkdir, readFile } from 'node:fs/promises';
const inputPath = process.argv[2] ?? 'urls.txt';
const outputDir = process.argv[3] ?? 'screenshots';
const urls = (await readFile(inputPath, 'utf8'))
.split(/\r?\n/)
.map((line) => line.trim())
.filter((line) => line && !line.startsWith('#'));
await mkdir(outputDir, { recursive: true });
const browser = await chromium.launch({ headless: true });
const context = await browser.newContext({
viewport: { width: 1440, height: 1000 },
deviceScaleFactor: 1,
});
const page = await context.newPage();
try {
for (let i = 0; i < urls.length; i += 1) {
const url = urls[i];
const filename = `${String(i + 1).padStart(4, '0')}.png`;
try {
const response = await page.goto(url, {
waitUntil: 'domcontentloaded',
timeout: 45_000,
});
if (!response) {
throw new Error('Navigation returned no main-document response');
}
if (!response.ok()) {
throw new Error(`HTTP ${response.status()} ${response.statusText()}`);
}
// Replace this generic readiness check with a permitted, page-specific
// selector when you know which listing content must be present.
await page.locator('body').waitFor({ state: 'visible', timeout: 15_000 });
await page.screenshot({
path: `${outputDir}/${filename}`,
fullPage: true,
type: 'png',
animations: 'disabled',
scale: 'css',
});
console.log(`Saved ${filename} from ${url}`);
} catch (error) {
console.error(`Failed ${url}: ${error.message}`);
}
}
} finally {
await context.close();
await browser.close();
}
Put authorized URLs in urls.txt, one per line, then run:
node capture.mjs urls.txt screenshots
domcontentloaded waits for the initial HTML document to be parsed without requiring every image, analytics request, or other resource to finish. The body visibility check is only a generic starting point; for a known page and authorized use, wait for a selector that identifies the content you actually need. There is no universal wait duration for property portals.
3. Change the capture scope and output
Capture only the visible viewport
Omit fullPage or set it to false. The screenshot uses the current viewport dimensions:
await page.screenshot({ path: 'listing-viewport.webp', type: 'webp', quality: 85 });
Capture one listing element
Use a selector that identifies the intended content. Inspect the page only in ways permitted by the portal and your authorization. If the selector is absent, fail clearly rather than silently saving an unrelated page.
const listing = page.locator('[data-testid="listing-details"]');
await listing.waitFor({ state: 'visible', timeout: 15_000 });
await listing.screenshot({ path: 'listing-details.png', type: 'png' });
The selector above is an example, not a claim about a particular portal’s markup. Replace it with a selector appropriate to the permitted page.
Capture the full scrollable page
await page.screenshot({ path: 'listing-full.jpg', fullPage: true, type: 'jpeg', quality: 85 });
Full-page mode captures the scrollable page as one image. Some pages load images or content only when scrolled. If those areas are missing, and the page’s terms and your permission allow it, scroll through the page or wait for known content before taking the screenshot. Do not treat this as permission to defeat access controls or bot protections.
Choose viewport size and pixel scale
Set a fixed viewport for repeatable dimensions. With scale: 'css', output dimensions correspond to CSS pixels. With scale: 'device', output dimensions account for the device scale factor and can be larger. A larger scale increases image dimensions and file size.
const context = await browser.newContext({
viewport: { width: 1365, height: 900 },
deviceScaleFactor: 2,
});
await page.screenshot({ path: 'retina.png', fullPage: true, scale: 'device' });
Make filenames useful
For repeat runs, derive filenames from a stable internal record ID or a sanitized URL path, and add a date or run identifier if old captures must be retained. Do not use raw URLs as paths: query strings can contain sensitive tokens and characters that are invalid in filenames. For example:
const filename = `${recordId}-${runDate}.png`;
4. Capture multiple listing URLs safely and predictably
- Confirm the portal’s current terms and your authorization for the automation and intended use.
- Prepare a reviewed URL list. Exclude URLs that require bypassing a login, access restriction, bot check, or CAPTCHA.
- Keep the browser version, operating system, viewport, scale, and screenshot settings fixed across runs you intend to compare.
- Wait for the specific content required for each capture, using a bounded timeout and a clear failure record.
- Save each result under a stable filename and keep a log mapping the filename to its source URL and capture time.
- Review failed pages and representative outputs. A saved image alone does not prove the intended listing loaded.
Use sequential navigation for a modest list as in the runnable example. If you increase concurrency, create separate pages or contexts per worker, limit the number of workers, and respect portal terms and any written authorization. More parallel browsers consume more memory and can create more requests in a short period.
5. Wait for the right page state
Navigation completion and content readiness are different. A page can finish its initial document load while a listing image, price, or other client-rendered content is still pending. Prefer a meaningful, authorized selector over an arbitrary long sleep:
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 45_000 });
await page.locator('YOUR_AUTHORIZED_CONTENT_SELECTOR').waitFor({
state: 'visible',
timeout: 15_000,
});
await page.screenshot({ path: 'ready.png', fullPage: true });
Replace the placeholder with a selector you are permitted to use. If you cannot identify a reliable readiness signal, use a short, bounded delay only as a fallback and record that the capture may be incomplete. Avoid relying on network-idle as a universal signal: pages with ongoing requests may never become idle, while a quiet network does not guarantee that the desired content rendered.
6. Repeatability, performance, and reliability
Keep the capture environment stable
Playwright notes that screenshot output can differ with host operating system, browser version, settings, hardware, power supply, headless mode, and other factors. Keep these conditions and the viewport consistent when comparing images; do not promise pixel-identical output across unrelated machines. See [Playwright’s visual comparison guidance](https://playwright.dev/docs/test-snapshots).
Control timeouts and failures
Set navigation and selector timeouts that fit your workflow, catch errors per URL, and retain the URL and error in a log. A timeout should mark that item as failed, not produce a misleading success record. Retry only transient failures, with a small bounded retry count and spacing; repeated retries against an access denial or bot check do not make the capture authorized.
Manage resource use
Browser rendering uses more memory and CPU than downloading an image file. Reusing one browser and context reduces startup overhead, while a modest number of concurrent pages can improve throughput at a resource cost. Full-page and device-scale captures can be much larger than viewport captures. Choose PNG only when its quality characteristics matter; JPEG or WebP can reduce output size when accepted by the workflow.
Protect captured material
Listings and images may have use restrictions even when visible in a browser. Limit access to output files, set retention rules, and avoid sharing or commercial reuse unless your permission and applicable terms allow it. Do not capture private account information or store session credentials in source code.
7. Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| Navigation timeout | The portal is slow, a resource is stuck, or navigation is blocked. | Use a bounded timeout appropriate to your run, check the URL and response, and record the failure. Do not repeatedly retry access restrictions. |
| Screenshot is blank or missing listing content | The page has not rendered the target content, or the selector/readiness check is too broad. | Wait for an authorized content-specific selector and inspect the resulting page state. Treat an absent target as a failed capture. |
| Element locator times out | The selector does not match this page, the markup changed, or the element is not visible. | Recheck the selector on an authorized page and wait for the correct state. Do not substitute a broad selector that captures unrelated content. |
| Images are absent in a full-page capture | Images may load lazily as the page scrolls or may not have completed loading. | Where permitted, scroll through the page and wait for the relevant images or content before capturing. Verify that the added wait is bounded. |
| Output differs between runs | Browser, OS, viewport, device scale, page state, dynamic content, or headless settings changed. | Pin the runtime and capture settings, and compare captures made under the same conditions. Dynamic page changes can still produce differences. |
| HTTP error or unexpected redirect | The URL is stale, the server returned an error, or the page redirected. | Log the final URL and response status, verify the input, and handle the item as an error when it did not reach the intended listing. |
| Bot check, CAPTCHA, or access denial | The portal is restricting automated access or requires an authorized access path. | Stop automation for that URL. Request permission or use an official, authorized route; do not bypass the control. |
| Files overwrite one another | The naming scheme is not unique across URLs or runs. | Use a stable record identifier and, when needed, a run date or unique run ID. Sanitize all filename components. |
8. Or skip the browser setup
For a permitted URL, ScreenshotNeo offers a one-request screenshot API. The [ScreenshotNeo documentation](https://screenshotneo.com/docs/) describes its request options. This example saves the response body as a WebP image:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
Replace the example target with a URL you are authorized to capture. ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before the shot. Bot checks, blank pages, and failed loads are never billed; responses include page-verdict and billing headers. Its MCP server lets AI agents use screenshot, page-info, and PDF-capture tools. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. See [ScreenshotNeo](https://screenshotneo.com) and [its API docs](https://screenshotneo.com/docs/).
Sign up free for 1,000 screenshots a month, no card required.
9. FAQ
How do I take screenshots of multiple property listing URLs automatically?
Put one permitted URL per line in a text file and run the Playwright loop above. It saves a separate file for each successful URL and logs failures so you can review them.
Can I save a full-page screenshot of a property listing?
Yes. In Playwright, set fullPage: true in page.screenshot(). Check that the intended content and lazy-loaded images are present, and confirm you are allowed to automate and use the capture.
Does a screenshot mean I can reuse the listing or its photos?
No. Capturing visible content does not establish permission to extract, reuse, redistribute, or commercially exploit it. Check the portal’s terms and obtain any required authorization for the specific use.
Is there one permission rule for all Indian property portals?
No such blanket rule is established here. Check each portal’s current official terms and your use case; the cited Magicbricks restriction should not be generalized to other portals.


