ScreenshotNeo

BlogHow-to

Scraping Real Estate Data with Python in 2026

Get real estate data with Python by starting with the right permissions and source. Learn when to use an authorized API, MLS feed, browser workflow, or Census data.

By the ScreenshotNeo team29 September 202611 min read

Scraping Real Estate Data with Python in 2026

Start with permission and the kind of data you need. For individual listings, use an API or feed you are authorized to access, such as a provider’s documented API or an MLS feed licensed for your use. For neighborhood-level housing and demographic context, use Census datasets. Then use Python’s requests library for a documented HTTP/JSON interface, or Playwright when an authorized source genuinely requires browser rendering.

A page being visible in a browser does not establish permission to collect it automatically, store its data, or redistribute it. This matters especially for real estate listings, where the source’s terms and license can govern access, display, retention, refresh, and reuse. This guide focuses on responsible, source-authorized collection—not bypassing a portal’s access controls.

1. Decide what data you need and whether you may collect it

“Real estate data” can mean several different things. Decide which one you need before choosing a tool:

Need Likely source Questions to settle
Current individual listings Authorized listing API or licensed MLS feed Are you eligible? May your use display, store, or redistribute records? What refresh frequency and attribution are required?
Property attributes or transaction records Authorized data provider or applicable public dataset Which fields and geography are covered? What do the field definitions mean? What are the license and update terms?
Neighborhood housing or demographic context U.S. Census datasets through the Census Data API Which dataset and geography answer the question? Remember these statistics are context, not individual property listings.

For every source, read the current terms and document the permitted use. Check whether your planned use is personal, research, internal business, public display, or commercial redistribution; whether records may be retained; and whether you need to show attribution or refresh data on a schedule. Applicable law and contractual terms can depend on the source, jurisdiction, and use, so do not treat this article as a universal legal conclusion.

2. Choose an authorized data route

Listings: obtain permission from the provider or MLS

Zillow’s general Terms of Use prohibit automated queries against its consumer services, including screen and database scraping, spiders, robots, and crawlers. Its separate API terms describe a route for preapproved licensees and constrain approved API and data use. Those are distinct access routes: an approved API does not make automated collection from the consumer website permissible, and API use remains subject to its terms. Check Zillow’s current [Terms of Use](https://www.zillow.com/z/corp/terms/) and [API terms](https://www.zillowgroup.com/developers/terms/) before relying on either route. The terms can change.

Choose the source and permission model based on whether you need listings, property records, or area-level context.
Choose the source and permission model based on whether you need listings, property records, or area-level context.

For MLS data, RESO Web API is a data transport standard, not a universal license or open-access permission. RESO explains that recipients agree to an MLS’s data-use and licensing policies, then work with that MLS’s software provider or technical staff to obtain credentials and access instructions. Zillow also says its listings are published through MLS IDX feeds. Ask the relevant MLS or authorized provider about eligibility, permitted display, retention, redistribution, refresh frequency, fees, and attribution; do not infer these from another MLS’s rules. See [RESO’s Web API overview](https://www.reso.org/reso-web-api/) and Zillow’s [IDX information](https://www.zillowgroup.com/developers/).

Market context: use Census data

If the question is about an area rather than an individual home—such as comparing housing or demographic context—look for an appropriate official Census dataset. The [Census Data API](https://www.census.gov/data/developers/data-sets.html) exposes U.S. Census datasets, and the Bureau offers free API-key registration. Census statistics are not a substitute for listing-level feeds, property records, or a licensed MLS dataset. Confirm that a dataset’s geography, period, and variables fit your analysis.

3. Collect authorized JSON data with Python Requests

Use Requests for a documented endpoint you are allowed to call. The example below is deliberately provider-neutral: set the endpoint, parameter names, and expected JSON structure from the provider’s documentation. It requests only the fields and geography your analysis needs. Requests documents query parameters, timeouts, JSON handling, and response errors in its [official documentation](https://requests.readthedocs.io/en/latest/).

Use a documented HTTP interface when available, and preserve provenance and validation alongside the response.
Use a documented HTTP interface when available, and preserve provenance and validation alongside the response.
import json
import os
import sys
from datetime import datetime, timezone
from pathlib import Path

import requests
from requests.exceptions import RequestException

API_URL = os.environ.get("REAL_ESTATE_API_URL")
API_KEY = os.environ.get("REAL_ESTATE_API_KEY")

if not API_URL:
    sys.exit("Set REAL_ESTATE_API_URL to your provider's documented endpoint.")
if not API_KEY:
    sys.exit("Set REAL_ESTATE_API_KEY in your environment; do not put secrets in source code.")

params = {
    # Replace these with documented parameter names and authorized values.
    "state": "CA",
    "city": "Oakland",
    "fields": "listing_id,address,price,updated_at",
}
headers = {"Authorization": f"Bearer {API_KEY}", "Accept": "application/json"}

try:
    response = requests.get(API_URL, params=params, headers=headers, timeout=(5, 30))
    response.raise_for_status()
    payload = response.json()
except requests.exceptions.Timeout:
    sys.exit("The provider did not respond before the timeout; retry later within its terms.")
except requests.exceptions.HTTPError as exc:
    # Avoid printing credentials or sensitive provider response bodies to logs.
    sys.exit(f"Provider returned HTTP {response.status_code}: {exc}")
except requests.exceptions.JSONDecodeError:
    sys.exit("The response was not valid JSON; check the endpoint and provider status.")
except RequestException as exc:
    sys.exit(f"Request failed: {exc}")

# Adapt this validation to the provider's documented response schema.
if not isinstance(payload, (dict, list)):
    sys.exit("Unexpected top-level JSON type; compare with the provider's schema.")

record = {
    "source": API_URL,
    "retrieved_at": datetime.now(timezone.utc).isoformat(),
    "geographic_scope": {"state": params["state"], "city": params["city"]},
    "license_or_use_reference": "RECORD YOUR APPLICABLE PROVIDER TERMS HERE",
    "data": payload,
}
Path("real_estate_data.json").write_text(json.dumps(record, indent=2), encoding="utf-8")
print("Saved real_estate_data.json")

Install the dependency in a virtual environment with python -m pip install requests. The documentation currently lists Requests 2.34.2 and Python 3.10+ support; check the project documentation for current compatibility. Keep credentials in environment variables or a secret manager, restrict access to saved data, and never commit keys to version control.

Adapt the client to the provider

  1. Read the API schema and authentication instructions. Replace the placeholder URL and parameters; do not guess endpoint paths or fields.
  2. Ask for only the fields and area needed, if the API supports field selection. Smaller responses use less bandwidth and make validation easier.
  3. Set explicit connect and read timeouts. The tuple (5, 30) bounds connection setup and response waiting; tune it to the provider’s documented behavior and your job’s requirements.
  4. Check the response status before parsing JSON. Handle documented pagination, rate limits, and retry guidance. Do not hammer an endpoint or retry indefinitely.
  5. Validate the schema and field meanings. A response can be valid JSON yet still omit a field, change a type, or represent a price or date differently than expected.
  6. Store source, retrieval time, geographic scope, applicable license reference, and the provider’s update timestamp alongside the records.

Address cleanup deserves its own validation step. Preserve the source value, normalize a separate comparison key, and review ambiguous unit numbers, directional prefixes, abbreviations, and postal codes. Do not assume that two differently formatted addresses are different properties, or that matching text guarantees the same property. Keep the source’s stable identifier when its terms permit it.

4. Use Playwright only for an authorized browser-rendered source

Some authorized sources require JavaScript rendering and provide no suitable documented API. In that case, browser automation can help observe the page’s request and response lifecycle. It does not grant access rights. Before automating, confirm that the source permits it and follow its authentication, rate, and use rules. If the source blocks automation, requires a CAPTCHA, or denies a request, stop and contact the provider or use an authorized interface. Do not evade the restriction.

Playwright’s Python documentation describes page request and response events. The following small example observes events while opening a page you are authorized to access; it does not extract records or bypass access controls. Install with python -m pip install playwright, then playwright install chromium.

import asyncio
from playwright.async_api import async_playwright

async def main():
    async with async_playwright() as p:
        browser = await p.chromium.launch(headless=True)
        page = await browser.new_page()
        page.on("request", lambda request: print("request", request.method, request.url))
        page.on("response", lambda response: print("response", response.status, response.url))
        try:
            response = await page.goto(
                "https://example.com/authorized-resource",
                wait_until="domcontentloaded",
                timeout=30000,
            )
            if response is None:
                print("Navigation produced no main-document response")
            elif response.status >= 400:
                print("Main document returned HTTP", response.status)
            else:
                print("Loaded authorized page; inspect only data your terms permit")
        finally:
            await browser.close()

asyncio.run(main())

Replace the example URL only with a resource you have permission to access. Keep logs minimal: request URLs can contain identifiers or sensitive query values. Treat browser failures, navigation timeouts, and HTTP error responses as different outcomes, and avoid collecting more than the approved scope. See the [Playwright Python page API](https://playwright.dev/python/docs/api/class-page) for event details.

5. Keep collection reliable, auditable, and within scope

  • Record provenance. Store the provider, retrieval timestamp, geographic scope, schema or dataset release when available, and license/use reference with each export.
  • Respect the update cadence. Use the provider’s approved refresh frequency. Avoid duplicate requests when a previously retrieved record is still current under the permitted retention rules.
  • Make retries bounded. Retry only transient failures according to provider guidance, with a limit and delay. Do not retry authorization failures, access denials, or CAPTCHA challenges as though they were network glitches.
  • Validate before analysis. Check required fields, data types, timestamps, duplicate IDs, units, missing values, and geographic coverage. Compare field semantics before comparing prices or attributes across sources.
  • Protect the output. Limit who can access the dataset and remove or expire it when required by its terms. Confirm rights before publishing a derived table or redistributing source records.

For larger jobs, use the provider’s pagination or bulk interface if documented, checkpoint completed pages, and make each run resumable. Keep a small audit log with request time, endpoint name, status, record count, and a non-secret job identifier. Avoid logging API keys, cookies, authorization headers, or full records when a count and status will do.

6. Common errors and how to fix them

Symptom Likely cause Responsible next step
401 or 403 response Missing, expired, or insufficient credentials; the account may not be approved for this dataset or use. Check the provider’s authentication and eligibility instructions, then contact its support or MLS channel. Do not try to bypass access checks.
404 response Wrong endpoint, retired version, or a resource outside the account’s scope. Copy the endpoint from current provider documentation and verify the resource and account permissions.
429 response Provider rate limit or request volume beyond the allowed cadence. Pause, consult documented limits and retry guidance, and reduce request frequency. Do not rotate identities to evade limits.
Timeout Network delay, slow provider response, or an operation taking longer than the configured timeout. Separate connection and read timeouts, retry a transient failure only within provider rules, and ask the provider about expected response times.
JSON decoding error An HTML error page, empty response, or unexpected content type arrived instead of JSON. Check status and content type before parsing; verify endpoint, credentials, and provider status. Do not log a sensitive response body.
Valid JSON but no expected records Filters, geography, pagination, or schema differ from assumptions. Inspect the documented schema and a permitted sample response; verify every filter and follow pagination instructions.
Playwright navigation fails or is challenged Network issue, browser incompatibility, denial, or an access control. Separate transport errors from HTTP status. Fix local setup for ordinary errors; for denials or challenges, stop automation and ask the source for an authorized route.
Duplicate or mismatched properties Address formatting, units, identifiers, or field semantics differ across records. Retain original values, normalize cautiously, use authorized stable IDs where available, and review ambiguous matches rather than merging automatically.

7. Performance, reliability, and cost

Use the simplest authorized interface that provides the needed data. A documented JSON API usually avoids the overhead of rendering a browser, while a licensed feed may be designed for an approved recurring workflow. Browser automation launches and runs a browser, so reserve it for cases where rendering is genuinely necessary and permitted. The sources here do not establish a universal speed benchmark, coverage level, fee, request quota, or update interval; confirm those details with the relevant provider or MLS.

Bound work by geography, date range, and fields; use documented pagination or bulk access; and schedule only the permitted refresh frequency. Cache or retain data only as allowed by the license. Budget for any provider, MLS, storage, and compute costs from the terms you actually receive—there is no single price that applies to all real estate data. For resilience, checkpoint progress, validate each page or batch, and make failures visible without silently treating incomplete results as complete.

8. Capture a permitted page for review

If your approved workflow needs a visual record of a publicly accessible page—for example, documenting a page you are authorized to review—a screenshot is a separate artifact from structured listing data and does not grant permission to collect or redistribute page contents. Use a browser locally if that suits your workflow, or an API designed to capture a page. ScreenshotNeo is a website screenshot API and MCP server; it is not an MLS feed or a source of listing records.

Or skip the browser setup

One GET request can return an image or PDF. For a permitted page, this cURL example saves a WebP image; see the [ScreenshotNeo API documentation](https://screenshotneo.com/docs/) for parameters and response details.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo accepts cookie or consent banners like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers say the page verdict and billing status. Its MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. See [ScreenshotNeo](https://screenshotneo.com) for plan details.

Use it only for pages you are permitted to capture. A screenshot is not structured listing data, an MLS license, or permission to automate a source that prohibits it. Create a free account for 1,000 screenshots a month with no card.

9. Short FAQ

There is no universal answer for every source, use, and jurisdiction. Check the site’s current terms, applicable license, and relevant law before automating or reusing data.

Can I use RESO Web API without asking an MLS?

RESO is a standard, not open permission. Access follows the relevant MLS’s policies and credentials process.

Can Census data tell me which homes are currently for sale?

No. Census datasets provide area-level statistical context; use an authorized listing source for individual current listings.

Does a screenshot API give me listing data?

No. It produces a visual capture. Use a permitted structured data API or licensed feed for records.

Primary sources