ScreenshotNeo

BlogHow-to

How to Scrape Websites in Real Time

Choose between direct HTTP, browser rendering, and recurring crawls. Build a bounded workflow that checks access rules, handles failures, and validates fresh data.

By the ScreenshotNeo team4 October 202610 min read

Direct answer: Start with an ordinary HTTP request if the content you need is already in the page’s returned HTML. Use a browser to render the page when its content depends on JavaScript or browser state. For recurring collection across a site, use a crawler that can discover pages from links or sitemaps and run asynchronously. In every case, define how fresh the data must be, check the target host’s access instructions and terms, bound your request rate, and validate what you collect.

“Real time” is a latency requirement, not a guarantee that you will notice every source-page change immediately. Measure the delay your application can tolerate, then choose a fetch interval or job workflow that fits it.

1. Define what “real time” means for your project

Write down the maximum acceptable time between a change on the source page and the moment your application can use the updated value. Include the whole pipeline: scheduling or polling, retrieval, rendering if needed, extraction, validation, and downstream processing.

  • One-time lookup: Fetch a known URL when a user or process needs it.
  • Frequent updates: Poll at a measured interval, with limits and backoff for errors.
  • Recurring site collection: Run a crawl job that discovers pages and can skip recently fetched or unchanged content where supported.

No reviewed source establishes one universal polling interval or a guarantee of instant change detection. Choose an interval based on the target’s rules, your freshness need, and the capacity of your own system.

2. Check the site before collecting data

  1. Identify the exact scheme and host you will access, such as https://www.example.com.
  2. Inspect that host’s root /robots.txt. Look at the applicable user-agent group, path rules, and any sitemap references. Rules apply to the protocol, host, and port where the file is published; a subdomain or alternate protocol can have different rules. See Google’s robots.txt guidance.
  3. Review the site’s current terms, API documentation, authentication requirements, and any data-use or rate restrictions that apply to your project.
  4. Do not treat a robots.txt allowance as permission to ignore other requirements. Robots.txt is advisory; it is not an access-control mechanism or a legal ruling. Cloudflare describes it as “robots.txt is advisory, not enforceable” in its Browser Run robots.txt and sitemaps documentation.
  5. If the server returns an access denial, bot challenge, or CAPTCHA, stop and use an authorized access path, such as a documented API or permission from the site owner. Do not try to defeat the challenge.

Rules and terms depend on the target and your use case. This guide cannot determine whether a particular collection activity is permitted in a specific jurisdiction.

3. Choose the lightest retrieval method that works

Method Use it when Latency and work to plan for Common limitation
Direct HTTP fetch The required content is in the returned HTML or a documented data endpoint. One request and parsing step. You handle scheduling, timeouts, retries, and validation. It may return an application shell without client-rendered content.
Browser rendering The content appears only after JavaScript runs or browser state is established. Add browser startup, page loading, and a wait for a meaningful selector or state. Rendering uses more resources and can still time out or encounter a block.
Managed asynchronous crawl You need recurring collection across pages discovered from links or sitemaps. Submit a job, receive a job ID, then retrieve results as pages are processed. You must configure scope and handle job status and results. A crawl cannot be assumed to bypass blocks.

Try a direct request first and inspect the response. If the target data is missing because it is populated in the browser, move to rendering. A managed crawler is a different choice: it is for page discovery and multi-page collection, not just rendering one known URL.

4. A bounded Python workflow for a known page

This runnable example fetches one page with Python’s standard library, checks the response, extracts the title, records the retrieval time, and uses a small retry limit with backoff for transient failures. It does not execute JavaScript. Set the target URL only after checking that host’s access instructions.

from datetime import datetime, timezone
from html.parser import HTMLParser
from time import sleep
from urllib.error import HTTPError, URLError
from urllib.request import Request, urlopen

URL = "https://example.com/"
TIMEOUT_SECONDS = 15
MAX_ATTEMPTS = 3

class TitleParser(HTMLParser):
    def __init__(self):
        super().__init__()
        self.in_title = False
        self.parts = []

    def handle_starttag(self, tag, attrs):
        if tag.lower() == "title":
            self.in_title = True

    def handle_endtag(self, tag):
        if tag.lower() == "title":
            self.in_title = False

    def handle_data(self, data):
        if self.in_title:
            self.parts.append(data)

request = Request(URL, headers={"User-Agent": "ExampleResearchBot/1.0"})

for attempt in range(1, MAX_ATTEMPTS + 1):
    try:
        with urlopen(request, timeout=TIMEOUT_SECONDS) as response:
            status = response.status
            content_type = response.headers.get("Content-Type", "")
            body = response.read()

        if status != 200:
            raise RuntimeError(f"Unexpected HTTP status: {status}")
        if "html" not in content_type.lower():
            raise RuntimeError(f"Expected HTML, received: {content_type}")

        charset = "utf-8"
        for part in content_type.split(";")[1:]:
            if "charset=" in part.lower():
                charset = part.split("=", 1)[1].strip()
        html = body.decode(charset, errors="replace")

        parser = TitleParser()
        parser.feed(html)
        title = " ".join(" ".join(parser.parts).split())
        if not title:
            raise RuntimeError("Page returned no title; check whether content is client-rendered")

        print({
            "url": URL,
            "retrieved_at": datetime.now(timezone.utc).isoformat(),
            "title": title,
        })
        break
    except HTTPError as exc:
        # Do not retry permanent access or not-found responses.
        if exc.code in (401, 403, 404, 410, 429):
            raise
        if attempt == MAX_ATTEMPTS:
            raise
        sleep(min(2 ** (attempt - 1), 8))
    except (URLError, TimeoutError) as exc:
        if attempt == MAX_ATTEMPTS:
            raise
        sleep(min(2 ** (attempt - 1), 8))

Replace the example title parser with extraction for the fields you need, and validate those fields before storing them. For production use, also set a response-size limit, respect published request limits, and distinguish a valid empty result from a failed retrieval.

5. Render a page when the data needs JavaScript

Use a browser only after confirming that a direct response does not contain the required content. Prefer a specific selector or other content signal as the readiness condition; an arbitrary delay can be too short on a slow page and waste time on a fast one.

Managed browser and extraction APIs vary. For example, WebscrapingAPI.dev documents static and JavaScript-rendered modes; its request limits, timeouts, and credit model are vendor-published settings that can change. Cloudflare’s crawl service documents static mode as well as browser rendering for its managed workflow. Avoid treating one provider’s quotas or costs as universal performance figures.

For a browser you operate yourself, make these choices explicit in your implementation:

  • Set a page-navigation timeout and an overall job deadline.
  • Wait for the selector that contains the data, or a documented application state.
  • Close browser pages and processes even after exceptions.
  • Limit concurrent pages to a level your host can sustain.
  • Capture the final URL, retrieval time, and relevant response status with the extracted record.

6. Use a recurring crawler for multi-page collection

When you need many related pages, a crawler can find URLs through links or sitemaps, apply scope and depth controls, and support incremental runs. Cloudflare documents an asynchronous Browser Rendering /crawl flow: start a job, receive a job ID, then check results as pages are processed. Its March 10, 2026 changelog describes the endpoint as open beta. Review the current Cloudflare crawl endpoint documentation for its current API shape, configuration, and limits before adopting it.

Configure a crawler conservatively:

  1. Choose a start URL and define allowed hostnames and URL paths.
  2. Set maximum depth, page count, and any other available crawl bounds.
  3. Choose static retrieval or browser rendering based on the pages’ actual needs.
  4. Use incremental or freshness options where supported to avoid repeatedly processing unchanged pages.
  5. Submit the job and store its ID. Poll status with a measured interval or use a documented completion mechanism if available.
  6. Validate each result and make downstream writes idempotent so processing a result twice does not duplicate records.

Directives are not interpreted identically by every crawler. Cloudflare documents support for crawl-delay in its managed crawl endpoint, while Amazon states that its named crawler agents do not support that directive. This is an example of crawler-specific behavior, not a rule for all tools; see Amazon’s AmazonBot documentation.

7. Bound rates, retries, and data quality

Use limits and backoff

  • Follow rate limits published by the target or the service you use. Keep concurrency modest until you know the allowed and sustainable rate.
  • Set connection and read timeouts. Cap response size so one unexpectedly large page cannot consume unbounded memory.
  • Retry only transient errors, with a small attempt limit and exponential backoff. Do not retry access denials or challenges in a tight loop.
  • For recurring jobs, add jitter to schedules so multiple workers do not all poll or crawl at the same instant.

Service limits differ and can change. As an example of vendor-published values, WebscrapingAPI.dev’s reviewed documentation lists request and credit limits, a default timeout, and a response-body cap; consult its current documentation rather than copying those figures as general rules.

Validate every result

  • Store the source URL and retrieval timestamp alongside each record.
  • Check required fields and value types against an expected schema.
  • Flag unexpectedly empty pages and structural changes instead of silently treating them as valid empty results.
  • Deduplicate or upsert using a stable key so repeated runs remain idempotent.
  • Track successful empty results, access failures, timeouts, parse failures, and challenges as separate outcomes.
  • Retain only data your project is allowed to collect and use.

8. Troubleshooting

Symptom Likely cause What to do
HTML is an app shell; expected text is missing The page fills in content with JavaScript, or the data comes from a separate documented endpoint. Inspect the page response and authorized network/API options. Use browser rendering if the content requires browser execution.
Browser wait times out The selector is wrong, the page failed, or the chosen wait condition never occurs. Confirm the selector in the rendered DOM, check navigation status, and wait for a concrete content signal. Keep an overall deadline.
HTTP 401 or 403 The resource requires authorization or denies the request. Use an authorized API or credentials if the site provides them. Do not attempt to bypass an access control.
HTTP 404 or 410 The URL is absent or has been removed. Check the source of the URL and record it as unavailable; do not retry indefinitely.
HTTP 429 A rate or quota limit was reached. Stop or slow requests, honor any retry guidance, and reduce concurrency or polling frequency.
CAPTCHA or bot challenge appears The site or an intermediary is challenging automated access. Stop collection for that route and seek an approved access method. Cloudflare says its documented crawl endpoint cannot bypass its bot detection or captchas.
Timeouts happen intermittently Network delay, a slow page, or an unsuitable timeout may be involved. Set a realistic deadline, retry transient failures a limited number of times with backoff, and record the failure separately.
Results are malformed or suddenly empty The source structure changed, the response is not HTML, or extraction selected the wrong element. Check status and content type, validate schema, and review a bounded sample of the returned content.
Repeated crawl results or duplicate records Runs overlap or ingestion is not idempotent. Use stable record keys, upserts, and job-state tracking. Avoid starting a new full crawl while a prior one is still processing unless intended.

9. Performance, reliability, and cost

There is no independent general-purpose benchmark in the reviewed sources proving that one scraping method is always fastest or cheapest. Direct fetching avoids browser work when it returns the needed data; rendering adds browser startup and page execution; an asynchronous crawl adds job submission and result handling but can automate discovery across many pages.

  • Latency: Measure end-to-end freshness in your own workflow, including scheduling, render waits, retries, and downstream processing.
  • Reliability: Use bounded retries, timeouts, idempotent ingestion, schema checks, and separate failure metrics. A successful HTTP response alone does not prove the extracted record is correct.
  • Capacity: Browser rendering consumes more resources than parsing a returned HTML response. Limit concurrency and choose rendering only where it is needed.
  • Cost: Compare the provider’s current pricing and quotas with your request volume, render requirements, response sizes, and retry rate. Vendor limits and plans can change; do not extrapolate from one provider to all others.

Or skip the browser setup

If your task is to capture a page as an image or PDF rather than extract structured records from many pages, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. It accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and the response identifies the page verdict and billing status in headers.

ScreenshotNeo also offers MCP tools for AI agents, including Claude, Cursor, and other MCP clients. One thousand screenshots a month are free with no card; paid plans start at $5 for 3,000. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Sign up for 1,000 free screenshots a month with no card.

Frequently asked questions

Does “real time” mean I will see every page change immediately?

No. It describes the freshness target you choose. Polling intervals, crawl schedules, processing time, and source availability all affect when a change reaches your application.

No. It communicates crawler preferences for a host and path; it does not enforce access controls or resolve terms, contracts, privacy, or legal requirements.

Should I use a browser for every page?

No. First check whether the returned HTML contains what you need. Browser rendering is useful when the data depends on JavaScript or browser state.

Can a managed crawler get past a CAPTCHA?

Do not assume so. Cloudflare says its documented crawl endpoint cannot bypass Cloudflare bot detection or captchas. Treat a challenge as a reason to stop and find an approved access path.