ScreenshotNeo

BlogHow-to

How to Monitor Websites with a Crawler API

Build reliable website change monitoring with a crawler API: scope crawls, render JavaScript, diff snapshots, respect robots.txt, and alert safely.

By the ScreenshotNeo team1 October 20269 min read

Direct answer: Schedule a crawl from a known starting URL, constrain its scope, honor robots.txt, fetch pages through an API, normalize the returned content, compare it with the previous snapshot, and alert only on meaningful changes or failures. Use a single-page fetch when one URL matters; use a crawl job when linked pages must be discovered. For JavaScript-rendered content, choose an API that runs a browser.

1. Decide what you are monitoring

Write down the target before choosing an API or interval. A useful monitor record contains:

  • Starting URL: the page or sitemap entry where discovery begins.
  • Scope: allowed hostnames, path prefixes, depth, and maximum pages.
  • Signal: HTTP availability, a CSS selector, visible text, structured data, screenshots, or the full rendered document.
  • Schedule: how often to run and any quiet hours.
  • Alert destinations: email, chat, incident system, or a webhook.
  • Retention: how long to keep raw responses, normalized snapshots, hashes, and crawl logs.

If the requirement is “tell me when this pricing page changes,” fetch one URL. If it is “find changes anywhere under /docs/,” submit a bounded crawl.

2. Respect robots.txt and site policies

Fetch and parse /robots.txt before crawling. Apply the rules for the user agent you send, including Disallow, Allow, and sitemap declarations. Robots.txt and robots meta tags are published controls that site owners use to communicate crawler access rules; Google documents the standard in its robots.txt guidance.

Re-check robots.txt periodically because it can change. Treat an unavailable robots file conservatively according to your provider’s policy, and record the response and decision in your crawl log. Do not use monitoring as a reason to bypass authentication, bot checks, paywalls, or explicit access controls.

3. Configure crawl scope and rendering

Unbounded discovery is expensive and can create accidental load. Set all of these where the API supports them:

Control Purpose
Allowed host and path Prevents links from sending the job to unrelated domains.
Maximum depth Limits how many link levels are followed from the start URL.
Maximum pages Caps work and cost for a run.
URL filters Exclude logout links, calendars, search results, tracking parameters, and duplicate URL forms.
JavaScript rendering Loads content that appears only after scripts execute.
Incremental parameters When available, use a modified-since or maximum-age boundary to avoid revisiting unchanged pages.

Cloudflare’s Browser Rendering crawl endpoint documents robots.txt compliance, depth and page limits, browser rendering, incremental crawling parameters such as modifiedSince and maxAge, and asynchronous job retrieval. Verify current limits and pricing in the provider documentation before production use.

4. Submit an asynchronous crawl job

Large crawls should be treated as state machines:

  1. Submit the starting URL and scope.
  2. Persist the returned job ID and configuration.
  3. Poll a status endpoint or subscribe to completion events.
  4. Handle queued, running, completed, partial, failed, and timed-out states.
  5. Retrieve results page by page and store the raw response.

Make submission idempotent. Generate a run key from the target, scope, and scheduled time, and do not create a second job when a retry receives an unknown response. Use exponential backoff with jitter for polling and API retries.

5. A self-managed monitoring implementation

The following Python example shows the parts that remain yours even when a provider supplies crawling: scheduling, normalization, hashing, comparison, and alerting. Replace the placeholder API calls with the crawler vendor’s documented endpoints.

import hashlib
import json
import re
import time
from datetime import datetime, timezone
from urllib.parse import urlparse, urlunparse

import requests

CRAWLER_API = "https://crawler.example.test/v1"
API_KEY = "YOUR_API_KEY"
START_URL = "https://example.com/docs"


def canonical_url(url):
    parts = urlparse(url)
    # Remove fragments and common tracking parameters before comparison.
    query = re.sub(r"(^|&)(utm_[^=]+|gclid|fbclid)=[^&]*", "", parts.query)
    query = query.strip("&")
    return urlunparse((parts.scheme, parts.netloc.lower(), parts.path or "/", "", query, ""))


def normalize_html(html):
    # Remove elements that change every run. Tune this for the site you monitor.
    html = re.sub(r"<time[^&]*>.*?</time>", "", html, flags=re.I | re.S)
    html = re.sub(r"]*>.*?", "", html, flags=re.I | re.S)
    html = re.sub(r"]*>.*?", "", html, flags=re.I | re.S)
    html = re.sub(r"\s+", " ", html).strip()
    return html


def digest(value):
    return hashlib.sha256(value.encode("utf-8")).hexdigest()


def submit_crawl():
    response = requests.post(
        f"{CRAWLER_API}/crawl",
        headers={"Authorization": f"Bearer {API_KEY}"},
        json={
            "url": START_URL,
            "maxDepth": 2,
            "maxPages": 100,
            "render": True,
            "respectRobotsTxt": True,
        },
        timeout=30,
    )
    response.raise_for_status()
    return response.json()["id"]


def wait_for_results(job_id):
    delay = 2
    for _ in range(30):
        response = requests.get(
            f"{CRAWLER_API}/crawl/{job_id}",
            headers={"Authorization": f"Bearer {API_KEY}"},
            timeout=30,
        )
        response.raise_for_status()
        job = response.json()
        if job["status"] in {"completed", "partial", "failed", "timed_out"}:
            return job
        time.sleep(delay)
        delay = min(delay * 2, 60)
    raise TimeoutError("crawl did not finish within the polling window")


def compare(previous, pages):
    changes = []
    current = {}
    for page in pages:
        url = canonical_url(page["url"])
        normalized = normalize_html(page.get("html", ""))
        item = {"hash": digest(normalized), "status": page.get("status"), "crawledAt": datetime.now(timezone.utc).isoformat()}
        current[url] = item
        old = previous.get(url)
        if old is None:
            changes.append({"url": url, "kind": "new", "current": item})
        elif old["hash"] != item["hash"] or old["status"] != item["status"]:
            changes.append({"url": url, "kind": "changed", "previous": old, "current": item})
    for url in previous.keys() - current.keys():
        changes.append({"url": url, "kind": "missing", "previous": previous[url]})
    return current, changes


if __name__ == "__main__":
    job_id = submit_crawl()
    job = wait_for_results(job_id)
    if job["status"] in {"failed", "timed_out"}:
        raise RuntimeError(json.dumps(job))
    previous = {}  # Load this from durable storage in production.
    current, changes = compare(previous, job.get("pages", []))
    print(json.dumps({"changes": changes, "snapshot": current}, indent=2))

Store the previous snapshot in a database or object store rather than process memory. Keep the raw provider response alongside the normalized representation so an alert can be investigated later.

6. Normalize before diffing

Raw HTML is noisy. Remove or isolate rotating timestamps, navigation chrome, analytics IDs, randomized tokens, ad slots, and tracking query parameters. Prefer one of these signals:

  • Selector text: best for a known price, status label, or heading.
  • Structured data: compare selected JSON-LD fields.
  • Rendered text: useful when markup changes but the visible content does not.
  • Content hash: efficient for stable documents after normalization.
  • Screenshot or PDF: useful for visual regressions, but compare with a tolerance for rendering differences.

Keep both old and new values for an alert. A hash alone tells you that something changed; it does not explain what changed.

7. Rate limits, retries, and scheduling

Use a per-domain concurrency limit and a delay between requests. Google reports that increased latency, 5xx responses, and 429 responses reduce crawl capacity. AWS Prescriptive Guidance gives 1–2 requests per second as a potentially appropriate rate for larger sites when you have explicit crawl permission. That is guidance, not a universal limit: follow the site’s policy and the API provider’s limits first.

  • Retry connection resets, DNS failures, and 502/503/504 responses with exponential backoff and jitter.
  • Do not blindly retry 401, 403, 404, or a robots denial.
  • Honor Retry-After on 429 responses.
  • Use a timeout for connect, read, and total job duration.
  • Stagger schedules so thousands of URLs do not hit one origin at the same second.

Choose frequency from the change you need to detect. Run hourly for rapidly changing availability data, daily for documentation, and less often for stable marketing pages. A shorter interval increases requests, storage, and provider usage.

8. Alerts that people can act on

Include the URL, change type, HTTP status, crawl timestamp, previous and current hashes, a short diff or selector value, and a link to the stored snapshot. Separate severity:

  • Availability: DNS errors, timeouts, repeated 5xx responses, or a missing page.
  • Policy: robots.txt changes, authentication failures, or a provider quota limit.
  • Content: a selected field or meaningful page region changed.
  • Visual: a screenshot difference above your chosen threshold.

Require confirmation across two runs for noisy content signals, while paging immediately for a production availability failure.

9. Troubleshooting common failures

Symptom Likely cause Fix
Only the app shell is returned JavaScript was not rendered or the page needed more wait time. Enable browser rendering and wait for a selector or network idle.
Many URLs leave the target site No host/path boundary or URL filter. Allow only the intended host and path; remove tracking and calendar links.
Every run reports a change Volatile timestamps, ads, tokens, or navigation are included. Normalize those regions or compare a stable selector.
429 responses Request pressure exceeded the origin or provider limit. Reduce concurrency, honor Retry-After, add backoff, and lower frequency.
403 or bot-check page The origin requires an approved access path or blocks automated traffic. Obtain permission, use documented credentials, or exclude the URL. Do not bypass controls.
Crawl job never completes Unbounded links, slow pages, or a provider timeout. Set depth/page caps, exclude expensive paths, and handle partial results.
Missing pages in a later run Temporary failure was treated as deletion. Require repeated absence or a confirmed 404 before marking a page removed.
Duplicate alerts Retries or overlapping schedules created multiple runs. Use an idempotency key and deduplicate by URL, signal, and snapshot hash.

10. Performance, reliability, and cost

Cost is driven by pages fetched, browser-rendered work, frequency, retries, storage, and any screenshot or PDF artifacts. Page and depth caps provide the clearest upper bound. Incremental crawling can reduce repeat work when the provider supports it.

Measure submission latency, queue time, per-page latency, status-code distribution, render failures, robots failures, pages discovered, pages changed, and quota consumption. Keep a dead-letter queue for failed jobs and replay them after correcting configuration. Store provider request IDs so support investigations can connect your logs to a run.

11. Or skip the browser setup

If your monitor needs screenshots or PDFs rather than a custom crawler, ScreenshotNeo provides a website screenshot API and MCP server. It accepts one GET request and returns PNG, JPEG, WebP, or PDF. Cookie and consent banners are accepted and removed before capture, along with 60+ known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

See the ScreenshotNeo API documentation for all options, including full-page capture with lazy images, CSS selectors, device presets, dark mode, custom CSS and JavaScript, waits, blocked resources, headers, cookies, geolocation, caching, signed links, asynchronous jobs, webhooks, bulk capture, and usage reporting.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

There are 1,000 screenshots a month free with no card. Paid plans start at $5 for 3,000 shots, and every feature is on every plan. Create a free ScreenshotNeo account.

12. Provider comparison checklist

Before selecting a crawler API, verify:

  • robots.txt behavior and user-agent identity
  • JavaScript rendering and wait controls
  • depth, page, URL, and response-size limits
  • incremental crawling support
  • scheduled runs, asynchronous jobs, and webhooks
  • retry semantics, rate limits, and timeout behavior
  • geographic execution and authentication support
  • snapshot retention, export, and deletion controls
  • usage metrics, request IDs, and failure observability
  • pricing for rendered pages, retries, storage, and overages

Cloudflare Browser Rendering is a managed option for browser-rendered, bounded, asynchronous crawls. CrawlZilla’s documentation lists scheduled crawls, Page Monitor change detection, analytics, and webhooks; verify its current limits, pricing, retention, and program terms directly before adopting it. A self-managed crawler gives maximum control but leaves you responsible for robots parsing, throttling, rendering, retries, storage, and alerting.

13. FAQ

Should I monitor one URL or crawl a site?

Use a single-page request for a known URL. Use a bounded crawl when discovery of linked pages is part of the requirement.

How do I monitor a page that changes after load?

Use browser rendering and wait for a selector, a delay, or network idle. Then compare the rendered region rather than the initial HTML.

How often should a crawler run?

Match the interval to the business impact and change rate, then lower it until request pressure and cost remain acceptable.

Can a hash explain what changed?

No. A hash detects a difference. Keep normalized snapshots or selector values to produce an explanation.

What should happen when a crawl partially fails?

Store successful pages, mark failed URLs separately, retry transient failures, and do not interpret an incomplete run as mass deletion.