ScreenshotNeo

BlogHow-to

How to fix Urlwatch when a site blocks automated requests

Find out whether Urlwatch is receiving a block, missing JavaScript content, or monitoring the wrong source—and choose a permitted fix.

By the ScreenshotNeo team4 October 20269 min read

If Urlwatch reports an error or stops detecting changes, first determine what the job actually received. A 403, a challenge page, missing JavaScript-rendered content, and an incorrect source can look similar but need different responses. If the site intentionally blocks automated access, do not try to disguise Urlwatch or defeat the block: check the site’s rules, request permission, or use an authorized alternative.

For content that is permitted to monitor, choose the simplest job that returns the data you need: a url job for a direct response, an authorized API endpoint when available, or a Browser job when the needed content only appears after JavaScript runs. A browser does not guarantee access if the site still denies or challenges the request.

1. Inspect what Urlwatch received

Before changing the job, collect the evidence from a run:

  1. Record the job type, target URL, time of the run, and the complete error or HTTP status shown in the output.
  2. Where the output makes it available, inspect the response content. Is it the expected page, an incomplete server response, an access-denied message, or a challenge page?
  3. Compare the response with what the monitored job is meant to track. A successful request can still return the wrong page or omit content that appears only after rendering.
  4. If you operate the site or can ask its operator, have them check relevant security event logs. Those may identify which security feature acted.

A 403 Forbidden says the request was refused; by itself it does not identify the rule, establish why it was refused, or prove that changing a request header is appropriate. Cloudflare, for example, documents rules that can block or challenge requests based on user-agent. The relevant remedy depends on the rule and the site’s policy. See [Cloudflare’s user-agent blocking documentation](https://developers.cloudflare.com/waf/tools/user-agent-blocking/) and [security feature interoperability guidance](https://developers.cloudflare.com/waf/feature-interoperability/).

2. Identify which problem you have

What you see What to investigate Next step
An HTTP error, denial page, CAPTCHA, or security challenge A site-side access or security rule, rate behavior, or request configuration Review the site’s published policy and the response. Ask the operator about permission or an allowed feed. Do not treat a different user-agent, browser job, proxy, or cookie replay as a guaranteed or appropriate bypass.
The URL job succeeds but its output lacks content visible in a normal browser JavaScript rendering or a mismatch between the page response and the data you need Look for an authorized API or direct source first. If the page itself must be rendered and access is allowed, consider a Browser job.
The output is valid but tracks the wrong content or too much content The selected URL, endpoint, or monitored data Choose a source that provides the needed content and use Urlwatch filters to focus the output.
The job succeeds sometimes and fails at other times Temporary connection problems, changing access policy, or request frequency Use the evidence from each run, avoid aggressive retries, and reduce or stop polling if the site’s instructions or behavior call for it. There is no universal retry interval that fixes an access denial.

3. Check an ordinary URL job’s request settings

Urlwatch URL jobs accept request headers and a user-agent value. These options are for configuring a permitted request—for example, supplying an API token to an authorized endpoint or using a truthful, stable client identity. They do not compel the destination to accept the request.

A minimal URL job in the Urlwatch jobs configuration can look like this:

- url: https://example.com/status
  name: Example status
  useragent: Urlwatch status monitor (contact: ops@example.com)
  headers:
    Accept: text/html

Replace the example URL and contact with values appropriate to your use. Add only headers required by the site’s documented interface or an agreement with its operator. Do not copy browser session credentials into a monitoring job unless the site explicitly authorizes that use and you can protect those secrets appropriately.

Urlwatch configuration details can change by version. Check the [current Urlwatch jobs documentation](https://urlwatch.readthedocs.io/en/latest/jobs.html) for the accepted job fields and syntax for your installed version.

When a header change is reasonable

  • The site documents a required header, such as an API authorization header, and you are permitted to use it.
  • The request is missing a format or content negotiation header that the documented endpoint requires.
  • You have verified a simple configuration error, such as a malformed URL or missing documented credential.

Do not cycle through user-agents to find one that passes a security rule. Cloudflare documents that user-agent rules can apply across a domain and take block or challenge actions; changing this value is not a general-purpose fix. See [Cloudflare’s user-agent blocking documentation](https://developers.cloudflare.com/waf/tools/user-agent-blocking/).

4. Prefer an authorized API or direct data source

If a page is displaying data from an API, a documented or otherwise authorized API response may be a better Urlwatch input than the full page. Urlwatch’s jobs documentation recommends considering an API-backed URL job in many such cases; it can be faster and avoids depending on page layout or browser rendering. Technical reachability alone does not establish permission to monitor an endpoint.

For example, if a site publishes an endpoint for a public status feed and its terms allow monitoring, the job can target that endpoint instead of the rendered status page:

- url: https://status.example.com/api/v1/summary.json
  name: Published status feed
  headers:
    Accept: application/json

This is illustrative configuration, not a real endpoint. Use the endpoint and authentication method the site actually documents. Confirm that the response contains the field or records you care about; a JSON response may include volatile timestamps or unrelated data that create noisy change alerts.

5. Use a Browser job only when rendering is needed

Urlwatch separates direct URL retrieval from browser navigation. A URL job retrieves a server response. A Browser job uses Playwright to render a page, which can help when JavaScript creates the content after the initial response. Browser jobs are resource-intensive and require the optional Playwright dependency and installed browsers described by Urlwatch 2.29’s documentation.

Install the optional dependency and browser components according to the instructions for your Urlwatch version. Then configure a navigation job using the documented fields. The exact setup can depend on the installed release, so consult the [Urlwatch Browser job documentation](https://urlwatch.readthedocs.io/en/latest/jobs.html) rather than assuming a config from a different version will work.

Browser jobs are a rendering option, not a block bypass. A rendered page can still receive an HTTP denial or challenge. If that happens, stop and ask the site operator for an authorized route instead of adding stealth settings or replaying visitor cookies.

6. Check monitoring frequency and change detection

Urlwatch retrieves job output, applies configured filters, compares it with previous output, and notifies you about changes. A job can therefore appear broken when the fetch succeeded but its output changed shape, includes dynamic values, or no longer contains the relevant content. Inspect the fetched output before rewriting filters.

The Urlwatch introduction recommends scheduling checks no more often than every 30 minutes. Treat that as a general upper-frequency recommendation from the documentation, not permission to poll every site at that rate. Site terms, published API limits, and the operator’s instructions may require a longer interval or no automated polling. Cloudflare’s [rate-limiting guidance](https://developers.cloudflare.com/waf/rate-limiting-rules/best-practices/) describes controls site operators can apply; its configuration examples are not universal Urlwatch retry rules.

  • Use a longer interval when the content changes infrequently or the site asks for one.
  • Avoid tight retry loops after denials, challenges, or repeated failures.
  • Filter out irrelevant changing fields only after confirming the job is allowed and returns the intended data.
  • For browser rendering, account for the additional runtime and resource use versus a direct URL job.

Urlwatch’s [introduction](https://urlwatch.readthedocs.io/en/latest/introduction.html) explains its jobs and notification workflow.

7. Troubleshoot common Urlwatch failures

Symptom Likely cause What to do
403 Forbidden or an access-denied response The destination refused the request; the status alone does not reveal which rule acted. Save the status and response details. Check the site’s terms and contact its operator about permission, an API, or a feed. Do not assume another user-agent will solve it.
A challenge or CAPTCHA page is recorded as the page content A security system is requiring interaction or denying automation. Do not automate challenge completion or reuse clearance cookies. Ask for an allowed monitoring route, or stop the job.
The URL job has no content that appears in a browser The content may be rendered by JavaScript, or the job may target the wrong source. Check for an authorized API first. If rendering is necessary and permitted, configure a Browser job with the required Playwright installation.
Browser job fails because Playwright or a browser is unavailable The optional dependency or browser installation required by the Urlwatch version is missing or incompatible. Follow the installation instructions for that version and verify its documented browser setup. A Browser job adds dependencies and runtime cost.
Job output changes on every run The page may contain timestamps, rotating content, or other irrelevant dynamic fields. Inspect the output and use appropriate Urlwatch filters to compare the stable content you need. Do not filter away an access-denial page and mistake it for a successful check.
Intermittent failures or throttling responses Temporary connectivity, site policy, request frequency, or a security control may be involved. Keep the run evidence, reduce frequency or stop polling, and ask the operator if needed. The available documentation does not establish a universal retry policy.
Changing the user-agent has no effect The cause may be unrelated to user-agent, or the site may intentionally block or challenge the request. Restore a truthful stable identity, inspect the response, and seek an authorized source. User-agent changes cannot guarantee acceptance.

8. Or skip the browser setup

If your goal is to capture a visual snapshot of an allowed page rather than keep a change-monitoring history, [ScreenshotNeo](https://screenshotneo.com) is a website screenshot API and MCP server for developers. Its screenshot API can return an image or PDF from one GET request; the [API documentation](https://screenshotneo.com/docs/) lists the supported parameters.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo request failed: ${res.status}`);
await Bun.write('shot.webp', res);

Replace YOUR_API_KEY with your key and the example URL with a page you are allowed to capture. The Node.js example uses Bun’s file-writing helper; in Node.js, save the returned response bytes using your preferred filesystem method. Keep API keys out of public code and client-side pages.

  • Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
  • Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers report the page verdict and billing status.
  • An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
  • The Free plan includes 1,000 screenshots a month with no card. Paid plans start at $5 for 3,000 screenshots.

Start with 1,000 free screenshots a month, with no card.

FAQ

Can Urlwatch monitor a website that blocks bots?

Only if the site provides or approves a route for that automated access. If it intentionally denies monitoring, ask the operator for permission or an alternative feed; otherwise stop the job.

Does a 403 prove Cloudflare blocked Urlwatch?

No. A 403 is an HTTP status, not an explanation of which provider or rule caused it. Inspect the response and, where available to the site operator, security event information.

Will a Browser job fix a Cloudflare block?

There is no such guarantee. Browser jobs render JavaScript content; they can still receive a block or challenge.

Should I use a proxy or copy browser cookies?

No. Those are not appropriate fixes for an intentional access denial. Use a documented, authorized endpoint or contact the site operator.

Is Urlwatch a screenshot tool?

Urlwatch monitors job output for changes. For visual captures, ScreenshotNeo provides a screenshot API and MCP server; see the [docs](https://screenshotneo.com/docs/).