ScreenshotNeo

BlogGuides

Urlwatch Review: Setup, Alerts, and Limitations

A practical Urlwatch review covering installation, jobs, filters, scheduled checks, alerts, troubleshooting, and when a self-managed monitor is the right fit.

By the ScreenshotNeo team4 October 20267 min read

Urlwatch is an open-source change monitor that retrieves configured job output, applies filters, compares the result with the previous run, and reports detected differences. It can watch ordinary web responses, pages that need a headless browser, and shell commands. You run it on a schedule you manage; it is not a hosted monitoring service that polls continuously on your behalf. Notifications can include a unified diff showing what changed. Urlwatch handbook: introduction and quick start.

Urlwatch suits developers who want control over what counts as a change, where the monitor runs, and how reports are delivered. Its trade-offs are the setup and ongoing care: configure jobs and filters, arrange recurring execution, and set up any external reporter you need.

What Urlwatch does

Each invocation retrieves each job’s output, applies its configured filters, compares the filtered result with the previous stored version, and invokes enabled reporters when it detects a difference. The result is a change report, not a judgment about whether a change matters to you. Dynamic timestamps, rotating recommendations, or unrelated page content can trigger noisy differences unless you narrow the content with filters.

The documented job types are:

  • url: fetch a web server response, using HTTP GET by default.
  • navigate: load a page with a headless browser when client-side rendering is required.
  • command: run a shell command and compare its output.

Each job uses one of these defining keys. The job editor validates configuration before activating it. See the jobs handbook.

Setup and first scheduled check

  1. Install Urlwatch. The current handbook documents installation with Python’s package installer. Check the official installation instructions for platform requirements and verification steps.
  2. Run urlwatch once. The quick start says this initializes or migrates its data.
  3. Define jobs and filters. Run urlwatch --edit to edit urls.yaml.
  4. Configure settings and reports. Run urlwatch --edit-config to edit urlwatch.yaml.
  5. Run it on a recurring schedule. Use cron on Unix-like systems or Task Scheduler on Windows. The interval is determined by how often the scheduler launches the command. The handbook recommends not checking more frequently than every 30 minutes.
  6. Verify the command and alert route. Run Urlwatch manually and use the reporter test option described below before relying on scheduled notifications.

The quick-start workflow and 30-minute recommendation are documented in the official introduction. Exact installation prerequisites can vary, so use the current installation page rather than assuming a particular Python or operating-system setup.

Example cron entry

*/30 * * * * /path/to/urlwatch

Use the executable path available to the account that owns the crontab. If the environment used for the interactive install differs from cron’s environment, the scheduled invocation may not find Urlwatch or its configuration. A scheduler controls timing; Urlwatch itself does not stay resident between invocations.

Define a job and reduce noisy changes

Edit jobs with urlwatch --edit. A minimal URL job looks like this:

name: Example release page
url: https://example.com/releases

That watches the retrieved response. To monitor a specific part of the page, chain filters: select content, convert it to readable text, then apply text operations. The following illustrates the documented filter pattern; adjust the XPath and matching text for the page you actually want to watch.

name: Example release notes
url: https://example.com/releases
filter:
  - xpath: '//main'
  - html2text:
      method: pyhtml2text
      unicode_snob: true
      body_width: 0
      inline_links: false
      ignore_links: true
      ignore_images: true
      pad_tables: false
      single_line_break: true
  - grep: 'Version|Released'
  - sort:

Filters can select HTML by CSS or XPath, convert HTML to text, process PDF or JSON data, and manipulate text with operations such as grep, strip, and sort. Chaining is useful for removing irrelevant page material, but it is configuration you maintain. It does not infer semantic importance automatically. The filters handbook describes available filters and examples.

JavaScript-rendered pages and other job options

Use a navigate job when the relevant content is rendered in a browser and is not present in the ordinary server response. Use command when the value to watch comes from a local or external command. Urlwatch’s jobs handbook documents optional keys for job types and shared settings; consult it for the exact options supported by your installed version instead of copying settings from an unrelated job.

Test filter changes against current content

Use urlwatch --test-filter to inspect a filter against the current page contents. This matters because stored output was filtered when it was retrieved: if you change a filter, an unchanged report can still show the older content processed with the earlier filter. Editing the filter does not retroactively re-filter that stored snapshot. This behavior is called out in the configuration handbook.

Configure alerts

Urlwatch displays reports on standard output by default. To send notifications elsewhere, enable and configure a reporter in urlwatch.yaml. Documented integrations include SMTP email, Mailgun, Slack, Discord, Pushbullet, Telegram, Matrix, Pushover, XMPP, and shell commands. Reporter availability and extra dependencies can differ; use urlwatch --features to see the reporters available in your installation. Most reporters need additional configuration, and integrations are disabled unless enabled.

Test a configured reporter with its name, for example:

urlwatch --test-reporter stdout
urlwatch --test-reporter email

The reporter handbook explains that the test report can include configured notification types and that --verbose can help diagnose delivery configuration. Provider-side service availability and terms are separate from Urlwatch. See reporters documentation and configuration documentation.

Limitations and practical trade-offs

  • You own the schedule. Checks happen only when a scheduler invokes Urlwatch. If the host is off, the job cannot run then. A server is an optional deployment choice, not a documented requirement.
  • Filters need maintenance. Page markup and content can change. A selector may stop matching or begin including unrelated material, so review the extracted output after target-page changes.
  • Changes are output differences. Urlwatch reports differences in retrieved and filtered output. It does not decide whether a price, announcement, or layout change is meaningful.
  • External alerts require setup. Reporter credentials and provider configuration add work, and delivery depends on the external destination.
  • No performance or reliability figures are established here. The reviewed documentation does not provide benchmarks, accuracy rates, or uptime statistics. Actual request time and reliability depend on the pages, browser jobs, host, scheduler, and network.

Performance, reliability, and cost

Urlwatch’s runtime depends on the number and type of jobs, target response times, and whether browser navigation is needed. Narrow filters help make reports more useful, but do not by themselves reduce the cost or duration of retrieving the page. Choose a schedule that respects the handbook’s recommendation of at least 30 minutes between runs, and consider how many requests your job list creates per run.

Reliability depends on the machine being available, the scheduler running correctly, network access, target-site behavior, and reporter configuration. Check scheduler logs and run Urlwatch manually when a notification is missing. The cited material does not establish a hosted service, service-level guarantee, or measured operating cost. The software can run on an existing computer; optional server hosting is a deployment decision, not a prerequisite in the docs.

Or skip the browser setup

If your goal is to capture a webpage as an image or PDF rather than monitor changes over time, ScreenshotNeo is a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. Its API supports URL capture and browser settings such as waits, selectors, and full-page capture; see the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot. Bot checks, blank pages, and failed loads are never billed; responses identify page verdict and billing status. Its MCP server lets AI agents take screenshots. The Free plan includes 1,000 screenshots per month with no card, and paid plans start at $5 for 3,000. Screenshot capture is a point-in-time image or PDF, not a scheduled change-monitoring workflow like Urlwatch.

Sign up for ScreenshotNeo’s free plan: 1,000 screenshots a month, no card required.

Troubleshooting

Symptom Likely cause What to do
No recurring checks The scheduler is not invoking Urlwatch, or the scheduled environment cannot find the command. Run the command manually as the scheduler’s user; check the scheduler entry, executable path, and logs.
Too many change reports The job includes volatile or irrelevant page content. Use CSS or XPath selection and chain text filters; inspect the extracted content with --test-filter.
Changed filter appears ineffective in an unchanged report The stored snapshot retains output filtered at retrieval time. Use urlwatch --test-filter against current content; do not treat an unchanged report as a re-filtered snapshot.
JavaScript content is missing The url job reads the server response, which may not include client-rendered content. Use a navigate job and consult the jobs handbook for its browser configuration.
Email or messaging alert does not arrive The reporter may be disabled, incomplete, unavailable in the installation, or rejected by its provider. Check reporter settings and installed features, run urlwatch --test-reporter NAME, then add --verbose for diagnostic output.
Job configuration does not activate Invalid YAML or an unsupported option can prevent the edited job from validating. Review indentation and spelling, use the job editor’s validation feedback, and verify options against the documentation for the installed version.

Frequently asked questions

Does Urlwatch notify me instantly?

No. It checks when invoked, so notification timing follows the schedule you configure and the time needed to fetch and report results.

Can Urlwatch watch something other than a webpage?

Yes. Its documented command job compares shell-command output, while URL and browser jobs cover web content.

Does it know whether a change is important?

No. It reports output differences after filtering. You choose the monitored content and decide what the diff means.

Is Urlwatch a hosted alternative to screenshot APIs?

No. Urlwatch is a self-managed change monitor. ScreenshotNeo captures images or PDFs through an API or MCP server; it does not replace Urlwatch’s scheduled diff workflow.