How to Configure Urlwatch to Monitor a Page for Text Changes Only
Configure Urlwatch to compare only the page text you care about. Learn how to select content, build a filter pipeline, test changes, and troubleshoot noisy diffs.
To make Urlwatch monitor text changes only, define a job in urls.yaml and add a filter pipeline that converts the retrieved HTML to text before Urlwatch compares it with the previous result. A typical pipeline is html2text, optionally followed by grep to retain relevant lines, then strip to remove surrounding whitespace. If the page has a stable content region, select it with a CSS or XPath filter before converting it to text. Preview the output with urlwatch --test-filter and adjust it until it contains only the text whose changes matter.
1. Create a text-only Urlwatch job
Install Urlwatch, then create or edit its job list with urlwatch --edit. For a page whose relevant content is present in the HTTP response, use a regular URL job:
name: "Watch page text"
url: "https://example.com/page"
filter:
- html2text
- grep: "Text I care about"
- strip
Replace the example URL and grep expression with values for your page. The grep filter retains matching lines, so use it when the text you care about has a stable phrase, label, or pattern. Remove the grep step if you want to compare all converted page text.
2. Isolate the relevant page region
Converting the whole page to text may still include navigation, footer links, timestamps, or other changing material. When the target copy sits inside a stable element, select that element first:
name: "Watch article text"
url: "https://example.com/page"
filter:
- css: "main article"
- html2text
- strip
main article is only an example selector; inspect the target page and substitute its actual markup. Urlwatch documentation also supports XPath and built-in element filters. Apply selection before html2text so that unrelated HTML never reaches the text conversion step.
Choose the narrowest stable region that includes all the copy you need. If the page changes its markup, the selector may stop matching; a filter preview helps reveal an empty or incomplete result.
3. Understand the filter pipeline
| Filter | Purpose | When to use it |
|---|---|---|
css or XPath |
Selects a part of the HTML document | Exclude page chrome and monitor a specific content block |
html2text |
Converts HTML to readable plaintext | Most text monitoring jobs |
grep |
Keeps lines matching a pattern | Track a stable labeled value or a small set of lines |
strip |
Removes leading and trailing whitespace | Ignore whitespace at the edges of the output |
striplines |
Removes empty or whitespace-only lines | Normalize blank-line differences |
sort |
Sorts lines | Ignore changes in order when line order is irrelevant |
Filters run in order. Select first, convert the selected HTML to text, then normalize or narrow the text as needed. Do not use sort when ordering conveys meaning, and do not use grep if the changing line might cease to match the expression: the resulting output could become empty and hide the change you meant to catch.
4. Preview exactly what Urlwatch will compare
- Run
urlwatch --test-filter <job-index-or-url>for the job. - Read the complete output. Confirm that the target text is present and unrelated page content is absent.
- Adjust the selector, grep expression, or whitespace filters and preview again.
- Run Urlwatch normally after the preview is correct, then schedule it at the interval appropriate for the page.
Urlwatch compares each job’s filtered output with the previously retrieved output and invokes enabled reporters when it detects differences. The schedule is controlled by how often Urlwatch is run; its current introduction recommends running no more often than every 30 minutes. See the Urlwatch documentation and its filter documentation for version-specific details.
5. Handle JavaScript-rendered content
A regular URL job is suitable when the server response already contains the text. If a browser normally displays the target only after JavaScript runs, a direct HTTP retrieval may not include it. First check whether the site exposes an underlying API that provides the relevant data. Otherwise, use the Urlwatch browser or navigation job approach described in its documentation and apply filters to the resulting content.
Browser rendering adds setup and runtime cost compared with fetching an HTML response directly. Keep the monitored region and wait condition focused so the job captures the intended state. If the page relies on client-side loading, verify the final filtered text rather than assuming the browser produced the expected content.
6. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| The diff contains menus, footer links, or other noise | The job converts the whole page | Add a CSS, XPath, or element-selection filter before html2text. |
| The filter output is empty | The selector does not match, or grep removed every line | Preview with --test-filter; verify the actual markup and pattern, and temporarily remove grep to locate the failure. |
| Text visible in a browser is missing | The content is injected by JavaScript after the HTTP response | Look for an underlying API or configure a browser/navigation job; then verify the rendered filter output. |
| Urlwatch reports a large change after editing filters | The new run uses the new pipeline while stored output was produced with old settings | Inspect and interpret the first diff after a filter change carefully. Confirm the new output is correct before treating it as a page-content change. |
| Whitespace-only differences trigger changes | Leading, trailing, or blank-line whitespace remains in the output | Use strip and, if appropriate, striplines; preview to ensure meaningful line boundaries remain. |
| Changes disappear after adding grep or sort | The grep pattern excludes the changed line, or sorting hides meaningful order changes | Use a broader matching pattern or remove the narrowing filter; preserve line order if it matters. |
7. Reliability, performance, and cost considerations
For the most reliable text-only comparison, target a stable content container, normalize only inconsequential whitespace, and inspect the filter output after any site markup or filter change. A direct HTTP job avoids browser rendering when the response already contains the text. JavaScript-heavy pages can require browser navigation and may be more sensitive to load timing.
Run Urlwatch on a schedule that meets your freshness needs while respecting the target site’s expectations; the Urlwatch introduction recommends no more often than every 30 minutes. The retrieved documentation does not specify a universal runtime or price, so both depend on your execution environment, job count, and whether browser rendering is needed.
Or skip the browser setup
If your page needs a rendered capture alongside monitoring, ScreenshotNeo is a website screenshot API and MCP server. One GET request returns an image or PDF; its capture options include waiting for selectors or network idle, custom JavaScript, and element capture. For text change monitoring specifically, Urlwatch’s filters perform the comparison; a screenshot is useful when the visual state matters too.
See the ScreenshotNeo API documentation. Example request:
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.com/page \
-o shot.webp
ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, with no card.
FAQ
Does Urlwatch compare the original HTML?
With a filter pipeline configured, it compares the output after those filters have been applied.
Can I watch one value on a page?
Yes. Select its containing element and use a text filter such as grep when the value has a stable matching pattern.
Will changing filters reprocess the old stored result?
Do not assume so. The next comparison can be between newly filtered output and previously stored output produced with the old filters, so review that first diff.


