ScreenshotNeo

BlogHow-to

How to Monitor a PDF Link on a Website with URLwatch

Track a PDF link’s destination with URLwatch, or monitor changes to the PDF’s extracted text. Set up filters, test them, schedule checks, and troubleshoot common issues.

By the ScreenshotNeo team4 October 20268 min read

To monitor a PDF link with URLwatch, create a job for the web page that contains the link, then filter the page output to the link’s destination. URLwatch compares that filtered result with the previous run and reports changes. If you mean changes to the PDF’s text rather than changes to the link, create a separate job for the PDF URL and use the pdf2text filter.

This guide covers both meanings, because they detect different changes: a link job detects when a link is added, removed, or points somewhere else; a PDF job detects changes in extracted document text. A text comparison does not establish that every visual or layout change in the PDF was detected.

1. Decide what you want to monitor

Goal Job URL What URLwatch compares
Detect a changed PDF link The page containing the link The selected link destination (href)
Detect changed PDF text The PDF itself Text extracted from the document

Use the first method if you need to know when a page’s PDF destination changes. Use the second if the URL remains the same but the file’s text may be revised. You can configure both jobs when both changes matter.

2. Install URLwatch and prepare its job list

Follow the official URLwatch installation instructions for your platform. Run URLwatch once, then open the job list with:

urlwatch --edit

Use urlwatch --edit-config if you want to configure reporters such as email. Without a configured reporter, you can still run the job and inspect its console output.

Set the job’s url to the HTML page that contains the link. Add an XPath filter that selects the matching anchor’s href. For example, this job selects links whose href contains .pdf:

name: "PDF link on example page"
url: "https://example.org/documents/"
filter:
  - xpath: '//a[contains(@href, ".pdf")]/@href'
  - sort:

Save the job in URLwatch’s job list. Replace the example page with the real page you want to watch. The XPath is only an example: a page may have multiple PDF links, use a different URL pattern, or structure its markup differently. Make the selector specific enough to select the intended link.

The XPath above selects every anchor whose href contains the literal string .pdf. If the page has multiple PDFs, narrow the XPath using a stable attribute, a parent element, or link text. For example, if the target link has a stable ID, use an expression such as:

filter:
  - xpath: '//a[@id="annual-report"]/@href'

Use only attributes and structure that actually exist on the target page. XPath and CSS selection, along with chained filters, are documented in the URLwatch filter documentation.

The sort filter in the example makes a list of selected destinations deterministic if more than one matching link is found. If your selector should return exactly one value, you may not need it.

Preview the filter before relying on it

Run URLwatch’s filter test command and check that the output is the destination you intend to track:

urlwatch --test-filter

Use the filter preview to catch selectors that return nothing, select the wrong link, or return several links unexpectedly. Confirm the output after changing the page URL or filter. A successful preview verifies what the current job extracts; it cannot guarantee the site’s markup will remain unchanged.

4. Monitor changes to the PDF’s text

To compare document text, make a separate job whose URL is the PDF itself. Put pdf2text first in the filter chain because it consumes binary PDF data and produces text:

name: "PDF document text"
url: "https://example.org/documents/report.pdf"
filter:
  - pdf2text
  - strip

The strip filter is optional; it can remove surrounding whitespace from the extracted output. The pdf2text filter requires pdftotext and its operating-system dependencies. Install those dependencies for your environment if PDF text extraction fails, following the URLwatch filter documentation for the supported requirements.

This method compares extracted text. It may not report a change that affects only layout, images, typography, or other visual details without changing the extracted text.

5. Schedule recurring checks and notifications

URLwatch checks jobs when it runs; the schedule is determined by how often you invoke it. The official introduction recommends running it no more frequently than every 30 minutes. For a simple recurring setup, schedule the command with cron on systems that use cron. On Windows, URLwatch’s quick start points to Windows Task Scheduler.

Choose an interval that matches how quickly you need to know about changes and the site’s acceptable request frequency. Configure a reporter with urlwatch --edit-config if you want changes delivered somewhere beyond the console. The URLwatch introduction lists email and third-party messaging reporters; each requires its own configuration.

URLwatch retrieves a job’s output, applies its filters, compares the result with the prior run, and invokes enabled reporters when it detects differences. A regular url job retrieves the web-server response. If the page requires JavaScript to render or reveal the PDF link, URLwatch also documents a navigate job that uses a headless browser.

6. Complete runnable examples

name: "PDF link on example page"
url: "https://example.org/documents/"
filter:
  - xpath: '//a[contains(@href, ".pdf")]/@href'
  - sort:

Save this job in the URLwatch job list, replace the example URL and adjust the XPath to match the actual page. Run the filter preview and confirm the extracted destination before scheduling checks.

URLwatch YAML: watch the PDF’s extracted text

name: "PDF document text"
url: "https://example.org/documents/report.pdf"
filter:
  - pdf2text
  - strip

Install pdftotext and its required system dependencies, then verify the filter output. The first filter must remain pdf2text so the binary response is converted to text before later filters run.

7. Or skip the browser setup

If you also need screenshots of the page or PDF, ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. Its API returns a screenshot or PDF from one GET request. See the ScreenshotNeo API documentation for the available parameters and response details.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.org/documents/ -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.org/documents/"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.org/documents/'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo returned HTTP ${res.status}`);
await Bun.write('shot.webp', new Uint8Array(await res.arrayBuffer()));

These requests capture the page; they do not replace URLwatch’s recurring change detection. ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture, and each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account to get started.

8. Troubleshooting

Symptom Likely cause What to check or change
The link job extracts no value The XPath does not match the page’s markup, the href does not contain .pdf, or the page has not exposed the link in the retrieved HTML. Inspect the page’s actual markup, make the selector match its anchor, and preview the filter output with urlwatch --test-filter.
The job selects the wrong PDF or several links The selector is too broad. Narrow the XPath to a stable attribute, containing element, or other page-specific structure. Check whether sort is appropriate if multiple values are intentional.
The link appears in a browser but not in the job output The page may render or reveal the link with JavaScript, or the site may respond differently to automated requests. Check the retrieved page and filter preview. If JavaScript rendering is required, consider URLwatch’s documented navigate job; investigate site-specific access restrictions separately.
The PDF text job errors during filtering pdftotext or an operating-system dependency is missing, or the response is not a PDF that the local setup can process. Install the required dependencies, confirm the URL returns the intended PDF, and ensure pdf2text is first in the filter list.
No notification arrives No reporter is enabled, its configuration is incomplete, or no filtered change was detected. Review reporter settings with urlwatch --edit-config, check the command’s output, and verify the filter returns the value you intend to compare.
Changes are missed between runs The target changed and changed back between scheduled checks, or the chosen filter omits the changing value. Choose an interval appropriate to the change window, within the documented recommendation of no more often than every 30 minutes, and test that the filter includes the destination or text of interest.

9. Performance, reliability, and cost considerations

  • Request frequency: URLwatch’s introduction recommends scheduling runs no more frequently than every 30 minutes. Select an interval based on how quickly you need change alerts.
  • Keep the comparison small: Filtering to a single href or extracted document text focuses the diff on the change you care about and avoids unrelated page content changes for the link-monitoring job.
  • JavaScript and site behavior: A normal URL job retrieves the web-server response. JavaScript-dependent links may require a navigate job. Automated requests can also encounter site-specific blocks or differences.
  • PDF text limitations: Text extraction needs pdftotext and OS dependencies. It compares extracted text, not every visual or layout property.
  • Reliability checks: Preview filters before scheduling, and revisit the extracted output if the site changes. URLwatch’s filter preview validates the current extraction, not future selector stability.
  • Cost: The documented workflow uses URLwatch and local PDF text extraction dependencies. The cited documentation does not state a service price or establish a universal cost for hosting or notifications.

10. Frequently asked questions

Use separate jobs for clarity: one job targets the containing page and extracts the href; another targets the PDF and extracts text with pdf2text.

Not necessarily. The link job compares the extracted destination. To detect text revisions at a stable URL, monitor the PDF itself with pdf2text.

Does a PDF text diff prove that the document looks identical?

No. The documented filter extracts text. A text comparison does not establish that visual or layout changes were detected.

The regular url job retrieves the web-server response. URLwatch documents a navigate job using a headless browser for pages that require JavaScript.