AI Tools for Journalists: Tracking Website Updates and Sources
Use feeds to find new publications and page monitors to catch edits to known sources. Learn how to configure alerts, preserve evidence, and verify changes before reporting them.
Use RSS feeds and search alerts to find new publications; use a website change monitor to detect edits to a page you already rely on. Treat every alert as a reporting lead, not proof: preserve the before-and-after evidence, inspect the original source, and verify consequential changes with the publisher or another authoritative record.
Choose the right kind of monitoring
“Website updates” can mean two different things: a new item has been published, or an existing page has changed. The distinction matters because a tool that finds new articles may not detect edits to an old policy page.
| Reporting need | Start with | What it can catch | Limit to remember |
|---|---|---|---|
| New articles, releases, or feed posts | RSS reader | New entries published in a feed | A feed does not necessarily report edits to existing pages. Feeder’s guide explains the distinction and feed options. |
| New items on a page with no feed | Custom feed, if available | New entries in a listing such as jobs, events, or headlines | Some services limit custom-feed features by plan; check current terms. |
| New indexed mentions of a subject | Search alerts | Newly found content matching a query | Search indexing is not a complete change log for a specific page. Tune source, language, region, and frequency where offered. The Media Rights Agenda guide describes Google Alerts options. |
| Changed wording or content on a known page | Page-change monitor | Scheduled comparisons of a whole page or selected area | Dynamic page elements can generate noise. Select a relevant section when possible and inspect the comparison. Visualping’s help documentation describes page and element monitoring. |
| New values in a frequently updated dataset | Dataset monitoring workflow | Criteria applied to updated data, with alerts for matches | A paper describes Datastringer’s approach; that paper is not evidence of current commercial availability. See the Datastringer paper. |
Investigative targets often include procurement records, court dockets, environmental action pages, terms of service, regulator notices, and open-data portals. GIJN discusses these use cases and monitoring selected portions of pages in its guide to searching text and tracking website changes.
Build a useful source-monitoring workflow
- Make a source list. Record the canonical URL, publisher, topic, why the page matters, and the type of change that could affect your reporting. Include official statements, public records, company policies, court and procurement pages, regulator pages, data portals, and recurring source pages.
- Check for an official update channel first. Subscribe to a publisher’s RSS feed or native alert for new releases. Add search alerts for new indexed mentions of a subject. Neither should be treated as a reliable record of every edit to an existing page.
- Add page monitoring for important pages without useful feeds. Choose whole-page monitoring if any wording could matter. Choose a specific element or text area when unrelated navigation, timestamps, ads, or rotating content would make whole-page alerts noisy. If the service supports criteria, describe the change you care about in plain language.
- Set a practical check interval. Match it to your deadline and the source’s likely update rate. A page that changes rarely may not need frequent checks; an imminent filing deadline may justify closer monitoring if the service offers it. Do not call a service “real time” unless its documentation for your plan supports that claim.
- Route alerts to a place you review. Use email, a team channel, or another supported destination. Keep the source URL and enough context in the alert so that you can tell which watch triggered it.
- Review every alert as a lead. Open the comparison and original page. Determine which text or element changed, whether the page loaded correctly, and whether the change affects the reporting question. Check for an official update, a second authoritative record, or confirmation from the publisher before making a consequential claim.
- Preserve evidence. Save dated before-and-after text or screenshots, the URL, the alert time, and the monitor settings or criteria. Follow newsroom rules for storing sensitive material and documenting how it was collected.
- Maintain the watch list. Revisit it after publication. Remove stale monitors, narrow noisy ones, and note any source pages the service could not access.
Reduce false alerts without hiding meaningful changes
- Monitor a stable page region. A selected article body or document section can avoid alerts caused by unrelated page furniture. Confirm that the selected region still includes the information you need.
- Distinguish content from page behavior. Rotating headlines, live counters, current dates, and personalized content may change on every visit. Where the tool allows it, exclude those areas or focus on the text that matters.
- Use more than one discovery path for high-value sources. A feed can surface new releases while a page monitor watches revisions to a standing policy page. These methods cover different kinds of change.
- Keep the original context. A snippet or AI-generated summary can help triage an alert, but it cannot establish why wording changed or whether the change is significant. Read the surrounding page and, where needed, the linked source document.
- Record access failures. A monitor that cannot load a page is a gap, not evidence that the page stayed unchanged. Some pages require a login, block automated access, or expose different content to different visitors.
What AI can and cannot do in this workflow
AI features can help sort alerts, summarize a before-and-after comparison, or answer questions about a monitored page. Visualping documents AI summaries and questions alongside comparison evidence in its help center. Use that output to decide what to inspect first; verify the change in the source itself.
Google Journalist Studio describes Pinpoint as a tool for analyzing large collections of PDFs, images, handwritten notes, emails, and audio using Google Search, AI, and machine learning. That supports document research; it does not establish that Pinpoint tracks arbitrary revisions to websites. See Google Journalist Studio.
Research on dataset monitoring offers another pattern: define conditions to check against changing datasets and alert when those conditions are met. The Datastringer paper describes such a system. Treat the paper as a description of the research system, not a claim that a current service is available or suitable for a newsroom.
Verify an alert before citing a change
- Open the alert’s source URL directly and note the time you reviewed it.
- Compare the preserved old and new versions. Identify the exact passage, number, image, or field that changed.
- Check that the page was not blank, partially loaded, or displaying a temporary error during either capture.
- Look for a dated release, linked document, public record, or other authoritative source that explains or confirms the change.
- Contact the publisher or responsible office when the change is material and its timing or meaning is unclear.
- In your reporting notes, distinguish what the archived evidence shows from what you infer about intent, cause, or significance.
An alert establishes that a monitoring system observed a difference under its settings. It does not by itself establish that the revision is accurate, complete, intentional, or newsworthy.
Capture a page for a dated comparison
A browser screenshot or saved page can help document what was visible at a particular time. For a manual capture, open the source in a browser, wait for the relevant content to load, capture the page or the relevant region, and record the URL and capture time with the file. A screenshot is a visual record, not a substitute for retaining text or checking the original.
Capture with a browser
- Open the canonical source URL and confirm you are looking at the intended page and access level.
- Wait for the text or data you need to appear. If the page loads content as you scroll, scroll through the relevant area first.
- Capture the whole page when page context matters, or the relevant region when the evidence is a specific panel or table.
- Keep the original file and record the URL, date and time, and any relevant monitor alert or notes.
Capture with a browser automation library
For repeatable captures in a Node.js project, Playwright can open a page and save a screenshot. Install it in a project with npm install playwright, then save the following as capture.mjs. Run it with node capture.mjs https://example.org. Replace the example URL with a page you are authorized to access.
import { chromium } from 'playwright';
const url = process.argv[2];
if (!url) throw new Error('Usage: node capture.mjs https://example.org');
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage({ viewport: { width: 1440, height: 1000 } });
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 45000 });
await page.screenshot({ path: 'page.png', fullPage: true });
console.log(`Saved page.png for ${page.url()} at ${new Date().toISOString()}`);
} finally {
await browser.close();
}
This basic example captures the rendered page after the initial document is ready. It does not handle authentication, consent interactions, page-specific readiness, or archival requirements. Add those only as appropriate for the source and your newsroom’s policies. A full-page screenshot may be long, and lazy-loaded content may need scrolling before capture.
Capture with cURL, Python, or Node.js using ScreenshotNeo
For a one-request screenshot of a public page, ScreenshotNeo is a website screenshot API and MCP server for developers. The API returns an image or PDF from a URL. This is useful for capturing a page snapshot for review or comparison; it does not monitor a URL on a schedule or establish that a revision is true. See the ScreenshotNeo API documentation for request options. Set YOUR_API_KEY to your API key and use a source URL you are authorized to access.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.org \
-o source.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.org"},
timeout=90,
)
r.raise_for_status()
with open("source.webp", "wb") as f:
f.write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://example.org'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('source.webp', res);
Or skip the browser setup
ScreenshotNeo takes a screenshot from one API request. Its capture flow accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each of those steps can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies page verdict and billing status in headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. It is a capture tool, not a scheduled page-change monitor.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo includes full-page capture with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets and custom viewports, retina scale, PDF settings, custom CSS and JavaScript, click-before-capture, hide selectors, waits, request and resource blocking, custom headers and cookies, user agent and Authorization, timezone and geolocation, transparent backgrounds, resizing, configurable cache TTL, signed links for public image tags, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, a usage API, and an OpenAPI spec. Parameters used by other screenshot APIs also work, which can make migration easier.
There are 1,000 screenshots a month free with no card. Paid plans start at $5 for 3,000 screenshots; higher listed plans are Growth at $15 for 15,000, Pro at $39 for 60,000, Scale at $99 for 250,000, and Business at $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan. Sign up for 1,000 free screenshots a month with no card.
Performance, reliability, and cost considerations
- Alert speed depends on the schedule and access. A scheduled monitor can only detect a change after a successful check. A fast interval may be useful for a deadline, but do not assume checks are instantaneous or continuous unless the service documents that for your plan.
- Dynamic pages need tuning. A page with rotating modules or client-side content can create noisy comparisons. Select a stable region and verify that the monitor loads the content you need.
- Access can fail silently unless you watch for it. Login walls, paywalls, bot protections, and network failures can prevent a monitor from seeing the source. Feeder notes that its monitoring cannot access logged-in or paywalled pages, and that feeds may be unavailable. Treat failed access as a monitoring gap. See the Feeder guide.
- Retain independent evidence. Store the alert, comparison, capture timestamp, source URL, and relevant original file under newsroom retention practices. A screenshot records visible appearance; it may not preserve searchable text, metadata, or the complete underlying record.
- Budget for coverage, not just alert volume. Compare how many pages, checks, or feeds you need and whether the plan supports the alert destinations and evidence retention your workflow requires. Vendor prices and limits change, so verify them before choosing.
- Do not infer accuracy from AI summaries. A summary can omit context or misread a layout change. Inspect the source evidence before citing a revision.
GIJN reported that Visualping’s journalist plan had a commercial value of $120 per year, up to 1,000 checks per month, and 25 pages at a time, compared with 150 checks and five pages for its public free level. These are vendor-program terms reported by GIJN, not independent performance measures; verify current eligibility and limits before relying on them. See GIJN’s report. The reviewed sources provide no cross-tool statistic for monitoring accuracy, detection latency, or journalism outcomes.
Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| No alert arrived | The page did not change in a monitored area, the next scheduled check has not run, the feed is unavailable, or the service could not access the page. | Check service status and the last successful check, confirm the URL and interval, and open the page as the monitor sees it if that feature exists. Add a second discovery method for critical sources. |
| Too many alerts | Ads, counters, dates, navigation, or rotating page elements are changing. | Monitor the relevant element or text, exclude routine regions where supported, and adjust criteria. Review several alert examples to ensure the filter does not hide meaningful edits. |
| Alert says changed but the wording looks the same | Formatting, hidden elements, a timestamp, or a small layout change may trigger a visual comparison. | Inspect both versions and the exact diff. Switch to text or element monitoring if available, or narrow the selected area. |
| Page monitor sees a blank or error page | The site may block automation, require login, time out, or have been temporarily unavailable. | Check the original manually, record the failed check, and do not interpret it as an unchanged page. Follow source access rules and newsroom policy. |
| Feed misses a page revision | RSS commonly announces new feed items and may not include edits to older pages. | Add a page-change monitor for the existing page and use the feed for new publications. |
| Browser capture omits lower-page content | Content may load only after scrolling or after client-side rendering completes. | Scroll to the relevant area before capturing, wait for a page-specific element, or use a suitable full-page capture option. Confirm the saved image includes the target content. |
| ScreenshotNeo returns an unexpected result | The URL may be inaccessible, the page may be blank, a bot check may intervene, or the request may have timed out. | Inspect the response’s X-Page-Verdict and X-Billed headers, verify the target URL, and retry only after checking whether the source is available. Consult the API documentation for supported parameters. |
| API request fails | The key, URL encoding, network response, or timeout may be wrong. | Confirm the API key is present, encode the URL, check the HTTP status and response headers, and allow an appropriate timeout for slow pages. Keep keys out of public code and logs. |
FAQ
Can an alert prove when a source changed?
It can document when a monitoring service observed a difference. The true edit time may be earlier, and the alert does not independently establish who made the edit or why.
Should I archive every alert?
Preserve alerts and versions that may affect a story, public record, or later verification. Apply your newsroom’s retention and security rules, especially for restricted sources.
Can Pinpoint replace a website change monitor?
The cited Google description covers analysis of document and media collections. It does not establish Pinpoint as an arbitrary website revision monitor.
Is a screenshot enough to establish what a page said?
It is useful visual evidence of what a capture displayed. Pair it with the URL, timestamp, and other available records, and retain text or source documents when the exact wording matters.
Sources and further reading
- GIJN: Toolbox on searching text and tracking website changes
- Visualping help: monitoring pages and reviewing changes
- Feeder: website change alerts and RSS guidance
- Media Rights Agenda: digital tools for journalism practice
- Google Journalist Studio
- Datastringer paper and survey of change detection and notification systems


