ScreenshotNeo

BlogUse cases

How to Automate Competitive Intelligence Across Websites

Build a competitor monitoring workflow that tracks pricing, product, and messaging changes, filters noise, and preserves evidence for review.

By the ScreenshotNeo team4 October 202612 min read

To automate competitive intelligence across websites, create a focused list of competitor pages tied to specific questions, check them on a schedule, compare either the whole page or relevant sections, and route changes with their source URL and timestamp to a place an analyst reviews. Treat an automated change as a lead to verify, not as a conclusion.

This guide covers a maintainable monitoring workflow, a runnable Python implementation, ways to choose hosted or developer-managed approaches, and how to preserve visual evidence with ScreenshotNeo.

1. Decide what you need to learn

Start with questions rather than a long list of URLs. For example:

  • Did a competitor change plan limits, prices, or packaging?
  • What product capability or integration was added?
  • Did positioning, claims, or target-customer language change?
  • Are hiring patterns signaling a new investment area?

Choose pages that can answer those questions: pricing and packaging, product or feature pages, changelogs and release notes, marketing or positioning pages, and selected careers pages. A page monitor will not cover every market signal. Add appropriate public news sources or filings when those matter.

2. Choose an implementation pattern

Approach Good fit Trade-off to consider
Hosted page monitor A team wants scheduled checks, page or element selection, and before-and-after alerts with little infrastructure. Notification options, integrations, and monitoring controls vary by service and plan. Visualping documents scheduled cloud checks, whole-page or selected-area monitoring, and visual, text, and code change detection. Visualping: What is Visualping?
Self-hosted monitor and API You want to manage deployment, watch configuration, fetchers, filters, and notifications yourself. You operate the installation and its configuration. changedetection.io documents a REST API for watches, groups, and notifications as well as schedules, browser fetchers, selectors, text filters, and change processors. API reference · Project documentation
Developer-oriented extraction pipeline You need structured fields from selected pages, or want to feed collected data into a warehouse, dashboard, CRM, or intelligence brief. Extraction and monitoring capabilities are vendor-described product claims, not independent performance results. Firecrawl describes competitor-page monitoring and structured extraction workflows. Firecrawl: competitive intelligence and market monitoring
Custom script You have a manageable list of public pages, clear comparison rules, and the capacity to maintain fetching and storage. You own browser behavior, retries, scheduling, alert delivery, data retention, and selector maintenance.

The sources describe different capabilities; they do not provide an independent, like-for-like benchmark. Select based on the pages you monitor, required output, and operating burden rather than an unsupported accuracy ranking.

3. Build a reliable monitoring workflow

  1. Keep the watchlist bounded. Store one row per question and page, with the competitor, category, URL, important region or selector, and owner.
  2. Set the comparison scope. Use a whole-page comparison when layout or broad content matters. Select a region or filter text when navigation, rotating promotions, or unrelated content create noise. Visualping documents page and element selection; changedetection.io documents CSS/XPath selectors and text filters.
  3. Match frequency to the signal. Check pricing or release pages more often only when the decision requires it. Less time-sensitive messaging can be checked less often. Frequent checks consume more service capacity and increase the volume of routine changes to review.
  4. Keep evidence with each event. Store the URL, check time, prior and current extracted text (or snapshots), and a concise diff. Keep an explicit result for failures so a broken fetch does not silently become a monitoring blind spot.
  5. Route to a review queue. Send alerts to a shared inbox, team channel, dashboard, or intelligence brief. Include enough context to verify the change without opening a separate system.
  6. Verify and label interpretation. An analyst checks the live page and saved evidence before updating a battlecard. Write “the pricing page now lists an annual plan” as an observation; label “this may target larger customers” as an interpretation.
  7. Review the system itself. Periodically check stale monitors, failed fetches, selectors that return no content, and pages whose structure changed. Remove watches that no longer answer a useful question.

4. A runnable Python monitor for a small watchlist

This example checks static or server-rendered HTML pages, optionally extracts a CSS selector, stores each successful text snapshot locally, and prints a unified diff when content changes. It also records fetch failures as errors instead of treating them as page changes. Use it as a starting point; it does not implement email or chat delivery.

Install dependencies

python -m venv .venv
. .venv/bin/activate
python -m pip install requests beautifulsoup4

Save the watchlist

Create watchlist.json. Set selector to a CSS selector for the section worth tracking, or use null for the whole page. Choose selectors based on the target page and review them when a page redesign changes its markup.

[
  {
    "name": "Example competitor pricing",
    "url": "https://example.com/pricing",
    "selector": "main"
  },
  {
    "name": "Example competitor release notes",
    "url": "https://example.org/releases",
    "selector": null
  }
]

Save the monitor

#!/usr/bin/env python3
"""Fetch a watchlist, save normalized text snapshots, and report diffs."""
import difflib
import hashlib
import json
import os
import re
import sys
from datetime import datetime, timezone
from pathlib import Path

import requests
from bs4 import BeautifulSoup

WATCHLIST = Path(os.environ.get("WATCHLIST", "watchlist.json"))
STATE_DIR = Path(os.environ.get("STATE_DIR", "monitor-state"))
TIMEOUT_SECONDS = 30
USER_AGENT = "CompetitorMonitor/1.0 (+contact your-team@example.com)"


def normalized_text(html, selector):
    soup = BeautifulSoup(html, "html.parser")
    for node in soup(["script", "style", "noscript", "svg"]):
        node.decompose()
    selected = soup.select_one(selector) if selector else soup.body or soup
    if selected is None:
        raise ValueError(f"CSS selector matched no element: {selector}")
    text = selected.get_text(" ", strip=True)
    return re.sub(r"\\s+", " ", text).strip()


def state_path(url):
    key = hashlib.sha256(url.encode("utf-8")).hexdigest()
    return STATE_DIR / f"{key}.txt"


def check(watch):
    name, url, selector = watch["name"], watch["url"], watch.get("selector")
    path = state_path(url)
    try:
        response = requests.get(
            url,
            headers={"User-Agent": USER_AGENT},
            timeout=TIMEOUT_SECONDS,
        )
        response.raise_for_status()
        current = normalized_text(response.text, selector)
        if not current:
            raise ValueError("extracted page text is empty")
    except (requests.RequestException, ValueError) as error:
        print(f"ERROR | {name} | {url} | {error}", file=sys.stderr)
        return False

    timestamp = datetime.now(timezone.utc).isoformat(timespec="seconds")
    if not path.exists():
        path.write_text(current + "\\n", encoding="utf-8")
        print(f"BASELINE | {timestamp} | {name} | {url}")
        return True

    previous = path.read_text(encoding="utf-8").rstrip("\\n")
    if current != previous:
        diff = difflib.unified_diff(
            previous.split(), current.split(),
            fromfile="previous", tofile="current", lineterm="",
        )
        print(f"CHANGE | {timestamp} | {name} | {url}")
        print("\\n".join(diff))
        path.write_text(current + "\\n", encoding="utf-8")
    else:
        print(f"UNCHANGED | {timestamp} | {name} | {url}")
    return True


def main():
    STATE_DIR.mkdir(parents=True, exist_ok=True)
    watches = json.loads(WATCHLIST.read_text(encoding="utf-8"))
    failed = 0
    for watch in watches:
        if not all(key in watch for key in ("name", "url")):
            print(f"ERROR | invalid watch entry: {watch!r}", file=sys.stderr)
            failed += 1
            continue
        if not check(watch):
            failed += 1
    return 1 if failed else 0


if __name__ == "__main__":
    raise SystemExit(main())

Run one check with python monitor.py. The first successful check for each URL creates a baseline. Later runs report a diff and replace that baseline. The sample prints the full diff to standard output; redirect output to a log or adapt the change branch to call an approved notification endpoint. Keep secrets out of the watchlist and source code if you add authenticated access.

Schedule checks

On a Unix-like host, a cron entry runs the monitor every six hours and appends output to a log:

0 */6 * * * cd /path/to/monitor && /path/to/monitor/.venv/bin/python monitor.py >> monitor.log 2>&1

Use a scheduler appropriate to your environment, and stagger large watchlists to avoid bursts. A production job should also alert on repeated execution failures and retain enough history to diagnose missed or noisy changes.

5. Browser-rendered pages and visual evidence

The Python example uses ordinary HTTP requests. It cannot run page JavaScript, interact with controls, or see content that only appears after a browser action. If a page relies on client-side rendering, use a monitoring system with a documented browser fetcher or a browser automation workflow that you operate. Check whether the page is publicly accessible and within the site’s permitted access rules before adding a browser, credentials, or a proxy.

For visual before-and-after evidence, capture the same URL with consistent viewport and capture settings, then retain each image with the URL, timestamp, and textual diff. A screenshot can help an analyst see layout or visual changes, but an image comparison alone does not establish why the page changed.

Or skip the browser setup for screenshot evidence: ScreenshotNeo is a website screenshot API, not a scheduled competitive-change monitor. Your scheduler can request a clean screenshot when you need an image of a page. The one-call request below uses the documented API pattern; see the ScreenshotNeo API documentation for options and response details.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

These examples request a screenshot of the example URL. Replace it with the page you are authorized to monitor, and keep the API key on the server rather than in browser code or a public repository. ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response includes X-Page-Verdict and X-Billed headers. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000, and every feature is on every plan.

Sign up for ScreenshotNeo’s free plan to capture up to 1,000 screenshots a month with no card.

6. Filtering changes and routing useful alerts

  • Ignore predictable churn: exclude navigation, footers, timestamps, rotating recommendations, or other regions unrelated to your question.
  • Use conditions carefully: text triggers can focus alerts on terms or values, but a condition that is too narrow can hide a significant change.
  • Preserve the raw signal: keep the extracted content or snapshot alongside any filtered alert so the analyst can inspect what the filter suppressed.
  • Record failures separately: log status codes, timeouts, empty extractions, selector misses, and browser errors as operational events.
  • Route with context: include competitor, page category, source URL, observation time, and a short diff. Send repeated fetch failures to the monitor owner.
  • Separate observed facts from analysis: automate collection and triage; require review before conclusions enter an executive brief or battlecard.

For programmatic watch management, Visualping documents an API for creating, updating, and deleting monitors and retrieving changes. Visualping API documentation. changedetection.io documents API management for watches, groups, and notifications, with the live API schema available in its installation. The appropriate notification channels and integration options depend on the service and plan.

7. Reliability, performance, cost, and access

Reliability

A scheduled check can fail because a site is unavailable, the response differs from expected content, a selector stops matching, or a page requires browser rendering. Keep the last known good snapshot, record failed checks separately, and retry transient network failures with a limit and delay. Do not replace a good baseline with an error page or empty response. If a page redesign invalidates a selector, treat it as a monitor maintenance task rather than a competitor signal.

Performance and scale

Keep the list limited to pages tied to decisions. Extracting a specific section reduces irrelevant text in a diff, though selectors require maintenance. Browser rendering and broad crawling generally demand more resources than a simple HTTP fetch; use them only for pages that need them. For a large watchlist, spread checks over time, set explicit timeouts, and monitor the queue and failure rate. The cited product documentation does not establish an independent accuracy or speed comparison.

Cost and retention

Estimate service cost from the number of monitored URLs, check frequency, browser or extraction needs, retained snapshots, and notification or API usage. For a self-hosted monitor, account for hosting, storage, backups, and maintenance time. Keep only the evidence needed for review and audit, following your organization’s retention rules. No topic-level independent statistic for accuracy, alert quality, time saved, or return on investment is established by the sources cited here.

Terms and data handling

Before monitoring, review the target site’s terms, robots.txt directives, access policies, and applicable law. changedetection.io explicitly places responsibility for compliant use on the person operating the software. Also consider what sensitive information an authenticated page might expose and whether it is appropriate to send that content to a third-party service.

8. Troubleshooting

Symptom Likely cause Fix
The first run reports a baseline, with no diff. No prior snapshot exists. This is expected. Keep the baseline and compare on a later successful check.
Every check reports a change. Dynamic content, rotating promotions, timestamps, or unstable whitespace is included. Compare a smaller section, remove known volatile content, normalize extracted text, and inspect the actual diff before suppressing it.
The monitor reports empty content or a missing selector. The selected section is absent, loaded with JavaScript, or changed in a redesign. Check the page and selector in a browser. Use a browser-capable fetcher if rendering is required; update the selector only after confirming the new target.
The request returns 403, 429, or a challenge page. The site denied or rate-limited the request, or requires an access path this fetcher does not provide. Reduce check frequency and review the site’s access rules. Do not treat a challenge page as a content update or attempt to evade an access control.
The script reports a timeout or connection error. Transient network or origin failure, slow page response, or an overly short timeout. Keep the last good snapshot, record the error, retry with bounded backoff, and investigate repeated failures before changing the baseline.
The alert arrives late or not at all. The scheduler stopped, the job exited nonzero, or notification delivery failed. Inspect scheduler and application logs, alert on repeated failures, and test notification routing independently.
Two URLs seem to share one saved snapshot. The storage key or watch identity is not unique in a modified implementation. Key snapshots by normalized URL or stable watch ID, and include the URL and timestamp in event records.
A ScreenshotNeo capture is not a useful visual record. The page may be blank, still loading, or showing a bot check rather than the content you intended to review. Review the response’s X-Page-Verdict and X-Billed headers, check the target URL, and confirm the page’s access behavior. Failed loads and bot checks are not billed under the stated ScreenshotNeo facts.

9. A weekly intelligence review checklist

  • Are all watched URLs still relevant and accessible under the applicable site rules?
  • Did a check fail, return empty content, or stop matching its selector?
  • Does every reported change have a timestamp, source URL, and inspectable evidence?
  • Have material changes been checked against the live page?
  • Are observations and interpretations clearly labeled separately?
  • Should the check schedule, filtering rules, or watchlist change based on current questions?
  • Have verified findings been summarized in a dated change log or battlecard?

A public-sector example in an OECD Global Forum on Competition submission describes Pakistan’s Competition Commission Market Intelligence Unit combining keyword-based Google Alerts with a Python scraper for stock-exchange announcements and metadata. It illustrates a useful design principle: combine broad discovery with structured monitoring of authoritative sources. It is a description of one implementation, not a controlled evaluation of impact. OECD Global Forum on Competition submission.

10. Frequently asked questions

How often should I check competitor pages?

Choose an interval based on how quickly the information can change and how quickly your team needs to respond. Start conservatively, then adjust based on verified changes and operational noise.

Should I monitor full pages or just prices?

Monitor the smallest region that answers the question when unrelated page changes create noise. Keep a whole-page or screenshot record too if broader presentation changes matter.

Can a website monitor prove a competitor’s strategy?

No. It can preserve evidence that a public page changed. The business meaning of that change is an interpretation that needs context and human review.

Can I use a screenshot API as the whole monitoring system?

A screenshot API returns page images. You still need a scheduler, storage, comparison logic, alert routing, and review process if you are building a monitoring workflow. ScreenshotNeo can provide screenshot evidence for that workflow.