How to Build a Competitor Tracking Tool
Build a focused competitor tracker for prices, availability, and website changes, with a practical collection pipeline, runnable Python starter, and operating checklist.
A useful competitor tracking tool starts with a bounded job: a short list of competitors, specific source pages, a small set of fields, a permitted checking cadence, and alerts tied to decisions. The basic pipeline is discover sources → fetch pages → extract fields → normalize observations → compare with history → notify people. Start with one source type, such as product prices, and add coverage only when the data is useful and maintainable.
This guide builds a small Python tracker for public product pages. It stores timestamped observations in SQLite, detects price or availability changes, and prints events you can route to email, a team channel, or a webhook. It is a foundation for a monitored workload, not a guarantee that arbitrary pages expose reliable or comparable data.
1. Define what the tracker must answer
Write down the decision the data supports before choosing a crawler or parser. For example: “Alert the pricing team when a named competitor’s listed monthly price changes by at least 5%.” A tool that collects pages without an action in mind tends to produce noisy alerts and maintenance work.
| Decision | Candidate sources | Fields to track | Useful alert |
|---|---|---|---|
| Reprice a product | Product or pricing page | Product match, amount, currency, billing period, availability | Price change beyond a chosen threshold |
| React to a plan change | Pricing page, public changelog | Plan name, included limits, feature wording, effective date | Plan or limit changed |
| Know when a product is available | Product page or marketplace listing | Availability state, region, variant | Out of stock or back in stock |
| Review public positioning | Changelog, product or company page | Headline, feature names, announcement date | A meaningful text change for human review |
Keep competitors separate from sources. One competitor can have several pages, each with its own source type, market or locale, extraction method, cadence, and enabled or paused state. A pricing page does not necessarily represent every product, region, or billing option.
- Begin with a handful of high-value sources and one or two fields.
- Record the canonical URL, market, expected update frequency, and why the source matters.
- Choose the decisions and thresholds that justify immediate alerts; send lower-priority changes in a digest.
- Prefer official APIs when they are available and appropriate for the data and terms.
- Set a review owner for every parser. A page change should have someone responsible for noticing and fixing extraction drift.
2. Design the collection pipeline
Source discovery
Maintain an explicit source registry instead of crawling an entire domain. Store the competitor, source URL, source type, locale or market, extraction strategy, cadence, and whether collection is enabled. Start with stable, public pages. If a site provides an appropriate official API, use it instead of parsing HTML.
Fetching and crawler rules
Use a descriptive user agent and contact information where practical, set request timeouts, limit concurrency per host, and back off when a source errors. Check the target’s robots.txt and applicable site terms before collection. RFC 9309 defines the Robots Exclusion Protocol for crawlers; it says its rules are not access authorization. The standard directs crawlers to follow parseable rules after successful retrieval and to assume complete disallow when robots.txt is unreachable because of server or network errors. These crawler rules do not decide whether a collection is legally permitted. Review relevant terms and obligations independently. RFC 9309.
The starter code below checks robots.txt before fetching a page and stops if the policy cannot be retrieved. For a production scheduler, cache robots policies according to the standard, use per-host rate limits, and pause sources that disallow access or repeatedly return errors.
Extraction and normalization
Convert the page into typed observations, not strings sprinkled through alert code. A useful price record includes competitor and source identifiers, product identifier, observed name, numeric amount, currency, availability, timestamp, parser version, and an evidence reference. Normalize locale-specific decimal separators, currencies, product identifiers, and availability values before comparison. Preserve raw responses or evidence snapshots only when your retention, rights, and storage policies permit it.
Price extraction is particularly error-prone: a page may show a sale price, a struck-through list price, a per-seat amount, an annual price billed monthly, tax-exclusive pricing, or a region-specific offer. Store enough context to review a match. Do not alert on a number until the parser has identified what the number represents.
History, comparison, and notification
Append observations to history rather than overwriting the current value. Compare the latest accepted observation with the previous accepted one, then emit an event containing the old and new values, observation time, source link, and evidence. Deduplicate repeated events. Track fetch or parser failures separately from competitor changes so a broken selector cannot look like a price drop.
Begin with a console or email alert. Include the competitor, changed field, old and new values, observation time, and source link. Send immediate notifications only for changes that require action; group other changes into a digest. A webhook can connect the event stream to an internal dashboard or workflow.
3. Run a minimal Python tracker
This starter tracks price and availability from explicitly configured product pages. It uses JSON-LD Product data when present and falls back to configurable CSS selectors. It stores observations in SQLite, fetches only sources allowed by the retrieved robots policy, and reports changes to the console. The example uses one request at a time; add a scheduler and host-level concurrency controls before expanding the source list.
Install dependencies
python -m venv .venv
. .venv/bin/activate
python -m pip install requests beautifulsoup4
Save the following as tracker.py. Replace the sample URL and selectors with pages you are permitted to monitor. Selectors are source-specific; the sample selectors will not fit every site.
import json
import os
import re
import sqlite3
import sys
import time
from datetime import datetime, timezone
from urllib.parse import urljoin, urlparse
from urllib.robotparser import RobotFileParser
import requests
from bs4 import BeautifulSoup
USER_AGENT = "CompetitorTracker/1.0 (+mailto:ops@example.com)"
DB_PATH = os.environ.get("TRACKER_DB", "tracker.sqlite3")
TIMEOUT_SECONDS = 20
# Keep one explicit entry per monitored source. Selectors are examples.
SOURCES = [
{
"competitor": "Example Competitor",
"product_id": "example-plan",
"url": "https://example.com/pricing",
"market": "US",
"currency": "USD",
"price_selector": "[data-testid='price']",
"availability_selector": "[data-testid='availability']",
},
]
session = requests.Session()
session.headers.update({"User-Agent": USER_AGENT, "Accept": "text/html,application/xhtml+xml"})
def utc_now():
return datetime.now(timezone.utc).isoformat(timespec="seconds")
def robots_allows(url):
parts = urlparse(url)
robots_url = f"{parts.scheme}://{parts.netloc}/robots.txt"
response = session.get(robots_url, timeout=TIMEOUT_SECONDS)
# RFC 9309 says unreachable robots.txt means assume complete disallow.
if response.status_code >= 500 or response.status_code in (408, 429):
raise RuntimeError(f"Cannot retrieve robots policy ({response.status_code}); source paused")
if response.status_code == 404:
return True
if not response.ok:
raise RuntimeError(f"Robots policy returned HTTP {response.status_code}; source paused")
parser = RobotFileParser()
parser.set_url(robots_url)
parser.parse(response.text.splitlines())
return parser.can_fetch(USER_AGENT, url)
def jsonld_product(soup):
"""Return the first JSON-LD Product node found, including in @graph."""
def walk(value):
if isinstance(value, dict):
kind = value.get("@type", [])
kinds = [kind] if isinstance(kind, str) else kind
if "Product" in kinds:
yield value
for child in value.values():
yield from walk(child)
elif isinstance(value, list):
for child in value:
yield from walk(child)
for tag in soup.select("script[type='application/ld+json']"):
try:
data = json.loads(tag.string or tag.get_text())
except (json.JSONDecodeError, TypeError):
continue
product = next(walk(data), None)
if product:
return product
return None
def clean_text(node):
return " ".join(node.get_text(" ", strip=True).split()) if node else None
def extract(source, html):
soup = BeautifulSoup(html, "html.parser")
product = jsonld_product(soup)
offers = (product or {}).get("offers", {})
if isinstance(offers, list):
offers = offers[0] if offers else {}
name = (product or {}).get("name") or clean_text(soup.select_one("h1"))
raw_price = offers.get("price") if isinstance(offers, dict) else None
currency = (offers.get("priceCurrency") if isinstance(offers, dict) else None) or source["currency"]
availability = offers.get("availability") if isinstance(offers, dict) else None
if raw_price is None and source.get("price_selector"):
raw_price = clean_text(soup.select_one(source["price_selector"]))
if availability is None and source.get("availability_selector"):
availability = clean_text(soup.select_one(source["availability_selector"]))
# This deliberately accepts only a simple decimal representation.
# Localized values need a source-specific parser; do not guess separators.
price = None
if raw_price is not None:
match = re.search(r"\d+(?:\.\d+)?", str(raw_price).replace(",", ""))
if match:
price = float(match.group(0))
if price is None and availability is None:
raise ValueError("No price or availability extracted; check markup and selectors")
return {
"name": str(name) if name else None,
"price": price,
"currency": str(currency) if currency else None,
"availability": str(availability) if availability else None,
}
def initialize(db):
db.execute("""CREATE TABLE IF NOT EXISTS observations (
id INTEGER PRIMARY KEY,
competitor TEXT NOT NULL,
product_id TEXT NOT NULL,
source_url TEXT NOT NULL,
market TEXT NOT NULL,
observed_at TEXT NOT NULL,
name TEXT,
price REAL,
currency TEXT,
availability TEXT,
parser_version TEXT NOT NULL
)""")
db.commit()
def latest(db, source):
return db.execute(
"SELECT name, price, currency, availability FROM observations "
"WHERE competitor=? AND product_id=? ORDER BY id DESC LIMIT 1",
(source["competitor"], source["product_id"]),
).fetchone()
def collect(db, source):
url = source["url"]
if not robots_allows(url):
print(f"PAUSED robots disallows {url}", file=sys.stderr)
return
response = session.get(url, timeout=TIMEOUT_SECONDS)
response.raise_for_status()
current = extract(source, response.text)
previous = latest(db, source)
fields = ("name", "price", "currency", "availability")
old = dict(zip(fields, previous)) if previous else None
db.execute(
"INSERT INTO observations "
"(competitor, product_id, source_url, market, observed_at, name, price, currency, availability, parser_version) "
"VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?)",
(source["competitor"], source["product_id"], url, source["market"], utc_now(),
current["name"], current["price"], current["currency"], current["availability"], "1"),
)
db.commit()
if old is None:
print(f"INITIAL {source['competitor']} {source['product_id']}: {current}")
return
changes = {key: (old.get(key), current.get(key)) for key in fields if old.get(key) != current.get(key)}
if changes:
print(json.dumps({
"event": "competitor_change",
"competitor": source["competitor"],
"product_id": source["product_id"],
"market": source["market"],
"source_url": url,
"observed_at": utc_now(),
"changes": changes,
}, ensure_ascii=False))
def main():
with sqlite3.connect(DB_PATH) as db:
initialize(db)
for index, source in enumerate(SOURCES):
try:
collect(db, source)
except (requests.RequestException, RuntimeError, ValueError) as exc:
print(f"ERROR {source['url']}: {exc}", file=sys.stderr)
# A small gap is not a substitute for a proper per-host scheduler.
if index + 1 < len(SOURCES):
time.sleep(2)
if __name__ == "__main__":
main()
Run it with python tracker.py. Each run adds a timestamped row to tracker.sqlite3. The first run establishes a baseline; later runs print JSON change events. Keep secrets out of source files if you extend the collector to authenticated APIs.
What to change before relying on it
- Replace the placeholder source with a page you are allowed to fetch. The user agent includes a contact address; change it to a real monitored mailbox.
- Use a source-specific parser for price formats, billing period, discounts, shipping, taxes, variants, and market. The sample numeric parser is intentionally simple and is not safe for every locale.
- Add a stable product-match rule. Do not assume the first Product JSON-LD node or first offer always represents the product and region you care about.
- Store parser version and extraction status with each run. Keep failures out of the observation history used to detect business changes.
- Replace console output with an alert sink only after deduplication, thresholds, and a human review path are defined.
4. Add useful change rules
Exact comparisons are appropriate for simple availability states, but prices often need context. A rule might alert when the same product, market, currency, and billing period changes by more than a percentage or absolute amount. Compare only like with like: monthly against monthly, matching variants, and consistent tax treatment.
def price_changed(old_price, new_price, minimum_percent=5, minimum_amount=1):
if old_price is None or new_price is None:
return old_price != new_price
delta = abs(new_price - old_price)
percent = delta / old_price * 100 if old_price else (100 if delta else 0)
return delta >= minimum_amount and percent >= minimum_percent
For richer tracking, emit separate events for fetch_failed, parse_failed, match_uncertain, and competitor_changed. Keep the raw evidence link or permitted snapshot alongside the event so an analyst can verify what changed. Promotions, shipping costs, region mismatches, and page experiments commonly require manual review.
5. Schedule collection and keep it reliable
Run the script from cron, a task queue, or a scheduled job runner. The appropriate cadence depends on how quickly the source changes and how often you can responsibly fetch it. Use per-source intervals rather than one global frequency. Add a bounded timeout, limited retries with exponential backoff for transient failures, per-host concurrency limits, and a circuit breaker that pauses repeated failures. Do not retry access denials or robots disallow rules as if they were temporary outages.
Conditional requests can save bandwidth when a source supports HTTP caching validators such as ETag or Last-Modified; Google documents these mechanisms for its crawler, but support and behavior vary by site. Do not assume every target supports them. Google crawler documentation.
Monitor the collector as a product of its own. Track fetch success, extraction completeness, source age, parser failure rate, duplicate alerts, and cost per useful observation. A page that returns HTTP 200 can still have changed its markup or served a challenge page. Validate extracted fields and flag implausible jumps for review rather than silently accepting them.
6. Troubleshoot common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| Robots check fails or source is paused | robots.txt is disallowed, unavailable, or returns an error. | Respect disallow rules. For unreachable policy, fail closed; investigate the source configuration and retry only after a reasonable interval. |
| HTTP 403, 429, or challenge page | The site denied or rate-limited the request, or served bot protection. | Pause or reduce collection, review terms and access rules, and use an official API if appropriate. Do not try to evade access controls. |
| Timeouts or intermittent 5xx responses | Network instability, source overload, or slow page response. | Use bounded retries with backoff for transient errors, lower host concurrency, and alert on stale data. Avoid retry storms. |
| No price extracted | The selector changed, price is rendered by client-side JavaScript, JSON-LD is absent, or the chosen offer is nested differently. | Inspect permitted page markup, update the parser, or use an appropriate official API or browser rendering method. Record a parser failure rather than a price change. |
| Price is off by a factor or has the wrong decimals | Locale separators, currency, tax, period, or a sale/list price was misinterpreted. | Write a locale-aware parser and keep currency and billing context. Add validation bounds and manual review for unusual changes. |
| Every run emits the same alert | The code compares against a wrong baseline, a field is unstable, or the event has no deduplication key. | Compare with the last accepted observation and deduplicate on source, product, field, old value, and new value. Separate transient fetch failures. |
| False availability changes | Text labels, variants, or regional inventory states differ. | Map source-specific labels into a controlled state set and retain the original value for review. |
| History has gaps | The scheduler failed, the source was paused, or writes failed. | Track last successful observation per source and alert when it exceeds the configured freshness window. |
7. Build or buy: compare the real workload
A custom tracker gives control over schema, source logic, storage, and workflow, but fetchers and parsers need ongoing care as sites change. Managed services can reduce that operating work, though their coverage, extraction output, history, retention, alert controls, integration, and terms need validation against your actual sources. Compare services with the same representative workload rather than treating advertised features as proof of accuracy.
| Option | Advertised focus in the reviewed material | Questions to verify |
|---|---|---|
| TrackBase | Scheduled price and page monitoring through an API | Are your sources and fields covered? What history and evidence are retained? What does the API return when extraction fails? |
| Ahrefs Firehose | Web-index streams and URL watches | Does its coverage fit page-level price and availability monitoring, or is your need web mentions and URL changes? |
| Scrapewise | Data APIs and managed price monitoring | Can it match the products, markets, cadence, output schema, and retention you require? |
| Custom system | Your own source list, schema, rules, and integrations | Who maintains parsers, fetch controls, evidence, alerts, and operational monitoring? |
Vendor descriptions establish advertised capabilities, not independently verified performance. Measure extraction quality and useful-alert rate on representative sources before committing. Recheck volatile plans, features, and terms directly with each provider.
8. Capture page evidence without operating a browser
For visual before-and-after evidence or pages that need browser rendering, a screenshot capture can complement structured extraction. ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. Its API can return an image or PDF from a URL, and its MCP tools let AI agents take screenshots, inspect page information, and capture PDFs.
Or skip the browser setup
One GET request captures a page. See the ScreenshotNeo API documentation for options and configuration.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
Before the shot, ScreenshotNeo accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server supports Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Use screenshots as review evidence, not as a substitute for validating structured price extraction. Create a free ScreenshotNeo account.
9. Performance, reliability, and cost
The main operating costs are engineering time maintaining source-specific parsers, fetch and storage infrastructure, alert review, and any managed data services. Keep the source set bounded, schedule per source, avoid needless duplicate fetches, and retain only evidence you have a reason and permission to keep. Estimate cost per useful observation, not just requests made: a cheap fetch that yields stale or misparsed data is not useful.
Reliability depends on source behavior and parser fit. Do not claim accuracy from a successful HTTP status or a vendor feature list. Track freshness and parser health, retain enough provenance to explain each event, and route uncertain matches to human review. Test new extraction rules against saved, permitted examples before turning on alerts.
10. Launch checklist
- Each source has an owner, business purpose, canonical URL, market, source type, cadence, and enabled state.
- Robots policy and applicable site terms have been reviewed; the collector identifies itself and honors denials.
- Timeouts, bounded retries, backoff, and per-host concurrency controls are configured.
- Observations include timestamps, normalized values, parser version, source, and evidence reference where permitted.
- Extraction failures are recorded separately from competitor changes.
- Price comparisons account for currency, locale, tax, billing period, variant, and promotion context.
- Alerts are deduplicated, thresholded, actionable, and include a source link.
- Stale sources and parser failures have an owner and a review path.
- A representative workload has been used to compare build and buy options.
FAQ
How many competitor pages should I monitor first?
There is no universal number. Start with the few sources that can change a real decision, then expand after you know the parser and alert maintenance cost.
Can I monitor any public page?
Public visibility alone does not answer whether collection is permitted. Follow crawler rules, review terms and applicable obligations, and do not evade access controls.
Should I store full page snapshots?
Only when they are useful and your retention, rights, and data policies allow it. A timestamped excerpt or source URL may be enough for some workflows.
Is a screenshot enough to track a price?
No. A screenshot helps a person review visual evidence, while a dependable price history requires structured extraction, normalization, and validation.


