How to Build a Competitor Price Monitoring Tool
Build a reliable competitor price monitor with source discovery, browser capture, product matching, historical data, alerts, and compliance controls.
Direct answer: Build the monitor as a pipeline that discovers permitted product URLs, captures each offer on a schedule, preserves raw provenance, parses and normalizes price data, matches equivalent products, stores timestamped observations, and alerts only on validated changes. Keep discovery separate from repeat monitoring, and treat product matching, currency, market, availability, and promotion context as part of the price—not optional metadata.
1. Define the monitoring scope
Write down these decisions before collecting anything:
- Products: your catalog, competitor categories, brands, or a manually curated list.
- Sources: official APIs, licensed feeds, retailer pages, or marketplace listings that your organization is permitted to access.
- Markets: country, region, language, currency, tax treatment, and delivery location.
- Offer context: base price, sale price, coupon, seller, fulfillment, stock status, pack size, and variant.
- Cadence: hourly, daily, or another schedule based on how quickly the business decision must react.
- Retention: how long to keep raw pages, screenshots, extracted records, and audit logs.
“Real time” is not a property you can assume. Freshness is limited by your schedule, page changes, regionalization, blocking, and source coverage.
2. Separate discovery from monitoring
Discovery finds candidate listings; monitoring repeatedly refreshes known URLs. Store the source URL or an explicit discovery method for every candidate. Possible discovery inputs include an official catalog API, a licensed feed, a merchant-provided export, or a human-reviewed search process. Do not silently turn search results into tracked products without recording how each URL was found.
3. Choose an allowed collection method
Prefer an official API or licensed feed when one exists. Use a basic HTTP client when the relevant price is present in the response. Use a browser only when the target legitimately requires client-side rendering and the access rules allow it. A browser adds latency, resource use, and more failure modes, so it should solve a real rendering requirement.
Apply a versioned policy before each outbound request: allowed host, request rate, fields collected, authentication method, and retention period. Do not bypass authentication or collect personal information merely because it appears beside a public listing. Robots.txt is a crawler protocol, not legal authorization. RFC 9309 states that its rules are “not a form of access authorization”; it also specifies that crawlers follow parseable rules and generally avoid using a cached robots file for more than 24 hours unless it is unreachable.
4. Design the observation record
Append observations instead of overwriting the latest value. A practical record contains:
| Field | Purpose |
|---|---|
| internal_product_id | Your stable catalog identity |
| competitor | Source organization or merchant |
| source_url | Exact requested listing URL |
| observed_title | Title as returned by the source |
| external_sku, GTIN, model | Identifiers when present |
| variant | Size, color, capacity, pack count, or configuration |
| amount, currency | Normalized numeric amount plus currency code |
| original_price_text | Original value retained for audit |
| availability | In stock, out of stock, preorder, or unknown |
| promotion | Sale, coupon, membership, or other offer context |
| seller | Seller or fulfillment context where relevant and permitted |
| market_context | Country, region, language, timezone, and delivery assumptions |
| observed_at | Timestamp of the observation |
| parse_status | Success, missing, malformed, or ambiguous |
| match_confidence | Confidence that the listing is the intended catalog item |
| parser_version, policy_version | Reproducibility when code or policy changes |
| raw_artifact_uri | Pointer to a retained HTML response, screenshot, or equivalent evidence |
5. Build product matching as its own subsystem
Similar titles do not prove equivalence. Match in this order:
- Use a stable identifier such as GTIN, model number, or the source SKU when available.
- Corroborate with brand, model, variant attributes, and pack size.
- Normalize units and quantities before comparing.
- Keep uncertain matches pending human review instead of forcing a match.
A wrong size, bundle, or variant creates a fictitious price gap. Store the matching evidence and confidence with every observation.
6. A minimal relational schema
CREATE TABLE products (
id TEXT PRIMARY KEY,
brand TEXT,
model TEXT,
variant_json TEXT,
active INTEGER NOT NULL DEFAULT 1
);
CREATE TABLE sources (
id INTEGER PRIMARY KEY,
competitor TEXT NOT NULL,
url TEXT NOT NULL,
market_json TEXT,
policy_version TEXT NOT NULL,
active INTEGER NOT NULL DEFAULT 1
);
CREATE TABLE observations (
id INTEGER PRIMARY KEY,
product_id TEXT NOT NULL,
source_id INTEGER NOT NULL,
observed_title TEXT,
external_sku TEXT,
amount_minor INTEGER,
currency TEXT,
original_price_text TEXT,
availability TEXT,
promotion_json TEXT,
seller TEXT,
observed_at TEXT NOT NULL,
parse_status TEXT NOT NULL,
match_confidence REAL,
parser_version TEXT NOT NULL,
raw_artifact_uri TEXT,
FOREIGN KEY(product_id) REFERENCES products(id),
FOREIGN KEY(source_id) REFERENCES sources(id)
);
CREATE INDEX observations_product_time
ON observations(product_id, observed_at);
Store money as an integer in the smallest currency unit plus an ISO currency code. Never compare amounts from different currencies until you have an explicit conversion policy and timestamp for the conversion.
7. Browser collection example with Python
The following example uses Playwright to load a page, wait for a price selector, and return a structured observation. Replace selectors for each permitted source. It deliberately fails closed when a price is missing or ambiguous.
from datetime import datetime, timezone
from decimal import Decimal
import json
from playwright.sync_api import sync_playwright
URL = "https://example.com/product"
PRICE_SELECTOR = "[data-price]"
TITLE_SELECTOR = "h1"
AVAILABILITY_SELECTOR = "[data-availability]"
def parse_amount(text: str) -> Decimal:
cleaned = "".join(ch for ch in text if ch.isdigit() or ch in ".,")
if not cleaned:
raise ValueError("price text contains no number")
# Adapt this parser to the source's locale; do not guess silently.
if cleaned.count(",") == 1 and cleaned.count(".") == 0:
cleaned = cleaned.replace(",", ".")
else:
cleaned = cleaned.replace(",", "")
return Decimal(cleaned)
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
page.goto(URL, wait_until="domcontentloaded", timeout=45_000)
page.wait_for_selector(PRICE_SELECTOR, timeout=15_000)
price_texts = page.locator(PRICE_SELECTOR).all_inner_texts()
if len(price_texts) != 1:
raise RuntimeError(f"expected one price, found {len(price_texts)}")
title = page.locator(TITLE_SELECTOR).first.inner_text()
availability = page.locator(AVAILABILITY_SELECTOR).first.inner_text() if page.locator(AVAILABILITY_SELECTOR).count() else None
amount = parse_amount(price_texts[0])
observation = {
"source_url": page.url,
"observed_title": title.strip(),
"amount": str(amount),
"currency": "USD", # Set from source metadata or a configured market policy.
"availability": availability.strip() if availability else "unknown",
"observed_at": datetime.now(timezone.utc).isoformat(),
"parse_status": "success",
"parser_version": "product-v1",
}
print(json.dumps(observation, indent=2))
browser.close()
Install the runtime with pip install playwright followed by playwright install chromium. Keep source-specific selectors and locale rules in version-controlled configuration rather than scattering them through the worker.
8. HTTP and API collection
If a permitted source exposes the price in its response, an HTTP client is cheaper and more reliable than a browser. Use connection timeouts, bounded retries, a descriptive user agent, and rate limits. Save the response status and parser result even when extraction fails.
import requests
response = requests.get(
"https://example.com/api/products/123",
headers={"Accept": "application/json", "User-Agent": "price-monitor/1.0"},
timeout=(5, 30),
)
response.raise_for_status()
data = response.json()
price = data["offer"]["price"]
currency = data["offer"]["priceCurrency"]
9. Normalize, validate, and reject bad records
- Parse the numeric amount and preserve the original text.
- Require a known currency and market context.
- Normalize units, decimal separators, and pack quantities.
- Classify availability separately from price.
- Distinguish base price from sale, coupon, and membership price.
- Reject missing, negative, malformed, or multiply matched prices.
- Flag implausible jumps for review; do not automatically overwrite a trusted value.
Validation should produce a visible status such as success, missing_price, ambiguous_price, blocked, or stale_source. Silent parser failure is one of the most dangerous failure modes because it can create confident-looking but false price decisions.
10. Schedule collection and preserve history
A scheduler creates jobs for active source records. Workers fetch and parse them, then append observations. At larger scale, separate ingestion from transformation with a queue or broker and worker pools. Validate before analytical storage, and expose matched pairs, deltas, and trends through an API or dashboard.
For each run, record:
- job ID and attempt number;
- requested URL and final URL;
- start and finish timestamps;
- HTTP or browser status;
- parser and policy versions;
- raw artifact location or an explicit reason it was not retained;
- failure class and retry decision.
11. Compute comparable price changes
Compare only observations with the same matched product and variant, currency and market, offer type, availability meaning, seller context, and a known observation time. Useful derived values include:
absolute_change = current_amount - previous_amount
percent_change = (absolute_change / previous_amount) * 100
Handle a zero or missing previous amount explicitly. A stockout is not a price decrease. A coupon ending is not necessarily a base-price change. Keep both the raw observations and the derived event so analysts can investigate the decision.
12. Alert design
Alert from validated observations rather than raw text differences. A useful alert includes the product match, old and new amounts, currencies, market, availability, promotion context, timestamps, source URL, parser status, and evidence link. Suppress duplicate alerts for the same observation pair and mark uncertain matches for review instead of paging a team.
Operational signals to monitor include last successful observation, age of the latest value by source, parse-validation failure rate, abrupt distribution changes, unresolved product matches, and incomplete collection.
13. Or skip the browser setup
ScreenshotNeo provides a website screenshot API and MCP server when visual evidence is useful in a price-monitoring workflow. One GET request returns a PNG, JPEG, WebP, or PDF. Before capture it accepts consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing state.
See the ScreenshotNeo API documentation for all options. The basic calls are:
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const bytes = await res.arrayBuffer();
For monitoring, relevant options include full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets or custom viewport, retina scale, custom CSS and JavaScript, clicks, selector or network-idle waits, blocked ads and trackers, custom headers, cookies, user agent and Authorization, timezone and geolocation, transparent backgrounds, resizing, selectable cache TTL, signed links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, and the usage API. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
ScreenshotNeo is useful when a screenshot is your provenance artifact or when browser setup and cleanup would otherwise be maintenance work. It has 1,000 free shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
14. Performance, reliability, and cost
Performance
- Use HTTP extraction where permitted and reserve browsers for rendered pages.
- Bound page waits and avoid loading unnecessary resource types when policy allows.
- Reuse browser processes carefully, but isolate jobs when pages leak state.
- Use queues and worker pools for many sources; enforce per-source concurrency.
- Cache only when the chosen TTL still meets the freshness requirement.
Reliability
- Retry transient network failures with exponential backoff and a maximum attempt count.
- Do not retry policy denials, authentication failures, or deterministic parser errors indefinitely.
- Keep raw evidence or an equivalent replayable record according to your retention policy.
- Alert on stale data and parser drift, not only on price changes.
Cost
Your main cost drivers are request volume, browser execution time, storage, retries, and engineering maintenance. A vendor’s SKU threshold is not a universal cutoff. Build is more plausible for a small, stable URL set with maintenance capacity; managed extraction or licensed feeds become more attractive as catalogs, markets, and retailer changes grow.
15. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| No price found | Selector drift, delayed rendering, or a different variant | Inspect the current response, wait for a stable selector, version the parser, and fail visibly. |
| Two or more prices found | Sale, list, installment, or hidden mobile markup | Classify price types and require an explicit selection rule. |
| Prices differ by region | Locale, currency, cookie, IP, or delivery context | Set market context explicitly and store it with every observation. |
| Sudden 0 or extreme value | Parser accepted malformed text or a placeholder | Validate ranges and compare against prior observations before alerting. |
| Browser timeout | Slow page, blocked resource, or overly short wait | Set bounded navigation and selector timeouts, capture the failure class, and review permitted resource blocking. |
| Repeated access denial | Collection is not permitted or the source is blocking requests | Stop retries, review terms and policy, and prefer an official API or licensed feed. |
| False competitor match | Similar title but different size, bundle, or model | Require identifiers and corroborating attributes; route uncertain matches to review. |
| No alert after a change | Observation failed validation or was marked stale | Inspect parse status, freshness, match confidence, and alert deduplication logs. |
| Too many duplicate alerts | Every run emits the same delta | Deduplicate by product, source, old observation, and new observation. |
16. Implementation checklist
- Define products, sources, markets, cadence, and retention.
- Record a permitted discovery method for every URL.
- Prefer official APIs or licensed feeds.
- Version target policy, parser, and matching rules.
- Append observations with provenance instead of overwriting values.
- Normalize money, units, currency, availability, and promotions.
- Match variants with identifiers and review uncertain cases.
- Validate before calculating deltas or sending alerts.
- Monitor freshness, parser failures, stale sources, and unresolved matches.
- Test the system against page changes and regional variants.
17. FAQ
How do I discover new competitor products to track?
Use an official catalog API, licensed feed, merchant export, or a documented human-reviewed search process. Store the discovery method and review the candidate before adding it to repeat monitoring.
Should I save every page?
Retain raw pages, screenshots, or an equivalent replayable artifact when auditability requires them, but make retention a documented policy with a defined duration and access controls.
Is robots.txt permission to collect prices?
No. Follow the parseable crawler rules, then separately evaluate terms, contracts, authentication, privacy, and applicable law.
When should I buy instead of build?
Compare source coverage, market support, match quality, cadence, reliability, provenance, permitted access methods, integration effort, support, and maintenance burden. Scale alone does not decide the question.
Can a screenshot prove a price?
It can preserve visual evidence of what was rendered at a time and market context, but it does not replace structured parsing, product matching, currency handling, or permission review.


