ScreenshotNeo

BlogEngineering

How to Track Prices and Product Catalogs at Scale

Build a reliable pipeline for tracking prices, availability, offers, and catalog changes across thousands of products and sources.

By the ScreenshotNeo team30 September 20269 min read

How to Track Prices and Product Catalogs at Scale

To track prices and product catalogs at scale, build a source-aware pipeline with five layers: product identity, compliant collection, normalization, immutable history, and change delivery. Every observation should retain its source, marketplace, currency, seller, condition, timestamp, and raw payload.

This design prevents the failures that make price monitoring unreliable: joining two different products, treating a currency conversion as a price change, missing seller or availability changes, and overwhelming an API or retailer. It also lets you combine scheduled snapshots with push notifications when a source supports both.

1. Define what you are tracking

A “price” is rarely one number. Decide which commercial facts matter before choosing an API or crawler.

Signal Examples Why it matters
Identity ASIN, SKU, GTIN, MPN, retailer ID, canonical URL Prevents false joins and duplicate products
Offer Seller, condition, fulfillment method, buy box or featured offer Explains why the visible price changed
Price Item price, shipping, tax, discount, landed price Separates displayed and comparable totals
Availability In stock, preorder, backorder, unavailable Price without availability can mislead
Catalog Title, images, attributes, variation, category, sales rank Detects merchandising and product-data changes
Context Marketplace, locale, currency, timestamp Makes histories comparable and auditable

Store both the observed components and your derived values. For example, keep item price, shipping, tax treatment, and discount separately, then calculate a normalized landed price. Never overwrite the original observation.

2. Choose an authorized collection method

First-party APIs

Use an official API whenever the source provides one and your application is authorized. Amazon describes its Product Pricing API as a way to retrieve pricing and offer information for products in the Amazon catalog. Its repricing guidance also covers monitoring competitor prices and related signals. The API documents getFeaturedOfferExpectedPriceBatch for up to 40 SKUs, competitive summaries, and notifications such as PRICING_HEALTH and ANY_OFFER_CHANGED. See the Amazon Product Pricing API reference.

A source-aware pipeline preserves identity, normalized offers, history, and change events.
A source-aware pipeline preserves identity, normalized offers, history, and change events.

Amazon’s catalog and listing APIs cover more than price. Catalog Items can return identifiers, attributes, images, product types, ranks, summaries, variations, and vendor details. Listings APIs expose listing state, offers, issues, and fulfillment availability. Notifications can report product-type, listing-status, listing-issue, and definition changes. Amazon distinguishes individual, batch, and bulk workflows; a batch operation can contain up to 20 similar requests. Plan for documented usage limits and throttling.

Historical data and tracking providers

For Amazon marketplace history, Keepa’s API provides price histories, product data, offers, deals, best-seller lists, seller information, and tracking notifications. Product search returns up to 10 results per result page. Keepa uses a token bucket: plans generate tokens continuously, and unused tokens expire after 60 minutes. Read the Keepa API documentation and token plans documentation before designing request volume.

Compliant crawlers

For sites without an authorized API, crawl only pages and data you are permitted to access. Check terms of use, robots rules, authentication requirements, contracts, and applicable law for every domain. A production crawler needs a per-domain queue, bounded concurrency, caching, retries with exponential backoff, parser versioning, and raw-response retention. AWS Prescriptive Guidance includes downloading and processing robots.txt as part of a scalable crawling system.

3. Use a source-aware data model

Do not use one global product table keyed only by title or URL. Separate a canonical product from source-specific listings and observations.

products
- product_id
- gtin, mpn, brand
- canonical_title
- created_at, updated_at

source_listings
- listing_id
- product_id
- source
- marketplace
- source_product_id
- canonical_url
- seller_id
- condition

observations
- observation_id
- listing_id
- observed_at
- currency
- item_price
- shipping_price
- tax_amount
- landed_price
- availability
- title
- image_hash
- variation_data
- raw_payload_uri
- parser_version

change_events
- event_id
- listing_id
- event_type
- old_value
- new_value
- detected_at

Keep source and marketplace on every row. A product can have different sellers, conditions, currencies, taxes, and availability in different markets. Preserve the raw payload in object storage or a document column so that a parser fix can replay historical data.

4. Collect with snapshots and notifications

Snapshots answer “what is true now?” Notifications answer “what changed?” Use both when a source supports both.

  1. Backfill: import your product identifiers and fetch an initial full record.
  2. Subscribe: enable source notifications for supported offer, pricing, listing, or catalog events.
  3. Schedule: poll high-value or fast-moving products more often than stable long-tail items.
  4. Reconcile: run periodic snapshots even when notifications are enabled; events can be delayed, duplicated, or missed.
  5. Record freshness: store the last successful observation and the next scheduled attempt.

Use a scheduler that creates jobs by source and domain. A queue worker should claim a job with a lease, apply the source’s rate limit, and extend or release the lease on retry. Keep separate budgets for backfills, routine refreshes, and urgent checks so a large import cannot starve monitoring.

5. Normalize prices and catalog fields

Normalization is where most misleading comparisons are created. Define explicit rules and keep the original values alongside normalized ones.

  • Currency: store the observed currency and conversion rate timestamp. Do not report a converted value as a source price.
  • Tax and shipping: record whether each is included, excluded, estimated, or unavailable. Calculate landed price only when the rule is consistent.
  • Units: normalize quantity, weight, volume, and pack size. A $10 two-pack is not equivalent to a $7 single item.
  • Condition: separate new, used, refurbished, rental, and collectible offers.
  • Seller: keep seller identity and fulfillment method. The cheapest offer may not be the featured offer.
  • Availability: use a controlled vocabulary and retain the source status.
  • Titles and attributes: normalize whitespace and obvious formatting while retaining the original text for audit.

6. Detect real changes without false alerts

Compare observations by stable listing identity, not by title. A practical detector emits separate event types for price, stock, offer, seller, condition, title, image, attribute, variation, category, and listing-state changes.

def classify_change(previous, current, fx_tolerance=0.01):
    events = []
    if previous.currency == current.currency:
        if abs(previous.landed_price - current.landed_price) > fx_tolerance:
            events.append("price")
    elif previous.landed_price != current.landed_price:
        events.append("currency_or_price")

    if previous.availability != current.availability:
        events.append("availability")
    if previous.seller_id != current.seller_id:
        events.append("seller")
    if previous.condition != current.condition:
        events.append("condition")
    if previous.title_hash != current.title_hash:
        events.append("title")
    if previous.image_hash != current.image_hash:
        events.append("image")
    return events

Apply tolerances for rounding and foreign-exchange movement. Debounce alerts: if three observations arrive within a short window, send one event containing the first and last values. Keep every raw observation even when you suppress a notification.

7. Handle duplicates and product matching

Prefer deterministic identifiers in this order: source product ID, GTIN, MPN plus brand, then a reviewed canonical URL. Use title and image similarity only to propose a match for review. Never merge records solely because titles look alike; variations, regional editions, and multipacks commonly share words.

Maintain a mapping table with confidence, match method, reviewer, and timestamps. When a source changes an identifier, create a new mapping event rather than silently rewriting history. For marketplaces with parent and child variations, model the parent catalog item and each purchasable child separately.

8. Rate limits, retries, and reliability

  • Read each provider’s usage plan and batch limits. Amazon documents throttling and batch limits; Keepa’s token bucket makes token consumption a scheduling constraint.
  • Use exponential backoff with jitter for transient 429 and 5xx responses.
  • Honor Retry-After when present.
  • Do not retry authentication failures, invalid identifiers, or permanent 4xx responses without changing the request.
  • Use idempotency keys or deterministic job IDs so a retry cannot duplicate an observation.
  • Record response status, latency, token or quota usage, parser version, and error category.
  • Alert on queue age, stale listings, rising parser errors, source coverage, and notification lag.

9. Storage, dashboards, and alerts

Append-only observations belong in a time-series or columnar store when volume is high. Keep a relational catalog for identity and mappings. Put raw payloads in durable object storage with retention and encryption controls appropriate to your data.

Useful dashboards show freshness by source, successful observations per hour, error rates, queue age, parser-version distribution, notification-to-snapshot reconciliation, and API token consumption. Alert only after applying a debounce window and minimum materiality rule, such as a percentage or absolute landed-price threshold.

10. Performance and cost planning

Estimate work as tracked listings × refreshes per period, then add retries and reconciliation snapshots. Batch where the provider supports it: Amazon’s featured-offer batch accepts up to 40 SKUs, while Amazon’s general batch operations can contain up to 20 similar requests. Keepa token cost and refill rate determine sustainable polling volume.

Use adaptive schedules. Refresh products with frequent changes, active promotions, or high business value more often. Slow down stable products. Cache unchanged responses where permitted, but never let cache age violate the freshness requirement you promised to users.

For crawlers, concurrency is limited per domain, not just globally. A high global worker count can still overload one retailer. Cost includes API calls, proxy or browser infrastructure where permitted, storage, parsing, alert delivery, and engineering maintenance. Historical depth and low latency usually increase one or more of those costs.

11. Troubleshooting checklist

Symptom Likely cause Fix
Many false price changes Currency, tax, shipping, or rounding differs Compare normalized landed values and retain conversion timestamps
Duplicate products Matching on title or URL alone Use source IDs, GTIN or MPN plus brand, then review uncertain matches
429 responses Exceeded quota or token rate Throttle per source, batch requests, back off, and monitor remaining quota
Stale catalog rows Notifications were assumed to be complete Run periodic reconciliation snapshots and track freshness
Parser suddenly returns nulls Source markup or schema changed Retain raw responses, version parsers, add field-level null alerts, and replay samples
Unexpected seller switch Featured offer changed Store seller, condition, and fulfillment as separate dimensions
Crawler blocked Robots, terms, authentication, or bot defenses Stop unauthorized collection; use an approved API or obtain permission
Alert storm No debounce or materiality threshold Group events by listing and send one summarized notification
A clean capture removes consent banners and overlays before visual comparison.
A clean capture removes consent banners and overlays before visual comparison.

Or skip the browser setup

If your workflow needs screenshots of product pages for visual audits, listing review, or evidence alongside structured price data, ScreenshotNeo provides a single GET request that returns PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing status. The MCP server includes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. You get 1,000 screenshots a month free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

FAQ

Should I poll or use notifications?

Use notifications for low-latency changes and scheduled snapshots for backfills and reconciliation. Notifications alone should not be treated as a complete historical record.

How often should prices be refreshed?

Base the schedule on observed volatility, business value, provider limits, and the freshness requirement. Adaptive schedules reduce cost while preserving coverage.

What is the best key for deduplication?

Use a provider’s stable product or listing identifier. Fall back to GTIN or MPN plus brand, and send ambiguous matches for review.

Do catalog changes matter if price is unchanged?

Yes. Availability, seller, condition, variation, title, images, attributes, category, and listing state can change customer outcomes without changing the displayed price.

Can I compare prices across countries?

Only after retaining marketplace, currency, tax, shipping, and conversion context. Report observed and normalized values separately.