How to Detect Price Changes in Screenshots of Product Pages with OCR
Build a price-change detector that captures product pages, extracts candidate prices with OCR, normalizes them, and flags meaningful changes for review.
To detect price changes in screenshots, capture the same product page and variant over time, run OCR on each screenshot, identify the likely offer price, normalize its currency and formatting, then compare it with the previous observation. OCR only reads visible text: it does not know which amount is the current product price. A page may also show a crossed-out list price, installment amount, shipping cost, coupon, or another unrelated number, so preserve the screenshot and flag ambiguous results for review.
This guide builds that workflow in Python with Google Cloud Vision, then covers equivalent request patterns in cURL and Node.js. It also explains capture consistency, candidate selection, normalization, operating costs, and troubleshooting.
1. Design the comparison before writing code
A comparison is useful only when both observations describe the same product, variant, currency, and comparable offer conditions. A different size, color, subscription term, locale, or logged-in offer can change the displayed amount without indicating a price change for the item you meant to track.
- Choose the observation identity. Record a stable product ID, canonical URL, variant, locale, and any relevant offer context.
- Capture on a schedule. Use a consistent viewport, device scale, locale, and page state. Wait for the offer area to appear and for prices to settle after client-side rendering.
- Keep the evidence. Store the original screenshot, capture time, URL, and OCR output alongside each extracted candidate.
- Extract text with OCR. Retain word-level geometry when available so candidate amounts can be related to the title and offer area.
- Select and normalize candidates. Parse plausible amounts, keep the raw strings, and map them to a currency-aware representation.
- Compare observations. Compare only matching identities and conditions; route uncertain or multiple-candidate cases to review.
Google Cloud Vision’s TEXT_DETECTION mode is designed to detect text in images and returns the full extracted string plus individual words and bounding boxes. Its DOCUMENT_TEXT_DETECTION mode is optimized for dense text and returns page, block, paragraph, word, and break structure. A typical product screenshot is closer to general image text detection, but validate the choice against screenshots from the sites, languages, and price formats you intend to monitor. These are OCR capabilities; candidate-price selection and comparison remain application logic. Google Cloud Vision text detection documentation.
2. Prepare credentials and a representative screenshot
Create a Google Cloud project, enable the Cloud Vision API, and configure application credentials using Google’s supported authentication setup. Install the client library:
python -m pip install google-cloud-vision
Save a representative product-page screenshot as product.png. Include the kinds of pages you expect to process: different themes, image quality, languages, currencies, and pages with sale prices or multiple offers. The example below uses the Google Cloud Python client and general image text detection.
3. Runnable Python pipeline
import json
import re
from datetime import datetime, timezone
from pathlib import Path
from google.cloud import vision
IMAGE_PATH = Path("product.png")
PRODUCT_ID = "example-product-blue-medium"
PRODUCT_URL = "https://shop.example/products/example"
LOCALE = "en-US"
CURRENCY = "USD"
# This deliberately recognizes common visible formats, not every global price format.
# Keep the raw OCR text and review candidates this parser cannot confidently interpret.
AMOUNT_RE = re.compile(
r"(?:(?:US\\$|\\$|€|£|¥)\\s?\\d[\\d,.]*(?:\\s?(?:USD|EUR|GBP|JPY))?"
r"|\\d[\\d,.]*\\s?(?:USD|EUR|GBP|JPY))",
re.IGNORECASE,
)
def detect_text(path: Path) -> dict:
client = vision.ImageAnnotatorClient()
image = vision.Image(content=path.read_bytes())
response = client.text_detection(image=image)
if response.error.message:
raise RuntimeError(response.error.message)
annotations = response.text_annotations
full_text = annotations[0].description if annotations else ""
words = []
# The first annotation is the combined text; subsequent entries are detected words.
for item in annotations[1:]:
vertices = item.bounding_poly.vertices
xs = [v.x for v in vertices]
ys = [v.y for v in vertices]
words.append({
"text": item.description,
"box": [min(xs), min(ys), max(xs), max(ys)],
})
return {"full_text": full_text, "words": words}
def normalize_amount(raw: str, currency: str):
"""Return integer minor units for a deliberately narrow example locale set.
Production parsers should use the known locale and currency and explicitly
handle that locale's decimal and grouping separators. Return None when unsure.
"""
value = re.sub(r"\\s", "", raw.upper())
for token in ("USD", "EUR", "GBP", "JPY", "US$", "$", "€", "£", "¥"):
value = value.replace(token, "")
value = value.strip()
if not re.fullmatch(r"[0-9][0-9,.]*", value):
return None
# Example conventions: comma thousands + dot decimals for USD/EUR/GBP;
# yen has no minor unit. This is not a universal locale parser.
if currency == "JPY":
digits = value.replace(",", "").replace(".", "")
return int(digits) if digits.isdigit() else None
if "." in value:
whole, fractional = value.rsplit(".", 1)
if len(fractional) not in (1, 2) or not fractional.isdigit():
return None
whole = whole.replace(",", "")
if not whole.isdigit():
return None
return int(whole) * 100 + int(fractional.ljust(2, "0"))
# In this example, a comma-only value is treated as a thousands grouping.
digits = value.replace(",", "")
return int(digits) * 100 if digits.isdigit() else None
def extract_candidates(full_text: str, currency: str) -> list[dict]:
candidates = []
for match in AMOUNT_RE.finditer(full_text):
raw = match.group(0)
candidates.append({"raw": raw, "minor_units": normalize_amount(raw, currency)})
return candidates
def main():
result = detect_text(IMAGE_PATH)
candidates = extract_candidates(result["full_text"], CURRENCY)
observation = {
"product_id": PRODUCT_ID,
"url": PRODUCT_URL,
"locale": LOCALE,
"currency": CURRENCY,
"captured_at": datetime.now(timezone.utc).isoformat(),
"screenshot_path": str(IMAGE_PATH),
"ocr_text": result["full_text"],
"ocr_words_with_boxes": result["words"],
"price_candidates": candidates,
"needs_review": len(candidates) != 1 or any(c["minor_units"] is None for c in candidates),
}
print(json.dumps(observation, indent=2, ensure_ascii=False))
if __name__ == "__main__":
main()
The script is runnable after authentication and installation, but its parser is intentionally a small example, not a universal money parser. It prints all recognized candidate amounts and requests review unless exactly one parseable candidate exists. In a real monitor, use the known locale’s number conventions, account for the expected currency, and select candidates using page context and geometry. Do not silently treat the first amount OCR finds as the offer price.
4. Select the likely offer price and compare it
A practical selector should combine signals rather than rely on a single regular expression:
- Restrict candidates to the expected currency and a plausible range for the tracked product.
- Use bounding boxes to prefer the price near the product title, selected variant, or offer area.
- Consider nearby labels such as “sale,” “from,” “per month,” or “shipping,” while treating labels as clues rather than guarantees.
- Detect multiple candidates and distinguish a current price from a crossed-out reference price when the layout makes that possible.
- Keep confidence and ambiguity explicit. If there are conflicting candidates, missing currency, or suspicious OCR substitutions such as
Ofor0, send the observation for review.
For each accepted observation, store a normalized integer amount in the currency’s minor unit (for example, cents where applicable), the currency code, raw matched text, screenshot reference, OCR text, timestamp, locale, product/variant identity, and selector decision. Compare only observations with matching identity and currency. Decide separately whether to alert on any numeric change or only changes above a chosen threshold; record that policy so alerts can be interpreted later.
OCR output does not establish that the page was fully loaded or that the displayed offer was available to the visitor. Keep capture status with the observation, and treat blank, timed-out, bot-check, or otherwise incomplete captures as missing data rather than as a price of zero.
5. Make the OCR request with cURL
For direct REST use, send the image bytes as base64 in a Vision API images:annotate request. The following shell example assumes GOOGLE_APPLICATION_CREDENTIALS is configured and gcloud can obtain an access token. It requests general text detection:
ACCESS_TOKEN="$(gcloud auth application-default print-access-token)"
IMAGE_B64="$(base64 -w 0 product.png)"
curl -sS -X POST \
-H "Authorization: Bearer ${ACCESS_TOKEN}" \
-H "Content-Type: application/json" \
"https://vision.googleapis.com/v1/images:annotate" \
-d "{\"requests\":[{\"image\":{\"content\":\"${IMAGE_B64}\"},\"features\":[{\"type\":\"TEXT_DETECTION\"}]}]}" \
-o vision-response.json
Inspect responses[0].textAnnotations for the combined text and word annotations with bounding polygons. If you are processing dense document-like material and need its structure, evaluate DOCUMENT_TEXT_DETECTION and inspect fullTextAnnotation. The Cloud Vision documentation describes the request and response structures: OCR text detection.
6. Make the OCR request with Node.js
This example uses the official Google Cloud client package. Run npm install @google-cloud/vision, configure Google Cloud application credentials, then save as detect-price.js and run node detect-price.js:
const vision = require('@google-cloud/vision');
const fs = require('node:fs');
async function main() {
const client = new vision.ImageAnnotatorClient();
const [result] = await client.textDetection('product.png');
const annotations = result.textAnnotations || [];
const text = annotations[0]?.description || '';
const words = annotations.slice(1).map(item => ({
text: item.description,
vertices: item.boundingPoly?.vertices || [],
}));
const amountPattern = /(?:US\\$|\\$|€|£|¥)\\s?\\d[\\d,.]*|\\d[\\d,.]*\\s?(?:USD|EUR|GBP|JPY)/gi;
const candidates = [...text.matchAll(amountPattern)].map(match => ({ raw: match[0] }));
console.log(JSON.stringify({
capturedAt: new Date().toISOString(),
screenshot: 'product.png',
text,
words,
candidates,
needsReview: candidates.length !== 1,
}, null, 2));
}
main().catch(error => {
console.error(error);
process.exitCode = 1;
});
This example extracts candidates but deliberately leaves locale-aware normalization and candidate selection visible as application responsibilities. Preserve the original match string and avoid guessing when punctuation could mean either a decimal separator or a thousands separator.
7. Capture consistently, or use an API
If you manage a browser capture pipeline yourself, keep the browser version, viewport, device scale, locale, timezone, cookies, authentication state, and wait condition stable where the site allows it. Wait for the specific offer selector or for the page to settle; a fixed delay alone can be wasteful and can still miss a slow price widget. Avoid comparing screenshots captured with different variants or different consent/login states.
For scheduled monitoring, track capture failures separately from OCR failures and price ambiguity. Retry transient navigation or service errors with bounded backoff, but avoid rapid retries that can trigger bot defenses or increase request volume. Keep the last valid observation and mark a missed interval; never convert a failed capture into a price-change event.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. Its API returns an image or PDF from one GET request. Cookie banners, newsletter popups, and chat widgets are removed before the shot; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. Responses identify the page verdict and billing status in headers. The MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
Install the Python dependency with python -m pip install requests, then run this example after setting your API key:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Use the ScreenshotNeo API documentation for request options and response details. For this workflow, save each returned image with the product identity and capture timestamp, then submit it to your OCR step. ScreenshotNeo offers 1,000 screenshots a month free with no card; paid plans start at $5 for 3,000. Every feature is available on every plan. Sign up for 1,000 free screenshots a month, with no card.
8. Accuracy, performance, reliability, and cost
Accuracy and review
The cited sources do not report a benchmark or error rate for extracting product prices from web screenshots, and they do not compare Vision with Textract for this task. Evaluate a representative sample from your target sites and measure your own candidate-selection errors. Include sale layouts, small text, localized separators, multiple currencies, and screenshots with overlays. Do not treat OCR confidence or a successful API response as proof that the selected amount is the correct offer price.
Throughput and batching
Google describes immediate OCR processing for up to 16 images and asynchronous batch processing for up to 2,000 images per request. These are documented service capabilities, not a recommended capture cadence or a guarantee of end-to-end completion time. Choose batch size and concurrency around your latency needs, quotas, and retry policy. Google Cloud OCR use case.
Cost
Google Cloud lists the first 1,000 units per month as free and Text Detection and Document Text Detection at $1.50 per 1,000 units for units 1,001 through 5,000,000 on the cited pricing page. Rates can depend on currency and tier and may change, so check the live pricing page before budgeting. The OCR use-case page gives a $27.36 monthly estimate for an example that includes 15,000 monthly Vision API calls and related infrastructure; that is a worked example under its stated assumptions, not a universal cost estimate for this pipeline. Cloud Vision pricing.
Estimate the full workflow, not just OCR units: screenshot capture, storage and retention, database or queue operations, retries, and human review can also contribute cost. Deduplicate unchanged captures only if that fits your monitoring objective; do not mistake caching behavior for a new observation.
Data handling
Screenshots can contain account details, addresses, or other information beyond the price. Restrict access, store only what the workflow needs, set a retention period, and review the OCR provider’s current data handling terms before sending sensitive pages. The sources cited here establish OCR capabilities and pricing, not comparative retention or privacy terms.
9. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| No text or no candidate price | The price is rendered after capture, too small, obscured, or present only in an image with poor quality. | Wait for the offer element, capture at a readable scale, inspect the screenshot, and log the full OCR response before changing parsing rules. |
| Several prices are detected | The page shows list price, sale price, financing, shipping, or prices for other variants. | Use page context and bounding boxes, bind observations to a variant, and mark unresolved cases for review. |
| Wrong amount after parsing | Locale separators were interpreted with the wrong convention, or OCR confused similar characters. | Parse with the known locale and currency, keep raw OCR text, reject ambiguous punctuation, and require review for suspicious substitutions. |
| Authentication or permission error | Credentials are missing, expired, or lack permission to call Vision. | Check the active project, API enablement, application credentials, and service-account permissions using Google’s current setup guidance. |
| Request rejected for image or payload | The image encoding or request structure is invalid, or the payload exceeds service constraints. | Confirm the base64 contains only image bytes, send valid JSON, and consult the current Vision API request limits. Consider using a supported cloud-storage image reference for larger inputs. |
| Capture appears blank or has a challenge page | The page did not finish loading or served a bot check rather than the product page. | Record the capture as invalid or incomplete, apply a bounded retry policy, and do not compare its OCR result as a price observation. |
| Frequent duplicate alerts | Equivalent prices have different formatting, or repeated captures are being treated as new changes. | Compare normalized amounts and currency, store observation IDs and timestamps, and define whether alerting is based on value change or every observation. |
10. FAQ
Can OCR tell which number is the actual selling price?
No. OCR returns recognized text and, in Vision’s general text mode, word locations. Your application must identify the offer price using page context and handle ambiguity.
Should I use text detection or document text detection?
Start by evaluating general text detection for ordinary web screenshots. Consider document text detection for unusually dense, document-like content. Validate both against your own pages; the documentation does not establish which is more accurate for product-price monitoring.
Can I compare screenshots from different locales?
Only after explicitly accounting for locale, currency, separators, and offer conditions. In most monitoring systems, separate observations by locale and currency so a formatting or conversion difference is not mistaken for a price change.
Is AWS Textract proven better for this use?
No such conclusion follows from the cited material. AWS describes Textract as extracting text and layout information from documents, but the sources do not provide a head-to-head product-screenshot benchmark. Evaluate it on your own representative images if it is a candidate. Amazon Textract.


