Scrape Product Prices and Stock from Any Store
Build a reliable price and stock scraper with robots.txt checks, structured data, browser rendering, variant handling, validation and history.
Use a domain-aware extraction pipeline: discover product URLs, read each host’s robots.txt, load static pages with HTTP and JavaScript pages with a browser, extract typed price and availability fields, validate the result, and store an auditable history. There is no universal price or stock selector that works across every store.
1. What to collect
For every observation, keep the product identity, selected variant, price, currency, stock label, normalized status, source URL, region and timestamp.
| Field | Purpose |
|---|---|
title, brand |
Human-readable identity |
sku, product_id, gtin |
Stable matching when shown |
variant |
Size, color, pack, seller or other selected dimensions |
price, currency |
Numeric amount and ISO-style currency kept separately |
availability_raw, availability |
Original store wording plus a controlled status |
quantity |
Nullable quantity only when explicitly exposed |
observed_at, parser_version |
Freshness and audit trail |
2. Respect crawl controls first
Fetch https://host/robots.txt separately for each host, protocol and port. Apply the matching user-agent group before requesting product pages. A rule served by one host, protocol or port does not automatically apply to another. Robots.txt is a crawler signal, not a complete answer to terms, privacy or copyright questions.
- Review terms of service and avoid login-only pages or personal data.
- Rate-limit requests and identify your crawler where appropriate.
- Collect factual fields needed for the use case instead of copying descriptions or images.
- Re-read robots.txt when rules change; Amazon documents that its ProductDiscoverybot can take up to 24 hours to reflect a change.
3. A runnable Python scraper
The example below uses ordinary HTTP first, JSON-LD and visible fallbacks, and a controlled stock vocabulary. It returns null for missing values so a failed extraction cannot silently become stale data.
#!/usr/bin/env python3
import json
import re
import sys
from datetime import datetime, timezone
from urllib.parse import urljoin, urlparse
from urllib.robotparser import RobotFileParser
import requests
from bs4 import BeautifulSoup
UA = "price-stock-monitor/1.0 (+https://example.com/bot-info)"
TIMEOUT = 30
def robots_allowed(url: str) -> bool:
p = urlparse(url)
robots_url = f"{p.scheme}://{p.netloc}/robots.txt"
rp = RobotFileParser(robots_url)
try:
rp.read()
except Exception:
# Choose your policy explicitly when robots.txt is unavailable.
return False
return rp.can_fetch(UA, url)
def number(value):
if value is None:
return None
text = str(value).strip()
# Keep digits, separators and a possible minus sign; adapt this for locale.
text = re.sub(r"[^0-9,.-]", "", text)
if not text:
return None
if text.count(",") == 1 and text.count(".") == 0:
text = text.replace(",", ".")
elif text.count(",") > 0 and text.count(".") > 0:
text = text.replace(",", "")
try:
return float(text)
except ValueError:
return None
def first_text(soup, selectors):
for selector in selectors:
node = soup.select_one(selector)
if node:
value = node.get("content") or node.get_text(" ", strip=True)
if value:
return value
return None
def normalize_stock(raw):
if not raw:
return "unknown"
value = raw.lower()
if any(x in value for x in ("out of stock", "sold out", "unavailable", "out-of-stock")):
return "out_of_stock"
if any(x in value for x in ("low stock", "only ", "left in stock", "limited")):
return "low_stock"
if any(x in value for x in ("in stock", "available", "ships today", "ready to ship")):
return "in_stock"
return "unknown"
def extract(url, html):
soup = BeautifulSoup(html, "html.parser")
jsonld = []
for tag in soup.select('script[type="application/ld+json"]'):
try:
data = json.loads(tag.string or tag.get_text())
jsonld.extend(data if isinstance(data, list) else [data])
except (json.JSONDecodeError, TypeError):
continue
product = next((x for x in jsonld if isinstance(x, dict) and x.get("@type") in ("Product", ["Product"])), {})
offers = product.get("offers") or {}
if isinstance(offers, list):
offers = offers[0] if offers else {}
raw_price = offers.get("price") or first_text(soup, [
'[itemprop="price"]', '[data-testid*="price"]', '.price', '[class*="price"]'
])
currency = offers.get("priceCurrency") or first_text(soup, [
'[itemprop="priceCurrency"]', '[data-currency]'
])
raw_stock = offers.get("availability") or first_text(soup, [
'[itemprop="availability"]', '[data-testid*="stock"]', '[class*="stock"]', '[class*="availability"]'
])
return {
"url": url,
"canonical_url": (soup.select_one('link[rel="canonical"]') or {}).get("href") if soup.select_one('link[rel="canonical"]") else url,
"title": product.get("name") or first_text(soup, ["h1", '[itemprop="name"]']),
"brand": (product.get("brand") or {}).get("name") if isinstance(product.get("brand"), dict) else product.get("brand"),
"sku": product.get("sku") or first_text(soup, ['[itemprop="sku"]']),
"price": number(raw_price),
"currency": currency,
"availability_raw": raw_stock,
"availability": normalize_stock(raw_stock),
"observed_at": datetime.now(timezone.utc).isoformat(),
}
def scrape(url):
if not robots_allowed(url):
raise RuntimeError(f"robots.txt disallows or could not be read for {url}")
response = requests.get(url, headers={"User-Agent": UA}, timeout=TIMEOUT)
response.raise_for_status()
result = extract(url, response.text)
if result["price"] is None and result["availability"] == "unknown":
result["extraction_alert"] = "no price or stock rule matched"
return result
if __name__ == "__main__":
for target in sys.argv[1:]:
print(json.dumps(scrape(target), ensure_ascii=False))
Install dependencies with python -m pip install requests beautifulsoup4, then run python scraper.py https://store.example/product. Replace the example selectors with selectors maintained per domain. The sample deliberately fails closed when robots.txt cannot be read; choose a documented alternative only after reviewing your policy.
4. Structured data before CSS selectors
Prefer Schema.org Product and Offer JSON-LD, then itemprop="price"/priceCurrency, then visible selectors. Keep fallbacks ordered and return null when a value is absent or invalid. Numeric typing prevents calculations on strings such as “$1,299.00”, while a separate currency field prevents mixing dollars, euros and pounds.
<script type="application/ld+json">
{
"@type": "Product",
"name": "Example",
"sku": "ABC-123",
"offers": {
"@type": "Offer",
"price": "1299.00",
"priceCurrency": "USD",
"availability": "https://schema.org/InStock"
}
}
</script>
5. JavaScript-rendered pages and browser interaction
Use a real browser when the price appears only after JavaScript, a location is selected, a size or color changes the offer, or products load through infinite scroll. Wait for the price or stock element rather than sleeping for an arbitrary fixed time.
import asyncio
from playwright.async_api import async_playwright
async def read_variant(url, size, color):
async with async_playwright() as p:
browser = await p.chromium.launch()
page = await browser.new_page()
await page.goto(url, wait_until="domcontentloaded", timeout=60000)
await page.select_option("select[name=size]", label=size)
await page.click(f"[data-color='{color}']")
await page.locator("[data-testid='price']").wait_for(state="visible", timeout=15000)
price = await page.locator("[data-testid='price']").inner_text()
stock = await page.locator("[data-testid='stock']").inner_text()
await browser.close()
return {"size": size, "color": color, "price_raw": price, "stock_raw": stock}
print(asyncio.run(read_variant("https://store.example/product", "Medium", "Black")))
Install with pip install playwright and playwright install chromium. For infinite scroll, repeatedly scroll, wait for new product cards, and stop when the count no longer increases. Record the selected variant with every observation; a default variant is not evidence that all variants share its price or stock.
6. Handling variants, sellers and regions
- Discover every permitted option from size, color, pack, seller and subscription controls.
- Select one combination at a time and wait for the offer to update.
- Read the price, currency and stock after the update, not before it.
- Store variant dimensions as structured fields and preserve the raw labels.
- Repeat with the required locale, timezone, location or shipping region.
Do not infer exact inventory from “limited time” or merchandising copy. If the store exposes a quantity, keep it as a nullable quantity field and retain the original wording.
7. Normalize availability safely
| Normalized value | Examples |
|---|---|
in_stock |
In stock, available, ready to ship |
out_of_stock |
Out of stock, sold out, unavailable |
low_stock |
Only 2 left, low stock, limited availability |
unknown |
Missing, ambiguous or contradictory label |
Always retain availability_raw. Store-specific wording changes, and an unknown result should raise an extraction alert rather than become an assumed in-stock or out-of-stock value.
8. cURL, Python and Node.js request patterns
curl -L --max-time 30 -A "price-stock-monitor/1.0" "https://store.example/product"
import requests
r = requests.get("https://store.example/product", headers={"User-Agent": "price-stock-monitor/1.0"}, timeout=30)
r.raise_for_status()
html = r.text
const res = await fetch('https://store.example/product', {
headers: { 'User-Agent': 'price-stock-monitor/1.0' },
signal: AbortSignal.timeout(30000)
});
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
const html = await res.text();
9. Validation and history
- Reject negative prices, impossible currency codes and malformed numbers.
- Compare the selected currency with the expected region.
- Detect anti-bot pages by title, status code, challenge markers and missing product identity.
- Alert on sudden increases in null values or selector misses.
- Keep the raw HTML or a content hash where permitted, plus parser version and observed timestamp.
- Compare repeated runs before publishing a price change.
A minimal record might be:
{
"store": "example",
"url": "https://store.example/product",
"sku": "ABC-123",
"variant": {"size": "M", "color": "Black"},
"price": 1299.0,
"currency": "USD",
"availability_raw": "Only 2 left in stock",
"availability": "low_stock",
"quantity": 2,
"observed_at": "2026-01-01T12:00:00Z",
"parser_version": "example-v4"
}
10. Performance, reliability and cost
- Use HTTP for static pages; browsers consume more CPU, memory and time.
- Cache robots.txt briefly and cache product responses only when freshness permits.
- Use bounded concurrency per host, exponential backoff for transient failures and connection reuse.
- Schedule high-change products more often than stable catalogs.
- Separate discovery, rendering, extraction and persistence so one failure does not erase prior history.
- Budget engineering time, browser minutes, proxy use, storage and vendor fees—not only request counts.
For a small, controlled set of stores, custom extraction gives maximum control. A hosted Actor can reduce browser and selector maintenance for many stores. A managed feed is useful when scheduled delivery, regional rendering and extractor repair matter. Verify current commercial terms before choosing a vendor.
11. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Price is null | Client-rendered value or changed selector | Inspect JSON-LD, wait for the price element, then update the domain fallback and alert on the miss. |
| Price is wrong by 1,000× | Minor units or locale separators | Check the store’s currency and decimal convention; parse locale explicitly. |
| Stock always unknown | Availability is variant-specific or hidden in an offer API | Select each variant, inspect JSON-LD/network responses where permitted, and preserve the raw label. |
| Browser sees a challenge | Bot check, rate limit or blocked automation | Stop retrying aggressively, respect terms and rate limits, and mark the observation as an anti-bot failure. |
| Different users see different prices | Region, cookies, account or location | Record locale, timezone, cookies and location assumptions; compare like with like. |
| Infinite-scroll items are missing | Extraction ran before new cards loaded | Scroll, wait for the card count to increase, and stop after several unchanged counts. |
| robots.txt cannot be fetched | Network error or unavailable file | Fail closed or apply your documented policy; do not silently crawl. |
12. Or skip the browser setup
ScreenshotNeo can render a store page and return a clean PNG, JPEG, WebP or PDF through one request. Use the API documentation for all options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Cookie banners, newsletter popups and chat widgets are removed before the shot. Bot checks, blank pages and failed loads are never billed, and response headers identify the page verdict and billing result. An MCP server lets AI agents use take_screenshot, get_page_info and capture_pdf. You get 1,000 screenshots each month free with no card; paid plans start at $5 for 3,000.
Create a free ScreenshotNeo account.
13. FAQ
Can one CSS selector scrape every store?
No. Store templates, labels and rendering differ. Use per-domain rules with structured-data fallbacks.
Should I store price as text?
Store a numeric amount and currency, plus the original text for audit and debugging.
Is robots.txt permission to scrape?
No. It is a crawler control signal. Also review terms, privacy, rate limits and applicable law.
How often should prices be checked?
Choose a schedule based on product volatility, freshness requirements, site limits and cost. Record the schedule with each source.
What should happen after an extraction fails?
Store null, retain the failure reason, alert on repeated misses and never overwrite a fresh value with an unmarked stale or guessed value.


