ScreenshotNeo

BlogEngineering

How to Scrape Amazon ASIN Data at Scale With Python

Build a reliable Python pipeline for Amazon ASIN discovery and retrieval with batching, signed requests, throttling, retries, checkpoints, and migration planning.

By the ScreenshotNeo team1 October 202610 min read

Short answer: use an authorized Amazon API for production-scale ASIN collection. Discover identifiers with SearchItems, retrieve details with GetItems in batches of up to 10 ASINs, request only the resources you need, sign every request with AWS Signature Version 4, throttle requests, retry transient failures, and checkpoint progress. Treat inaccessible ASINs as a normal result that must be stored separately from successful records.

Amazon defines an ASIN as a 10-character alphanumeric item identifier. Use the normalized ASIN plus marketplace as the stable key in your data model. HTML scraping can be technically possible, but it requires a marketplace-specific compliance review: robots.txt behavior does not grant permission to collect or reuse catalog data.

1. Decide what you are collecting

Write down the marketplace, discovery method, fields, refresh interval, and retention policy before writing a crawler.

  • Marketplace: choose one explicitly. API resource availability varies by locale.
  • Discovery: use SearchItems with keywords, search index, and marketplace parameters.
  • Retrieval: use GetItems for known ASINs.
  • Fields: request only what you need, such as ItemInfo, Images, BrowseNodeInfo, Offers, OffersV2, and ParentASIN.
  • Freshness: store retrieval timestamps. Do not describe price or availability as current unless you retrieved and timestamped it.
  • Storage: keep raw responses only when your agreement and retention rules allow it.

2. Model ASINs as marketplace-scoped identifiers

An ASIN can occur in more than one marketplace, and parent-child relationships matter for variations. A practical record includes:

marketplace
asin
parent_asin
title
brand
images
offers
browse_nodes
retrieved_at
source_request_id
status
error_code
raw_response_reference

Normalize input with strip().upper(), reject values that are not 10 alphanumeric characters, and deduplicate on (marketplace, asin). Keep invalid, missing, and inaccessible identifiers in a separate status table so one bad item does not hide a batch failure.

3. Use SearchItems for discovery

SearchItems is the normal starting point when you have keywords or a category rather than ASINs. Persist every returned ASIN and its parent relationship before requesting details.

Discovery should be restartable. Store the search parameters, marketplace, page or token, retrieval time, and the last successful response. If a process stops after page 7, resume from the stored token instead of repeating the entire search.

4. Retrieve details with GetItems batches

Send no more than 10 ASINs in one GetItems request. The response has separate Items and Errors containers, so a successful HTTP response can still contain inaccessible ASINs.

  1. Normalize and deduplicate ASINs.
  2. Split them into groups of 10.
  3. Send one signed request per group.
  4. Write successful items and per-ASIN errors independently.
  5. Checkpoint the batch only after both writes succeed.

5. Python implementation

The following worker uses botocore to produce AWS Signature Version 4 headers. It deliberately reads the endpoint, host, region, partner tag, and marketplace from environment variables because these values vary by account, locale, and API generation. Install dependencies with pip install requests botocore.

import json
import os
import re
import time
from datetime import datetime, timezone
from pathlib import Path

import requests
from botocore.auth import SigV4Auth
from botocore.awsrequest import AWSRequest
from botocore.credentials import Credentials

ASIN_RE = re.compile(r"^[A-Z0-9]{10}$")
API_ENDPOINT = os.environ["AMAZON_API_ENDPOINT"]
AMAZON_HOST = os.environ["AMAZON_HOST"]
AMAZON_REGION = os.environ["AMAZON_REGION"]
AMAZON_ACCESS_KEY = os.environ["AMAZON_ACCESS_KEY"]
AMAZON_SECRET_KEY = os.environ["AMAZON_SECRET_KEY"]
AMAZON_PARTNER_TAG = os.environ["AMAZON_PARTNER_TAG"]
AMAZON_MARKETPLACE = os.environ["AMAZON_MARKETPLACE"]
CHECKPOINT = Path("asin-checkpoint.json")


def normalize_asin(value: str) -> str | None:
    value = value.strip().upper()
    return value if ASIN_RE.fullmatch(value) else None


def chunks(values, size=10):
    for start in range(0, len(values), size):
        yield values[start:start + size]


def signed_post(operation: str, payload: dict) -> dict:
    body = json.dumps(payload, separators=(",", ":"))
    headers = {
        "content-type": "application/json; charset=UTF-8",
        "host": AMAZON_HOST,
        "x-amz-target": f"com.amazon.paapi5.v1.ProductAdvertisingAPIv1.{operation}",
        "user-agent": "asin-pipeline/1.0",
    }
    request = AWSRequest(method="POST", url=API_ENDPOINT, data=body, headers=headers)
    credentials = Credentials(AMAZON_ACCESS_KEY, AMAZON_SECRET_KEY)
    SigV4Auth(credentials, "ProductAdvertisingAPI", AMAZON_REGION).add_auth(request)
    response = requests.post(API_ENDPOINT, data=body, headers=dict(request.headers), timeout=30)
    response.raise_for_status()
    return response.json()


def get_items(asins: list[str], resources: list[str]) -> dict:
    payload = {
        "ItemIds": asins,
        "ItemIdType": "ASIN",
        "Resources": resources,
        "PartnerTag": AMAZON_PARTNER_TAG,
        "PartnerType": "Associates",
        "Marketplace": AMAZON_MARKETPLACE,
    }
    return signed_post("GetItems", payload)


def load_checkpoint() -> set[str]:
    if not CHECKPOINT.exists():
        return set()
    return set(json.loads(CHECKPOINT.read_text()).get("completed", []))


def save_checkpoint(completed: set[str]):
    temporary = CHECKPOINT.with_suffix(".tmp")
    temporary.write_text(json.dumps({"completed": sorted(completed)}))
    temporary.replace(CHECKPOINT)


def fetch_all(raw_asins: list[str]):
    normalized = sorted({a for value in raw_asins if (a := normalize_asin(value))})
    completed = load_checkpoint()
    resources = [
        "ItemInfo.Title",
        "ItemInfo.ByLineInfo",
        "Images.Primary.Large",
        "BrowseNodeInfo.BrowseNodes",
        "ParentASIN",
    ]
    for batch in chunks([a for a in normalized if a not in completed]):
        try:
            response = get_items(batch, resources)
        except requests.RequestException as exc:
            print(f"transport failure for {batch}: {exc}")
            time.sleep(2)
            continue

        retrieved_at = datetime.now(timezone.utc).isoformat()
        for item in response.get("ItemsResult", {}).get("Items", []):
            record = {"retrieved_at": retrieved_at, "status": "ok", "item": item}
            print(json.dumps(record))
        for error in response.get("Errors", []):
            record = {"retrieved_at": retrieved_at, "status": "error", "error": error}
            print(json.dumps(record))

        completed.update(batch)
        save_checkpoint(completed)
        time.sleep(1.0)


if __name__ == "__main__":
    fetch_all([line.strip() for line in open("asins.txt") if line.strip()])

Use the exact operation names, resource names, endpoint, and authentication requirements documented for the API version available to your account. Amazon’s documentation currently carries the statement: “PA-API will be deprecated on May 15th, 2026. Please migrate to Creators API.” Make migration a release dependency rather than assuming this payload will remain valid.

Discovery request shape

For discovery, replace the operation and payload with the current SearchItems schema for your API generation. Persist the returned ASINs and pagination token. Do not silently discard a response that contains both results and errors.

payload = {
    "Keywords": "wireless keyboard",
    "SearchIndex": "All",
    "ItemCount": 10,
    "Resources": ["ItemInfo.Title", "ParentASIN"],
    "PartnerTag": AMAZON_PARTNER_TAG,
    "PartnerType": "Associates",
    "Marketplace": AMAZON_MARKETPLACE,
}
response = signed_post("SearchItems", payload)
items = response.get("SearchResult", {}).get("Items", [])
next_token = response.get("SearchResult", {}).get("SearchURL")
for item in items:
    print(item.get("ASIN"))

6. cURL and Node.js request patterns

Signature Version 4 must be generated for each request. The examples below assume your signing layer has already produced the required Authorization, x-amz-date, and related headers.

curl -X POST "$AMAZON_API_ENDPOINT" \
  -H "content-type: application/json; charset=UTF-8" \
  -H "x-amz-target: com.amazon.paapi5.v1.ProductAdvertisingAPIv1.GetItems" \
  -H "x-amz-date: $AMZ_DATE" \
  -H "authorization: $AWS_AUTHORIZATION" \
  --data '{"ItemIds":["B000123456"],"ItemIdType":"ASIN","Resources":["ItemInfo.Title"],"PartnerTag":"YOUR_PARTNER_TAG","PartnerType":"Associates","Marketplace":"YOUR_MARKETPLACE"}'
const payload = {
  ItemIds: ["B000123456"],
  ItemIdType: "ASIN",
  Resources: ["ItemInfo.Title"],
  PartnerTag: process.env.AMAZON_PARTNER_TAG,
  PartnerType: "Associates",
  Marketplace: process.env.AMAZON_MARKETPLACE
};
// Sign method, URL, headers, region, service, and body with AWS Signature Version 4.
const response = await fetch(process.env.AMAZON_API_ENDPOINT, {
  method: "POST",
  headers: signedHeaders,
  body: JSON.stringify(payload)
});
if (!response.ok) throw new Error(`${response.status}: ${await response.text()}`);
const data = await response.json();
console.log(data);

7. Throttling, retries, and concurrency

Amazon documents TPS and TPD limits that depend on account history and shipped revenue. One current Associates help page lists an initial rate of one request per second, with one additional request per second per $4,600 in shipped revenue, capped at 10 requests per second. Treat those values as account-dependent and verify the current limit before deployment.

  • Use a token bucket or leaky bucket shared by all workers.
  • Start with one request per second and increase only after observing account limits.
  • Use bounded concurrency; concurrency does not remove the need for a rate limiter.
  • Retry throttling and temporary 5xx responses with exponential backoff and jitter.
  • Do not retry malformed requests, invalid credentials, or permanently inaccessible ASINs without changing the request.
  • Honor Retry-After when supplied.
  • Persist checkpoints after each completed batch.

A simple schedule is 1, 2, 4, 8 seconds plus random jitter, capped at a practical maximum, with a finite retry count. Record the final failure and continue with later batches.

8. Pagination and checkpoints

Search results are paginated. Store the marketplace, normalized search parameters, page token, request timestamp, and response status. Checkpoint after durable writes, not immediately after sending a request. If a process is interrupted, replaying the last batch is safer than skipping it; deduplicate on the marketplace and ASIN key.

9. Resource selection and payload cost

Requesting every resource increases response size, latency, parsing work, and the chance of field-level errors. Begin with title and identifier fields, then add images, browse nodes, offers, or parent information only when a downstream use requires them. Keep resource lists version-controlled so a schema change is visible in code review.

10. HTML scraping versus an authorized API

Question Authorized API HTML collection
Authorization and terms Defined by your API agreement Requires marketplace-specific review
Field coverage Documented resources, varying by locale Page markup can change without notice
Freshness Controlled by API limits and retrieval time Depends on page access and rendering
Failure handling Structured errors and throttling signals Selectors, bot checks, and layouts can fail
Reproducibility Request and resource lists can be logged Rendered output varies by session and geography
Migration risk PA-API is scheduled for deprecation Terms and page structure can change

Robots.txt is not a license to collect, store, or redistribute catalog data. Review Amazon terms, marketplace policy, privacy obligations, affiliate requirements, and data-retention limits before collecting pages or publishing derived data.

11. Validation and storage checklist

  • Normalize ASINs to uppercase.
  • Reject values that are not exactly 10 alphanumeric characters.
  • Deduplicate by marketplace and ASIN.
  • Store retrieval time and API generation.
  • Separate successful items from inaccessible IDs.
  • Preserve parent-child relationships.
  • Redact credentials and signed authorization values from logs.
  • Do not claim price or availability is current without a timestamp.
  • Retain raw responses only where policy permits.

12. Troubleshooting

Signature or authorization errors

Cause: wrong region, service name, host, timestamp, canonical headers, or body hash. Fix: sign the exact bytes sent over the wire, keep marketplace host and region explicit, synchronize the system clock, and never reuse an old authorization header.

Throttling responses

Cause: request rate or daily volume exceeds the account limit. Fix: lower the shared limiter rate, add exponential backoff with jitter, reduce concurrency, and request a documented rate increase when eligible.

HTTP success but missing items

Cause: inaccessible ASINs are returned in the response’s Errors container. Fix: process Items and Errors independently and persist the per-ASIN reason.

Invalid ASINs

Cause: whitespace, lowercase input, parent identifiers from another marketplace, or a non-ASIN product code. Fix: normalize, validate, deduplicate, and confirm the marketplace before calling GetItems.

Empty search pages

Cause: an overly narrow keyword, unsupported search index, exhausted pagination token, or locale mismatch. Fix: log the complete search parameters, verify the marketplace, and handle an empty page as a valid terminal result.

Large or slow responses

Cause: too many resources or image and offer fields in one request. Fix: request only required resources and split downstream enrichment into separate stages.

Duplicate records after restart

Cause: a crash between the API response and checkpoint write. Fix: use an idempotent upsert keyed by marketplace and ASIN, then checkpoint after the transaction commits.

13. Performance, reliability, and cost notes

Batching up to 10 ASINs per GetItems request reduces request overhead, but it does not change account quotas. Payload minimization improves latency and memory use. Durable checkpoints let you resume long jobs without repeating discovery. A bounded worker pool is safer than unbounded async tasks because it keeps rate, memory, and retry pressure predictable.

Measure request count, throttle count, error count by code, batch latency, response bytes, successful items, inaccessible items, and checkpoint age. Alert when error rates or checkpoint age rises, rather than only when the process exits.

14. Plan the PA-API to Creators API migration

Amazon’s indexed documentation says: “PA-API will be deprecated on May 15th, 2026. Please migrate to Creators API.” Before releasing a new integration, verify Creators API enrollment, quotas, authentication, marketplace coverage, field mappings, affiliate attribution, and retention rules. Keep the discovery and storage layers separated from the transport layer so you can replace signed request code without rewriting validation and persistence.

15. Or skip the browser setup

If your workflow also needs visual evidence of product pages, ScreenshotNeo provides a single screenshot request. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf.

See the ScreenshotNeo API documentation for options such as full-page capture, CSS element capture, custom headers and cookies, JavaScript, waits, blocked resources, caching, signed links, async jobs, bulk capture, and usage reporting.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://www.amazon.com/dp/B000123456 -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://www.amazon.com/dp/B000123456"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://www.amazon.com/dp/B000123456' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo includes 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

FAQ

What is the safest stable key?

Use the normalized ASIN together with the Amazon marketplace. Store parent ASIN separately for variation analysis.

How many ASINs fit in one GetItems call?

Up to 10 ASINs per request.

Should I scrape product HTML instead?

Use an authorized API for production workloads unless a marketplace-specific legal and policy review approves HTML collection.

Why did an HTTP 200 response contain no item?

Inaccessible identifiers can appear in the response’s Errors container. Process that container explicitly.

When should I migrate?

Before shipping new PA-API-dependent work. The documented deprecation date is May 15, 2026, so verify current Creators API requirements now.