How to Scrape, Monitor, and Download Supplier Product Data
Build a refreshable supplier catalog by choosing an approved data route first, then collecting, matching, validating, monitoring, and exporting product records.
To collect supplier product data, first ask for an approved supplier feed, API, portal export, or applicable data-pool connection. Use website scraping only when the source permits your intended access and no suitable structured route is available. Then preserve each source record, match products by stable identifiers where possible, validate fields, compare each refresh with the previous accepted version, and export a documented catalog with retrieval timestamps.
GS1 GDSN supports product-data exchange through data pools and synchronization between participating trading partners; participation and item coverage are not universal. GS1 US also documents API and bulk-data workflows, but the available fields and access conditions depend on the service and subscription. GS1 GDSN · GS1 data services
1. Choose the right supplier-data route
Start with the route that supplies the fields you need under terms that allow your intended storage and use. Compare coverage, field completeness, update latency, identifier quality, permitted use and redistribution, integration work, and total cost. There is no universal best route.
| Route | Useful when | Check before you build |
|---|---|---|
| Supplier feed or GS1 GDSN data pool | A supplier participates and recurring synchronization matters. | Supplier and item coverage, schema, attributes, update behavior, subscription, and terms. |
| Supplier or registry API | You need structured queries and repeatable integration. | Authentication, rate limits, fields, bulk support, price, license, geography, and storage or redistribution rights. |
| Supplier portal export | A one-time or periodic download is sufficient. | Format, field selection, record limits, refresh process, and terms. |
| Website extraction | No suitable approved structured source is available and the intended access is permitted. | Current terms, robots.txt instructions, technical restrictions, request volume, content rights, and applicable law. |
GDSN is built for product information exchange through certified data pools. Ask the supplier whether it participates and whether the items you need are published. GS1’s GDSN overview describes supplier data being shared and synchronized with trading partners.
Ask suppliers these questions
- Can you provide a feed, API, portal export, or data-pool connection for these products?
- Which identifiers and fields are included, and what does each field mean?
- How often is the data updated, and how are deletions or corrections represented?
- May we store, transform, display, and redistribute the data for our intended use?
- What authentication, rate limits, record limits, fees, and notification mechanisms apply?
- Can you provide a sample payload and a versioned schema?
2. Define the catalog fields and identity rules
Decide what “a product record” means in your system before ingestion. A retail item, a case, and a pallet may have related names but distinct identifiers and packaging levels. Matching on a title alone can merge variants or package sizes.
| Field | Purpose | Validation examples |
|---|---|---|
| Supplier ID and supplier SKU | Preserve the supplier’s identity and key. | Non-empty; uniqueness scoped to supplier if that is how the supplier defines it. |
| GTIN or other stable identifier | Help match the same trade item across systems. | Preserve as a string, including leading zeroes; validate format and source. |
| Brand, title, description | Support search, listing, and human review. | Normalize whitespace; retain original source text as evidence. |
| Variant and packaging attributes | Distinguish size, color, count, and packaging level. | Validate units and compare like packaging levels. |
| Dimensions and net content | Support logistics and catalog display. | Store numeric value and unit separately; do not infer missing units. |
| Images and availability | Support merchandising and purchasing workflows. | Record source URL and retrieval outcome; define what availability values mean. |
| Source, retrieval time, schema version | Make values traceable and staleness visible. | Required for each accepted record or field update. |
Keep source values distinct from normalized values. A useful record includes supplier, original source identifier, source URL or feed name, retrieval time, schema version, raw response or snapshot where permitted, normalized fields, and validation result. This is practical pipeline guidance rather than a required GS1 record format.
GS1’s Verified by GS1 can help check whether a GS1 identifier is structured correctly and which company is associated with it. An identity lookup is not a full product catalog or proof that every product field is complete and current.
3. Set access and crawl boundaries before scraping
For a website source, read its current terms and robots.txt, use a clear crawler identity, limit request rates, and stop if access is blocked or requires credentials you do not have. Do not bypass authentication, CAPTCHAs, bot checks, or other access controls. When the permission or permitted use is unclear, ask the supplier or obtain legal review.
RFC 9309 standardizes robots.txt rules that crawlers are requested to honor. It explicitly says those rules “are not a form of access authorization.” A robots.txt check is one input to crawler behavior, not permission to collect, store, or reuse data. RFC 9309 also says crawlers should not use a cached robots.txt version for more than 24 hours unless the file is unreachable; this is not a recommended product-data refresh interval.
Pre-crawl checklist
- Confirm the source and intended use are permitted under current terms and applicable requirements.
- Check robots.txt for the relevant paths and identify your crawler honestly.
- Set a conservative request rate and concurrency limit; honor server errors and Retry-After when present.
- Fetch only the pages and fields necessary for the catalog.
- Stop on access denial, authentication requirements, CAPTCHA, or blocking rather than attempting to evade it.
- Record retrieval time and source URL for every collected record.
4. Example: extract product cards from an authorized page
The following Python example is deliberately limited to a page you are authorized to access. It uses placeholder selectors because every supplier site has a different markup; inspect the permitted page and replace them. It checks robots.txt for a disallow rule as a basic guard, identifies the crawler, makes one request, extracts product fields, and writes JSON Lines. It is not a complete robots.txt parser or a legal authorization check.
python -m pip install requests beautifulsoup4
import json
import time
from datetime import datetime, timezone
from urllib.parse import urljoin, urlparse
from urllib.robotparser import RobotFileParser
import requests
from bs4 import BeautifulSoup
START_URL = "https://supplier.example/catalog"
USER_AGENT = "SupplierCatalogBot/1.0 (contact: data-team@example.com)"
OUTPUT = "supplier-products.jsonl"
session = requests.Session()
session.headers.update({"User-Agent": USER_AGENT})
robots_url = urljoin(START_URL, "/robots.txt")
robots = RobotFileParser()
robots.set_url(robots_url)
robots.read()
if not robots.can_fetch(USER_AGENT, START_URL):
raise SystemExit(f"robots.txt disallows this URL: {START_URL}")
response = session.get(START_URL, timeout=(5, 30))
response.raise_for_status()
if "text/html" not in response.headers.get("Content-Type", ""):
raise SystemExit("Expected an HTML page; use the supplier's documented format instead.")
soup = BeautifulSoup(response.text, "html.parser")
retrieved_at = datetime.now(timezone.utc).isoformat()
# Replace these selectors and attributes after inspecting the authorized page.
records = []
for card in soup.select(".product-card"):
link = card.select_one("a.product-card__link")
title = card.select_one(".product-card__title")
sku = card.select_one("[data-sku]")
price = card.select_one(".product-card__price")
if not title or not link:
continue
records.append({
"supplier": "supplier.example",
"supplier_sku": sku.get("data-sku") if sku else None,
"title_source": title.get_text(" ", strip=True),
"product_url": urljoin(START_URL, link.get("href", "")),
"price_source": price.get_text(" ", strip=True) if price else None,
"source_url": START_URL,
"retrieved_at": retrieved_at,
"validation_status": "needs_review",
})
with open(OUTPUT, "w", encoding="utf-8") as f:
for record in records:
f.write(json.dumps(record, ensure_ascii=False) + "\\n")
print(f"Wrote {len(records)} records to {OUTPUT}")
The sample collects one page only. For pagination, follow only links that are in scope, validate each URL against the host and permitted paths, deduplicate records by source key, and add a delay and retry policy appropriate to the supplier’s documented limits. The Python standard library robots parser offers a basic check; if precise RFC-compliant robots handling is a requirement, use a maintained implementation and verify its behavior against the applicable standard.
5. Prefer structured feeds and APIs when available
When a supplier provides CSV, XML, JSON, or an API, ingest that documented format instead of reverse-engineering its product pages. Keep the original file or response where permitted. Parse into a staging table first; validate required columns, encodings, identifiers, units, and schema version before updating the production catalog.
GS1 US describes API-based product, location, and company data workflows, including automated ingestion and bulk search or export capabilities in its official materials. Exact functions and access depend on the selected service and subscription. Ask for current documentation and credentials from the provider; do not assume a public endpoint or a particular schema.
Staging and acceptance sequence
- Download or receive the supplier artifact and record source, retrieval time, and declared schema version.
- Parse into staging without overwriting accepted data.
- Validate field types, required values, identifier format, unit consistency, and duplicate keys.
- Compare the staged record with the last accepted record.
- Automatically accept only changes covered by clear rules; send ambiguous identity or high-impact changes for review.
- Commit the accepted version and produce a dated export with a documented schema.
6. Match, normalize, and validate records
Use the strongest shared identifier available, preserving its source and exact value. Prefer a supplier SKU scoped to that supplier or a verified product identifier over fuzzy title matching. Treat name matching as a candidate-generation aid, not definitive identity. A check that an identifier is valid or associated with a company does not establish that all attributes describe the same package or are current.
Practical validation rules
- Identifiers: store GTINs and SKUs as strings so leading zeroes survive. Keep both the supplied value and any normalized form.
- Units: store amount and unit separately. Convert only when the source unit is clear and preserve the original.
- Variants: compare size, count, flavor, color, and packaging level before merging.
- Images: retain the supplier image URL and retrieval status; do not assume a broken link means the product was removed.
- Missing values: distinguish unknown, not supplied, not applicable, and explicitly cleared values.
- Duplicates: detect duplicate supplier keys and conflicting identifiers before loading.
- Provenance: retain source, retrieval time, and validation outcome so a reviewer can trace a field.
7. Monitor supplier changes and download reliable exports
Monitor by comparing each new accepted candidate with the last accepted record. Track source changes separately from your own normalization changes. The supplier’s update cadence and the cost of stale data should determine refresh frequency; there is no universal polling interval.
| Change | Suggested handling |
|---|---|
| Description or image URL changed | Record the delta and accept or review according to merchandising policy. |
| GTIN or supplier SKU changed | Do not silently overwrite identity; investigate whether it is a correction, replacement, or distinct item. |
| Unit, pack count, or dimensions changed | Flag for review because the packaging level or logistics meaning may have changed. |
| Record absent from this refresh | Do not infer deletion from one missing response; check feed semantics and source coverage. |
| Parse or validation failure | Keep the prior accepted value, record the failure, and alert the owner. |
For downloads, publish a stable schema and include retrieval time, source, and schema version. Common options include CSV for spreadsheet workflows, JSON Lines for record-by-record pipelines, and a database or API for downstream systems. Document whether an export contains the latest source values, last accepted values, or both. Preserve a change log so consumers can tell what changed between exports.
Minimum operational signals
- Last successful retrieval time by supplier and source.
- Records received, accepted, rejected, changed, and missing.
- Schema, parsing, authentication, and rate-limit errors.
- Age of the latest accepted data and whether it exceeds your freshness target.
- Manual review queue size and unresolved identity conflicts.
8. Or skip the browser setup
If your authorized workflow needs page screenshots as visual evidence or for a catalog review, ScreenshotNeo is a website screenshot API and MCP server. A screenshot is not structured product data, so use the supplier’s feed or API for catalog ingestion when available. For pages you are permitted to capture, one GET request returns an image or PDF; see the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await require('node:fs/promises').writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, and failed loads are never billed; responses identify the page verdict and billing status. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, no card required.
9. Troubleshooting
| Symptom | Likely cause | Response |
|---|---|---|
| 403, CAPTCHA, or bot-check page | The site blocks automated access or requires authorization. | Stop the crawl. Ask the supplier for an approved feed, API, or export. Do not evade the control. |
| robots.txt disallows the path | The crawler-facing rules do not permit that path for your agent. | Do not crawl that path; request an approved data route. Robots rules do not grant legal authorization for paths they allow. |
| 200 response but no products parsed | Selectors changed, content is rendered client-side, or the response is a different page. | Inspect the permitted response and markup; prefer the supplier’s structured endpoint. Do not add access-control bypasses. |
| Duplicate or mismatched catalog items | Records were matched by title or ignored packaging and variant fields. | Match with supplier-scoped identifiers or GTIN where available; review variants and packaging levels. |
| Leading zeroes disappear | An identifier was parsed as a number by a spreadsheet or ingestion tool. | Store identifiers as text from the first import, then reimport the original source. |
| Encoding or unit corruption | Wrong character encoding, locale-specific number parsing, or ambiguous units. | Use the supplier’s documented encoding and schema; retain original values and validate conversions. |
| Supplier export no longer parses | Schema or column names changed without an expected version. | Quarantine the file, alert the owner, inspect the schema change, and retain the prior accepted catalog. |
| Timeouts or rate limits | Requests are too frequent, source is slow, or a service limit was reached. | Reduce concurrency, honor Retry-After, use bounded retries with backoff, and ask for bulk access. |
| Missing record mistaken for deletion | The source export is partial, filtered, or temporarily incomplete. | Check the source’s deletion semantics and coverage before removing a catalog item. |
10. Performance, reliability, and cost
- Prefer bulk delivery: a supplier feed or bulk API can reduce per-page requests and parsing work. Check its actual coverage and refresh semantics first.
- Limit crawl load: set a low, documented request rate and concurrency; use timeouts and bounded retries. Do not turn retries into a way to overcome blocking.
- Make ingestion restartable: checkpoint pages or batches, deduplicate by source key, and keep raw permitted inputs so a parser fix does not require refetching everything.
- Protect the accepted catalog: stage and validate before committing. On source failure, retain the last accepted values and mark freshness rather than replacing good records with blanks.
- Estimate total cost: include supplier or data-pool subscriptions, API access, integration and maintenance, storage, review time, and the operational cost of stale or incorrect data. No universal cost or refresh benchmark applies.
- Plan for drift: websites and feeds change. Monitor schema and selector failures, record provenance, and assign an owner for supplier and parser changes.
11. Frequently asked questions
Can I use Verified by GS1 as my supplier catalog?
No. It can help verify GS1 identity information, such as whether an identifier is structured correctly and which company is associated with it. The returned information is not necessarily a complete product record. Verified by GS1
Does a permissive robots.txt mean I can scrape a site?
No. RFC 9309 defines crawler instructions and says they are not access authorization. Check the applicable terms, access controls, contracts, and legal requirements for your situation.
How often should supplier data refresh?
Use the supplier’s stated update cadence and the business impact of stale values. Separate source refresh frequency from the urgency with which your team reviews changes.
Should I overwrite the catalog when a supplier field is blank?
Only if the source defines blank as an intentional clearing operation. Otherwise preserve the accepted value, record the missing input, and investigate.
Is a screenshot enough to extract product data?
No. A screenshot can serve as visual evidence, but it does not provide reliable structured fields for matching, validation, or catalog export. Use a supplier feed, API, portal export, or permitted structured extraction.


