How to Scrape Cdiscount Product Pages
Learn how to access Cdiscount product data responsibly, parse JSON-LD with Python, handle rendering failures, and choose the right API route.
Short answer: If you sell on Cdiscount, start with the documented Octopia Marketplace API for seller operations. For an authorized research or analysis workflow, inspect product pages, prefer structured data such as JSON-LD when present, parse only fields you are allowed to collect, and retain the URL and retrieval time with every record. The available sources do not establish whether independent automated scraping or downstream reuse is permitted, so check the current terms and obtain authorization before collecting data.
1. Choose the right access route
For sellers: use the official Marketplace API
Cdiscount’s Marketplace API documentation describes an Octopia-provided API for seller activities including listing creation, offer updates, orders, customer relations, and financial statements. Seller credentials are obtained through the seller account and settings. Cdiscount’s seller FAQ also describes catalogue management and stock updates through Seller Space or the API. This evidence concerns seller workflows; it does not establish a general public, read-only product-data API for arbitrary pages.
Start with the Cdiscount Marketplace API documentation and the seller FAQ if your goal is to manage your own catalogue or offers.
For authorized page analysis: inspect and parse
If you have authorization to collect rendered product-page fields, use a narrow workflow:
- Fetch only URLs and fields covered by your authorization.
- Inspect the response for JSON-LD before depending on presentation markup.
- Parse values that are present; do not silently guess missing prices, currencies, or availability.
- Store the source URL and collection timestamp with each record.
- Detect consent pages, block pages, errors, and empty results. Stop or back off when access is refused.
- Recheck the parser against current pages because page markup is not a stable API contract.
2. Check permission before automating
Public visibility and a successful HTTP response do not by themselves answer whether automated collection, request rates, or reuse are allowed. The official Cdiscount conditions page lists current conditions and version history, but the material reviewed here does not resolve independent scraping rights. Read the terms that apply to your account and use case, and obtain appropriate authorization where needed.
Cdiscount’s privacy and cookie notice describes consent-dependent tracking that can include pages viewed, product references, searches, and cart contents. Do not collect or reuse personal, account, or tracking data merely because it is technically observable.
3. Inspect a product page before writing a parser
Crawlbase reports that Cdiscount product pages usually include a JSON-LD block containing fields such as name, price, currency, and availability. Treat this as a vendor observation and a starting hypothesis, not a guarantee from Cdiscount. A page may instead return a consent screen, a bot check, an error page, or markup with missing fields.
For each representative URL, save a response sample in an authorized environment and answer:
- Is the response the product page or an interstitial?
- Which JSON-LD objects are present?
- Are prices numeric, strings, or localized text?
- Is availability represented as a schema.org value, visible text, or absent?
- Does the page require JavaScript rendering?
4. A tolerant Python parser
The following example fetches one URL, scans every JSON-LD block, and extracts values when available. It intentionally fails safely when a field is absent. Install dependencies with python -m pip install requests beautifulsoup4. Review the target site’s applicable terms and your authorization before running it.
import json
from datetime import datetime, timezone
from decimal import Decimal, InvalidOperation
import requests
from bs4 import BeautifulSoup
URL = "https://www.cdiscount.com/"
HEADERS = {"User-Agent": "authorized-research-client/1.0"}
response = requests.get(URL, headers=HEADERS, timeout=30)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
records = []
for tag in soup.select('script[type="application/ld+json"]'):
try:
value = json.loads(tag.string or tag.get_text())
except json.JSONDecodeError:
continue
values = value if isinstance(value, list) else [value]
for item in values:
if isinstance(item, dict):
records.append(item)
product = next((x for x in records if x.get("@type") in ("Product", ["Product"])), None)
if not product:
raise RuntimeError("No Product JSON-LD found; inspect the response for a consent or block page.")
offers = product.get("offers") or {}
if isinstance(offers, list):
offers = offers[0] if offers else {}
def decimal_or_none(value):
try:
return str(Decimal(str(value)))
except (InvalidOperation, TypeError, ValueError):
return None
result = {
"url": URL,
"collected_at": datetime.now(timezone.utc).isoformat(),
"name": product.get("name"),
"price": decimal_or_none(offers.get("price")),
"currency": offers.get("priceCurrency"),
"availability": offers.get("availability"),
"sku": product.get("sku"),
}
print(json.dumps(result, ensure_ascii=False, indent=2))
5. Equivalent requests with cURL and Node.js
cURL: inspect the raw response
curl --fail-with-body --max-time 30 \
-A 'authorized-research-client/1.0' \
'https://www.cdiscount.com/' \
-o cdiscount-page.html
Node.js: fetch and scan JSON-LD
const res = await fetch('https://www.cdiscount.com/', {
headers: { 'user-agent': 'authorized-research-client/1.0' },
signal: AbortSignal.timeout(30000)
});
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const html = await res.text();
const scripts = [...html.matchAll(/<script[^>]+type=["']application\/ld\+json["'][^>]*>([\s\S]*?)<\/script>/gi)];
const jsonLd = scripts.flatMap(([, text]) => {
try { const value = JSON.parse(text); return Array.isArray(value) ? value : [value]; }
catch { return []; }
});
const product = jsonLd.find(x => x && (x['@type'] === 'Product' || x['@type']?.includes?.('Product')));
console.log(product ?? { warning: 'No Product JSON-LD found' });
6. Rendering, consent pages, and refusals
Crawlbase reports that all successful calls in its observed August 2026 sample used a JavaScript token, and that failures included HTTP 403 responses plus consent or block pages. These are provider-specific observations from a stated period, not a universal success rate or a recommendation to evade access controls.
When a response is not a product page:
- Classify it as consent, bot check, login, error, or unknown.
- Record the status and stop if your authorization does not cover the next action.
- Use a slower, bounded schedule rather than increasing concurrency.
- Do not attempt to bypass a refusal or challenge without explicit authorization.
7. Data normalization and provenance
| Field | Recommended handling |
|---|---|
| name | Keep the original string; trim surrounding whitespace. |
| price | Parse as a decimal, never binary floating point; retain the original value if conversion fails. |
| currency | Require an explicit currency code; do not infer one from a symbol alone. |
| availability | Store the source value and map it to your own enum separately. |
| url | Store the requested URL and, if available, the final URL after redirects. |
| collected_at | Use UTC. Prices and stock are time-dependent observations. |
8. Performance, reliability, and cost
- Bound concurrency: use a small worker pool and an explicit timeout. Increase only when your authorization and service limits allow it.
- Retry selectively: retry transient network failures with exponential backoff; do not retry authentication failures, 403 responses, or consent pages indefinitely.
- Cache carefully: cache only when freshness requirements permit it, and record the cache time.
- Validate samples: periodically compare parsed fields with the rendered page and alert on sudden missing-field rates.
- Control cost: avoid fetching the same URL repeatedly, request only required fields, and stop on repeated interstitials.
- Preserve evidence: retain status, headers needed for debugging, parser version, URL, and timestamp without storing unnecessary personal data.
9. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| 403 Forbidden | Request refusal or access policy | Stop, review authorization and terms, and contact the site owner if needed. |
| HTML contains no Product JSON-LD | Markup changed, JavaScript rendering, or an interstitial | Save the response, classify it, and update the parser only within the allowed workflow. |
| JSON decode error | Malformed or escaped script content | Skip that block, log its location, and inspect a current sample. |
| Price is null | Offer missing, multiple offers, or localized markup | Preserve null, inspect all offers, and require an explicit currency. |
| Timeouts | Slow rendering or network conditions | Use a bounded timeout, limited retries, and lower concurrency. |
| Repeated consent page | Consent state is required | Follow the authorized consent process or stop; do not bypass it. |
10. Or skip the browser setup
ScreenshotNeo provides a website screenshot API and MCP server. One GET request captures a clean PNG, JPEG, WebP, or PDF. Before capture it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and responses identify the page verdict and billing status.
See the ScreenshotNeo API documentation for the full option list.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://www.cdiscount.com/ -o cdiscount.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://www.cdiscount.com/"}, timeout=90)
open("cdiscount.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://www.cdiscount.com/' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Options include full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets and custom viewports, retina scale, custom CSS and JavaScript, click and wait actions, request blocking, headers, cookies, user agents, timezone and geolocation, resizing, caching, signed links, asynchronous jobs, bulk capture, and PDF output. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
Plans include 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
11. FAQ
Does Cdiscount have a product API?
The documented API evidence covers seller operations. It does not establish a general public read-only API for arbitrary product pages.
Can I treat JSON-LD as a permanent schema?
No. It is a useful inspection target, but page markup can change and fields can be absent.
Should I scrape prices as numbers?
Store a decimal value together with an explicit currency, the original representation, URL, and timestamp.
What should I do when Cdiscount returns a bot check?
Classify the response and stop or follow an explicitly authorized access process. Do not escalate around a refusal.


