How to Capture Full-Page Screenshots of Multiple Indian Ecommerce Product Pages in Bulk
Capture full-page screenshots of Indian ecommerce product pages in bulk with Playwright, consistent filenames, page-specific waits, and a failure manifest.
To capture full-page screenshots of many Indian ecommerce product pages, prepare a URL list with stable product IDs, use a browser automation script to visit each authorized page, wait for the specific content you need, save a full-page image, and record each result in a manifest. Playwright supports this with page.screenshot({ fullPage: true }). Full-page capture includes the scrollable page, but it does not guarantee that lazy-loaded images or dynamic content have finished loading.
This guide uses Playwright with Node.js, then covers cURL, Python, and a hosted API option. Use only pages you are authorized to access, and check each marketplace’s current terms and access controls before running a batch.
1. Choose what to capture
| Mode | What it captures | Use it when |
|---|---|---|
| Viewport | The currently visible browser area | You need comparable above-the-fold previews. |
| Full page | The whole scrollable page as one tall image | Below-the-fold product details, reviews, or layout matter. |
| Element | A selected element, such as the product details region | The entire page is unnecessary or too long for the review. |
Playwright documents page screenshots, full-page capture, and locator screenshots. It also supports PNG, JPEG, and WebP output, CSS- or device-pixel scale, masking, and disabling animations. Choose settings that serve the comparison; the sources do not establish one universally best format. See the Playwright screenshots documentation.
2. Prepare a stable input list
Keep a stable row ID or SKU beside every URL. Use the ID for filenames because product titles can change or contain characters unsuitable for paths. Remove accidental duplicates, check that URLs are well-formed, and include only targets you are authorized to visit.
id,url
SKU123,https://www.example.in/product-one
SKU456,https://www.example.in/product-two
The example domains and SKUs above are placeholders. Replace them with your own authorized product URLs and identifiers.
3. Run a Playwright batch capture
The following runnable Node.js script reads a CSV with id,url columns, visits URLs sequentially, waits for a page-specific readiness selector, saves full-page PNG files, and writes a JSON Lines manifest containing success or failure for every input row.
Install and set up
mkdir bulk-capture
cd bulk-capture
npm init -y
npm install playwright
npx playwright install chromium
Save the input as urls.csv in this directory. Save this script as capture.mjs:
import fs from 'node:fs';
import path from 'node:path';
import { chromium } from 'playwright';
const inputPath = process.argv[2] ?? 'urls.csv';
const outputDir = process.argv[3] ?? 'screenshots';
const manifestPath = path.join(outputDir, 'manifest.jsonl');
const readinessSelector = process.env.READY_SELECTOR ?? 'body';
const navigationTimeoutMs = Number(process.env.NAVIGATION_TIMEOUT_MS ?? 45000);
const selectorTimeoutMs = Number(process.env.SELECTOR_TIMEOUT_MS ?? 15000);
function parseCsvLine(line) {
// This compact parser is suitable for simple id,url files without quoted commas.
const comma = line.indexOf(',');
if (comma < 1) throw new Error(`Expected id,url row: ${line}`);
return { id: line.slice(0, comma).trim(), url: line.slice(comma + 1).trim() };
}
const lines = fs.readFileSync(inputPath, 'utf8').split(/\r?\n/).filter(Boolean);
if (lines.length < 2 || lines[0].trim() !== 'id,url') {
throw new Error('CSV must start with the header id,url and contain at least one row');
}
const rows = lines.slice(1).map(parseCsvLine);
const seen = new Set();
for (const row of rows) {
if (!row.id || !row.url) throw new Error(`Missing id or URL: ${JSON.stringify(row)}`);
if (seen.has(row.id)) throw new Error(`Duplicate id: ${row.id}`);
seen.add(row.id);
const parsed = new URL(row.url);
if (!['http:', 'https:'].includes(parsed.protocol)) throw new Error(`Unsupported URL protocol: ${row.url}`);
}
fs.mkdirSync(outputDir, { recursive: true });
fs.writeFileSync(manifestPath, '');
const browser = await chromium.launch({ headless: true });
const context = await browser.newContext({ viewport: { width: 1440, height: 1000 }, deviceScaleFactor: 1 });
const page = await context.newPage();
try {
for (const { id, url } of rows) {
const startedAt = new Date().toISOString();
const file = path.join(outputDir, `${id.replace(/[^a-zA-Z0-9_-]/g, '_')}.png`);
const record = { id, url, startedAt, file, status: 'failed' };
try {
const response = await page.goto(url, { waitUntil: 'domcontentloaded', timeout: navigationTimeoutMs });
record.httpStatus = response?.status() ?? null;
await page.locator(readinessSelector).first().waitFor({ state: 'visible', timeout: selectorTimeoutMs });
// Optional site-specific wait can be added here for a product image or price selector.
await page.screenshot({ path: file, fullPage: true, type: 'png', animations: 'disabled' });
record.status = 'success';
} catch (error) {
record.error = error instanceof Error ? error.message : String(error);
}
record.finishedAt = new Date().toISOString();
fs.appendFileSync(manifestPath, `${JSON.stringify(record)}\n`);
console.log(`${record.status}: ${id}${record.error ? ` — ${record.error}` : ''}`);
}
} finally {
await context.close();
await browser.close();
}
Run it with:
node capture.mjs urls.csv screenshots
The default readiness selector is body, which only confirms that the document body is visible. For meaningful product readiness, set a selector that appears when the content under review is present. The selector varies by site and page template:
READY_SELECTOR='main' node capture.mjs urls.csv screenshots
READY_SELECTOR='[data-testid="product-title"]' node capture.mjs urls.csv screenshots
Those selectors are examples, not claims about any particular marketplace’s markup. Inspect the authorized page and choose a selector that exists there. If the site uses several templates, make readiness configurable per input row or group the URLs by template.
4. Wait for the content that matters
fullPage: true controls the captured area. It does not scroll through the page to trigger every lazy-loaded element or prove that dynamic content is complete. Wait for the specific product image, price, or other relevant element, or add a measured delay when the page provides no reliable signal. Verify a sample from each site and page type before processing the full list.
Possible readiness strategies include:
- Selector: wait for the product title, main image, or price element to become visible.
- Response: inspect navigation status and handle error pages or redirects in your workflow.
- Short delay: allow deferred rendering time if no stable selector is available, while recognizing that a delay is less reliable than a page-specific signal.
- Network idle: use only when appropriate for the site. Pages with analytics or persistent requests may never become idle.
Do not assume a successful browser navigation means the intended product page rendered. Login gates, bot checks, consent dialogs, unavailable listings, or regional variants can produce a screenshot of a different state.
5. Make the batch repeatable
- Keep settings fixed. Use one viewport, device scale, capture mode, and image format for comparable outputs.
- Name by stable ID. Keep the mapping between SKU, URL, and filename in the manifest.
- Record outcomes. The example manifest records ID, URL, timestamps, output path, HTTP status when available, and errors.
- Review representative pages. Check at least one result per site and template for missing deferred content, blocked navigation, unusual page state, or excessive image height.
- Rerun failures selectively. Use the manifest to build a retry list instead of repeating successful work.
For large lists, keep concurrency low at first. Sequential capture makes failures easier to trace and reduces simultaneous browser load. If you introduce parallel pages, tune concurrency to your machine and the target sites’ allowed request rates; the research does not establish a universal safe or fastest value.
6. Capture a selected product element instead
When the full page is too long or includes irrelevant recommendations and footer content, capture a locator. Replace the screenshot line in the script with a locator that matches the relevant element:
const productPanel = page.locator('[data-testid="product-details"]');
await productPanel.waitFor({ state: 'visible', timeout: selectorTimeoutMs });
await productPanel.screenshot({ path: file, type: 'png', animations: 'disabled' });
The selector is illustrative and must be adapted to the page. Locator screenshots capture the selected element rather than the entire scrollable document.
7. cURL, Python, and Node.js with ScreenshotNeo
If you prefer a hosted screenshot API, ScreenshotNeo accepts one GET request per URL and returns an image or PDF. For a batch, loop over your URL and ID list, save each response using its stable ID, and write a manifest as in the browser workflow. See the ScreenshotNeo API documentation for request options.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
These one-request examples target stripe.com as provided. Replace the target with a product URL you are authorized to capture. For bulk use, loop over your rows, check each response, save to an ID-based filename, and record the response outcome. ScreenshotNeo accepts the parameter names used by other screenshot APIs, which can make an existing integration easier to adapt.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. One GET request returns a PNG, JPEG, WebP, or PDF. Cookie banners and consent prompts are accepted like a visitor and more than 60 known consent platforms, newsletter popups, and chat widgets are removed before the shot; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. AI agents using Claude, Cursor, or another MCP client can use its take_screenshot, get_page_info, and capture_pdf tools.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Loop this call over your authorized URL list and use each stable ID for the output filename. ScreenshotNeo provides 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000 screenshots. Every feature is available on every plan. Read the API documentation and sign up for 1,000 free screenshots a month with no card.
8. No-code batch capture
FullPage Capture’s vendor page describes a Chrome and Edge extension workflow that accepts addresses one per line or open tabs, then captures them in a queue. The vendor states that a run can include up to 200 pages, that captures are stored in capture history, and that downloads can be individual files or a combined PDF. The vendor page says batch capture is part of Pro and was updated September 29, 2026. These are vendor claims; recheck capacity, plan status, and behavior before relying on them. See the FullPage Capture vendor site.
A browser script offers more control over readiness, stable filenames, and per-page error handling. An extension may suit a one-off batch when its stated workflow meets the need. These are workflow tradeoffs based on documented capabilities, not independent product test results.
9. Performance, reliability, and cost
- Processing time: it depends on navigation time, page weight, readiness waits, screenshot height, and concurrency. No independent benchmark establishes a general pages-per-minute figure.
- Browser resources: full-page captures of long product pages can use more memory and produce large files. Start sequentially, close pages or contexts cleanly, and use element captures when the whole page is not needed.
- Reliability: navigation and screenshot errors should be recorded per row. A manifest makes missing captures visible and supports targeted retries. Inspect representative outputs because automated success does not prove the correct page state.
- Output size: PNG is lossless and can be large; JPEG or WebP may reduce storage, depending on content and quality settings. The sources do not establish a single best format.
- Hosted API cost: calculate volume from your expected clean captures and the provider’s published plan. ScreenshotNeo lists 1,000 free shots per month, then plans of $5 for 3,000, $15 for 15,000, $39 for 60,000, $99 for 250,000, and $249 for 1,000,000; yearly billing gives two months free. Every feature is on every plan. Only clean shots are billed, according to the product facts; use the response billing headers when reconciling requests.
10. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Screenshot shows only a loading state | The script waited for document navigation but not the product content. | Set a page-specific readiness selector and verify it on each page template. |
| Images or below-the-fold content are missing | Lazy-loaded elements did not load before capture. | Wait for the relevant image or content selector; inspect representative full-page output. |
| Selector wait times out | The selector is wrong, differs by template, or the intended page did not load. | Inspect the authorized page’s markup, make selector configuration template-specific, and check redirects or access gates. |
| Navigation times out | The page is slow, blocked, or has ongoing activity. | Check whether the target is reachable and authorized, prefer a useful readiness signal over waiting for every request, and adjust timeout only when justified. |
| Some rows have no file | Navigation or screenshot failed, or the output directory was unavailable. | Read the manifest error, confirm write permissions and disk space, then rerun only failed IDs. |
| Names collide or files overwrite | IDs were duplicated or unsafe characters collapsed to the same filename. | Validate unique IDs and use a collision-resistant filename mapping. |
| Very tall image is hard to review | The full page includes long recommendation, review, or footer sections. | Capture the product element or use a consistent viewport capture for the comparison. |
| API response is not the expected image | The request may have returned an error or a page verdict that needs inspection. | Check the response status and ScreenshotNeo’s X-Page-Verdict and X-Billed headers; consult the API documentation. |
11. Marketplace access and permissions
The available research does not establish whether bulk screenshotting is permitted by Amazon.in, Flipkart, or any other specific marketplace. Check each site’s current terms, access controls, and applicable authorization before capture. A successful browser or API response is not evidence of permission. Keep request volume within the access conditions that apply to your use.
12. Frequently asked questions
Does full-page capture scroll the page first?
Playwright describes full-page capture as capturing the scrollable page as if it fit on a very tall screen. That option alone does not guarantee that lazy content has been triggered or loaded.
Can I produce one PDF for all product pages?
The cited Playwright screenshot workflow saves individual image files. FullPage Capture’s vendor says its extension can download a combined PDF. ScreenshotNeo supports PDF output, but a single PDF assembled from a batch depends on your chosen workflow.
Should I use a fixed delay or a selector?
A selector that represents the content you need is usually a clearer readiness condition. A delay can be a fallback for pages without a stable signal, but it can still be too short or unnecessarily long.
How should I compare pages from different marketplaces?
Use the same viewport, device scale, capture mode, and output format where possible, and retain the URL and capture time in the manifest. Different sites may still render personalized or region-specific states.


