ScreenshotNeo

BlogHow-to

How to Bulk Capture Product Page Screenshots from Indian D2C Stores

Capture consistent product page screenshots across Indian D2C stores with a reusable Playwright workflow, reliable filenames, and clear retry handling.

By the ScreenshotNeo team4 October 202610 min read

To bulk capture product page screenshots from Indian D2C stores, prepare a list of product URLs, then use a browser automation script to visit each URL, wait for a consistent page condition, capture the viewport, full page, or a chosen element, and save each result under a stable filename. Playwright provides the screenshot operations; the loop, naming, retries, and results log are your batch workflow.

This guide uses Node.js and Playwright for local capture. It also covers cURL, Python, and Node.js calls to ScreenshotNeo for hosted capture. Store-specific behavior varies, so check representative captures for consent prompts, variant state, lazy-loaded images, and access restrictions before running a large collection.

1. Decide what each screenshot should show

Choose one capture scope and keep it consistent across stores and products:

  • Viewport: captures the visible first screen at the configured browser size. Use it to compare above-the-fold layout.
  • Full page: captures the full scrollable page as one image. Use it to review a long product detail page.
  • Element: captures a selected component, such as the product gallery or purchase panel. Use it when that component is the comparison target.

Playwright documents viewport, element, and full-page screenshots. Its screenshot API does not provide a turnkey URL-list batch feature; the script below supplies the iteration. Full-page capture and element targeting are separate modes: do not combine them in one screenshot call.

Also decide the viewport dimensions, image format, and pixel scale. For visual comparisons, use the same settings for every URL. PNG is lossless and can be large; JPEG and WebP are supported alternatives. Playwright uses CSS-pixel sizing by default; device scale captures at the device pixel ratio and can produce larger files.

2. Prepare and label your URL list

Put one absolute product-page URL on each line in urls.txt. Use a stable output name based on the store, product slug, and capture date. The example creates a sanitized name from each URL and keeps a separate JSON Lines results file, so failures are visible and can be retried.

https://example.in/products/linen-shirt
https://shop.example.in/products/face-serum

Before a batch, remove duplicates and confirm the URLs point to the product pages you intend to compare. If a site requires a particular variant, locale, or region, include the relevant page state in your capture plan; a URL alone may not select it.

3. Install Playwright and capture locally

Install Node.js, create a project, add Playwright, and install its Chromium browser:

mkdir d2c-captures
cd d2c-captures
npm init -y
npm install playwright
npx playwright install chromium

Save this as capture.mjs beside urls.txt. It captures each page at a fixed viewport, waits for the page load event and then a short configurable settling period, records success or failure, and continues after individual errors. The default is a viewport PNG. Set FULL_PAGE=1 for a full-page image, or set SELECTOR to capture one element instead.

import { chromium } from 'playwright';
import { readFile, mkdir, appendFile } from 'node:fs/promises';

const urls = (await readFile('urls.txt', 'utf8'))
  .split(/\r?\n/)
  .map((line) => line.trim())
  .filter((line) => line && !line.startsWith('#'));

const outputDir = 'captures';
const resultsFile = 'results.jsonl';
const width = Number(process.env.WIDTH || 1440);
const height = Number(process.env.HEIGHT || 1000);
const settleMs = Number(process.env.SETTLE_MS || 1500);
const fullPage = process.env.FULL_PAGE === '1';
const selector = process.env.SELECTOR || '';
const format = process.env.FORMAT || 'png';
const scale = process.env.SCALE === 'device' ? 'device' : 'css';

if (!['png', 'jpeg'].includes(format)) {
  throw new Error('FORMAT must be png or jpeg');
}
if (!Number.isFinite(width) || !Number.isFinite(height) || width < 1 || height < 1) {
  throw new Error('WIDTH and HEIGHT must be positive numbers');
}
if (fullPage && selector) {
  throw new Error('Choose FULL_PAGE or SELECTOR; they cannot be combined');
}

function safeName(url, index) {
  const parsed = new URL(url);
  const slug = parsed.pathname.split('/').filter(Boolean).join('-') || 'home';
  const base = `${parsed.hostname}-${slug}`.toLowerCase()
    .replace(/[^a-z0-9.-]+/g, '-')
    .replace(/-+/g, '-')
    .slice(0, 140);
  return `${String(index + 1).padStart(3, '0')}-${base}`;
}

await mkdir(outputDir, { recursive: true });
const browser = await chromium.launch({ headless: true });
const context = await browser.newContext({ viewport: { width, height }, deviceScaleFactor: 1 });

try {
  for (const [index, url] of urls.entries()) {
    const page = await context.newPage();
    const startedAt = new Date().toISOString();
    const file = `${outputDir}/${safeName(url, index)}.${format}`;
    try {
      const response = await page.goto(url, { waitUntil: 'load', timeout: 45000 });
      await page.waitForTimeout(settleMs);
      const options = {
        path: file,
        type: format,
        scale,
        fullPage,
      };
      if (selector) {
        const target = page.locator(selector).first();
        await target.waitFor({ state: 'visible', timeout: 15000 });
        await target.screenshot({ ...options, fullPage: undefined });
      } else {
        await page.screenshot(options);
      }
      await appendFile(resultsFile, JSON.stringify({
        url, file, startedAt, status: 'ok', httpStatus: response?.status() ?? null,
      }) + '\n');
    } catch (error) {
      await appendFile(resultsFile, JSON.stringify({
        url, file, startedAt, status: 'error', error: String(error),
      }) + '\n');
    } finally {
      await page.close();
    }
  }
} finally {
  await context.close();
  await browser.close();
}

Run the default viewport capture:

node capture.mjs

Examples of alternative settings:

# Full scrollable page
FULL_PAGE=1 node capture.mjs

# Capture just the product gallery
SELECTOR='.product-gallery' node capture.mjs

# A fixed viewport and device-pixel scale
WIDTH=1280 HEIGHT=900 SCALE=device node capture.mjs

# JPEG output
FORMAT=jpeg node capture.mjs

The script uses a fixed settling delay as a simple baseline, not proof that every image or interactive component has finished rendering. For pages with lazy content, inspect samples and use a site-appropriate readiness condition. If the desired content appears only after scrolling, add a deliberate scroll-and-wait step before taking the screenshot. Do not assume all Indian D2C stores use the same markup or loading behavior.

4. Choose output format, scale, and naming

Choice Use it when Trade-off
PNG You need crisp text and interface details Files can be larger
JPEG You prefer smaller photographic captures Compression can soften fine edges
WebP Your capture tool supports it and your downstream tools accept it Confirm that your image viewer and comparison pipeline handle it
CSS scale You want dimensions tied to the CSS viewport It may not match a high-density device capture
Device scale You need pixels at the device pixel ratio Images may be larger; keep the ratio consistent

Playwright documents PNG, JPEG, and WebP screenshot output and notes that an extension can determine the format, with PNG as the default when none determines it. The local sample explicitly supports PNG and JPEG to keep its filename and encoding settings aligned; add WebP only if your installed Playwright version and downstream workflow support it. Keep the original URL in the results log even when filenames are sanitized, since different URLs can reduce to similar names.

5. Make captures comparable across stores and dates

  • Use the same viewport width, height, pixel scale, browser version, and capture mode.
  • Choose and record a consistent page readiness rule. A page load event alone may not mean product images or client-rendered content are ready.
  • Decide how to handle cookie prompts, newsletter overlays, chat widgets, and variant selectors. Record whether the capture shows them; do not silently treat different states as directly comparable.
  • Review a sample from each store before scaling up. Check for blank images, incomplete galleries, unexpected overlays, or a wrong selected variant.
  • Keep the date and original URL with each image. Product pages change, and dated names help distinguish captures.
  • For automated access, check the relevant site’s rules and any access controls. The available research does not establish policies for specific stores.

Playwright notes that operating system, browser version, settings, hardware, power source, and headless mode can affect rendering. If you compare images over time, keep the environment stable and interpret pixel differences with those sources of variation in mind.

Consent dialogs, promotional popups, chat widgets, variant selectors, and lazy-loaded images can change what appears in a capture. There is no single selector or wait rule that applies to every store. For a local workflow, inspect the page and choose a deliberate interaction or visibility rule for the sites you are permitted to automate. If a site presents a bot check, CAPTCHA, access restriction, or terms notice, do not attempt to evade it; follow the site’s requirements and use an allowed capture path.

Keep a record of pages that failed or showed an unexpected state, then retry only after addressing the cause. A screenshot of an error page is not a successful product-page capture. For comparison sets, note whether a capture includes the consent state and selected variant so that later reviewers can understand what they are seeing.

7. Hosted capture with ScreenshotNeo

ScreenshotNeo is a website screenshot API and MCP server by Yorker Media. It accepts one GET request with a URL and returns an image or PDF. You can loop over a URL list in your own script, save each response, and retain the per-response verdict and billing headers. See the ScreenshotNeo API documentation for request options and setup, and ScreenshotNeo for the product overview.

Here is a cURL call for one product page; repeat it for each URL in your list, using a distinct output filename:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python example using requests:

import requests

url = "https://stripe.com"
r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": url},
    timeout=90,
)
r.raise_for_status()
with open("shot.webp", "wb") as f:
    f.write(r.content)
print(r.headers.get("X-Page-Verdict"), r.headers.get("X-Billed"))

Node.js example:

const target = 'https://stripe.com';
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: target });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
console.log(res.headers.get('X-Page-Verdict'), res.headers.get('X-Billed'));

Replace the example target with each product URL and generate a unique filename from your store/product identifier. For bulk capture, ScreenshotNeo also supports bulk capture of up to 100 URLs per call. It has full-page capture with lazy images loaded, element capture by CSS selector, custom CSS and JavaScript, click-before-capture, selector/delay/network-idle waits, viewport and device presets, dark mode, retina scale, request blocking, headers, cookies, user agent, authorization, timezone, geolocation, transparent backgrounds, resizing, configurable caching, signed image links, async jobs with signed webhooks, PDF options, a usage API, and an OpenAPI spec. The same parameter names used by other screenshot APIs also work, which can make switching easier.

ScreenshotNeo removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing. Responses identify the page verdict and billing state in X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

Plans are Free: 1,000 shots per month with no card; Starter: $5 for 3,000; Growth: $15 for 15,000; Pro: $39 for 60,000; Scale: $99 for 250,000; and Business: $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan. See the docs for current request details. Sign up for 1,000 free screenshots a month with no card.

8. Performance, reliability, and cost

Local Playwright avoids sending each capture request to a separate hosted screenshot API, but you manage the browser installation, execution, output files, and retries. Running several pages concurrently can reduce elapsed time, but it also increases browser and memory load; start sequentially, then increase concurrency carefully while checking that pages still render correctly. No benchmark is established here, so measure your own URL set and machine if throughput matters.

Hosted capture moves browser setup out of your script. ScreenshotNeo bills only clean shots and exposes verdict and billing headers, so your batch can distinguish billable captures from failed or unsuitable pages. Its bulk endpoint accepts up to 100 URLs per call. For either route, set timeouts, retain a result record for every URL, and retry selectively rather than losing track of partial failures.

9. Troubleshooting

Symptom Likely cause What to do
Browser executable missing Playwright package is installed but Chromium is not Run npx playwright install chromium.
Navigation timeout The page did not reach the selected load condition in time, or the site is slow or inaccessible Check the URL manually, increase the navigation timeout for that case, and log the failure. Avoid treating a timeout image as a valid product capture.
Screenshot is blank or incomplete Content may be client-rendered, lazy-loaded, blocked, or still loading Inspect the page state, choose a suitable readiness condition, and scroll or wait for the relevant content before capture.
Element selector not found The selector differs across stores or the component has not appeared Inspect that page’s markup, use a store-specific selector, and wait for visibility.
Full-page and selector options conflict They describe different capture scopes Choose either full page or one element for that capture.
Images vary between runs Viewport, device scale, browser environment, page state, or timing changed Keep the environment and settings stable and record capture conditions.
Hosted response is not a usable image The target may have returned a bot check, blank page, timeout, or failed load Read X-Page-Verdict and X-Billed, inspect the target and request settings, and retry only when appropriate.
Output was overwritten Multiple URLs produced the same sanitized filename Include a unique index or stable product identifier in each filename and retain the source URL in the log.

10. FAQ

How do I take screenshots of multiple product pages at once?

Use a URL list and a loop or a batch-capable endpoint. The local Playwright script above processes a list sequentially; ScreenshotNeo’s bulk capture accepts up to 100 URLs per call.

Should I use a viewport or full-page screenshot?

Use a viewport for first-screen layout comparisons, full page for overall page structure, and element capture for a specific component. Keep the choice consistent within a comparison set.

Will every Indian D2C store work the same way?

No uniform behavior is established. Check the intended URLs for their own loading state, overlays, variants, access requirements, and automated browsing rules.

Can I compare screenshots captured on different machines?

You can, but browser rendering can vary with the operating system, browser version, settings, hardware, power source, and headless mode. Keep the capture environment stable where pixel-level comparison matters.