ScreenshotNeo

BlogHow-to

How to Bulk Screenshot Pages from an Indian Ecommerce Product URL List

Capture product pages from a URL list with a paced Playwright script, traceable filenames, and clear failure handling—or send each URL to ScreenshotNeo.

By the ScreenshotNeo team4 October 202612 min read

To bulk screenshot product pages from an Indian ecommerce URL list, read the URLs from a file, open each one in a browser, wait for the page state you need, and save a screenshot with a stable filename. The Playwright example below does this sequentially, records each URL and outcome in a manifest, and supports either viewport or full-page captures. Check the rules for the actual destination sites before capturing or redistributing their pages.

This guide uses Node.js and Playwright for the do-it-yourself workflow. It also includes cURL, Python, and Node.js examples for sending individual URLs to ScreenshotNeo when you want a hosted screenshot API.

1. Choose a capture route and define the output

A local Playwright script gives you control over browser setup, waits, naming, and storage. You implement list parsing, pacing, retries, and reporting yourself. A hosted service can handle rendering for each URL; ScreenshotNeo accepts one URL per request, so a bulk workflow calls it once per list entry. Neither route guarantees every destination page will render: access restrictions, bot checks, network failures, and page-specific behavior can affect results.

Choice Use it when Trade-off
Playwright script You need browser-level control or want the capture loop in your own process. You operate the browser and build pacing, retries, storage, and failure reporting.
ScreenshotNeo API You want a hosted URL-to-image request and API options such as full-page capture. You still build the list loop and decide how to store and retry results.
Hosted list workflow You prefer a service workflow or spreadsheet-oriented input. Check the provider’s current documentation for supported list sources, batch limits, output handling, and rate limits before committing to it.

Use viewport capture when the visible first screen is the comparison you need. Use full-page capture when below-the-fold product information matters. Playwright describes full-page mode as capturing the entire scrollable page, rather than only the current viewport. Playwright screenshot documentation

2. Prepare the URL list and filename plan

Save one absolute URL per line in urls.txt. Blank lines and lines beginning with # are ignored by the script below.

https://shop.example.in/product/one
https://shop.example.in/product/two
# Add one product page URL per line

The script creates an output directory and assigns each input a numbered filename, such as 0001-product.png. The number preserves the input order even when two URLs share a path or product name. A JSON Lines manifest records the original URL, timestamp, status, output filename, and error when a capture fails. Keep that manifest with the images so each file remains traceable to its source.

3. Install Playwright

Use a supported Node.js installation, create a working directory, and install Playwright. Install its Chromium browser once for that environment.

npm init -y
npm install playwright
npx playwright install chromium

Save the following as bulk-screenshot.mjs in the same directory. It reads urls.txt by default and writes PNG files to screenshots/.

4. Run a sequential Playwright capture

import { chromium } from 'playwright';
import { mkdir, readFile, appendFile } from 'node:fs/promises';
import path from 'node:path';

const inputFile = process.argv[2] ?? 'urls.txt';
const outputDir = process.argv[3] ?? 'screenshots';
const fullPage = process.env.FULL_PAGE === '1';
const navigationTimeoutMs = Number(process.env.NAVIGATION_TIMEOUT_MS ?? 45000);
const settleDelayMs = Number(process.env.SETTLE_DELAY_MS ?? 1200);
const gapMs = Number(process.env.GAP_MS ?? 1500);
const maxAttempts = Number(process.env.MAX_ATTEMPTS ?? 2);

const urls = (await readFile(inputFile, 'utf8'))
  .split(/\r?\n/)
  .map(line => line.trim())
  .filter(line => line && !line.startsWith('#'));

await mkdir(outputDir, { recursive: true });
const manifestPath = path.join(outputDir, 'manifest.jsonl');
const browser = await chromium.launch({ headless: true });
const context = await browser.newContext({ viewport: { width: 1440, height: 1000 } });

try {
  for (let index = 0; index < urls.length; index++) {
    const url = urls[index];
    const filename = `${String(index + 1).padStart(4, '0')}-product.png`;
    const outputPath = path.join(outputDir, filename);
    const record = { index: index + 1, url, capturedAt: new Date().toISOString(), filename };
    let lastError;

    for (let attempt = 1; attempt <= maxAttempts; attempt++) {
      const page = await context.newPage();
      try {
        const response = await page.goto(url, {
          waitUntil: 'domcontentloaded',
          timeout: navigationTimeoutMs
        });
        // A short settle delay allows ordinary client-side rendering to begin.
        // Replace or supplement it with a site-specific locator when appropriate.
        if (settleDelayMs > 0) await page.waitForTimeout(settleDelayMs);
        await page.screenshot({ path: outputPath, fullPage });
        record.status = 'captured';
        record.httpStatus = response?.status() ?? null;
        record.attempts = attempt;
        lastError = undefined;
        break;
      } catch (error) {
        lastError = error;
        record.attempts = attempt;
      } finally {
        await page.close();
      }
    }

    if (lastError) {
      record.status = 'failed';
      record.error = String(lastError?.message ?? lastError);
    }
    await appendFile(manifestPath, `${JSON.stringify(record)}\n`);
    console.log(`${record.status}: ${url}${record.error ? ` — ${record.error}` : ''}`);
    if (index < urls.length - 1 && gapMs > 0) {
      await new Promise(resolve => setTimeout(resolve, gapMs));
    }
  }
} finally {
  await context.close();
  await browser.close();
}

Run it with the default viewport settings:

node bulk-screenshot.mjs urls.txt screenshots

For full-page captures, increase the timeout, or change pacing, set environment variables for the run:

FULL_PAGE=1 NAVIGATION_TIMEOUT_MS=60000 SETTLE_DELAY_MS=2000 GAP_MS=2500 MAX_ATTEMPTS=3 node bulk-screenshot.mjs urls.txt screenshots

On Windows PowerShell, set variables for the current session before running Node:

$env:FULL_PAGE = "1"
$env:NAVIGATION_TIMEOUT_MS = "60000"
$env:SETTLE_DELAY_MS = "2000"
$env:GAP_MS = "2500"
$env:MAX_ATTEMPTS = "3"
node bulk-screenshot.mjs urls.txt screenshots

The example uses domcontentloaded plus a short delay to avoid waiting indefinitely for pages that keep network connections open. It is a general starting point, not a guarantee that a product page’s images or client-rendered details are ready. For a known site, replace the delay with a relevant selector wait, such as await page.locator('[data-testid="product-title"]').waitFor(), using a selector that actually exists on that site. Navigation and screenshot behavior are documented in the Playwright Page API.

5. Choose the right capture options

Viewport or full page

The default screenshot is the current viewport. Set fullPage: true to capture the full scrollable page. Full-page images can be very tall and large; choose them only when below-the-fold content is relevant. A long page may also load images lazily as it scrolls, so inspect output before treating it as a complete record.

Wait strategy

domcontentloaded means the initial document was parsed; it does not mean every product detail, image, or recommendation widget has finished rendering. A fixed delay is simple but can be too short on a slow page and wasteful on a fast one. A selector wait is more specific when the target element is stable. Network-idle waits can be unsuitable on pages with ongoing analytics or polling. Pick the condition that matches the content you need to capture.

Viewport, device scale, and output

The example uses a 1440 by 1000 CSS-pixel viewport and PNG output. Change the context viewport to match the comparison scenario. If you need mobile product pages, use a mobile viewport and appropriate device settings consistently across the list. Playwright supports screenshot options such as image type and quality for supported formats; consult its screenshot API documentation for the current option details. Keep capture settings consistent if you plan to compare pages.

Cookies, sessions, and region

The script starts with a fresh browser context and no retailer-specific login state. If your authorized use requires a particular session or locale, configure the browser context deliberately and protect any credentials. Do not put secrets directly into the URL list or manifest. The destinations may serve different content based on location, cookies, account state, or stock status; record relevant capture conditions when that context matters.

6. cURL, Python, and Node.js with ScreenshotNeo

ScreenshotNeo is a website screenshot API and MCP server. Its API accepts a URL and returns a screenshot or PDF. For a URL list, loop over the entries and save one response per URL; the examples below show the individual request shape. See the ScreenshotNeo API documentation for the available parameters, including full-page capture and output options.

cURL: capture one URL

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

Replace the example target with a product URL. To process a list from a shell, call cURL once for each line and use a unique output name; keep request pacing and error handling explicit. Avoid putting an API key in a shared script or shell history where others can read it.

Python: capture one URL

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
with open("shot.webp", "wb") as f:
    f.write(r.content)

Node.js: capture one URL

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: HTTP ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));

Node.js: process a URL list with the API

import { readFile, mkdir, writeFile, appendFile } from 'node:fs/promises';

const accessKey = process.env.SCREENSHOTNEO_ACCESS_KEY;
if (!accessKey) throw new Error('Set SCREENSHOTNEO_ACCESS_KEY first');
const urls = (await readFile('urls.txt', 'utf8'))
  .split(/\r?\n/).map(s => s.trim()).filter(s => s && !s.startsWith('#'));
await mkdir('api-shots', { recursive: true });

for (let i = 0; i < urls.length; i++) {
  const url = urls[i];
  const filename = `${String(i + 1).padStart(4, '0')}-product.webp`;
  const q = new URLSearchParams({ access_key: accessKey, url });
  try {
    const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`, { signal: AbortSignal.timeout(90000) });
    if (!res.ok) throw new Error(`HTTP ${res.status}`);
    await writeFile(`api-shots/${filename}`, Buffer.from(await res.arrayBuffer()));
    await appendFile('api-shots/manifest.jsonl', `${JSON.stringify({ url, filename, status: 'saved', capturedAt: new Date().toISOString() })}\n`);
    console.log(`saved: ${url}`);
  } catch (error) {
    await appendFile('api-shots/manifest.jsonl', `${JSON.stringify({ url, status: 'failed', error: String(error), capturedAt: new Date().toISOString() })}\n`);
    console.error(`failed: ${url}: ${error}`);
  }
  if (i < urls.length - 1) await new Promise(resolve => setTimeout(resolve, 1000));
}

Set the key in the environment before running that script: SCREENSHOTNEO_ACCESS_KEY. The API loop is sequential and saves each response as returned; if you need a different format or capture mode, use the documented parameters. For very large lists, consider the API’s async jobs and bulk capture options described in its docs, and retain the URL-to-result manifest.

7. Or skip the browser setup

ScreenshotNeo can render each URL through one API request, so you do not need to install or manage a browser for the capture workflow. For a list, send one request per product URL. The API supports full-page capture and other options; check the docs for parameter names and response details.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Cookie and consent banners are accepted like a visitor and removed, along with known newsletter popups and chat widgets; each cleanup step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers indicate the page verdict and billing status. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.

8. Make bulk runs reliable and polite

  • Start sequentially. The examples process one URL at a time with a gap between requests. Avoid launching a large parallel burst against retailer sites.
  • Set a finite timeout and retry limit. A timeout prevents one stalled page from hanging the entire batch; retries should be bounded, because repeating a blocked or broken URL indefinitely adds load without fixing it.
  • Keep failures in the manifest. Do not silently skip URLs. The manifest lets you retry only failed entries and distinguish a capture failure from a missing input.
  • Make reruns safe. Numbered outputs preserve order, but a rerun overwrites same-named files. For archival runs, write each run to a dated directory or include a run identifier in the directory name.
  • Validate inputs. Confirm every line is an absolute HTTP or HTTPS URL. For production scripts, reject other schemes and malformed values before navigation.
  • Review a sample first. Check viewport, page length, image loading, consent state, and filename mapping before processing the full list.

Only capture pages you are allowed to access and use. Review the applicable terms and access restrictions for each retailer, especially if you plan to store, share, or republish screenshots. Browser automation capability and a service’s robots.txt behavior do not determine whether a particular use is permitted.

9. Performance, storage, and cost

Run time depends on the number of URLs, page load times, your wait conditions, retries, and the gap between requests. This guide provides no speed estimate or benchmark. Sequential processing is easier to control and diagnose, while parallel processing may increase load on destination sites and make failures harder to isolate. If you add concurrency, keep it low, respect the target site’s limits, and verify the provider’s current limits for hosted API calls.

Full-page PNGs can consume more storage than viewport captures, particularly for long pages. Choose viewport capture when only the first screen matters, and select a different supported format or image settings when smaller files are important. Keep the original URL and capture timestamp in the manifest regardless of storage format.

For local Playwright, account for the compute and storage used by the machine or environment running the browser; no vendor price applies to the script itself. For ScreenshotNeo, the stated plans are Free: 1,000 shots/month with no card; Starter: $5 for 3,000; Growth: $15 for 15,000; Pro: $39 for 60,000; Scale: $99 for 250,000; Business: $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan. Only clean shots are billed; consult the product docs for current billing details and response headers.

10. Troubleshooting

Symptom Likely cause What to do
Navigation timeout The site is slow, the timeout is too short, or the page never reaches the selected load condition. Raise the finite timeout for your use case, use domcontentloaded, and wait for a relevant selector or short settle period. Keep bounded retries.
Screenshot is blank or incomplete Client-side content has not rendered, a page failed to load, or access was restricted. Check the page manually and the manifest/error; wait on a meaningful element. Do not assume a successful navigation means useful content appeared.
Product images are missing Images may load lazily or after the screenshot was taken. Use full-page capture where appropriate, add a site-specific wait, and inspect a sample. Some pages require scrolling behavior or other site-specific handling.
Only part of a long page appears The capture used viewport mode or the page changes as it scrolls. Enable fullPage and check the resulting image. Consider whether a full-page image is the right output for a very long page.
Repeated HTTP error or bot check The destination rejected or challenged the request, or access is restricted. Stop repeated retries, check site rules and access, and use an authorized route. A different capture method does not grant access permission.
Images overwrite one another The filename scheme is based only on product name or URL path. Use the input index or a collision-resistant identifier and retain the manifest.
API response is an error The key, URL, request parameters, or service response may be invalid. Check the HTTP status and response details, verify the key and encoded URL, and consult the ScreenshotNeo docs. Do not save an error body with an image extension.
Batch stops after one bad URL An unhandled exception escaped the per-URL loop. Catch errors per item, write a failed manifest record, and continue. Retry failures separately after inspecting their causes.

11. Frequently asked questions

Can I screenshot every page on a retailer site from one URL list request?

This workflow processes the URLs you supply. It does not discover every product page on a site, and the ScreenshotNeo API examples make one request per URL; use its documented bulk options if they fit your list workflow.

Should I capture product pages as full-page images?

Only when details below the fold are part of the task. A viewport shot is smaller and focuses on the visible state; full-page mode includes the scrollable page and may produce a very tall image.

Will the same URL always produce the same screenshot?

No. Page content can vary with time, stock, location, cookies, account state, experiments, and other site behavior. Record when and under what capture settings each image was made.

Does robots.txt settle whether I can capture a page?

No. A service may describe how its own crawler treats robots.txt, but that does not answer the terms or access rules for every site or use. Check the rules that apply to your target pages and intended use.