ScreenshotNeo

BlogHow-to

Bulk Website Screenshot Generation in Python for Indian E-commerce Product Pages

Build a Python batch workflow with Playwright to capture Indian e-commerce product pages, save predictable files, and record failures for review.

By the ScreenshotNeo team4 October 20269 min read

Use Python with Playwright to visit each product URL, capture the viewport, full page, or a chosen element, and save the result under a stable filename. Put the URLs and item IDs in a CSV, record each outcome in a manifest, and handle failures per page so one inaccessible or slow page does not stop the batch. The script below is a starting point: the right readiness condition and access method depend on each target site.

1. Check access and prepare the input

Before automating a marketplace or store, check its current terms and use an authorized access method. The available documentation does not establish policies for Amazon.in, Flipkart, or other named Indian marketplaces. Site behavior can also vary with consent state, login, region, language, and other page conditions, so verify each target rather than assuming one setup works everywhere.

Create products.csv with a stable identifier and an absolute URL for each page:

id,url
sku-1001,https://shop.example.in/products/item-one
sku-1002,https://shop.example.in/products/item-two

Use IDs that are unique and stable. The script sanitizes them for filenames; it does not derive names from page titles, which may be missing or duplicated.

2. Install Playwright and a browser

Playwright for Python offers synchronous and asynchronous interfaces and supports Chromium, Firefox, and WebKit. This example uses the synchronous API with Chromium:

python -m venv .venv
# macOS or Linux:
source .venv/bin/activate
# Windows PowerShell:
# .venv\\Scripts\\Activate.ps1
python -m pip install playwright
python -m playwright install chromium

Playwright’s documented workflow is to launch a browser, navigate a page, and save a screenshot. See the Playwright screenshot documentation and Python getting-started guide.

3. Choose the screenshot scope

Scope Use it for Playwright option
Viewport Consistent preview of the initially visible layout Default screenshot behavior
Full page Below-the-fold product details and page layout full_page=True
Element A product card, price area, or other identified component Locator screenshot
Image bytes Post-processing or pixel comparison page.screenshot() without a path

Full-page captures can be very tall. An element capture is clipped to the element’s bounds, but content covered by another element can still appear obscured. Use the same viewport and scope across a comparison batch.

4. Run a batch and write an outcome manifest

Save this as capture_products.py. It reads the CSV, creates one browser page per URL, waits for navigation to reach domcontentloaded, then captures the selected scope. The readiness setting is an example, not a guarantee that product images, prices, or client-rendered content are ready. Adjust it for the target site.

import csv
import re
import sys
from pathlib import Path
from urllib.parse import urlparse

from playwright.sync_api import sync_playwright

INPUT_CSV = Path("products.csv")
OUTPUT_DIR = Path("screenshots")
MANIFEST = Path("manifest.csv")
SCOPE = "full_page"  # "viewport", "full_page", or "element"
ELEMENT_SELECTOR = "[data-testid='product-card']"  # change for the target site
VIEWPORT = {"width": 1365, "height": 900}
NAVIGATION_TIMEOUT_MS = 45_000


def safe_id(value: str) -> str:
    cleaned = re.sub(r"[^A-Za-z0-9_-]+", "-", value.strip()).strip("-_")
    return cleaned or "item"


def valid_http_url(value: str) -> bool:
    parsed = urlparse(value)
    return parsed.scheme in {"http", "https"} and bool(parsed.netloc)


def capture_one(browser, row):
    item_id = row["id"].strip()
    url = row["url"].strip()
    if not item_id:
        raise ValueError("missing id")
    if not valid_http_url(url):
        raise ValueError("url must be an absolute http or https URL")

    output_path = OUTPUT_DIR / f"{safe_id(item_id)}.png"
    page = browser.new_page(viewport=VIEWPORT, device_scale_factor=1)
    try:
        response = page.goto(
            url,
            wait_until="domcontentloaded",
            timeout=NAVIGATION_TIMEOUT_MS,
        )
        # A navigation response can be absent for some navigations. Record its
        # status when available; do not treat HTTP status alone as proof that
        # the intended product content rendered correctly.
        status = str(response.status) if response is not None else ""

        # Add a site-specific readiness condition here when appropriate, e.g.:
        # page.locator("[data-testid='product-title']").wait_for(timeout=15000)
        if SCOPE == "viewport":
            page.screenshot(path=str(output_path))
        elif SCOPE == "full_page":
            page.screenshot(path=str(output_path), full_page=True)
        elif SCOPE == "element":
            page.locator(ELEMENT_SELECTOR).screenshot(path=str(output_path))
        else:
            raise ValueError(f"unknown SCOPE: {SCOPE}")
        return {"id": item_id, "url": url, "file": str(output_path),
                "status": status, "outcome": "success", "error": ""}
    finally:
        page.close()


def main():
    OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
    required = {"id", "url"}
    with INPUT_CSV.open(newline="", encoding="utf-8-sig") as source:
        reader = csv.DictReader(source)
        if not required.issubset(reader.fieldnames or []):
            raise SystemExit("CSV must have id and url columns")
        rows = list(reader)

    results = []
    with sync_playwright() as playwright:
        browser = playwright.chromium.launch(headless=True)
        try:
            for row in rows:
                try:
                    results.append(capture_one(browser, row))
                except Exception as exc:
                    results.append({
                        "id": row.get("id", ""), "url": row.get("url", ""),
                        "file": "", "status": "", "outcome": "failed",
                        "error": f"{type(exc).__name__}: {exc}",
                    })
        finally:
            browser.close()

    columns = ["id", "url", "file", "status", "outcome", "error"]
    with MANIFEST.open("w", newline="", encoding="utf-8") as target:
        writer = csv.DictWriter(target, fieldnames=columns)
        writer.writeheader()
        writer.writerows(results)

    failed = sum(result["outcome"] == "failed" for result in results)
    print(f"Processed {len(results)} rows; {failed} failed. See {MANIFEST}.")
    return 1 if failed else 0


if __name__ == "__main__":
    sys.exit(main())

Run it with python capture_products.py. The output folder contains ID-based PNG files, and manifest.csv records the requested URL, output path, available response status, and any error. A failed page is recorded and the loop continues; the process exits with a nonzero status if any row failed.

Set a site-specific ready condition

domcontentloaded means the initial HTML document has been parsed; it does not establish that every product detail or image has rendered. When the page has a reliable product marker, wait for it before taking the screenshot:

page.locator("[data-testid='product-title']").wait_for(state="visible", timeout=15000)

Replace the selector with one verified on the target page. For lazy-loaded images or sections below the fold, a full-page capture may not by itself make every site load those resources. Check the resulting files and adapt the page-specific wait or scrolling procedure where needed; do not assume a universal wait condition.

5. Useful variations

Capture an element

Set SCOPE = "element" and configure ELEMENT_SELECTOR for an element present on the product page. If the selector matches nothing, Playwright will time out or report that the element cannot be found. Confirm the selector in the page and wait for visibility when rendering is delayed.

Capture to memory

For image processing or pixel comparison, omit the path and receive bytes:

image_bytes = page.screenshot(full_page=True)
# Pass image_bytes to your image-processing step.

Use another browser engine

Playwright documents Chromium, Firefox, and WebKit. Install the chosen browser with python -m playwright install firefox or python -m playwright install webkit, then replace playwright.chromium with playwright.firefox or playwright.webkit. Browser engines can render pages differently, so keep the engine consistent when comparing images.

Capture viewport only

Set SCOPE = "viewport" to capture the initial viewport. The example fixes the viewport dimensions and device scale factor to make captures more comparable. Those settings do not make layouts identical across browser engines or site states.

6. cURL, Python, and Node.js with ScreenshotNeo

For the browser-based workflow above, Python and Playwright handle the batch loop locally. If you want a managed screenshot request instead, ScreenshotNeo accepts a URL and returns an image or PDF. The following examples capture one page; to process a CSV, wrap the Python or Node.js request in your own loop and write each response to an ID-based filename.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo API documentation for request options and response details. ScreenshotNeo supports full-page captures, CSS element capture, custom CSS and JavaScript, waits, cookies and headers, device and viewport settings, image formats, caching, bulk capture up to 100 URLs per call, and more. Check the docs for the parameters relevant to your job.

Or skip the browser setup

ScreenshotNeo accepts a URL in one API call and can return a PNG, JPEG, WebP, or PDF. Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. See ScreenshotNeo and the API docs. Sign up for 1,000 free screenshots a month, no card required.

7. Performance, reliability, and cost

This local workflow launches one browser and processes pages sequentially, which keeps the example simple and limits simultaneous browser work. No throughput or failure-rate figure is established here. If you need more capacity, measure it on your target pages and environment before increasing concurrency: browser processes and page contents consume resources, and concurrent requests may interact with a site’s access controls or rate limits.

For reliability, keep the per-row error recording, retain the source CSV, and review the manifest and actual images. Consider rerunning only failed IDs after investigating their causes. Avoid treating an HTTP response status or an output file’s existence as proof that the intended product information is visible. Local browser operation has no per-screenshot service charge in this example, but it does require a machine, browser installation, and maintenance. A managed service has its own plan and request behavior; compare current terms and documentation for your workload.

8. Troubleshooting

Symptom Likely cause What to do
Browser executable missing The Playwright package is installed but its browser was not installed for this environment. Run python -m playwright install chromium in the active environment.
Navigation timeout The page did not reach the chosen lifecycle state before the timeout, or the site is slow or unreachable. Check the URL and network, inspect the page manually, and adjust the timeout deliberately. Avoid an unbounded wait.
Screenshot is blank or incomplete Content may render after navigation, require a site-specific readiness condition, or depend on lazy loading or access state. Wait for a verified content locator, inspect the saved image, and confirm consent, login, locale, and permitted access conditions.
Element selector not found The selector is wrong, changes by page, or the element has not appeared. Verify the selector for that page and wait for the element in a visible state before capturing.
Duplicate or overwritten output Two rows sanitize to the same filename or IDs are duplicated. Use unique IDs and check sanitized IDs for collisions before running a large batch.
Some rows fail but others save The workflow catches errors per row so the batch can continue. Review manifest.csv, investigate each failed row, and rerun the relevant IDs after fixing the cause.
Images differ between runs Page state, viewport, browser engine, dynamic content, or timing differs. Keep browser and viewport settings consistent, wait for a meaningful ready condition, and account for changing content.

9. FAQ

Can I use this for any Indian ecommerce site?

The code can visit valid web URLs, but that does not establish that every site permits automation or that every page renders the same way. Check current site terms and verify access and output for each target.

Should I choose full-page or viewport screenshots?

Choose viewport for comparable first-screen previews. Choose full-page when the below-the-fold page matters. Use an element capture when only a specific component is needed.

Does a successful screenshot prove the product price is current?

No. A screenshot records what appeared in that browser session. Verify the page state and any data freshness requirements independently.

Can I compare screenshots pixel by pixel?

Playwright can return screenshot bytes for later processing. For meaningful comparisons, keep scope, viewport, browser engine, and page readiness consistent, and account for dynamic page content.