ScreenshotNeo

BlogHTML to image & PDF

Convert Indian Ecommerce Product URLs to PDF Using Python Playwright

Save an Indian ecommerce product page as a PDF with Python Playwright. Choose print or screen styling, handle page readiness, and troubleshoot common capture problems.

By the ScreenshotNeo team4 October 20269 min read

Use Playwright’s Python library to open the product URL in headless Chromium, then call page.pdf(). The method works for pages you can access, but ecommerce layouts, login walls, consent banners, lazy-loaded images, and anti-bot checks can affect what appears in the PDF. Playwright uses print CSS by default; select screen media first if you want a closer match to the browser view.

This guide covers a local, repeatable workflow for saving a product page you are authorized to access. For product records you plan to share, review the retailer’s terms and consider copyright and privacy implications.

1. Install Playwright and Chromium

Install the Python package and its browser binaries as separate steps. Use a virtual environment for a project-specific installation:

python -m venv .venv

# macOS or Linux
source .venv/bin/activate

# Windows PowerShell
# .venv\Scripts\Activate.ps1

python -m pip install playwright
python -m playwright install chromium

Playwright supports multiple browser engines generally, but its PDF generation API is supported only in headless Chromium. The install workflow is documented in the Playwright Python getting-started guide; see also the browser installation guide.

2. Save a product page as an A4 PDF

Save this as product_to_pdf.py. Replace the example URL with the product page you want to capture. The script accepts a URL and output path as command-line arguments, uses Chromium, waits for the page’s load event, and closes the browser even if navigation or PDF generation fails.

import argparse
from pathlib import Path
from playwright.sync_api import sync_playwright, TimeoutError as PlaywrightTimeoutError


def main():
    parser = argparse.ArgumentParser(description="Save an accessible product page as PDF")
    parser.add_argument("url", help="Full product page URL, including https://")
    parser.add_argument("-o", "--output", default="product.pdf", help="Output PDF path")
    parser.add_argument(
        "--screen", action="store_true",
        help="Use screen CSS media instead of the default print media",
    )
    args = parser.parse_args()

    output = Path(args.output)
    output.parent.mkdir(parents=True, exist_ok=True)

    with sync_playwright() as p:
        browser = p.chromium.launch(headless=True)
        try:
            page = browser.new_page(viewport={"width": 1440, "height": 1000})
            response = page.goto(args.url, wait_until="load", timeout=60_000)

            if response is not None and response.status >= 400:
                raise RuntimeError(f"Navigation returned HTTP {response.status}")

            # Optional: wait briefly for client-rendered product details or images.
            page.wait_for_timeout(1500)

            if args.screen:
                page.emulate_media(media="screen")

            page.pdf(
                path=str(output),
                format="A4",
                print_background=True,
                prefer_css_page_size=True,
                margin={"top": "12mm", "right": "12mm", "bottom": "12mm", "left": "12mm"},
            )
            print(f"Saved PDF to {output.resolve()}")
        except PlaywrightTimeoutError as exc:
            raise SystemExit(
                "Navigation or a page wait timed out. Try a different readiness strategy "
                "or inspect the URL in a normal browser."
            ) from exc
        finally:
            browser.close()


if __name__ == "__main__":
    main()

Run it with:

python product_to_pdf.py "https://example.in/product" --output "out/product.pdf"

# Use screen CSS instead of print CSS
python product_to_pdf.py "https://example.in/product" -o product-screen.pdf --screen

The example uses A4 paper, prints background graphics, and gives the page 12 mm margins. You can remove or change these options. prefer_css_page_size=True tells Chromium to prefer a page size specified by the page’s CSS; omit it if you want the requested A4 format to take precedence.

3. Choose print styling or screen styling

page.pdf() renders with print CSS media by default. Print styles can hide navigation or rearrange product details for paper. To render with screen styles, call page.emulate_media(media="screen") before page.pdf(). That can preserve a more screen-like layout, but the result may paginate less neatly.

Print output can change colors. To request more exact colors in your own page or a page you control, the CSS property -webkit-print-color-adjust can affect print rendering:

@media print {
  html {
    -webkit-print-color-adjust: exact;
    print-color-adjust: exact;
  }
}

This CSS cannot be applied to a third-party product page unless you can change its styles; for a page you do not control, print_background=True asks Chromium to include background graphics, but does not guarantee identical screen colors.

See the Playwright Page API PDF documentation for the print-media behavior, Chromium limitation, and color guidance.

4. Handle page readiness and product images

There is no single wait condition that reliably fits every ecommerce page:

  • wait_until="load" waits for the page load event. It is a reasonable starting point, but client-rendered prices, variants, or images may appear later.
  • wait_until="domcontentloaded" returns earlier, after the initial HTML is parsed. Use it only when you add a targeted wait for the content you need.
  • wait_until="networkidle" waits for network activity to become idle. Analytics, ads, polling, or other ongoing requests can make this wait unsuitable or cause it to time out.
  • A short page.wait_for_timeout() can give delayed rendering a chance to finish, but it is a fixed delay rather than proof that the required content is ready.

When you know a stable product-detail selector on a site, wait for that element instead of guessing how long the page needs:

page.goto(url, wait_until="domcontentloaded", timeout=60_000)
page.locator("YOUR_PRODUCT_TITLE_SELECTOR").wait_for(state="visible", timeout=15_000)
page.pdf(path="product.pdf", format="A4", print_background=True)

Replace the placeholder selector with one that exists on the target page. Product images may be lazy-loaded as they enter the viewport; a PDF capture may omit images that never loaded. For a page you control, explicitly wait for image loading before generating the PDF. On a third-party page, inspect the resulting PDF and adjust the wait or page interaction only where permitted.

5. Adjust the PDF output

Common page.pdf() options include:

Option What it controls Typical use
path Where to save the PDF Use a per-product filename or output directory.
format Paper size, such as A4 or Letter A4 is a common choice for product records in India.
print_background Whether background graphics are printed Set to True when colored sections or backgrounds matter.
margin Page margins; values can use units such as mm Leave room for printed or annotated pages.
prefer_css_page_size Whether CSS page sizing takes priority over the requested format Use when the page defines its own print size.
landscape Switches page orientation Use when a wide product comparison layout is being captured.
scale Scales the rendered content Adjust when content overflows or prints too small; check readability.
page_ranges Selects PDF pages by range Useful when a long page produces more pages than needed.

Confirm exact option names and accepted values against the current Page API reference when adapting the script. Page CSS can also control printed page breaks and sizing; the final pagination depends on both those styles and your PDF options.

6. Capture several product URLs

For a small list of URLs, reuse one browser and create a fresh page for each capture. A fresh page helps avoid carrying page state such as cookies between products. This example writes one PDF per URL:

from pathlib import Path
from playwright.sync_api import sync_playwright

urls = [
    "https://example.in/product-one",
    "https://example.in/product-two",
]

Path("pdfs").mkdir(exist_ok=True)

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    try:
        for index, url in enumerate(urls, start=1):
            page = browser.new_page()
            try:
                response = page.goto(url, wait_until="load", timeout=60_000)
                if response is not None and response.status >= 400:
                    print(f"Skipping {url}: HTTP {response.status}")
                    continue
                page.pdf(
                    path=f"pdfs/product-{index}.pdf",
                    format="A4",
                    print_background=True,
                )
                print(f"Saved pdfs/product-{index}.pdf")
            except Exception as exc:
                print(f"Failed {url}: {exc}")
            finally:
                page.close()
    finally:
        browser.close()

For larger batches, add bounded concurrency and retries only for transient failures, and keep the number of simultaneous browser pages within the memory and CPU limits of the machine. Respect site access limits and do not use automation to bypass a login, CAPTCHA, or other access control.

7. Troubleshoot common problems

Symptom Likely cause What to try
Browser executable is missing The Python package was installed but Chromium’s browser binary was not. Run python -m playwright install chromium in the same environment.
PDF generation is unsupported in the selected browser The script launched Firefox or WebKit. Use p.chromium.launch(); PDF generation is documented for headless Chromium.
Navigation times out The page is slow, still making requests, or unreachable from the environment. Try wait_until="load" or "domcontentloaded", increase the navigation timeout if appropriate, and wait separately for the specific content you need.
PDF is blank or shows an error page The URL redirected, returned an error, or displayed a bot challenge or access wall. Check the final page in a regular browser and inspect the navigation response. Capture only pages you are authorized to access; do not attempt to bypass access controls.
Product price or variant is missing Client-side rendering may not have finished, or the product option needs a user interaction. Wait for the relevant visible selector. If interaction is permitted and needed, perform it before calling page.pdf().
Product photos are missing Images may be lazy-loaded or blocked. Wait for the page’s relevant image elements to load where possible, then inspect the PDF. A fixed delay may help but is not a guarantee.
Colors or layout differ from the browser Print CSS is active by default and print rendering may modify colors. Try screen media for a closer screen layout; enable background printing. On pages you control, use print color CSS and page-break rules.
PDF has unexpected page breaks The page’s print styles, paper size, margins, or long content affect pagination. Try A4, adjust margins or scale, or use CSS page-break rules if you control the page.
Access is denied or a CAPTCHA appears The retailer restricts automated access or requires an authorized session. Use the retailer’s supported access or download option where available. Do not try to evade the challenge.

8. Performance, reliability, and cost

Local Playwright does not charge per PDF through the code shown here, but it does use your machine’s CPU, memory, storage, and bandwidth. Browser startup and each page load contribute to runtime; for a small batch, reusing one Chromium process avoids repeated browser startups. For concurrent work, limit open pages, save outputs incrementally, and record failures so one inaccessible product does not stop the whole batch.

Capture reliability depends on the target site and your network: product pages can change, redirect, delay content, or refuse automated access. Set timeouts, check HTTP responses where present, close pages and browsers in finally blocks, and inspect generated PDFs when completeness matters. Avoid retrying permanent errors such as access denial as if they were transient.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. Its PDF endpoint can capture a product page in one request. See the ScreenshotNeo documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.in/product -o product.pdf
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.in/product", "format": "pdf"},
    timeout=90,
)
r.raise_for_status()
open("product.pdf", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.in/product',
  format: 'pdf'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo returned HTTP ${res.status}`);
await Bun.write('product.pdf', res);
  • Cookie banners are accepted and 60+ known consent platforms, newsletter popups, and chat widgets are removed before capture; each step can be turned off.
  • Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing. Responses include X-Page-Verdict and X-Billed headers.
  • An MCP server lets AI agents using Claude, Cursor, or another MCP client take screenshots with take_screenshot, inspect pages with get_page_info, and create PDFs with capture_pdf.
  • The Free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000 shots. Every feature is on every plan.

Sign up for ScreenshotNeo free: 1,000 screenshots a month, no card required.

FAQ

Can I use Playwright PDF generation with Firefox?

No. The documented PDF generation support is for headless Chromium, even though Playwright supports other browser engines for other tasks.

Does this make a PDF of the entire product page?

It prints the page as rendered, which can span multiple PDF pages. Print CSS and the page’s content determine what is included and how it breaks across pages.

Can I use this for any Indian ecommerce URL?

You can attempt it for URLs you are authorized to access, but no workflow guarantees that every retailer permits automation or will render the same way. Check retailer terms and use supported access methods.

Why does a PDF sometimes differ from a screenshot?

A PDF is laid out for print by default, while a screenshot captures pixels from a viewport. Print CSS, paper dimensions, margins, and page breaks can change both the composition and colors.