ScreenshotNeo

BlogHow-to

How to Save Selenium Screenshots of Indian Web Pages as PDFs

Use Selenium’s print-page command to save the current page as a PDF, configure its layout, and check Indian-language fonts in the result.

By the ScreenshotNeo team4 October 20263 min read

Short answer: To save the current Selenium page as a PDF, call WebDriver’s print-page command, decode the returned base64 data, and write the bytes to a .pdf file. In Chromium, run the browser in headless mode for printing. A WebDriver screenshot is an image; the print-page command creates a PDF with printed-page layout. Selenium’s print-page documentation and Chrome browser documentation describe these workflows.

1. Save the current page as a PDF with Selenium

The examples below use Python with Selenium 4 and Chrome. They navigate to a page, wait for its document to finish loading, print the current page, decode the returned PDF data, and save it. Replace the sample URL with the page you need. For a page requiring sign-in, add your normal Selenium authentication steps before printing.

from base64 import b64decode
from pathlib import Path

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait

url = "https://example.com"
output_path = Path("page.pdf")

options = Options()
options.add_argument("--headless=new")

# Selenium Manager can configure the driver for supported setups.
with webdriver.Chrome(options=options) as driver:
    driver.get(url)
    WebDriverWait(driver, 30).until(
        lambda browser: browser.execute_script("return document.readyState") == "complete"
    )

    pdf_base64 = driver.print_page()
    output_path.write_bytes(b64decode(pdf_base64))

print(f"Saved {output_path}")

The readiness check waits for the document load event, but it cannot know that every single-page app, lazy image, API request, or web font is ready. Add a wait for a page-specific element or condition before calling print_page(). Selenium documents print options for page size, orientation, margins, scale, background printing, and shrink-to-fit; configure them in the Selenium binding you use and check its version’s API documentation. Binding implementations and option names can differ, so record your Selenium and browser versions when reproducing output. See Selenium’s print-page options and examples.

Wait for page-specific content and fonts

For dynamic pages, wait until a known content element appears and, if it matters to the output, until fonts are ready. This JavaScript condition waits for the browser’s font-loading set to settle:

WebDriverWait(driver, 30).until(
    lambda browser: browser.execute_script(
        "return document.fonts ? document.fonts.status === 'loaded' : true"
    )
)

That condition does not prove that every glyph has an appropriate font or that a page’s application has finished rendering. Prefer a page-specific readiness signal, such as a results container becoming visible or a loading indicator disappearing. A fixed sleep can be useful for diagnosing a timing issue, but it is usually less reliable than waiting for an observable condition.

2. Choose PDF layout options

Printing follows browser print behavior, including the page’s print CSS. Use Selenium’s PrintOptions to set the options exposed by your binding. Common choices include:

Setting When to adjust it
Orientation Use portrait for typical articles and forms; landscape for wide tables or dashboards.
Page size Choose a paper size such as A4 or Letter when the document needs a predictable page format. The default may depend on the browser.
Margins Reduce margins for more content per page, or increase them for printed documents with room for notes.
Scale Adjust when content is too large or too small. Check text legibility after scaling.
Background Enable background printing if colored sections or background graphics are part of the intended document.
Shrink to fit Use when wide content must fit the printable area; inspect tables and small text for readability.

For example, Selenium’s Python API exposes options through PrintOptions. Verify the available properties against the Selenium version installed in your environment:

from selenium.webdriver.common.print_page_options import PrintOptions

print_options = PrintOptions()
print_options.orientation = "landscape"
print_options.scale = 0.9
# Set page size, margins, background, or shrink-to-fit as supported
# by the PrintOptions API in your installed Selenium version.

pdf_base64 = driver.print_page(print_options)

To control browser-specific print details such as header and footer templates, Chromium’s DevTools Protocol provides Page.printToPDF. Treat that as a Chromium-specific route: the protocol documents options including orientation, margins, scale, and header/footer templates. Selenium’s print command is the cross-binding abstraction; select the protocol route when you need a Chromium control that your binding does not expose. Read the Page.printToPDF protocol documentation.

3. Check Indian-language text and fonts

Indian web pages may use Devanagari, Tamil, Bengali, Telugu, Kannada, Malayalam, Gujarati, Gurmukhi, Odia, or more than one script. Correct rendering depends on the page’s CSS font stack, loaded web fonts, and fonts available to the browser. A screenshot that looks correct does not guarantee the PDF has the same glyph coverage, so inspect the PDF itself.

Noto’s browser font guidance explains that browsers fall back through the CSS font family list when a font lacks a character and recommends choosing fonts for the writing systems used. Its Hindi and Tamil example includes Noto Sans Devanagari, Noto Sans Tamil, Noto Sans, and Noto Sans Symbols 2. This is an example stack, not a universal requirement: use fonts that cover the scripts on the specific page.

For Linux Devanagari specifically, a Chromium source change dated March 4, 2026 identifies Noto Sans Devanagari as the default for standard and sans-serif, Noto Serif Devanagari for serif, and Noto Sans Mono for fixed-width text. That detail is specific to Devanagari and the described Linux defaults; do not assume it applies to every Indian script or every Linux image. See the Chromium source change.

  • If characters appear as boxes or blanks, identify the script and check whether an appropriate font is available to the browser process.
  • If conjuncts or marks look wrong, check whether the page’s intended web font finished loading and whether the browser has a suitable fallback.
  • If only the PDF is wrong, compare the saved PDF with the page and check the browser’s print output and font availability; do not infer PDF correctness from a screenshot.
  • For repeatable output, use a consistent browser environment and record its version, operating system, and installed fonts.

There is no universal font-install command for all Indian scripts and operating systems. Follow the documentation for your runtime image and the fonts required by the scripts in your pages.

4. PDF output versus screenshot images

driver.print_page() produces a PDF representation using print layout. It can paginate a long page, apply print CSS, and scale content to paper. By contrast, driver.save_screenshot("page.png") captures a raster image of the browser viewport. Use a screenshot when you need pixels that match a visible browser view; use print-to-PDF when you need a document. Selenium documents both printing and screenshots.

If you need an image embedded in a PDF, capture the image and convert it with an appropriate image-to-PDF tool. That conversion preserves pixels rather than reproducing browser print layout, and a long page may require multiple captures or a full-page capture strategy.

5. Troubleshooting

Symptom Likely cause Fix
Printing command is unsupported or errors in Chromium The browser is running with a visible UI or the browser/binding does not support the command in that configuration. Run Chromium headless, check the Selenium browser documentation, and confirm compatible Selenium and browser versions.
The PDF is empty or missing page content Printing began before the app rendered its content, or navigation ended on an error/redirect page. Wait for a page-specific ready condition; verify the final URL and page state before printing.
Images are missing Images or lazy-loaded content had not loaded when printing began, or the page excludes them in print CSS. Wait for relevant images or content, scroll/load lazy sections if needed, and inspect the site’s print styles.
Text is clipped or pages break badly Wide elements, fixed dimensions, print CSS, margins, or scale do not fit the selected paper. Try landscape, adjust margins or scale, enable shrink-to-fit if supported, and inspect the site’s print stylesheet.
Indian-language glyphs show as boxes No installed or loaded font covers those characters. Check the script, page font loading, CSS fallback list, and fonts available in the browser environment.
Glyphs or combining marks are shaped incorrectly The intended web font did not load, or the fallback environment lacks suitable script support. Wait for font loading, inspect browser console/network errors, and verify the exact PDF output.
Output differs across machines Browser builds, operating systems, fonts, print defaults, or Selenium binding implementations differ. Pin or record versions and runtime fonts; compare the same URL and options in the same environment.
PDF file cannot be opened The base64 return value was written as text instead of decoded, or the write was interrupted. Decode the base64 data to bytes as shown, then check that the file begins as a valid PDF and is non-empty.

6. Performance, reliability, and cost

PDF generation time depends on page loading, scripts, fonts, image size, and print layout. Avoid arbitrary long waits where a readiness condition can be used. Reuse a browser session when processing multiple pages if isolation requirements allow it, but verify each page’s final URL and ready state. For reliability, save each result to a temporary path and rename it after the write completes; log the URL, browser and Selenium versions, selected options, and any exception. For Indian-language documents, include a visual check of the generated PDF in the workflow.

Selenium’s direct costs depend on where and how you run the browser; the sources cited here do not provide a universal cost figure. Account for browser compute, storage, and any infrastructure or grid service you use. The method makes no promise of identical output across browser versions or operating systems.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. Its PDF endpoint takes one GET request with a URL. Use it when you want a hosted capture without installing and managing a browser. The output is a browser-rendered PDF, so inspect script coverage and pagination for your target pages just as you would with a local browser.

First create an API key and see the ScreenshotNeo API documentation. Then run:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o page.pdf
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com", "format": "pdf"},
    timeout=90,
)
r.raise_for_status()
open("page.pdf", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.com',
  format: 'pdf'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await (await import('node:fs/promises')).writeFile('page.pdf', bytes);

Cookie banners, newsletter popups, and chat widgets are removed before capture. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers say which outcome occurred. An MCP server lets AI agents use the take_screenshot, get_page_info, and capture_pdf tools. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month with no card.

FAQ

Does the PDF contain the whole page?

It prints the current document according to browser print layout and pagination. Check print CSS, lazy content, and page breaks when the result omits or rearranges material.

Can I save a PDF from a page that requires login?

Yes, if your Selenium session is authenticated and permitted to access the page. Complete the login flow before printing and avoid saving credentials in source code.

Will an Indian-language page always render correctly?

No single font setup covers every script and runtime. Verify the page’s script-specific fonts and inspect the generated PDF.

Can I add page numbers or a header?

Chromium’s DevTools Page.printToPDF documents header and footer templates. Use that Chromium-specific route when those controls are needed and unavailable through your Selenium binding.