ScreenshotNeo

BlogHow-to

How to Take a PDF of a Webpage with Selenium and Headless Chrome

Use Selenium and headless Chrome to save a webpage as PDF, wait for dynamic content, control print settings, and fix common rendering issues.

By the ScreenshotNeo team4 October 20267 min read

Use Selenium’s Python Chrome WebDriver to open the page, wait for the content you need, call print_page(), and decode its base64 result into a PDF file. Headless mode runs Chrome without a visible browser window. Selenium describes PDF printing as best effort, so inspect the output for missing content, pagination, clipping, fonts, and backgrounds.

1. Install Selenium and prepare Chrome

Install Selenium in your Python environment:

python -m pip install selenium

Recent Selenium versions include Selenium Manager, which handles driver setup in most supported cases. Chrome and ChromeDriver still need to match at the major-version level. If automatic setup fails, check the installed Chrome version and driver compatibility before debugging the page.

2. Save the current page as a PDF

This complete example waits for navigation to finish, prints the current page, writes the returned PDF bytes, and always closes the browser:

import base64
from pathlib import Path

from selenium import webdriver
from selenium.webdriver.support.ui import WebDriverWait

url = "https://example.com"

options = webdriver.ChromeOptions()
options.add_argument("--headless=new")

driver = webdriver.Chrome(options=options)
try:
    driver.get(url)

    # A baseline only: a page can still load data or images after this.
    WebDriverWait(driver, 20).until(
        lambda browser: browser.execute_script(
            "return document.readyState"
        ) == "complete"
    )

    pdf_base64 = driver.print_page()
    pdf_bytes = base64.b64decode(pdf_base64)
    Path("page.pdf").write_bytes(pdf_bytes)
finally:
    driver.quit()

Run it with python save_page.py. The PDF is written to the current working directory. print_page() prints the page currently open in that WebDriver session; navigate before calling it. The returned value is base64 data, so decode it before writing the file.

Wait for the content that matters

document.readyState == "complete" is a practical baseline, not proof that a single-page application has finished rendering. If the page loads its main content asynchronously, wait for a stable, page-specific element:

from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC

WebDriverWait(driver, 30).until(
    EC.visibility_of_element_located((By.CSS_SELECTOR, "main article"))
)

For lazy-loaded images, scroll through the page before printing and wait for the images you need. For an application that fetches data after navigation, wait for a selector or state that indicates the relevant data has appeared. There is no universal wait condition that proves every site is ready.

3. Set PDF layout options with Chrome DevTools Protocol

Use Selenium’s print_page() as the starting point. If you need explicit paper size, margins, page ranges, background printing, or headers and footers, invoke Chrome’s Page.printToPDF command through Selenium’s Chrome DevTools Protocol method. Decode its returned base64 data in the same way:

import base64
from pathlib import Path

result = driver.execute_cdp_cmd(
    "Page.printToPDF",
    {
        "landscape": False,
        "displayHeaderFooter": False,
        "printBackground": True,
        "paperWidth": 8.27,
        "paperHeight": 11.7,
        "marginTop": 0.4,
        "marginBottom": 0.4,
        "marginLeft": 0.4,
        "marginRight": 0.4,
        "scale": 1,
        "preferCSSPageSize": True,
    },
)
Path("page.pdf").write_bytes(base64.b64decode(result["data"]))

Paper dimensions and margins are expressed in inches. The protocol also documents page ranges and header/footer templates. Consult the installed browser’s supported protocol version if a parameter is rejected; protocol options can vary with browser versions. Relevant options include:

Option What it controls
landscape Print orientation.
paperWidth, paperHeight Paper dimensions in inches.
marginTop, marginBottom, marginLeft, marginRight Page margins in inches.
scale Content scaling for the printed page.
pageRanges Pages to include, using the protocol’s page-range format.
displayHeaderFooter Whether to print header and footer templates.
printBackground Whether to include background graphics and colors.
preferCSSPageSize Whether CSS @page size takes precedence over the paper dimensions.

For available parameters and their exact behavior, see the Chrome DevTools Protocol Page reference and the Selenium Python WebDriver API.

4. Choose between Selenium, CDP, and Chrome’s command line

Route Choose it when Trade-off
Selenium print_page() You already use WebDriver to navigate, interact, or wait for page content. Simple browser workflow; fewer explicit PDF controls than the protocol route.
CDP Page.printToPDF You need documented print settings such as margins, paper dimensions, or page ranges. Uses Chrome’s DevTools Protocol; confirm support in the browser version you run.
Chrome headless CLI The URL can be opened directly and no WebDriver interaction is needed. Does not provide Selenium interactions before printing.

Chrome’s headless command line can print directly to PDF:

google-chrome --headless --print-to-pdf=page.pdf https://example.com

To suppress Chrome’s default header and footer, use its documented --no-pdf-header-footer flag. The headless CLI also documents a --timeout maximum wait. A timeout is not proof that application-specific content has loaded; use WebDriver when you need to wait for a known element or interact with the page before printing. See Chrome’s headless command-line reference.

5. Handle dynamic pages and print layout

  • Single-page applications: wait for the rendered content, not only navigation completion.
  • Lazy images: scroll to trigger loading, then wait for the required images before printing.
  • Print-specific styling: pages may use CSS @media print or @page rules that change what appears on paper.
  • Backgrounds: if colors or background graphics are missing, enable printBackground with CDP and check whether the page’s print styles allow them.
  • Pagination: long pages can split content across pages. Adjust CSS print rules or paper and margin settings when you control the page; inspect page breaks in the generated file.
  • Authentication: if the page requires login, establish the authenticated browser session before printing. The PDF reflects the content available to that session.
  • Fonts and external resources: wait for important fonts and assets to load, and verify the output in the same environment used for deployment.

6. Troubleshoot common problems

Symptom Likely cause What to do
print_page() errors or is unavailable Older Selenium, unsupported driver/browser combination, or a session that is not using a compatible browser. Update Selenium, confirm the active browser is Chrome/Chromium, and check the installed Selenium API documentation.
Chrome fails to start or WebDriver cannot create a session Chrome and ChromeDriver major versions do not match, or Chrome is missing in the runtime. Check browser installation and major version compatibility; let Selenium Manager resolve the driver where supported.
PDF is blank or missing app content The page had not rendered its asynchronous content when printing began. Wait for a page-specific visible element or application state; do not rely only on a fixed delay or document readiness.
Images are missing Lazy loading or slow image requests had not completed. Scroll through the relevant page area and wait for required images to load before printing.
Background colors or images are absent Background printing is disabled or print CSS removes the backgrounds. Set CDP printBackground to true and inspect the page’s print styles.
Content is clipped or breaks awkwardly Paper size, margins, scale, or CSS print rules do not fit the content. Adjust dimensions and margins, try a different scale or orientation, and review @page and print media styles.
Unexpected headers or footers appear Browser print headers and footers are enabled. Disable displayHeaderFooter in CDP or use Chrome CLI’s documented suppression flag.
PDF output changes between machines Browser version, fonts, installed dependencies, or page load timing differ. Pin or record the browser environment, wait for needed assets, and inspect output in the deployment environment.
Chrome processes remain after an exception The WebDriver session was not closed during error handling. Put driver.quit() in a finally block.

7. Performance, reliability, and cost

Each WebDriver session starts a browser process, so reusing a session for multiple pages can avoid repeated startup work when your workflow allows it. Always navigate and wait for each page’s needed content before printing, and close the driver when the batch is done. Printing time depends on page complexity, assets, and readiness waits; there is no universal capture duration.

For reliability, record the Chrome and Selenium versions in the runtime, use explicit waits tied to page content, and treat the resulting PDF as an artifact to validate. Selenium documents the print operation as best effort rather than a guarantee of identical rendering across environments. Browser automation also consumes compute and memory while Chrome runs; plan worker limits around your deployment resources. Selenium and Chrome are software components rather than a per-screenshot API plan, so direct costs depend on the infrastructure where you run them.

8. Or skip the browser setup

If all you need is a PDF file from a URL, ScreenshotNeo provides a one-request website screenshot API that can return a PDF. See the API documentation for its parameters. For example, save a PDF response with cURL:

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com \
  -d format=pdf \
  -o page.pdf

Cookie banners, newsletter popups, and chat widgets are removed before the shot, and each step can be turned off. Bot checks, blank pages, failed loads, and cache hits are not billed; response headers report the page verdict and billing status. An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000 shots. Every feature is on every plan.

Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.

9. Frequently asked questions

Can Selenium save a webpage as a PDF without opening a visible window?

Yes. Start Chrome with --headless=new, navigate with WebDriver, and call print_page().

Does the PDF include the entire page?

It prints the page using browser print behavior and styles. Check the generated PDF for pagination, hidden print content, and page breaks; long-page output depends on the site’s print CSS and your selected settings.

Can I use Selenium’s print method with a browser other than Chrome?

The example here uses Selenium’s Chrome WebDriver and Chrome’s printing support. The available method and options depend on the driver and browser combination; consult that browser’s Selenium API documentation.

What if I need a visual image instead of a PDF?

Use a browser screenshot workflow for PNG, JPEG, or WebP output. For an API-based capture, ScreenshotNeo supports image formats as well as PDF.

Primary references