How to Save a Webpage as PDF Using Selenium and Chrome
Use Selenium and headless Chrome to print a rendered webpage to PDF, control print layout, and handle common loading and setup issues.
To save a webpage as a PDF with Selenium and Chrome, open the page in headless Chrome, call Selenium’s print-page operation, decode its base64 result, and write the bytes to a .pdf file. Selenium’s print operation returns PDF content for the current page. The examples below use Python; Chrome’s own headless command-line option is also covered for cases that do not need WebDriver automation.
Use Selenium to print the current page
Install Selenium with python -m pip install selenium. Selenium Manager can manage the driver for many standard setups. You still need Chrome or Chromium installed. Check the installed browser and driver versions if startup fails: Selenium’s Chrome guide advises matching their major versions and says Selenium 4 is compatible with Chrome 75 and newer. Confirm the current requirements for your environment in the Selenium Chrome documentation.
Save this as save_page_pdf.py and run it with python save_page_pdf.py https://example.com page.pdf:
import argparse
import base64
from pathlib import Path
from selenium import webdriver
from selenium.webdriver.common.print_page_options import PrintOptions
from selenium.webdriver.support.ui import WebDriverWait
def main():
parser = argparse.ArgumentParser(description="Save a webpage as a PDF using headless Chrome")
parser.add_argument("url", help="Webpage URL to print")
parser.add_argument("output", nargs="?", default="page.pdf", help="Output PDF path")
parser.add_argument(
"--ready-selector",
help="Optional CSS selector that must appear before printing",
)
args = parser.parse_args()
chrome_options = webdriver.ChromeOptions()
chrome_options.add_argument("--headless=new")
driver = webdriver.Chrome(options=chrome_options)
try:
driver.get(args.url)
# For dynamic pages, wait for a site-specific element that indicates
# the content to include in the PDF is present.
if args.ready_selector:
WebDriverWait(driver, 30).until(
lambda browser: browser.find_elements("css selector", args.ready_selector)
)
print_options = PrintOptions()
pdf_base64 = driver.print_page(print_options)
pdf_bytes = base64.b64decode(pdf_base64, validate=True)
output_path = Path(args.output)
output_path.parent.mkdir(parents=True, exist_ok=True)
output_path.write_bytes(pdf_bytes)
print(f"Saved {len(pdf_bytes)} bytes to {output_path}")
finally:
driver.quit()
if __name__ == "__main__":
main()
The official Selenium guide documents driver.print_page() returning base64-encoded PDF content. Decode it before writing: saving the base64 text itself creates a file that is not a valid PDF. The finally block closes Chrome even when navigation, printing, or file writing raises an error. The Python WebDriver API describes the result as a best-effort PDF based on the supplied parameters; inspect the output for your page and print settings.
Wait for the content you need
driver.get() returning does not guarantee that every site-specific component, image, or API-driven section is ready to print. For a page that fills in asynchronously, wait for a meaningful selector as shown above. Choose a selector tied to the content you need, such as the article body, rather than an element that appears before its data has loaded.
If content appears only after scrolling, clicking, or signing in, perform those actions in the WebDriver session before calling print_page(). Selenium prints the current browser page and state. There is no universal wait that guarantees every dynamic or lazy-loaded asset is ready; verify the generated PDF against the page state you intended to capture.
Set page size, margins, orientation, and other print options
Selenium’s PrintOptions supports orientation, page ranges, page size, margins, scale, background rendering, and shrink-to-fit. Set only the properties you need. The exact accepted values and units are defined by Selenium’s print-page documentation.
from selenium.webdriver.common.print_page_options import PrintOptions
print_options = PrintOptions()
print_options.orientation = "landscape" # or "portrait"
print_options.scale = 1
print_options.background = True
print_options.shrink_to_fit = True
print_options.page_ranges = ["1-3"]
print_options.margin_top = 0.4
print_options.margin_bottom = 0.4
print_options.margin_left = 0.4
print_options.margin_right = 0.4
pdf_base64 = driver.print_page(print_options)
This example shows the documented option categories; consult the Selenium guide for the property names and accepted values for the Selenium version you install. Page ranges let you limit which pages are printed. Margins and paper dimensions affect pagination, while scale and shrink-to-fit can help wide content fit. Background rendering is useful when color blocks or other background styling carry meaning. Compare output PDFs after changing layout settings because a change can alter page breaks and text size.
Use Chrome headless directly when Selenium is unnecessary
If you only need Chrome to navigate to a URL and write a PDF, use its headless command-line route:
chrome --headless --print-to-pdf https://example.com
Chrome writes output.pdf in the current working directory by default. To choose a path and omit the browser’s print header and footer:
chrome --headless --print-to-pdf=/absolute/path/page.pdf --no-pdf-header-footer https://example.com
Chrome documents --timeout as a maximum wait in milliseconds before capture, including when the page is still loading. For example:
chrome --headless --print-to-pdf=page.pdf --no-pdf-header-footer --timeout=5000 https://example.com
A fixed timeout is only a timing limit; it does not establish that a particular page’s content is ready. Check the generated file and use Selenium with an application-specific readiness wait when you need to coordinate with page state. See Google’s Chrome Headless command-line reference for the documented switches and behavior.
Choose between Selenium and Chrome’s command line
| Need | Use | Why |
|---|---|---|
| Interact with the page before printing | Selenium | Use a WebDriver session to navigate and prepare the current page. |
| Set print layout in application code | Selenium | PrintOptions covers orientation, ranges, size, margins, scale, background, and shrink-to-fit. |
| Print a URL to a local PDF with a simple command | Chrome headless CLI | Chrome writes the PDF directly, with documented switches such as header/footer omission and timeout. |
| Control PDF bytes and output path in a Python program | Selenium | Decode the returned data and write it to the path your program chooses. |
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Chrome will not start or WebDriver reports a session creation error | Chrome is missing, the driver cannot be found, or Chrome and ChromeDriver have incompatible major versions. | Install Chrome or Chromium, check the installed versions, and follow Selenium’s current Chrome setup guidance. Ensure the process can launch a headless browser. |
print_page() fails in a visible browser session |
Selenium’s print-page operation requires Chromium browsers to run headlessly. | Add a headless Chrome argument such as --headless=new when creating Chrome options, as in the example. |
| The saved file is unreadable or not recognized as a PDF | The returned base64 string was written as text, or decoding was skipped or corrupted. | Use base64.b64decode() and write the resulting bytes with Path.write_bytes(). Keep the output extension as .pdf. |
| The PDF is blank or misses a section | The page may not have loaded the required content when printing began, or navigation may have reached an unexpected page. | Check the current URL and page state, wait for a meaningful readiness selector, and inspect the PDF. A generic load event or fixed delay may not cover site-specific rendering. |
| The PDF has unexpected breaks, tiny text, or clipped wide content | Paper size, margins, scale, orientation, or shrink-to-fit settings do not match the page. | Adjust the relevant PrintOptions values and review the output after each layout change. Try landscape orientation for wide tables and diagrams. |
| The Chrome CLI saves the file somewhere unexpected | --print-to-pdf without an explicit path writes output.pdf in the current working directory. |
Run from the intended directory or provide an explicit path with --print-to-pdf=/path/to/file.pdf. |
| CLI output omits recent dynamic content | The capture timeout elapsed while the page was still loading. | Adjust --timeout and check the output. When readiness depends on a specific page element, use Selenium to wait for it before printing. |
| The output directory does not exist or writing fails | The chosen path is invalid or its parent directory is absent or unwritable. | Use a writable path. The Python example creates parent directories before writing; permissions can still prevent output. |
Performance, reliability, and cost
Both approaches run Chrome, so browser startup and page rendering are part of the work. Reusing a Selenium session for multiple pages can avoid starting a fresh browser for every PDF, but navigate and wait for each page deliberately, and always quit the driver when finished. For parallel jobs, account for the memory and CPU used by each browser process; the cited documentation does not provide universal throughput figures.
Reliability depends on the target site as well as the print call. Pages can render different content based on authentication, cookies, location, or timing. Network failures, bot checks, and page changes can affect output. Save to a controlled path, catch and log failures in production code, and validate PDFs when downstream processing depends on their contents. Neither a fixed sleep nor Chrome’s capture timeout proves that all desired content was included.
The Selenium and Chrome command-line methods do not add an API charge, but your own machine or hosted runner still uses compute, memory, storage, and network resources. The sources cited here do not state a general cost per PDF; infrastructure cost depends on where and how often you run Chrome.
Or skip the browser setup
If you want a hosted PDF capture without installing or managing Chrome, ScreenshotNeo provides a screenshot and PDF API. One GET request can return a PDF; see the API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o page.pdf
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com", "format": "pdf"},
timeout=90,
)
r.raise_for_status()
open("page.pdf", "wb").write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com',
format: 'pdf'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('page.pdf', Buffer.from(await res.arrayBuffer())));
- Cookie and consent banners are accepted and removed before capture; 60+ known consent platforms, newsletter popups, and chat widgets can be removed, and each step can be turned off.
- Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed. Response headers report the page verdict and billing status.
- An MCP server provides
take_screenshot,get_page_info, andcapture_pdftools for Claude, Cursor, and other MCP clients. - The Free plan includes 1,000 shots a month with no card. Paid plans start at $5 for 3,000 shots. Every feature is on every plan.
Sign up free for 1,000 screenshots a month with no card.
FAQ
Does Selenium save the PDF directly to a file?
No. The documented print-page workflow returns base64-encoded PDF content. Decode it and write the bytes to the destination path.
Can I print only selected pages?
Yes. Selenium’s print options include page ranges. Set the range using the syntax supported by your Selenium version and verify the resulting page count.
Will lazy-loaded images always appear?
There is no such guarantee in the cited documentation. Wait for the content your workflow needs, trigger any required page behavior, and inspect the PDF.
Can I use Selenium’s print operation with Firefox?
The print-page guide specifies that printing requires Chromium browsers in headless mode. This article’s Selenium instructions therefore use Chrome.


