How to Generate PDFs with Selenium
Print any rendered webpage to a PDF with Selenium, configure layout, troubleshoot headless Chrome, and automate reliable PDF generation.
Use Selenium’s print-page API to turn the page currently rendered in a browser into a PDF. In Python, navigate to the page, create PrintOptions, call driver.print_page(print_options), base64-decode the returned string, and write the bytes to a file. Chromium printing requires headless mode according to Selenium’s browser documentation.
This workflow prints the rendered HTML page. It does not download a PDF that is already hosted at a URL; that is a separate download or HTTP-response workflow.
1. Generate a PDF from the current page
Install Selenium and make sure a compatible Chromium browser and driver are available to your environment.
python -m pip install selenium
The following complete script opens a page, prints it, decodes Selenium’s base64 response, and saves page.pdf:
from base64 import b64decode
from selenium import webdriver
from selenium.webdriver.common.print_page_options import PrintOptions
url = "https://example.com"
output_path = "page.pdf"
options = webdriver.ChromeOptions()
options.add_argument("--headless")
driver = webdriver.Chrome(options=options)
try:
driver.get(url)
print_options = PrintOptions()
pdf_base64 = driver.print_page(print_options)
with open(output_path, "wb") as output:
output.write(b64decode(pdf_base64))
finally:
driver.quit()
print(f"Wrote {output_path}")
print_page() returns PDF data encoded as base64. Decode it before opening the destination in binary mode. The browser prints the page as it is rendered after navigation, including the current DOM and loaded styles.
2. Wait for dynamic content before printing
A navigation call can return before JavaScript has finished rendering charts, tables, images, or other application data. Wait for a meaningful element instead of relying only on a fixed sleep.
from base64 import b64decode
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.print_page_options import PrintOptions
options = webdriver.ChromeOptions()
options.add_argument("--headless")
driver = webdriver.Chrome(options=options)
try:
driver.get("https://example.com/report")
WebDriverWait(driver, 30).until(
EC.visibility_of_element_located((By.CSS_SELECTOR, "main.report"))
)
WebDriverWait(driver, 30).until(
EC.invisibility_of_element_located((By.CSS_SELECTOR, ".loading"))
)
pdf_base64 = driver.print_page(PrintOptions())
with open("report.pdf", "wb") as output:
output.write(b64decode(pdf_base64))
finally:
driver.quit()
For pages that load images lazily, scroll through the page or wait for the image elements you need. A wait for a selector proves that an element exists; it does not necessarily prove that every network request or animation has completed.
3. Configure print layout
Selenium’s PrintOptions exposes common print controls. Depending on the language binding and Selenium version, configure orientation, page dimensions, margins, background graphics, and selected pages or page ranges through that object. The official Selenium reference documents portrait and landscape orientation, page size, margins, backgrounds, and page ranges.
from base64 import b64decode
from selenium import webdriver
from selenium.webdriver.common.print_page_options import PrintOptions
options = webdriver.ChromeOptions()
options.add_argument("--headless")
driver = webdriver.Chrome(options=options)
try:
driver.get("https://example.com/invoice")
print_options = PrintOptions()
# Use the properties supported by your Selenium language binding:
# print_options.orientation = "landscape"
# print_options.page_width = 11
# print_options.page_height = 8.5
# print_options.margin_top = 0.25
# print_options.margin_bottom = 0.25
# print_options.margin_left = 0.25
# print_options.margin_right = 0.25
# print_options.background = True
# print_options.page_ranges = ["1-3"]
pdf_base64 = driver.print_page(print_options)
with open("invoice.pdf", "wb") as output:
output.write(b64decode(pdf_base64))
finally:
driver.quit()
Check the API reference for your binding before deploying these properties: names and value types differ between Python, Java, JavaScript, C#, Kotlin, and Ruby. Use CSS print rules for document-specific layout:
<style>
@media print {
.screen-only { display: none !important; }
@page { margin: 12mm; }
.invoice { break-inside: avoid; }
}
</style>
Background colors and images may be omitted unless background printing is enabled in the print options. Headers and footers are not a universal Selenium print option; if you need templated headers or footers, use Chromium’s DevTools Protocol method described below.
4. Chromium-specific control with Page.printToPDF
Chromium’s DevTools Protocol provides Page.printToPDF. It is useful when you need controls beyond Selenium’s portable print API, such as header and footer templates, CSS page-size preference, streaming, tagged PDFs, or detailed page ranges. It is Chromium-specific and protocol parameters can vary with browser versions.
from base64 import b64decode
from selenium import webdriver
options = webdriver.ChromeOptions()
options.add_argument("--headless")
driver = webdriver.Chrome(options=options)
try:
driver.get("https://example.com/report")
result = driver.execute_cdp_cmd("Page.printToPDF", {
"printBackground": True,
"preferCSSPageSize": True,
"landscape": False,
"displayHeaderFooter": True,
"headerTemplate": "<span></span>",
"footerTemplate": "<span class='pageNumber'></span> / <span class='totalPages'></span>",
"marginTop": 0.4,
"marginBottom": 0.4,
"marginLeft": 0.4,
"marginRight": 0.4
})
with open("report.pdf", "wb") as output:
output.write(b64decode(result["data"]))
finally:
driver.quit()
Choose Selenium’s print API when WebDriver portability and ordinary print settings matter. Choose Page.printToPDF when you deliberately target Chromium and require protocol-only features. Neither method downloads a PDF response that the site already serves.
5. Printing versus downloading an existing PDF
| Goal | Correct workflow |
|---|---|
| Make a PDF representation of an HTML page | Navigate with Selenium, wait for rendering, then call print_page() or Chromium Page.printToPDF. |
| Save a PDF linked by a page | Use a download or HTTP client workflow and wait for the download to complete. Printing the link page will print the HTML link, not the linked PDF. |
| Fetch an authenticated PDF endpoint | Reuse the authenticated HTTP session or configure the browser download flow. Handle response status, content type, and completion explicitly. |
The Selenium print documentation covers rendered-page printing. It does not establish one universal, browser-independent recipe for authenticated PDF downloads, so treat downloads as a separate implementation.
6. Reliability checklist for automation
- Run Chromium headless in CI and containers.
- Pin compatible browser and Selenium versions where reproducibility matters.
- Set an explicit page-load and element wait timeout.
- Wait for application data, fonts, images, and loading indicators that affect the PDF.
- Use a
try/finallyblock sodriver.quit()always runs. - Write to a temporary file, verify it begins with the PDF signature
%PDF-, then rename it into place. - Log the URL, elapsed time, browser version, and failure stage without putting secrets in logs.
- Retry transient browser startup or navigation failures with a fresh driver, rather than reusing a corrupted session indefinitely.
7. Troubleshooting common failures
“Print” returns an error in headed Chrome
Cause: Selenium’s Chromium print example requires headless mode. Fix: add --headless and verify that the browser and driver are compatible.
The PDF is blank or missing application data
Cause: printing happened before JavaScript finished rendering, or the page requires authentication. Fix: wait for a content selector and for loading indicators to disappear; establish the required cookies or login state before printing.
Images or colors are absent
Cause: images were still loading, or background printing was disabled. Fix: wait for image elements and enable the background option supported by your binding. CSS @media print rules can also hide or change content.
Only the first page appears
Cause: page ranges were configured unintentionally, content was hidden by print CSS, or the document was not laid out for printing. Fix: remove the range restriction, inspect @media print styles, and set page dimensions and margins deliberately.
WebDriverException during startup
Cause: missing browser binaries, an incompatible driver, or container restrictions. Fix: install matching browser and driver versions, confirm the executable is on the runtime path, and use the required container flags for your environment.
The file cannot be opened
Cause: base64 data was written as text or was not decoded. Fix: call b64decode() and write with "wb"; then check the first bytes are %PDF-.
8. Performance, reliability, and cost considerations
PDF generation time is dominated by browser startup, page navigation, JavaScript execution, fonts, images, and network dependencies. Reuse a driver for a controlled batch of pages when isolation requirements allow it, but restart after repeated crashes or memory growth. Keep waits targeted: a short selector wait is usually more predictable than a large fixed sleep.
For reliable output, make the page deterministic: freeze timestamps where appropriate, wait for charts to finish, use print-specific CSS, and avoid capturing while transitions are running. Test representative pages with long tables, images, custom fonts, and multiple page breaks.
Self-hosted Selenium costs come from the machines, browser processes, network traffic, and maintenance needed to keep browser and driver versions compatible. Selenium itself does not charge per PDF; infrastructure and any third-party services used by the page determine operating cost.
9. Or skip the browser setup
If you need a hosted screenshot or PDF capture endpoint instead of maintaining Selenium and Chromium, ScreenshotNeo accepts one GET request and returns a PDF or image. See the ScreenshotNeo API documentation for the available options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and the response identifies the page verdict and billing status in headers. Its MCP server lets AI agents call take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
10. FAQ
Does Selenium create a PDF from the current DOM?
Yes. The print API creates a PDF representation of the page currently rendered by the browser, after navigation and any waits you perform.
Can I use the same code for Firefox?
Firefox also exposes a print-page API, but browser support and output details are not identical. Follow the binding’s browser-specific documentation and verify the resulting layout.
Should I use Selenium or DevTools Protocol?
Use Selenium’s print API for the WebDriver-oriented path and common options. Use Chromium’s protocol when you need Chromium-only controls such as templates, streaming, or tagged-PDF settings.
Why is my downloaded PDF not produced by print_page()?
print_page() prints HTML. An already hosted PDF must be retrieved through a download or HTTP workflow.
Is headless mode required everywhere?
Selenium’s Chromium print example states that Chromium printing requires headless mode. Confirm the requirement for your selected browser and current Selenium release before deployment.


