How to Generate PDFs from HTML with Playwright and Java
Generate PDFs from HTML or a URL with Playwright for Java. Configure print styles, page size, margins, backgrounds, headers, and output options.
Use Playwright’s Java API to create a browser page, load HTML with Page.setContent() or navigate to a URL, then call Page.pdf(). PDF generation uses print CSS media by default. To render screen styles instead, call Page.emulateMedia() with SCREEN before exporting. Playwright’s Java Page API documents that page.pdf() generates a PDF using print CSS media: Playwright Java Page API.
The examples use Chromium. Playwright supports multiple browser engines generally, but do not assume PDF generation behaves identically across engines; check the documentation for the Playwright version and engine you deploy. See the browser documentation.
1. Set up a Java Playwright project
Add the Playwright Java dependency using the installation instructions for your build system and install the browser binaries required by your environment. The official Java introduction covers setup. The code below uses the documented Java API; it is illustrative and has not been run as part of this article.
Use a Playwright release that supports the PDF options you need. For example, tagged PDF and outline options are documented as added in v1.42. Consult the API page for version annotations and current signatures.
2. Generate a PDF from an HTML string
This minimal example launches Chromium, writes a small HTML document into a page, and saves the PDF as report.pdf. The browser and Playwright resources are closed even if an exception occurs.
import com.microsoft.playwright.Browser;
import com.microsoft.playwright.Page;
import com.microsoft.playwright.Playwright;
import java.nio.file.Paths;
public class HtmlToPdf {
public static void main(String[] args) {
try (Playwright playwright = Playwright.create()) {
Browser browser = playwright.chromium().launch();
try {
Page page = browser.newPage();
String html = "<!doctype html>"
+ "<html><head><meta charset='utf-8'>"
+ "<title>Quarterly report</title></head>"
+ "<body><h1>Quarterly report</h1>"
+ "<p>Generated from HTML with Playwright.</p>"
+ "</body></html>";
page.setContent(html);
page.pdf(new Page.PdfOptions().setPath(Paths.get("report.pdf")));
} finally {
browser.close();
}
}
}
}
Page.pdf() returns a PDF buffer. Supplying setPath(Paths.get(...)) writes the file to that path; without a path, consume the returned buffer yourself if you need to persist or transmit it. setContent() assigns markup to the page and uses document.write() internally, so supply a complete document and avoid depending on prior page state. See Page.setContent and Page.pdf API details.
3. Generate a PDF from a URL
Navigate before exporting when the source is an existing web page. Handle navigation failures and use an application-specific readiness condition when the page fills content asynchronously.
import com.microsoft.playwright.Browser;
import com.microsoft.playwright.Page;
import com.microsoft.playwright.Playwright;
import java.nio.file.Paths;
public class UrlToPdf {
public static void main(String[] args) {
try (Playwright playwright = Playwright.create()) {
Browser browser = playwright.chromium().launch();
try {
Page page = browser.newPage();
page.navigate("https://example.com");
page.pdf(new Page.PdfOptions().setPath(Paths.get("page.pdf")));
} finally {
browser.close();
}
}
}
}
navigate() waits according to its navigation behavior, and setContent() defaults to waiting for the load event. A load event does not prove that your app’s data, web fonts, lazy images, or client-side rendering are ready. Wait for a selector or other known signal that means the particular document is ready to print.
4. Configure PDF layout and output
Pass options to Page.PdfOptions. These are the main controls documented by the Java API:
| Option | Purpose and behavior |
|---|---|
setFormat("A4") |
Select a named paper format. The default is Letter; common named formats such as A4 are supported. |
setWidth("210mm"), setHeight("297mm") |
Set explicit dimensions with units such as px, in, cm, or mm. Choose either a named format or explicit dimensions to express the desired paper size. |
setLandscape(true) |
Use landscape orientation. Portrait is the default. |
setMargin(...) |
Set page margins. Margins default to none. Use the margin fields supported by your installed Java API version. |
setScale(1.0) |
Scale the page content; valid values are 0.1 through 2. |
setPageRanges("1-3, 5") |
Restrict the exported page range using the documented page-range syntax. |
setPrintBackground(true) |
Include background graphics, which are omitted unless enabled. |
setPreferCSSPageSize(true) |
Give CSS @page dimensions priority. Otherwise, content is scaled to fit the PDF paper size selected in options. |
setPath(path) |
Write the generated PDF to a filesystem path. Without it, the API returns a buffer but does not save a file. |
setDisplayHeaderFooter(true) |
Enable header and footer templates. Templates can use documented classes for date, title, URL, page number, and total pages. |
| Tagged output and outline | Options are available for tagged PDFs and document outlines; the API documents these as added in v1.42. They do not by themselves guarantee accessibility conformance. |
Option names and overloads can evolve. Check the Java Page API for the exact version in your project, especially when configuring margins, templates, or newer output options.
Example: A4, landscape, margins, and backgrounds
Page.PdfOptions options = new Page.PdfOptions()
.setPath(Paths.get("report-a4.pdf"))
.setFormat("A4")
.setLandscape(true)
.setPrintBackground(true)
.setScale(1.0);
page.pdf(options);
Add the margin configuration supported by your Playwright Java version when the print layout needs reserved space. If the HTML defines its own paper dimensions, use setPreferCSSPageSize(true) so CSS @page sizing takes priority.
5. Choose print or screen styling
By default, Page.pdf() renders with print media. This activates @media print styles and can hide navigation, alter colors, or change layout. If the output should match screen media, set it before calling pdf():
page.emulateMedia(new Page.EmulateMediaOptions()
.setMedia(com.microsoft.playwright.options.Media.SCREEN));
page.pdf(new Page.PdfOptions().setPath(Paths.get("screen-style.pdf")));
Use print media for documents designed for paper and screen media when the screen layout is the intended source. Print colors can be modified by default. When exact colors matter, the API points to the CSS property -webkit-print-color-adjust; for example, apply it to the relevant elements in the document’s print stylesheet. Also enable setPrintBackground(true) if background graphics need to appear.
CSS page size versus PDF paper settings
These controls solve different layout needs:
- Use
setFormator dimensions when the export should follow a known paper size such as A4. - Use CSS
@pagerules for document-controlled page geometry, andsetPreferCSSPageSize(true)to give those rules priority. - When CSS page size does not take priority, Playwright says content is scaled to fit the configured paper size.
6. Wait for content that loads after navigation
Page readiness is application-specific. A navigation or setContent() load wait does not guarantee a chart, API response, custom font, or lazy image is ready. Add a meaningful readiness signal before PDF generation. For example, if the page displays a known report container only after data is ready:
page.navigate("https://example.com/report");
page.locator("#report-ready").waitFor();
page.pdf(new Page.PdfOptions().setPath(Paths.get("report.pdf")));
For generated HTML, include the content and resources needed by the browser, and wait on an app-specific selector or state when scripts populate the page. Avoid treating a fixed delay as a universal readiness guarantee. If a required image or font is remote, verify it is reachable from the browser runtime and is loaded before capture.
7. Header and footer templates
Enable headers and footers with setDisplayHeaderFooter(true) and provide the documented header or footer template options for your Java API version. Templates can include classes for date, title, URL, page number, and total pages. Two limits matter: scripts in templates are not evaluated, and page styles are not visible inside the templates. Keep their markup self-contained and style them using the supported template behavior documented by Playwright.
8. Or skip the browser setup
If the goal is a screenshot of a web page rather than a Java-managed PDF workflow, ScreenshotNeo offers a website screenshot API and MCP server. A single GET request can return PNG, JPEG, WebP, or PDF. See the ScreenshotNeo site and API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.pdf
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.pdf", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.pdf', bytes));
Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents use screenshot tools. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, with no card.
9. Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| No PDF file appears | No path was supplied, or the process cannot write to the target directory. | Set setPath(Paths.get(...)), check the returned buffer path if saving it yourself, and confirm directory permissions. |
| Styles differ from the browser screenshot | PDF uses print media by default. | Use print CSS intentionally, or set screen media before pdf() when screen styling is required. |
| Backgrounds or exact colors are missing | Background printing is disabled or print color adjustment changes colors. | Enable setPrintBackground(true) and use -webkit-print-color-adjust for exact print colors where appropriate. |
| Content is clipped or scaled unexpectedly | Paper size, margins, scaling, or CSS @page rules conflict. |
Review format and dimensions, margins, scale bounds, and setPreferCSSPageSize. |
| Last section, chart, or image is absent | Content was still loading or lazy resources were not triggered. | Wait for a page-specific ready selector or state; ensure remote resources can load in the browser environment. |
| Navigation or browser launch fails | The URL is unreachable from the runtime, browser binaries are missing, or the selected browser cannot launch in its environment. | Follow Java installation guidance, install the required browser, check network access and launch environment, and inspect the thrown Playwright exception. |
| Header/footer has no expected styling or dynamic content | Page styles do not apply inside templates and template scripts are not evaluated. | Use self-contained template markup and the documented template classes rather than relying on page CSS or script execution. |
| PDF behavior differs by browser | PDF support and behavior may depend on engine and version. | Verify the chosen Java API and engine documentation; do not assume cross-engine parity. |
10. Performance, reliability, and cost
PDF export runs a browser, so include browser startup, navigation, resource loading, and PDF rendering in the job’s latency and memory budget. For repeated work, choose a lifecycle strategy that reuses browser processes safely for your service rather than launching one for every document; ensure pages and contexts are closed and that concurrent jobs have enough memory. The appropriate concurrency depends on document complexity and deployment resources, so measure it with representative pages instead of assuming a fixed throughput.
For reliability, set application-level timeouts around navigation and readiness waits, capture logs for failed navigations, and handle filesystem or returned-buffer errors. A generic load wait is not a substitute for confirming that the document’s data and resources are ready. Pin and review the Playwright version used by the deployment because the API documentation is rolling.
Playwright itself is an automation library rather than a per-PDF price in this workflow. Budget for the infrastructure that runs the browser, stores or transfers PDFs, and supports retries. Retries should be limited to transient failures; repeatedly rendering a permanently broken page adds load without fixing it.
11. Frequently asked questions
Can I return the PDF directly from a Java service?
Yes. Use the PDF buffer returned by Page.pdf() and write it to your HTTP response or storage layer. Set a path only when you want Playwright to save it to disk.
Does calling setContent() execute arbitrary scripts in the HTML?
setContent() assigns content using document.write(). Script behavior depends on the markup and browser execution; if scripts build the final document, wait for a defined ready condition before export.
Do tagged PDFs guarantee accessibility?
No. Tagged output is an available PDF option, but the option alone does not establish that the source document or resulting PDF conforms to an accessibility standard.
Can I use CSS to choose the paper dimensions?
Yes. Define CSS @page rules and enable setPreferCSSPageSize(true) when those dimensions should take priority over the PDF paper setting.


