Convert HTML Files to PDF in Java
Convert HTML files to PDF in Java with OpenHTMLtoPDF, browser-backed alternatives, complete code, troubleshooting, and deployment guidance.

Direct answer: For well-formed XHTML or deliberately authored HTML with supported CSS, use OpenHTMLtoPDF from Java. It is a pure-Java renderer, but it is not a browser: it does not execute JavaScript and does not implement many modern CSS features such as flexbox and grid. If your document depends on browser JavaScript or modern HTML5/CSS3, evaluate Flying Saucer’s Chrome PDF module, which delegates rendering to chrome-headless-shell.
1. Choose the rendering route
| Requirement | Route | What to expect |
|---|---|---|
| Controlled XHTML/XML templates and supported CSS | OpenHTMLtoPDF | Pure Java; CSS 2.1 and a subset of HTML5; no JavaScript; no flex or grid support. Read the project README before selecting a release. |
| XHTML/CSS with the Flying Saucer family | Flying Saucer PDF module | Pure-Java XML/XHTML renderer based on CSS 2.1. The README lists Java 11+ for 9.5.0, Java 17+ for 9.6.0, and Java 21+ for 10.0.0. See the official repository. |
| Modern HTML5/CSS3 or JavaScript-driven pages | Flying Saucer Chrome PDF module | Uses chrome-headless-shell. Package and operate the external browser and validate the exact deployment environment. |
| Low-level PDF creation or editing after conversion | Apache PDFBox | Creates and manipulates PDFs, but its official description does not make it an HTML/CSS renderer. |
There is no renderer that automatically behaves exactly like Chrome. Test representative templates for pagination, fonts, images, links, page breaks, and accessibility requirements before committing to a library. OpenHTMLtoPDF’s own documentation cautions that modern HTML5 should be crafted for its engine rather than passed in unchanged.
2. Convert an HTML file with OpenHTMLtoPDF
Prerequisites
- Java and Maven or Gradle.
- Well-formed XHTML/XML markup. Close every element and use valid nesting.
- CSS supported by the selected OpenHTMLtoPDF release.
- A writable output path and readable asset paths.
The exact artifact and version should come from the current project integration guidance. The example below uses the commonly published openhtmltopdf-pdfbox artifact at version 1.0.10; verify that version and its transitive dependencies in your build before production use. A Maven parent POM is not automatically the runtime module you need.

Maven
<dependency>
<groupId>com.openhtmltopdf</groupId>
<artifactId>openhtmltopdf-pdfbox</artifactId>
<version>1.0.10</version>
</dependency>
Gradle
implementation("com.openhtmltopdf:openhtmltopdf-pdfbox:1.0.10")
Complete Java program
import com.openhtmltopdf.pdfboxout.PdfRendererBuilder;
import java.io.File;
import java.io.FileOutputStream;
import java.io.IOException;
import java.nio.file.Path;
public final class HtmlToPdf {
private HtmlToPdf() {}
public static void main(String[] args) throws IOException {
if (args.length != 2) {
System.err.println("Usage: java HtmlToPdf input.html output.pdf");
System.exit(2);
}
File input = Path.of(args[0]).toFile();
File output = Path.of(args[1]).toFile();
if (!input.isFile()) {
throw new IOException("Input HTML file does not exist: " + input);
}
File parent = output.getParentFile();
if (parent != null) {
parent.mkdirs();
}
PdfRendererBuilder builder = new PdfRendererBuilder();
builder.useFastMode();
builder.withFile(input);
// The base URI lets relative CSS, images, and fonts resolve beside input.html.
builder.toStream(new FileOutputStream(output));
builder.run();
System.out.println("Wrote " + output.getAbsolutePath());
}
}
Run it with:
mvn compile exec:java -Dexec.mainClass=HtmlToPdf -Dexec.args="report.html report.pdf"
For a service, prefer a Path-based input policy and stream the result to the caller. Keep the output stream open until builder.run() returns, then close it with try-with-resources:
try (FileOutputStream pdf = new FileOutputStream("report.pdf")) {
new PdfRendererBuilder()
.useFastMode()
.withFile(new File("report.html"))
.toStream(pdf)
.run();
}
3. Prepare HTML and CSS the renderer can paginate
Use valid, self-contained markup
<!DOCTYPE html>
<html xmlns="http://www.w3.org/1999/xhtml">
<head>
<meta charset="UTF-8" />
<style>
@page { size: A4; margin: 18mm 15mm; }
body { font-family: "DejaVu Sans", sans-serif; font-size: 10pt; }
h1, h2 { page-break-after: avoid; }
.invoice { page-break-inside: avoid; }
.page-break { page-break-before: always; }
table { width: 100%; border-collapse: collapse; }
td, th { border: 0.2mm solid #bbb; padding: 2mm; }
</style>
</head>
<body>
<h1>Invoice</h1>
<section class="invoice">...</section>
</body>
</html>
OpenHTMLtoPDF works best when you author for its supported engine. The project recommends table layouts for predictable results and warns against floats near page breaks. Replace flexbox, grid, sticky positioning, browser-specific selectors, and JavaScript-generated content with server-rendered markup or supported table/block layouts.
Resolve relative assets
Relative CSS, images, and fonts need a base URI. withFile(input) supplies the input file location. When creating a document from a string, supply the base URI explicitly:
String html = Files.readString(Path.of("report.html"));
try (FileOutputStream pdf = new FileOutputStream("report.pdf")) {
new PdfRendererBuilder()
.withHtmlContent(html, Path.of("report.html").toUri().toString())
.toStream(pdf)
.run();
}
Use absolute, allow-listed asset locations in production. Missing assets commonly produce blank image areas or fallback fonts.
Fonts
Install the required fonts in the runtime image or register them through the renderer’s font APIs for your selected version. Verify that the license permits server redistribution. A PDF that looks correct on a laptop can change in a minimal container with different installed fonts.
4. When JavaScript or modern CSS is required
OpenHTMLtoPDF does not execute JavaScript and does not implement many modern standards. A page that fills a chart after window.onload, uses client-side templating, or relies on flex/grid may render empty or incorrectly. Flying Saucer documents a Chrome PDF module that delegates to chrome-headless-shell for modern HTML5/CSS3. That route adds a browser binary, process startup, sandboxing, patching, and container deployment requirements. Test the exact HTML, browser version, fonts, and operating system rather than assuming compatibility from the library name.
If you control the application, another reliable pattern is to render data into static HTML on the server first, then pass that normalized XHTML to a pure-Java renderer. This removes timing races and makes failures reproducible.
5. Security and licensing checklist
- Treat uploaded HTML, CSS, SVG, images, and URLs as untrusted input.
- Restrict file and network access so a document cannot read local secrets or call internal services.
- Set size, page-count, render-time, and output-size limits.
- Keep XML parsing hardened against XXE and review the exact release’s security notes. Flying Saucer’s changelog records recent
DocumentBuilderFactoryhardening, but that is not a guarantee for every version or application. - Review licenses for every artifact. OpenHTMLtoPDF states LGPL 2.1-or-later; its PDF/A testing module has a separate GPL exception and is not distributed to Maven Central. PDFBox is Apache License 2.0.
- Scan transitive dependencies and update according to the selected project’s security guidance.
6. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Blank or partly blank PDF | Malformed markup, unsupported CSS, or JavaScript-generated content | Validate and normalize XHTML; move content generation to Java; replace unsupported layout rules. |
| Images or CSS are missing | No usable base URI, blocked file access, or incorrect relative paths | Use withFile or withHtmlContent(html, baseUri); verify each asset path and permission. |
| Text wraps differently in production | Different fonts or font versions | Package and register the same fonts; inspect embedded fonts in the PDF. |
| Content overlaps a page boundary | Floats, fixed heights, or unsupported page-break behavior | Prefer tables and normal flow; remove fixed heights; add page-break-inside: avoid to small blocks. |
| Charts, menus, or totals are absent | They are created by JavaScript | Pre-render them, or use the Chrome-backed route and validate its browser runtime. |
| OutOfMemoryError or very slow renders | Very large images, long documents, or concurrent renders | Resize images, cap input size and page count, process with bounded concurrency, and profile representative documents. |
| Security scanner reports XML issues | Parser or dependency configuration | Upgrade the exact renderer line, apply its hardening guidance, and disable external entities and unrestricted URL access. |
7. Performance, reliability, and cost
The inspected project pages do not establish an empirical performance winner for typical workloads. Measure your own templates with cold and warm JVM runs, image-heavy pages, long tables, custom fonts, and concurrent requests. Track render duration, heap use, output size, page count, and failure reason.
- Reuse application infrastructure, but avoid sharing mutable renderer state between requests unless the selected version documents it as safe.
- Cache immutable CSS, fonts, and images. Downscale source images before rendering.
- Use a bounded worker pool and a queue so one pathological document cannot exhaust the JVM.
- Record input identity, renderer version, Java version, page count, duration, and output hash for reproducibility.
- For browser-backed rendering, account for browser startup and process supervision in capacity planning.
- Software license and infrastructure cost depend on the artifacts, browser runtime, hosting, and your usage; the research sources do not provide a universal benchmark or price.
8. Or skip the browser setup
If the HTML is already available at a public URL, ScreenshotNeo can return a PDF through one GET request. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. See the ScreenshotNeo API documentation for the current option names.

cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://screenshotneo.com -d format=pdf -o page.pdf
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://screenshotneo.com", "format": "pdf"},
timeout=90,
)
r.raise_for_status()
open("page.pdf", "wb").write(r.content)
Node.js
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://screenshotneo.com',
format: 'pdf'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
const fs = await import('node:fs/promises');
await fs.writeFile('page.pdf', Buffer.from(await res.arrayBuffer()));
ScreenshotNeo also has an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. Every plan includes the features; 1,000 screenshots per month are free with no card, and paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
9. FAQ
Can OpenHTMLtoPDF convert any HTML page?
No. It targets well-formed XML/XHTML and supported CSS. JavaScript, flexbox, grid, and other browser features may not render.
Should I use PDFBox instead?
Use PDFBox when you need to create or manipulate PDF objects directly. It is not presented by its official project page as an HTML/CSS renderer.
Which Java version should I install?
Follow the selected renderer release. Flying Saucer lists different minimums for 9.5.0, 9.6.0, and 10.0.0, so do not infer a single requirement for the whole project.
How do I guarantee identical output?
Pin the renderer, Java runtime, fonts, browser version if used, and input assets. Keep representative PDFs as regression fixtures and inspect page breaks after upgrades.


