ScreenshotNeo

BlogHTML to image & PDF

Convert HTML to PDF in Java: Code Examples

Convert HTML strings or files to PDFs in Java with iText pdfHTML and OpenHTMLtoPDF, including CSS, images, fonts, troubleshooting, and production guidance.

By the ScreenshotNeo team1 October 20267 min read

Use iText pdfHTML when you need broad HTML/CSS support, tagged or accessible PDFs, PDF/A, forms, or further iText processing. Use OpenHTMLtoPDF when an LGPL, PDFBox-based renderer is a better fit for controlled XHTML/CSS templates that do not require JavaScript, flexbox, or CSS grid.

This guide shows runnable Java code for HTML strings and files, explains relative assets, fonts, page layout, accessibility, alternatives, production limits, and common failures.

1. Choose a Java HTML-to-PDF library

Requirement Recommended choice Reason
Modern HTML/CSS, accessibility, tagging, PDF/A, forms, or iText post-processing iText pdfHTML Maintained iText add-on with a Java API for standards-oriented PDF generation.
LGPL licensing and controlled XHTML/CSS templates OpenHTMLtoPDF Pure Java, PDFBox-based rendering for a documented subset of XHTML and CSS.
Arbitrary web pages that depend on JavaScript, browser APIs, or dynamic consent UI Use a browser renderer or a screenshot/PDF service Neither library is a complete web browser.

Do not start new code with the old iText HTMLWorker or XML Worker examples. HTMLWorker was deprecated and removed; XML Worker expected predictable XHTML/CSS rather than arbitrary web pages.

2. Convert an HTML string with iText pdfHTML

Add the pdfHTML and iText Core dependencies using the versions approved for your project. The official repository contains the current dependency coordinates and examples.

package example;

import com.itextpdf.html2pdf.HtmlConverter;
import java.io.FileOutputStream;
import java.io.IOException;

public class HtmlStringToPdf {
    public static void main(String[] args) throws IOException {
        String html = """
            <!doctype html>
            <html>
              <head>
                <meta charset=\"UTF-8\">
                <style>
                  @page { size: A4; margin: 20mm; }
                  body { font-family: sans-serif; color: #222; }
                  h1 { color: #174ea6; }
                </style>
              </head>
              <body>
                <h1>Invoice 1001</h1>
                <p>Generated from an HTML string.</p>
              </body>
            </html>
            """;

        HtmlConverter.convertToPdf(html, new FileOutputStream("out.pdf"));
    }
}

The same conversion can write to any OutputStream, File, PdfWriter, or PdfDocument. Use convertToDocument when you need to append iText content after parsing, or convertToElements when you want to insert parsed elements into your own document flow.

3. Convert an HTML file and resolve CSS and images

Relative URLs such as img/logo.png cannot be resolved from an arbitrary stream unless you provide a base URI. Set the base URI to the directory that contains the HTML file and its assets.

package example;

import com.itextpdf.html2pdf.ConverterProperties;
import com.itextpdf.html2pdf.HtmlConverter;
import java.io.FileInputStream;
import java.io.FileOutputStream;
import java.io.IOException;

public class HtmlFileToPdf {
    public static void main(String[] args) throws IOException {
        String source = "./templates/invoice.html";
        String output = "./out/invoice.pdf";

        ConverterProperties properties = new ConverterProperties();
        properties.setBaseUri("./templates/");

        try (FileInputStream html = new FileInputStream(source);
             FileOutputStream pdf = new FileOutputStream(output)) {
            HtmlConverter.convertToPdf(html, pdf, properties);
        }
    }
}

When you pass a File, iText can use the file’s parent directory as the default base URI. With streams, set it explicitly. Prefer absolute, controlled asset paths in production and make sure the process can read every referenced image, stylesheet, and font.

4. Add fonts, metadata, tagging, and PDF/A

For accessible output, create a PDF document, enable tagging, and pass it to the converter. The exact font registration and PDF/A setup depend on your pdfHTML version and compliance target; validate the generated file with the validator required by your project.

import com.itextpdf.html2pdf.ConverterProperties;
import com.itextpdf.html2pdf.HtmlConverter;
import com.itextpdf.kernel.pdf.PdfDocument;
import com.itextpdf.kernel.pdf.PdfWriter;

ConverterProperties properties = new ConverterProperties();
properties.setBaseUri("./templates/");

try (PdfWriter writer = new PdfWriter("tagged.pdf");
     PdfDocument pdf = new PdfDocument(writer)) {
    pdf.setTagged();
    HtmlConverter.convertToPdf(html, pdf, properties);
}

The iText examples cover tagged PDFs, PDF/A-3B, custom fonts, forms, SVG, Arabic, and Hebrew. Treat those as version-specific capabilities: pin a library version, then verify your exact document with accessibility and archival validation before release.

5. OpenHTMLtoPDF example

OpenHTMLtoPDF is a pure-Java renderer built on PDFBox. It supports a reasonable subset of well-formed XML/XHTML and CSS 2.1 plus later features. It does not execute JavaScript and does not implement many modern browser standards, including flexbox and grid.

import com.openhtmltopdf.pdfboxout.PdfRendererBuilder;
import java.io.FileOutputStream;

public class OpenHtmlToPdf {
    public static void main(String[] args) throws Exception {
        String html = "<html><body><h1>Hello</h1></body></html>";

        try (FileOutputStream output = new FileOutputStream("out.pdf")) {
            PdfRendererBuilder builder = new PdfRendererBuilder();
            builder.useFastMode();
            builder.withHtmlContent(html, "file:///absolute/path/to/templates/");
            builder.toStream(output);
            builder.run();
        }
    }
}

Craft templates for this engine: use well-formed markup, prefer tables for layouts that cross page boundaries, and avoid floats near page breaks. The project documents Java 8 as the minimum runtime; verify the current release before pinning a dependency.

6. CSS, images, and fonts that work reliably

  • Use a correct <meta charset=\"UTF-8\"> and pass Java strings as UTF-8 data.
  • Give every relative URL a base URI. Confirm the process has permission to read local files.
  • Prefer HTTPS or local assets you control. Remote assets make output dependent on DNS, TLS, latency, and availability.
  • Register and embed the fonts required for brand or multilingual output; check licenses before distributing font files.
  • Design explicit print styles with @page, margins, page size, and page-break rules.
  • Test SVG, RTL text, forms, and complex selectors against the selected library version rather than assuming browser parity.

7. Page size, margins, and large documents

Define paper size and margins in print CSS, for example @page { size: A4 landscape; margin: 12mm; }. Break a large report into sections when memory usage becomes a concern, stream output where the API permits it, and avoid embedding unnecessarily large raster images. Measure conversion time and heap usage with your real templates; the supplied sources publish no general throughput benchmark.

8. Troubleshooting

Symptom Likely cause Fix
Images or CSS are missing No base URI or an incorrect one Set ConverterProperties.setBaseUri (or the renderer’s equivalent) to the asset directory.
Output is blank Malformed HTML, an unreadable stream, or unsupported markup Validate the HTML, close or flush streams, and reduce the template to a minimal reproducer.
Flexbox or grid layout collapses OpenHTMLtoPDF is not a browser and lacks these standards Rewrite the template using supported CSS, or choose a browser-based renderer.
JavaScript-generated content is absent Server-side converters do not execute page JavaScript Render the content before conversion or use a browser automation/service that runs JavaScript.
Fonts show as boxes or wrong glyphs Font is unavailable, not embedded, or lacks the needed glyphs Register an appropriate font, embed it where required, and test the target scripts.
Content overlaps or breaks at page boundaries Unsupported CSS, floats, or an element too large for a page Use print-specific CSS, table layouts for repeated structure, and explicit break rules.
Remote images intermittently fail Network or TLS dependency during conversion Download and validate assets first, or serve them from a reliable internal location.

9. Production checklist

  1. Pin and periodically review the renderer version.
  2. Validate HTML before conversion and reject untrusted input when templates can contain user data.
  3. Set conversion, network, and job time limits in the surrounding service.
  4. Bound input size, image dimensions, and number of external resources.
  5. Log the template version, renderer version, elapsed time, and output size.
  6. Keep golden PDFs or structural assertions for regression testing across library upgrades.
  7. Validate accessibility, PDF/A, signatures, or forms when those are requirements.

10. Or skip the browser setup

If the source is a live website rather than a controlled HTML template, ScreenshotNeo provides a single API request for a screenshot or PDF. Its capture flow accepts cookie and consent banners, removes more than 60 known consent platforms plus newsletter popups and chat widgets, and lets you turn each step off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and billing state.

See the ScreenshotNeo API documentation for the complete option set, including full-page capture, CSS selectors, waits, custom CSS and JavaScript, headers, cookies, user agents, geolocation, PDF paper settings, caching, signed links, asynchronous jobs, webhooks, bulk capture, and usage reporting.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const bytes = await res.arrayBuffer();
require('node:fs').writeFileSync('shot.webp', Buffer.from(bytes));

ScreenshotNeo includes an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

11. FAQ

Can Java convert a complete modern website to PDF?

Not reliably with a non-browser HTML converter. JavaScript, flexbox, grid, and browser-specific behavior require a browser renderer or a service that runs one.

Which library should a commercial application choose?

Compare HTML/CSS fidelity, accessibility, PDF/A, forms, fonts, runtime support, performance, post-processing needs, and license or support requirements. iText pdfHTML is the usual starting point for advanced PDF features; OpenHTMLtoPDF fits controlled templates where LGPL and PDFBox matter.

Why does a relative image work in one example but not another?

A file-based conversion can infer the source file’s parent directory. A stream or string has no location, so you must provide a base URI.

How do I append a signature page or other iText content?

Use convertToDocument or pass a managed PdfDocument, then append content through the iText APIs after HTML conversion.