ScreenshotNeo

BlogHTML to image & PDF

How to Convert XHTML to PDF with iText in Java

Convert XHTML and CSS to PDF in Java with iText pdfHTML, configure resources, avoid legacy APIs, and troubleshoot common rendering issues.

By the ScreenshotNeo team1 October 20267 min read

Use iText pdfHTML with iText Core. It is iText’s current HTML/XML-to-PDF add-on for Java. The high-level entry point is HtmlConverter.convertToPdf; for real XHTML files, configure a base URI so relative CSS, images and links can be resolved.

iText identifies XML Worker as the iText 5-era option and recommends pdfHTML for current iText Core projects. See the iText product overview and the official HTML-to-PDF tutorial.

1. Add pdfHTML to your Java project

Use the com.itextpdf:html2pdf dependency. Match the pdfHTML and iText Core versions covered by your project’s license and dependency policy; iText’s installation guide explains the supported setup.

<dependency>
  <groupId>com.itextpdf</groupId>
  <artifactId>html2pdf</artifactId>
  <version>6.3.3</version>
</dependency>

The current feature reference associates pdfHTML 6.3.3 with iText Core 9.7.0. Check the versioned support matrix before relying on a particular CSS property or HTML element.

2. Convert an XHTML string

This is the smallest complete Java example. The input must be valid enough for the parser and should use HTML/CSS that pdfHTML supports.

import com.itextpdf.html2pdf.HtmlConverter;
import java.io.FileOutputStream;

public class XhtmlStringToPdf {
    public static void main(String[] args) throws Exception {
        String xhtml = "<!doctype html>"
                + "<html xmlns=\"http://www.w3.org/1999/xhtml\">"
                + "<head><meta charset=\"UTF-8\" />"
                + "<style>body { font-family: sans-serif; } h1 { color: #17324d; }</style>"
                + "</head>"
                + "<body><h1>XHTML report</h1><p>Created with pdfHTML.</p></body>"
                + "</html>";

        HtmlConverter.convertToPdf(xhtml, new FileOutputStream("output.pdf"));
    }
}

The one-line conversion shown above follows iText’s official tutorial. It is suitable for self-contained markup. External stylesheets, images and fonts require resource configuration.

3. Convert an XHTML file with relative resources

Set a base URI to the directory that contains the XHTML document and its relative resources. The stream overload below keeps resource handling explicit and works well when the input and output are files.

import com.itextpdf.html2pdf.ConverterProperties;
import com.itextpdf.html2pdf.HtmlConverter;
import java.io.FileInputStream;
import java.io.FileOutputStream;
import java.nio.file.Path;

public class XhtmlFileToPdf {
    public static void main(String[] args) throws Exception {
        Path input = Path.of("src/main/resources/report.xhtml");
        Path output = Path.of("target/report.pdf");

        ConverterProperties properties = new ConverterProperties()
                .setBaseUri(input.toAbsolutePath().getParent().toString());

        try (FileInputStream in = new FileInputStream(input.toFile());
             FileOutputStream out = new FileOutputStream(output.toFile())) {
            HtmlConverter.convertToPdf(in, out, properties);
        }
    }
}

For a document containing <link href="css/report.css" rel="stylesheet" /> or <img src="images/logo.png" />, the base URI must point to the directory from which those relative paths resolve. Use absolute, controlled resource locations when input comes from an untrusted source.

4. Prepare XHTML that converts predictably

  • Close every element and quote every attribute. XHTML-style empty elements such as <br /> and <img ... /> avoid ambiguous parsing.
  • Declare the character encoding, preferably with <meta charset="UTF-8" />, and read the Java input as UTF-8.
  • Keep CSS in linked files or <style> blocks that pdfHTML supports. Browser-only behavior, JavaScript-driven layout and unsupported CSS are not automatically reproduced.
  • Give images a resolvable URL or file path and include dimensions when stable pagination matters.
  • Design for pages. Long unbreakable strings, very large tables and elements that cannot split can create overflow or unexpected page breaks.

5. Check HTML and CSS support before relying on it

pdfHTML is not a browser engine. The official feature matrix is the authority for supported and unsupported tags, properties and layout behavior for a specific release. Treat its details as versioned: support changes over time.

pdfHTML 6.3.3 lists support for CSS :is(), :where() and :not() selectors, improved tolerance of malformed CSS, and fixes related to CSS Grid pagination and list-rendering performance. Those release notes do not mean every browser CSS feature is supported.

6. Licensing and deployment checklist

  1. Choose versions of html2pdf and iText Core that are compatible with each other.
  2. Read iText’s installation and licensing guide before shipping.
  3. For noncommercial use, iText describes use under the AGPL subject to its terms.
  4. For closed-source commercial software, iText requires a commercial license for both iText Core and pdfHTML and describes the license-key library installation.
  5. Record the exact pdfHTML/Core versions in your build so later rendering changes can be traced.

7. iText 5, XML Worker and HTMLWorker migration

XML Worker belongs to iText 5, which iText describes as end of life. HTMLWorker was deprecated and removed from recent versions and was intended for simple snippets rather than complete CSS-driven pages. Migrating is more than changing an import: review dependencies, resource resolution, CSS coverage, pagination and licensing.

8. Troubleshooting

Symptom Likely cause Fix
Images or CSS are missing No base URI or an incorrect relative path Set ConverterProperties.setBaseUri to the resource root; verify the resolved files and permissions.
Output is blank or incomplete Malformed XHTML, an exception while loading a resource, or unsupported markup Validate the source, convert a minimal fragment, inspect the exception, then compare elements and CSS with the support matrix.
Fonts differ from the browser The required font is unavailable or not configured Install or register the font available to the Java process and test the PDF on the deployment machine.
CSS appears ignored The property is unsupported or the stylesheet was not resolved Confirm the stylesheet URL and check that property in the release-specific feature reference.
Content is clipped at page boundaries Large unbreakable elements, fixed heights or pagination-sensitive layout Remove unnecessary fixed heights, allow wrapping, split large tables or redesign the page for printable dimensions.
Old classes cannot be found Code still targets HTMLWorker/XML Worker or iText 5 Move to pdfHTML APIs and update dependencies systematically; do not treat a class-name replacement as a complete migration.
License errors at startup License-key setup does not match the deployment or license Follow the current installation guide and verify the commercial license covers Core and pdfHTML.

9. Performance, reliability and cost considerations

  • Reuse a JVM and avoid starting a new process for every document.
  • Keep resource paths local or on dependable internal storage when repeatability matters; remote resources add latency and failure points.
  • Measure conversion time and memory with your actual XHTML, images, fonts and table sizes. The release notes contain targeted fixes, not a universal benchmark.
  • Use bounded input sizes and timeouts around resource loading in services that accept user-supplied documents.
  • Log the pdfHTML/Core versions, input identifier and conversion exception without logging secrets embedded in URLs or headers.
  • Render representative documents in CI and compare PDFs or page images after dependency upgrades.

10. Or skip the browser setup

If your real task is obtaining a clean screenshot or PDF of a web page before another pipeline processes it, ScreenshotNeo provides a single HTTP request. It accepts cookie and consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each step can be turned off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and responses identify the result with X-Page-Verdict and X-Billed headers. Its MCP server exposes take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients.

See the ScreenshotNeo API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo includes full-page capture, CSS-selector element capture, dark mode, device presets, custom viewports, retina scale, PDF paper and margin controls, custom CSS and JavaScript, click and wait actions, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, configurable caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage reporting and an OpenAPI specification. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

11. FAQ

Is XHTML required?

No. pdfHTML handles HTML/XML input, but well-formed XHTML makes parsing and resource resolution easier to reason about.

Can pdfHTML reproduce any browser page?

No. Check the release-specific support matrix and test the exact markup and CSS you depend on.

Should a new project use XML Worker?

No. XML Worker is associated with end-of-life iText 5. Start with pdfHTML for current iText Core work.

Where should relative URLs point?

Set the converter’s base URI to the directory or URL root from which the XHTML document’s relative resources resolve.

What changed in pdfHTML 6.3.3?

The release notes mention support for several selector pseudo-classes, more tolerance for malformed CSS, and fixes for CSS Grid pagination and list-rendering performance. Verify the release notes and matrix for your exact needs.

Can I use iText in a closed-source commercial product?

iText’s licensing guide says closed-source commercial use requires a commercial license for iText Core and pdfHTML.

12. Primary references