ScreenshotNeo

BlogHTML to image & PDF

How to Convert HTML to PDF in Java with OpenPDF

Convert HTML to PDF in Java with OpenPDF’s HTML module. Follow the dependency setup, runnable example, rendering limits, security guidance, and fixes for common errors.

By the ScreenshotNeo team4 October 20267 min read

Use OpenPDF’s openpdf-html module to render HTML into a PDF in Java. Add the HTML module and matching OpenPDF core dependency, create an ITextRenderer, pass it the HTML, call layout(), and write the PDF to an output stream with createPDF().

The examples below use OpenPDF 3.0.5, the version listed in the research sources. Check the project’s current published versions before pinning dependencies. OpenPDF 3.x uses the org.openpdf namespace; older examples using com.lowagie may not compile against it.

1. Add the OpenPDF dependencies

For Maven, include both the HTML renderer and its matching core artifact:

<dependencies>
    <dependency>
        <groupId>com.github.librepdf</groupId>
        <artifactId>openpdf-html</artifactId>
        <version>3.0.5</version>
    </dependency>
    <dependency>
        <groupId>com.github.librepdf</groupId>
        <artifactId>openpdf</artifactId>
        <version>3.0.5</version>
    </dependency>
</dependencies>

Keep the two versions aligned. If your build uses Gradle, the equivalent declarations are:

dependencies {
    implementation 'com.github.librepdf:openpdf-html:3.0.5'
    implementation 'com.github.librepdf:openpdf:3.0.5'
}

These coordinates and the renderer API are documented by the OpenPDF project; verify the version against the project or artifact repository when upgrading.

2. Convert an HTML string to a PDF file

This complete Java example converts a small HTML document, writes output.pdf in the current working directory, and closes the output stream even if conversion fails:

import org.openpdf.pdf.ITextRenderer;

import java.io.FileOutputStream;

public class HtmlToPdf {
    public static void main(String[] args) throws Exception {
        String html = """
                <html>
                  <head>
                    <style>
                      body { font-family: sans-serif; margin: 2rem; }
                      h1 { color: #203864; }
                    </style>
                  </head>
                  <body>
                    <h1>Invoice</h1>
                    <p>Rendered from HTML with OpenPDF.</p>
                  </body>
                </html>
                """;

        try (FileOutputStream output = new FileOutputStream("output.pdf")) {
            ITextRenderer renderer = new ITextRenderer();
            renderer.setDocumentFromString(html);
            renderer.layout();
            renderer.createPDF(output);
        }
    }
}

Compile and run it with the dependencies on your classpath, for example through your IDE or a Maven project. The essential order is:

  1. Provide a complete HTML document with setDocumentFromString(html).
  2. Call layout() so the renderer calculates page layout.
  3. Call createPDF(output) to write the PDF to the stream.

Call layout() before creating the PDF. For server code, prefer writing to a caller-provided stream or a bounded temporary file rather than assuming a local output path is suitable.

3. Convert HTML that references CSS, images, or other resources

An HTML string that contains relative references needs a base URL so the renderer can resolve them. The base URL must point to a directory or document location appropriate for those references. Use the overload supported by your selected OpenPDF version, and confirm the method signature against that version’s API documentation:

String html = """
        <html>
          <head><link rel="stylesheet" href="styles/report.css"></head>
          <body><img src="images/logo.png" alt=""><h1>Report</h1></body>
        </html>
        """;

String baseUrl = "file:/srv/reports/";

try (FileOutputStream output = new FileOutputStream("report.pdf")) {
    ITextRenderer renderer = new ITextRenderer();
    renderer.setDocumentFromString(html, baseUrl);
    renderer.layout();
    renderer.createPDF(output);
}

Resource resolution is a common source of incomplete PDFs. Check that the base URL is correctly formed, the process can read the referenced files, and CSS and image paths resolve from that base. Do not assume that a browser-only feature or arbitrary modern CSS will render identically.

4. Understand rendering support and page layout

openpdf-html is an HTML-to-PDF renderer integrated with OpenPDF and derived from Flying Saucer. The project describes modern HTML5 support as in progress and notes improved CSS3 compatibility; that does not promise browser-equivalent rendering.

Before relying on a template in production, render representative documents and inspect:

  • CSS selectors, layout, and page-break behavior.
  • Fonts, glyph coverage, and font resource paths.
  • Images, their dimensions, and resource URLs.
  • Long tables, wide content, and content that spans multiple pages.
  • Any HTML5 or CSS3 features your template depends on.

Keep templates intentionally compatible with the renderer, and treat the generated PDF—not the browser preview—as the result to validate. The project’s HTML module README describes the module and its rendering scope.

5. Handle input and resource access safely

Do not pass arbitrary user-supplied HTML or resource URLs into a conversion service without controls. OpenPDF’s project documentation places responsibility on the application developer to ensure input is trusted, sanitized, and safe; it says the library does not validate input or provide sandboxing.

  • Use trusted templates and escape or sanitize data inserted into them.
  • Validate resource URLs and restrict which files or hosts the conversion process can access.
  • Avoid allowing untrusted HTML to reference local files, internal services, or unrestricted network resources.
  • Run conversion with the minimum filesystem and network permissions needed for the application.
  • Set application-level limits for input size, conversion time, and output size.

These controls belong in the surrounding application and deployment. The renderer itself should not be treated as an isolation boundary.

6. Check licenses before shipping

The OpenPDF project identifies the core library as dual-licensed under MPL 2.0 or LGPL 2.1. It identifies openpdf-html and openpdf-renderer as LGPL 2.1 only. Review the license texts for the exact artifacts you distribute and assess obligations for your application and distribution model. See the project’s license and project documentation.

7. Troubleshoot common conversion problems

Symptom Likely cause What to check or change
ITextRenderer or org.openpdf import cannot be resolved The HTML module is missing, dependencies are not on the classpath, or imports are from another OpenPDF generation. Add openpdf-html and the matching core dependency. For OpenPDF 3.x use the org.openpdf package names shown in the current module README.
Code using com.lowagie fails with OpenPDF 3.x OpenPDF 3.0 changed the package namespace. Update imports and code to the org.openpdf namespace, or use documentation matching the older version you intentionally selected. Consult the release notes for migration details.
Stylesheets or images are missing Relative resources have no usable base URL, an incorrect path, or are unreadable by the process. Set the document base URL where appropriate; verify each resource path and process permissions. Check logs for resource-loading errors.
PDF content differs from browser output The renderer does not implement every browser HTML or CSS behavior. Reduce the template to supported features, test the exact markup and CSS, and inspect the rendered PDF. HTML5 support is described as in progress.
Text is missing or appears as boxes The selected font may be unavailable or lack the required glyphs. Verify font availability and glyph coverage in the runtime environment; test representative text including accented and non-Latin characters.
The PDF is empty or layout seems incomplete The layout step may have been skipped, HTML may be malformed, or content/resources failed to load. Follow the documented sequence: set the document, call layout(), then createPDF(). Simplify the input and check resource paths.
Conversion fails only in production Production may have different filesystem access, fonts, network policy, dependency versions, or runtime limits. Compare deployed dependency versions and environment permissions; include required assets in deployment and test with the same runtime constraints.

8. Plan for performance, reliability, and cost

The cited OpenPDF material does not provide a benchmark or fixed resource estimate. Measure with your own document sizes and templates. Large images, long documents, complex styling, and remote resources can add work and memory pressure; constrain inputs and observe conversion time and output size in your application.

  • Reuse stable templates and local assets where practical, and avoid fetching unnecessary resources during each conversion.
  • Set a request deadline and handle conversion failures explicitly; do not return a partial or invalid PDF as success.
  • For concurrent conversions, set concurrency limits based on measurements in your deployment environment.
  • Write to managed streams or temporary files and clean them up on both success and failure.
  • Track failures by template and input class so rendering regressions are diagnosable.

OpenPDF is a Java library distributed through software artifacts, so your direct costs depend on your hosting, compute, storage, and operational choices. Review the licenses separately from infrastructure costs.

9. When the input is a web page, use a screenshot or PDF capture API

If the source is a live URL and your goal is a faithful capture of the page as a visitor sees it, a browser-based capture workflow may fit better than adapting the page into HTML for this Java renderer. ScreenshotNeo is a website screenshot API and MCP server for developers, with a one-request API that can return an image or PDF. Its options include PDF paper size, margins, landscape orientation, and page ranges. Read the ScreenshotNeo overview and API documentation for the available request parameters.

Or skip the browser setup

For a live webpage, one GET request can return a PDF. This example uses the ScreenshotNeo API base and a placeholder key; replace the target URL as needed. See the ScreenshotNeo API docs for PDF parameters and account setup.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -d format=pdf \
  -o page.pdf

ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server lets AI agents use screenshot, page-info, and PDF-capture tools. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up free for 1,000 screenshots a month, with no card required.

FAQ

Does OpenPDF convert an HTML string directly?

Yes. The documented renderer flow accepts a string with setDocumentFromString, lays it out, and writes the result with createPDF.

Can I expect the same result as Chrome or Firefox?

No browser-equivalence guarantee is established by the project documentation. Test the exact HTML, CSS, fonts, and assets you need with your selected version.

Can I safely convert HTML submitted by users?

Only with application-level validation, sanitization, and resource restrictions appropriate to your threat model. OpenPDF states that input validation and sandboxing are the application developer’s responsibility.

Which package namespace should I use?

For OpenPDF 3.x, use org.openpdf. Older examples may use com.lowagie; match imports and API documentation to the dependency version in your build.