ScreenshotNeo

BlogHTML to image & PDF

How to Convert HTML to PDF with PDFBox

PDFBox does not parse HTML itself. Use OpenHTMLtoPDF for HTML/CSS layout, then PDFBox for the PDF document and post-processing.

By the ScreenshotNeo team1 October 20268 min read

Direct answer: Apache PDFBox is a Java library for creating and manipulating PDF files; it is not an HTML or CSS layout engine. To convert HTML to PDF, use an HTML renderer such as OpenHTMLtoPDF with its PDFBox integration. The renderer lays out well-formed HTML/XHTML and supported CSS, while PDFBox supplies the PDF backend and APIs for later PDF operations.

This distinction explains most failed implementations. PDFBox cannot take an arbitrary web page and reproduce it like Chrome. OpenHTMLtoPDF supports a subset of HTML and CSS, does not run JavaScript, and does not implement many browser layout features such as modern flexbox and grid. Pages often need a print-oriented template.

1. Choose the PDFBox integration that matches your project

First identify the PDFBox major version already used by your application. OpenHTMLtoPDF publishes separate integration artifacts for PDFBox 2 and PDFBox 3:

Application dependency OpenHTMLtoPDF integration Use when
PDFBox 3 io.github.openhtmltopdf:openhtmltopdf-pdfbox Your application uses Apache PDFBox 3.x
PDFBox 2 com.openhtmltopdf:openhtmltopdf-pdfbox Your application uses Apache PDFBox 2.x

These are OpenHTMLtoPDF artifacts, not Apache PDFBox modules. Confirm the current compatible versions in your build before publishing or deploying. The PDFBox 3 getting-started documentation currently demonstrates org.apache.pdfbox:pdfbox:3.0.8; release numbers change over time.

2. Maven setup for PDFBox 3

The following Maven configuration shows the correct dependency shape. Set openhtmltopdf.version to a current release compatible with your application and verify the resolved dependency tree.

<properties>
  <maven.compiler.source>17</maven.compiler.source>
  <maven.compiler.target>17</maven.compiler.target>
  <pdfbox.version>3.0.8</pdfbox.version>
  <openhtmltopdf.version>REPLACE_WITH_COMPATIBLE_VERSION</openhtmltopdf.version>
</properties>

<dependencies>
  <dependency>
    <groupId>io.github.openhtmltopdf</groupId>
    <artifactId>openhtmltopdf-pdfbox</artifactId>
    <version>${openhtmltopdf.version}</version>
  </dependency>

  <dependency>
    <groupId>org.apache.pdfbox</groupId>
    <artifactId>pdfbox</artifactId>
    <version>${pdfbox.version}</version>
  </dependency>
</dependencies>

If your project is already using PDFBox 3 transitively, avoid forcing a second incompatible version. Run mvn dependency:tree and align the renderer, PDFBox, and any other PDF libraries.

3. Convert a self-contained HTML string

This example renders a small invoice-style document and writes target/invoice.pdf. It uses the OpenHTMLtoPDF builder for HTML layout and PDFBox APIs to inspect the resulting document.

package com.example.pdf;

import com.openhtmltopdf.pdfboxout.PdfRendererBuilder;
import java.io.ByteArrayOutputStream;
import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.nio.file.Files;
import java.nio.file.Path;
import org.apache.pdfbox.Loader;
import org.apache.pdfbox.pdmodel.PDDocument;

public final class HtmlToPdf {
  private HtmlToPdf() {}

  public static void main(String[] args) throws Exception {
    String html = """
        <!DOCTYPE html>
        <html>
        <head>
          <meta charset=\"UTF-8\">
          <style>
            @page { size: A4; margin: 18mm; }
            body { font-family: sans-serif; color: #222; }
            h1 { font-size: 22pt; }
            .total { margin-top: 24px; font-weight: bold; }
          </style>
        </head>
        <body>
          <h1>Invoice 1001</h1>
          <p>Rendered from well-formed HTML and supported CSS.</p>
          <p class=\"total\">Total: $125.00</p>
        </body>
        </html>
        """;

    byte[] pdf = render(html, "https://example.com/");
    Path output = Path.of("target", "invoice.pdf");
    Files.createDirectories(output.getParent());
    Files.write(output, pdf);

    try (PDDocument document = Loader.loadPDF(pdf)) {
      System.out.println("Pages: " + document.getNumberOfPages());
    }
  }

  static byte[] render(String html, String baseUri) throws IOException {
    ByteArrayOutputStream output = new ByteArrayOutputStream();
    PdfRendererBuilder builder = new PdfRendererBuilder();
    builder.withHtmlContent(html, baseUri);
    builder.toStream(output);
    builder.run();
    return output.toByteArray();
  }
}

The baseUri matters when your HTML references relative images, stylesheets, or fonts. Use a trusted absolute directory or URL that contains those resources. For untrusted input, restrict or replace external resource resolution.

4. Convert an HTML file and resolve relative assets

Path htmlFile = Path.of("src/main/resources/templates/report.html");
Path outputFile = Path.of("target/report.pdf");

try (OutputStream output = Files.newOutputStream(outputFile)) {
  PdfRendererBuilder builder = new PdfRendererBuilder();
  builder.withFile(htmlFile.toFile());
  builder.toStream(output);
  builder.run();
}

Keep the HTML, CSS, images, and fonts in a directory structure that matches the relative URLs in the document. A missing base URI is a common reason for blank images and missing styles.

5. HTML and CSS that work reliably

OpenHTMLtoPDF describes support for a reasonable subset of well-formed XML/XHTML, some HTML5, and CSS 2.1 plus later standards. Treat it as a print renderer, not a browser.

  • Use valid, closed elements and a declared character set.
  • Prefer normal block layout, tables, floats, and explicit widths.
  • Use @page for paper size and margins.
  • Use print-specific CSS instead of responsive browser breakpoints.
  • Use stable image dimensions and accessible local or absolute resource URLs.
  • Design explicit page breaks with page-break-before, page-break-after, or the renderer’s supported equivalents.
  • Embed or register fonts when the output must be consistent across machines.

Do not assume JavaScript will run. Client-side templates, chart libraries, lazy-loaded content, authentication redirects, and DOM mutations must be completed before conversion or replaced with server-rendered markup.

6. Images, fonts, and external resources

Relative URLs are resolved against the base URI passed to withHtmlContent or the location of the HTML file. Test:

  • PNG, JPEG, and other image formats used by your documents.
  • Images behind authentication or expiring URLs.
  • SVG files and data URLs if your design depends on them.
  • Font files, font weights, fallback characters, and non-Latin scripts.
  • HTTPS certificates and network access from the conversion host.

For repeatable builds, package CSS, fonts, and images with the application or serve them from controlled locations. Avoid depending on a live website whose content can change during a conversion.

7. PDFBox operations after conversion

Once the renderer has produced a PDF, use PDFBox for PDF-specific work such as metadata, merging, attachments, encryption, or page inspection. Always close documents and streams.

try (PDDocument document = Loader.loadPDF(Path.of("target/report.pdf").toFile())) {
  document.getDocumentInformation().setTitle("Monthly report");
  document.save("target/report-with-metadata.pdf");
}

Only one thread may access a single PDDocument at a time. For parallel jobs, create a separate document instance per job; do not share one loaded document between worker threads.

8. PDFBox 2 projects

If the application is still on PDFBox 2, use the PDFBox 2 OpenHTMLtoPDF integration coordinate listed by Maven Central: com.openhtmltopdf:openhtmltopdf-pdfbox. Keep every PDFBox-related dependency on the 2.x line and compile against the APIs used by that major version. Do not copy PDFBox 3 examples into a PDFBox 2 project without checking API differences.

PDFBox’s migration documentation also matters when rendering PDFs to images: older APIs such as PDPage.convertToImage were removed in PDFBox 2, and PDFRenderer is the replacement for rasterizing an existing PDF. That operation is separate from converting HTML into a PDF.

9. Common errors and fixes

Error or symptom Likely cause Fix
ClassNotFoundException or linkage errors PDFBox 2 and PDFBox 3 artifacts are mixed Inspect mvn dependency:tree; use the integration matching one PDFBox major version.
Blank or missing images Relative URLs have no usable base URI, or the process cannot fetch the resource Pass a correct base URI, use local packaged assets, and verify file permissions and network access.
CSS appears ignored Unsupported browser CSS, malformed markup, or a stylesheet that was not resolved Validate the HTML, simplify to supported CSS, and use print-oriented rules.
JavaScript-generated content is absent The renderer does not run JavaScript Render the data on the server first or use a browser capture workflow that executes JavaScript.
Text has boxes or wrong glyphs Missing font or unsupported character coverage Register and embed an appropriate font; test every required script.
Content is clipped or pagination is wrong Browser layout assumptions, oversized fixed elements, or missing page rules Set page size and margins, remove fixed heights, add explicit break rules, and inspect representative long documents.
Out-of-memory failures Large images, many pages, high render resolution, or retained documents Reduce image dimensions and render resolution, close documents promptly, and use PDFBox scratch-file loading where appropriate.
Concurrent modification or unreliable output One PDDocument is shared by multiple threads Give each job its own document instance and synchronize access to any shared state.

10. Validation checklist

  1. Test a short document and a document long enough to span many pages.
  2. Check headings, tables, lists, headers, footers, and intentional page breaks.
  3. Verify images at their actual production sizes.
  4. Test the fonts and languages your users need.
  5. Open the PDF in more than one viewer and extract text if searchability matters.
  6. Test missing, slow, or inaccessible external resources.
  7. Measure memory and conversion time with production-sized input instead of assuming browser-like performance.
  8. Confirm every PDDocument, stream, and temporary resource is closed.

11. Performance, reliability, and cost considerations

Conversion time depends on markup size, CSS complexity, image dimensions, font processing, page count, and resource access. PDFBox documentation notes that rendering memory depends on the PDF and render resolution. Reduce unnecessary image resolution, avoid retaining large document objects, and use scratch files when your workload requires them.

For reliable batch conversion, isolate jobs, impose input and resource timeouts at the application boundary, cache stable assets, and record the renderer and PDFBox versions with each release. Do not publish a universal pages-per-second or memory benchmark without measuring your own documents.

PDFBox and OpenHTMLtoPDF are software dependencies. Your cost is the infrastructure and engineering work needed to run them; licensing and version compatibility should be reviewed for your application before distribution.

12. Or skip the browser setup

If your input is already a public webpage and you need a clean capture, ScreenshotNeo provides a website screenshot API and MCP server. It accepts one GET request and returns PNG, JPEG, WebP, or PDF. The DIY PDFBox workflow above is appropriate when you control the HTML and need Java-side PDF processing; ScreenshotNeo is useful when the page must be captured as rendered by a browser.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const buffer = Buffer.from(await res.arrayBuffer());
require('fs').writeFileSync('shot.webp', buffer);

See the ScreenshotNeo API documentation for request options and PDF capture details. Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots, and the free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

FAQ

Can PDFBox convert HTML by itself?

No. PDFBox creates and manipulates PDF documents. An HTML/CSS renderer must perform the layout step.

Does OpenHTMLtoPDF behave like Chrome?

No. It supports a defined subset of HTML/XHTML and CSS, does not run JavaScript, and does not implement many modern browser standards.

Should I use PDFBox 2 or PDFBox 3?

Use the major version already required by your application unless you are planning a dependency migration. Match the OpenHTMLtoPDF integration artifact to that major version.

Can I convert a JavaScript-heavy website with this method?

Not directly. Produce server-rendered HTML first, or use a browser-based capture service for pages that require JavaScript execution.

Can multiple threads share one PDFBox document?

No. Give each concurrent operation its own PDDocument instance.

Primary references