ScreenshotNeo

BlogHTML to image & PDF

Convert a URL to PDF in Java Using Puppeteer

Puppeteer is a JavaScript library, not a Java API. Learn how Java can run a Puppeteer process or call a hosted PDF endpoint, with runnable examples and print settings.

By the ScreenshotNeo team4 October 202611 min read

Short answer: Puppeteer is a JavaScript library, not a native Java API. To use it in a Java application, either run a separate Node.js process that uses Puppeteer, or have Java send an HTTP request to a hosted browser service that returns a PDF. If the requirement is specifically Puppeteer, the local-process approach uses Puppeteer directly; the hosted approach is a Java-to-HTTP integration and may use a browser automation service behind the endpoint.

Puppeteer’s basic PDF workflow is to launch a browser, navigate to the page, call page.pdf(), then close the browser. Its PDF output uses print CSS by default. [Puppeteer PDF guide; Page.pdf() API]

1. Choose how Java will reach the browser

Approach What runs where Choose it when Trade-offs
Java launches a Node.js Puppeteer script Your deployment runs Java, Node.js, Puppeteer, and a compatible browser. You need browser interaction, custom readiness logic, and control of the browser environment. You manage browser installation, updates, process lifecycle, resource limits, and isolation.
Java calls a hosted PDF endpoint Java sends an HTTP request; the service operates the browser and returns PDF bytes. You prefer not to manage a browser process and the endpoint supports the required options. You depend on a provider, protect API credentials, check its current limits and costs, and review where page data is processed.

Browserless documents a Java HttpClient example for sending a URL and PDF options to its hosted endpoint. That is a service integration called from Java, not Puppeteer running inside the JVM. Check the provider’s current documentation for endpoint details, account limits, and pricing before adopting it. [Browserless documentation]

2. Run Puppeteer as a separate Node.js process

This option keeps Puppeteer’s browser control available while letting your Java application invoke it as a child process. The following example accepts a URL and output path, navigates to the URL, writes a PDF, and closes Chromium even if capture fails.

Install the script dependencies

npm init -y
npm install puppeteer

Save this as render-pdf.cjs:

const puppeteer = require('puppeteer');

async function main() {
  const [url, outputPath] = process.argv.slice(2);
  if (!url || !outputPath) {
    throw new Error('Usage: node render-pdf.cjs <url> <output.pdf>');
  }

  const parsed = new URL(url);
  if (!['http:', 'https:'].includes(parsed.protocol)) {
    throw new Error('URL must use http or https');
  }

  const browser = await puppeteer.launch({ headless: true });
  try {
    const page = await browser.newPage();
    await page.goto(url, { waitUntil: 'networkidle2', timeout: 60000 });
    await page.pdf({
      path: outputPath,
      format: 'A4',
      printBackground: true,
      margin: { top: '16mm', right: '14mm', bottom: '16mm', left: '14mm' }
    });
  } finally {
    await browser.close();
  }
}

main().catch((error) => {
  console.error(error.message);
  process.exitCode = 1;
});

Run it directly:

node render-pdf.cjs https://example.com ./example.pdf

Java 11 or later can invoke the script with ProcessBuilder. Pass each argument separately; this avoids shell quoting problems when URLs contain query parameters.

import java.io.IOException;
import java.nio.charset.StandardCharsets;
import java.nio.file.Path;
import java.util.concurrent.TimeUnit;

public class UrlToPdf {
    public static void main(String[] args) throws Exception {
        if (args.length != 2) {
            throw new IllegalArgumentException("Usage: UrlToPdf <url> <output.pdf>");
        }

        Process process = new ProcessBuilder(
                "node", "render-pdf.cjs", args[0], args[1])
                .redirectErrorStream(true)
                .start();

        boolean finished = process.waitFor(90, TimeUnit.SECONDS);
        if (!finished) {
            process.destroyForcibly();
            throw new IOException("PDF generation timed out");
        }

        String output = new String(process.getInputStream().readAllBytes(), StandardCharsets.UTF_8);
        int exitCode = process.exitValue();
        if (exitCode != 0) {
            throw new IOException("Puppeteer failed (exit " + exitCode + "): " + output);
        }

        if (!Path.of(args[1]).toFile().isFile()) {
            throw new IOException("Puppeteer exited successfully but PDF was not created");
        }
        System.out.println("PDF created at " + args[1]);
    }
}

Compile and run with the script and installed Node dependencies available in the working directory:

javac UrlToPdf.java
java UrlToPdf https://example.com ./example.pdf

For production, also drain process output while it runs or redirect it to a bounded log. If output is not consumed, a busy child process can fill its output pipe and stall. Set an application-level concurrency limit, enforce an overall deadline, and terminate child processes on cancellation or shutdown.

3. Configure PDF output and page readiness

page.pdf() renders using the print media type by default. This means print-specific CSS can hide navigation, change layout, or add page breaks. To use screen styles, call page.emulateMediaType('screen') before creating the PDF. Printed colors may be adjusted by the browser; for exact CSS colors, the Puppeteer API points to -webkit-print-color-adjust. [Puppeteer Page.pdf() API]

await page.emulateMediaType('screen');
await page.pdf({ path: outputPath, printBackground: true, format: 'A4' });

For print output, CSS can explicitly preserve colors:

@media print {
  html {
    -webkit-print-color-adjust: exact;
    print-color-adjust: exact;
  }
}

Page size, margins, orientation, and backgrounds

  • format selects a named paper format such as A4 or Letter. Use the PDF options documented for your installed Puppeteer version.
  • Set margin with top, right, bottom, and left CSS units such as mm or in.
  • Set landscape: true for a landscape page.
  • Set printBackground: true when background colors and images should appear.
  • Header and footer templates can add printed details when the selected Puppeteer version supports the relevant PDF options. Reserve enough bottom and top margin for them.

When page CSS defines its own paper size, consult the installed Puppeteer API’s preferCSSPageSize option and decide whether CSS or the PDF call should control the final dimensions.

Wait for the page that you actually need

The guide’s example uses waitUntil: 'networkidle2', which is a useful starting point, but many sites keep network connections open or load important content after navigation. Prefer a page-specific readiness signal where possible, such as waiting for a report container or a known state change. A fixed delay can help with a known animation or delayed widget, but it is not proof that content has finished loading. Puppeteer’s documented PDF flow waits for fonts to load by default. [Puppeteer PDF guide]

await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 60000 });
await page.waitForSelector('[data-report-ready="true"]', { timeout: 20000 });
await page.pdf({ path: outputPath, format: 'A4', printBackground: true });

Use networkidle0 or networkidle2 only when the site’s network behavior makes those conditions useful. For authenticated pages, set the required cookies or headers before navigation and handle credentials as secrets. Avoid accepting arbitrary URLs from untrusted users without network access controls: a browser can otherwise be induced to request internal services.

4. Call a hosted PDF endpoint from Java

A hosted endpoint is useful when you want the Java service to exchange HTTP requests and PDF bytes instead of operating Chromium locally. Browserless’s documented example uses java.net.http.HttpClient, sends JSON with a URL and PDF options, and consumes the response as a PDF. Adapt the endpoint and JSON fields to the provider’s current API documentation. [Browserless documentation]

import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.nio.file.Files;
import java.nio.file.Path;
import java.time.Duration;

public class HostedUrlToPdf {
    public static void main(String[] args) throws Exception {
        if (args.length != 2) {
            throw new IllegalArgumentException("Usage: HostedUrlToPdf <url> <output.pdf>");
        }

        String token = System.getenv("BROWSERLESS_TOKEN");
        if (token == null || token.isBlank()) {
            throw new IllegalStateException("Set BROWSERLESS_TOKEN in the environment");
        }

        String endpoint = "https://production-sfo.browserless.io/pdf?token=" + token;
        String json = "{\"url\":\"" + jsonEscape(args[0]) +","
                + "\"options\":{\"format\":\"A4\","
                + "\"printBackground\":true,\"displayHeaderFooter\":false}}";

        HttpClient client = HttpClient.newBuilder()
                .connectTimeout(Duration.ofSeconds(10))
                .build();
        HttpRequest request = HttpRequest.newBuilder(URI.create(endpoint))
                .timeout(Duration.ofSeconds(90))
                .header("Content-Type", "application/json")
                .POST(HttpRequest.BodyPublishers.ofString(json))
                .build();

        HttpResponse<byte[]> response = client.send(request, HttpResponse.BodyHandlers.ofByteArray());
        if (response.statusCode() < 200 || response.statusCode() >= 300) {
            throw new IllegalStateException("PDF service returned HTTP " + response.statusCode()
                    + ": " + new String(response.body()));
        }
        String contentType = response.headers().firstValue("Content-Type").orElse("");
        if (!contentType.toLowerCase().contains("application/pdf")) {
            throw new IllegalStateException("Expected application/pdf, received " + contentType);
        }
        Files.write(Path.of(args[1]), response.body());
    }

    private static String jsonEscape(String value) {
        return value.replace("\\", "\\\\").replace("\"", "\\\"")
                .replace("\n", "\\n").replace("\r", "\\r");
    }
}

This compact example demonstrates the request shape; for production, use a JSON library to serialize the body and parse error responses. Never place a real token in source control or logs. The service documentation says its PDF endpoint accepts a URL or raw HTML and returns application/pdf; verify the exact host, option names, authentication method, and current limits against the provider’s docs before deployment.

5. JavaScript and command-line reference

The Puppeteer library runs in Node.js. These examples are useful for checking the browser workflow independently of the Java wrapper.

cURL for a hosted endpoint

Use the provider’s documented endpoint and JSON request schema. The following shows the general POST shape; replace the host, token placement, and options with the current provider instructions.

curl -X POST 'https://production-sfo.browserless.io/pdf?token=YOUR_TOKEN' \
  -H 'Content-Type: application/json' \
  --data '{"url":"https://example.com","options":{"format":"A4","printBackground":true}}' \
  --output page.pdf

Node.js with Puppeteer

const puppeteer = require('puppeteer');

(async () => {
  const browser = await puppeteer.launch({ headless: true });
  try {
    const page = await browser.newPage();
    await page.goto('https://example.com', { waitUntil: 'networkidle2', timeout: 60000 });
    await page.pdf({ path: 'page.pdf', format: 'A4', printBackground: true });
  } finally {
    await browser.close();
  }
})();

There is no equivalent cURL command for a local Puppeteer process itself: Puppeteer is a library. cURL applies to an HTTP service endpoint.

6. Troubleshooting

Symptom Likely cause Fix
Java says Puppeteer is not found Puppeteer is a Node.js package, not a Java dependency. Install Node.js and the package, run the script directly first, then check the Java process working directory and executable path.
Chromium fails to launch Browser download is missing, runtime libraries are unavailable, or container permissions/configuration prevent launch. Install the browser dependencies required by the Puppeteer version and deployment image; inspect the child process error output.
Navigation times out The page is slow, keeps connections open, or never reaches the selected network-idle condition. Use a realistic timeout and a site-specific readiness selector or state check. Do not treat a longer fixed sleep as a general solution.
PDF is blank or missing dynamic content Capture began before the page rendered its data, or the content requires interaction. Wait for the actual content selector/state and perform required interactions before page.pdf().
PDF colors or layout differ from the browser PDF uses print media by default and print color adjustment can change appearance. Use print CSS intentionally, or emulate screen media before PDF generation; enable background printing and CSS color adjustment as needed.
Hosted response is not a PDF The service returned an error body, authentication failure, or a different response type. Check status code and content type before writing the file; inspect a safely redacted error body and verify endpoint and token.
Only some pages appear in a ranged PDF Requested page ranges do not cover every page. Ensure the ranges cover all desired pages. Browserless documents that uncovered pages can be silently omitted and out-of-range requests can error. [Browserless documentation]
PDF has no title or author metadata The documented Puppeteer PDF flow does not provide built-in metadata options. Post-process the generated file with a PDF library if metadata is required. [Browserless documentation]

7. Performance, reliability, and cost

  • Reuse carefully: launching a browser for every small request adds work. A managed worker can reuse a browser while creating a fresh page per job, with bounded concurrency and cleanup after failures.
  • Limit resource use: cap simultaneous pages, navigation time, PDF size, and input URL length. Close pages and browsers reliably, and clean up partial output files.
  • Retry selectively: retry transient connection failures or provider-side temporary errors with a small bounded backoff. Do not repeatedly retry invalid URLs, authentication failures, or deterministic page errors.
  • Keep output deterministic: page content, fonts, locale, viewport, cookies, and readiness conditions can affect the PDF. Fix these inputs where repeatability matters.
  • Protect data: a hosted service receives the URL or HTML and may fetch page resources; evaluate data handling and credentials. A local browser keeps browser execution in your environment but still needs network and runtime controls.
  • Estimate cost from actual usage: local operation consumes compute, memory, storage, and maintenance time. A hosted endpoint may charge by request or another usage measure; current prices and limits were not established by the documentation reviewed, so check the provider’s current plan terms.

8. Or skip the browser setup

If the job is simply to turn a public URL into a PDF, ScreenshotNeo provides a screenshot and PDF API. It accepts one GET request with the target URL and returns a PDF when configured for PDF output. The API supports paper size, margins, landscape orientation, and page ranges. See the ScreenshotNeo API documentation for the PDF parameters and current request syntax.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com \
  -d format=pdf \
  -o page.pdf

The same API can be called from Java with HttpClient, or from Python and Node.js:

// Java 11+: save the response bytes returned by the GET request
import java.net.URI;
import java.net.http.*;
import java.nio.file.*;

var query = "access_key=YOUR_API_KEY&url=https%3A%2F%2Fexample.com&format=pdf";
var request = HttpRequest.newBuilder(URI.create("https://api.screenshotneo.com/v1/shot?" + query)).GET().build();
var response = HttpClient.newHttpClient().send(request, HttpResponse.BodyHandlers.ofByteArray());
if (response.statusCode() < 200 || response.statusCode() >= 300) throw new RuntimeException("HTTP " + response.statusCode());
Files.write(Path.of("page.pdf"), response.body());
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com", "format": "pdf"},
    timeout=90,
)
r.raise_for_status()
open("page.pdf", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com', format: 'pdf' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
await require('node:fs/promises').writeFile('page.pdf', Buffer.from(await res.arrayBuffer()));
  • Cookie and consent banners are accepted or removed before capture; newsletter popups and chat widgets from known platforms are also removed.
  • Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Response headers identify the page verdict and billing status.
  • An MCP server lets AI agents use screenshot and PDF capture tools.
  • The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.

Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.

9. Frequently asked questions

Can I import Puppeteer into a Java class?

No. Puppeteer is a JavaScript library. Use a Node.js process or an HTTP service from Java.

Does the PDF exactly match what I see on screen?

Not by default. Puppeteer generates PDFs with print media styles. Emulate screen media when that is the intended output, and account for paper sizing and print color rules.

Can I add PDF title and author using Puppeteer’s PDF call?

The documented PDF flow does not expose metadata options for title or author. Generate the PDF and set metadata with a PDF library afterward if needed.

Is a hosted PDF endpoint still “using Puppeteer”?

It depends on the service implementation. From Java, you are calling an HTTP API; you are not using Puppeteer as a Java library. Confirm which browser technology and options the selected service provides.