ScreenshotNeo

BlogHow-to

How to Test PDF Files with Selenium

Test PDF downloads, browser viewing, and generated print output with Selenium, then validate the saved document with a PDF library.

By the ScreenshotNeo team4 October 20269 min read

Use Selenium to exercise the browser workflow, then use an HTTP client or PDF library to inspect the actual file. For downloads, Selenium can click the link and help obtain session cookies, but WebDriver does not report download progress. For PDFs generated from a webpage, use Selenium’s print API, save its returned PDF bytes, and validate the document separately.

1. Choose the PDF workflow you need to test

Keep these cases distinct because they fail in different places:

  • Download: the application links to a PDF file, possibly behind authentication.
  • Browser viewing: a PDF URL should open in the browser’s PDF viewer.
  • Print generation: a webpage is converted to a PDF through the browser’s print interface.

Decide what constitutes success before writing assertions. A successful click only proves that the browser action occurred; it does not prove a successful transfer, valid PDF structure, correct text, or visual layout.

2. Test a downloaded PDF

Selenium’s own guidance recommends using WebDriver to locate the download link and obtain browser cookies when needed, then fetching the file with an HTTP client. WebDriver does not expose download progress as a stable file-download API. Selenium file-download guidance

This example uses Selenium 4, Java’s built-in HttpClient, and a link with id pdf-download. Replace the selector and expected URL for the application under test. Add the Selenium Java dependency and a browser driver configured for your environment.

import java.net.CookieManager;
import java.net.CookiePolicy;
import java.net.CookieStore;
import java.net.HttpCookie;
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.nio.file.Files;
import java.nio.file.Path;
import java.time.Duration;
import java.util.List;
import org.openqa.selenium.By;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.WebElement;
import org.openqa.selenium.chrome.ChromeDriver;
import org.openqa.selenium.support.ui.ExpectedConditions;
import org.openqa.selenium.support.ui.WebDriverWait;

public class DownloadPdfTest {
  public static void main(String[] args) throws Exception {
    WebDriver driver = new ChromeDriver();
    try {
      driver.get("https://example.com/reports");
      WebDriverWait wait = new WebDriverWait(driver, Duration.ofSeconds(15));
      WebElement link = wait.until(ExpectedConditions
          .visibilityOfElementLocated(By.id("pdf-download")));
      String href = link.getAttribute("href");
      if (href == null || href.isBlank()) {
        throw new AssertionError("PDF link has no href");
      }

      // Browser-managed downloads are not monitored through WebDriver.
      // Fetch the link directly and forward the site's session cookies.
      CookieManager cookieManager = new CookieManager(null, CookiePolicy.ACCEPT_ALL);
      CookieStore store = cookieManager.getCookieStore();
      URI pdfUri = URI.create(href);
      for (org.openqa.selenium.Cookie c : driver.manage().getCookies()) {
        HttpCookie cookie = new HttpCookie(c.getName(), c.getValue());
        cookie.setPath(c.getPath());
        cookie.setSecure(c.isSecure());
        cookie.setHttpOnly(c.isHttpOnly());
        store.add(URI.create("https://" + c.getDomain()), cookie);
      }

      HttpClient client = HttpClient.newBuilder()
          .cookieHandler(cookieManager)
          .followRedirects(HttpClient.Redirect.NORMAL)
          .connectTimeout(Duration.ofSeconds(20))
          .build();
      HttpRequest request = HttpRequest.newBuilder(pdfUri)
          .timeout(Duration.ofSeconds(60))
          .GET()
          .build();
      HttpResponse<byte[]> response = client.send(request,
          HttpResponse.BodyHandlers.ofByteArray());
      if (response.statusCode() != 200) {
        throw new AssertionError("Expected HTTP 200, got " + response.statusCode());
      }
      byte[] bytes = response.body();
      if (bytes.length < 5 || bytes[0] != '%' || bytes[1] != 'P'
          || bytes[2] != 'D' || bytes[3] != 'F' || bytes[4] != '-') {
        throw new AssertionError("Response is not a PDF; check redirects, auth, and MIME/body");
      }
      Path output = Path.of("build", "test-output", "report.pdf");
      Files.createDirectories(output.getParent());
      Files.write(output, bytes);
      System.out.println("Saved " + bytes.length + " bytes to " + output);
    } finally {
      driver.quit();
    }
  }
}

Cookie domain handling may need adjustment for applications using several subdomains, non-default paths, or a different authentication mechanism. If the download requires a bearer token, CSRF header, or signed URL, reproduce the required request state in the HTTP client instead of assuming cookies are sufficient.

Equivalent retrieval with cURL

For a public or signed download URL, retrieve the file directly and inspect the HTTP status and content type. For an authenticated URL, supply the appropriate cookie or authorization header securely; do not commit secrets to source control.

curl --fail --location --show-error \
  --output report.pdf \
  --write-out 'HTTP %{http_code}; content type %{content_type}; bytes %{size_download}\n' \
  'https://example.com/reports/latest.pdf'

A 200 response and a %PDF- header are useful transport checks, but they do not establish that the document contains the expected pages or text.

3. Test a PDF generated from a webpage

Selenium’s print interface accepts options such as orientation, margins, scale, background printing, and shrink-to-fit. In Java, PrintsPage returns base64-encoded PDF data; decode and save it before document assertions. Selenium print-page documentation

Java example: print, save, and check the PDF

The following saves the browser’s print output. It uses PDFBox for document checks; add the PDFBox dependency to the project. The example checks that a required phrase appears in extracted text.

import java.nio.file.Files;
import java.nio.file.Path;
import java.time.Duration;
import java.util.Base64;
import org.apache.pdfbox.Loader;
import org.apache.pdfbox.text.PDFTextStripper;
import org.openqa.selenium.By;
import org.openqa.selenium.PrintsPage;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.chrome.ChromeDriver;
import org.openqa.selenium.print.PageSize;
import org.openqa.selenium.print.PrintOptions;
import org.openqa.selenium.support.ui.ExpectedConditions;
import org.openqa.selenium.support.ui.WebDriverWait;

public class PrintPdfTest {
  public static void main(String[] args) throws Exception {
    WebDriver driver = new ChromeDriver();
    try {
      driver.get("https://example.com/invoice/123");
      new WebDriverWait(driver, Duration.ofSeconds(15)).until(
          ExpectedConditions.visibilityOfElementLocated(By.cssSelector(".invoice-total")));

      PrintOptions options = new PrintOptions();
      options.setOrientation(PrintOptions.Orientation.PORTRAIT);
      options.setScale(1.0);
      options.setBackground(true);
      options.setShrinkToFit(true);
      options.setPageSize(new PageSize(8.27, 11.69)); // A4 in inches
      options.setPageMargin(java.util.Map.of(
          "top", 0.4, "bottom", 0.4, "left", 0.4, "right", 0.4));

      String base64 = ((PrintsPage) driver).print(options);
      byte[] pdf = Base64.getDecoder().decode(base64);
      Path output = Path.of("build", "test-output", "invoice.pdf");
      Files.createDirectories(output.getParent());
      Files.write(output, pdf);

      try (var document = Loader.loadPDF(pdf)) {
        if (document.getNumberOfPages() < 1) {
          throw new AssertionError("Generated PDF has no pages");
        }
        String text = new PDFTextStripper().getText(document);
        if (!text.contains("Invoice total")) {
          throw new AssertionError("Expected invoice text was not found");
        }
      }
    } finally {
      driver.quit();
    }
  }
}

Check the current Selenium binding documentation for exact option method signatures in your version. The print feature and interfaces vary by language binding; Selenium also documents a BiDi BrowsingContext printing path. Avoid assuming that every browser driver exposes identical viewer controls.

4. Validate the saved document

Choose assertions that match the requirement, and separate them from browser interaction assertions.

Requirement Useful check Limit
Download endpoint works HTTP status, redirect destination, content type, non-empty body A response can be an HTML login or error page despite a successful status.
File is a PDF PDF parser opens it; optionally check the file signature A signature alone does not prove the document is complete or correct.
Required values are present Extract text and assert stable labels or values Scanned images, custom fonts, and reading order can complicate extraction.
Page count or metadata matters Inspect page count, dimensions, and metadata with a PDF library Metadata may vary across otherwise equivalent output.
Print conformance is required Run the required profile validator, such as PDF/A-1b preflight Validate only a conformance profile the product actually promises.
Layout or visual appearance matters Render pages to images and compare with controlled tolerances Fonts, browser versions, and antialiasing can create harmless pixel changes.

Apache PDFBox provides Unicode text extraction, form handling, image export, and PDF/A-1b preflight validation. Use a parser for structure and text; use rendered page comparisons when layout fidelity is the requirement. Apache PDFBox project

Minimal Python document check

If a Python test suite already retrieves the bytes, a PDF-aware parser can check text independently of Selenium. For example, with pypdf installed:

from pathlib import Path
from pypdf import PdfReader

path = Path("build/test-output/report.pdf")
reader = PdfReader(str(path), strict=True)
assert len(reader.pages) > 0
text = "\n".join(page.extract_text() or "" for page in reader.pages)
assert "Invoice total" in text

5. Test the browser PDF viewer separately

If the requirement is that clicking a PDF opens it in the browser, test that presentation path as a browser-specific scenario. Firefox uses its built-in PDF viewer when configured to open PDFs in Firefox by default, with documented exceptions for incorrectly set MIME types. Mozilla’s PDF viewer guidance

Keep viewer assertions separate from the HTTP response and document-content assertions. A viewer may be implemented with browser-specific controls, so selectors and preferences should be taken from the current documentation for the selected browser and Selenium binding. Selenium supported browsers

6. Reliability, performance, and cost

  • Keep transfer and browser work separate: use Selenium for the user-visible action and an HTTP client for predictable download assertions.
  • Wait for application readiness: wait for a meaningful page element before printing; fixed sleeps make tests slower and still race with delayed content.
  • Use stable output paths: write into a per-test temporary directory and include the test identifier to prevent parallel jobs overwriting files.
  • Bound timeouts and file sizes: set connection/read timeouts and reject unexpectedly large responses to prevent a stalled or erroneous endpoint from consuming CI resources.
  • Control rendering inputs: pin browser and font environment where visual output matters; use tolerant comparisons for rasterized pages.
  • Control cost: browser startup, rendering, and storing artifacts consume CI time and storage. Reuse a driver within a test scope where appropriate, and retain PDFs only when they help diagnose failures.

7. Troubleshooting

Symptom Likely cause Fix
Click succeeds but no file is available WebDriver does not provide download progress; browser download location or permissions may also differ. Fetch the discovered URL through an HTTP client and assert the saved response, or configure and inspect browser-specific download behavior only when that behavior itself is under test.
HTTP client gets 401 or 403 Session cookies, authorization, CSRF state, or signed URL was not forwarded. Copy the required authentication state from the browser session or obtain a fresh signed URL; check cookie domain, path, secure flag, and expiration.
Saved file begins with HTML Login redirect, error page, proxy response, or wrong endpoint was saved as a PDF. Inspect final URL, status, content type, and a short safe diagnostic excerpt; correct authentication or endpoint selection.
Print returns blank or incomplete pages Content had not loaded, lazy content was absent, or print CSS changes visibility. Wait for application readiness and required assets; validate print styles and print options such as background and shrink-to-fit.
Text assertion fails although the page looks correct Text may be rasterized, encoded with unusual font mappings, or extracted in a different order. Inspect extracted text, assert robust labels, and use rendered-page checks for scanned or image-based content.
PDF parser reports corruption Partial transfer, non-PDF body, or interrupted write. Check response length/status and signature, write atomically, and retry only transient network failures with a bounded policy.
Viewer test breaks after a browser update Viewer controls and browser capabilities are implementation-specific. Prefer checking response and PDF bytes for product behavior; maintain a separate, version-aware viewer test when required.

8. Or skip the browser setup

If your goal is to capture a webpage as a PDF for inspection, ScreenshotNeo provides a website screenshot API and MCP server. A single request can return a PDF; use the API documentation for available parameters.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.pdf

python -c 'import requests; r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com", "format": "pdf"}, timeout=90); open("shot.pdf", "wb").write(r.content)'

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com', format: 'pdf' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo removes cookie banners, popups, and chat widgets before capture; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The Free plan includes 1,000 screenshots per month with no card, and paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month with no card.

FAQ

Can Selenium verify that a download finished?

WebDriver does not expose download progress. Use Selenium to identify the link or session, then use an HTTP client to retrieve and verify the resource.

Can Selenium create a PDF without opening a print dialog?

Selenium’s print interface returns PDF data for supported drivers and bindings, allowing the test to save and inspect it programmatically.

Is checking for the PDF header enough?

No. It helps detect an obvious HTML response, but parse the document and assert the content or conformance your application requires.