How to Do Visual Testing for PDF Files
Compare rendered PDF pages against a reviewed baseline, catch layout regressions in CI, and choose thresholds without hiding real changes.
Visual testing for PDF files means rendering each page of a new PDF and comparing it with a trusted reference version. A repeatable test uses the same renderer and settings for both files, reports differences page by page, checks page counts and dimensions, and requires review before accepting a baseline update. It catches visual layout changes; it does not prove that wording, metadata, accessibility structure, or embedded resources are correct.
1. Build a repeatable PDF visual regression workflow
- Create a trusted baseline. Generate the reference PDF from an approved version of your document output and keep it in version control or another controlled artifact store. Treat baseline changes as reviewed code changes, not routine cleanup.
- Fix the rendering conditions. Use the same renderer and version, resolution, font environment, page selection, page dimensions, and print-mark trimming for expected and actual PDFs. Record these choices so a renderer or font change does not look like an unexplained document regression.
- Compare corresponding pages. Generate a per-page report and, where available, difference PNGs. Check page count and dimensions explicitly: a missing page or changed page size is significant even if the remaining pages compare well.
- Review every flagged change. Decide whether each difference is intended. Update the baseline only after review and alongside the change that caused it.
- Add nonvisual checks for nonvisual requirements. Assert important text, metadata, accessibility structure, or embedded files separately. A visually similar page can still contain incorrect or inaccessible content.
Keep the baseline, renderer configuration, comparator configuration, and test output together as build artifacts. That makes a CI failure actionable instead of a bare pass/fail result.
2. JavaScript and TypeScript: compare PDFs with pdf-visual-compare
pdf-visual-compare provides a JavaScript/TypeScript library and CLI for comparing actual and expected PDF files. Its documented capabilities include page selection, pixel and percentage thresholds, exclusion regions, per-page mismatch details, optional diff PNGs, and CLI output suitable for failing CI and generating JUnit results. Check the project documentation for current installation and CLI syntax before pinning it in a build.
Install and run the CLI
npm install --save-dev pdf-visual-compare
npx pdf-visual-compare --help
Use the help output for the installed version to confirm its exact flags. A CI invocation should name the expected baseline and actual output, select the pages you intend to test, set conservative thresholds, request diff images or JUnit output if supported by that version, and return a failing exit code when differences exceed the configured rule.
# Example workflow shape; substitute the flags shown by your installed version's --help.
npx pdf-visual-compare \
--expected artifacts/baseline/invoice.pdf \
--actual artifacts/current/invoice.pdf \
--diff-output artifacts/diff \
--junit-output artifacts/pdf-visual-results.xml
The command above shows the inputs and outputs a CI step needs; flag spelling and availability can vary by release, so use the installed CLI help rather than copying unverified flags into a pipeline. Store the actual PDF, report, and diff images when the command fails.
Library use and safe inputs
The project also exposes library functionality for applications that need structured per-page results or custom reporting. Follow the API shown in the documentation for the version you pin. Prefer passing PDF bytes or buffers to the comparison API where practical. The project documentation recommends binary inputs, or restricting file paths to a trusted input root when paths are used; avoid allowing untrusted input to choose arbitrary filesystem paths.
Understand threshold semantics before tuning
The documentation describes a page as passing when it is within either the configured pixel threshold or the configured percentage threshold. That is an OR rule: meeting either limit can pass the page. Do not assume the thresholds must both be met. Begin with strict settings, inspect repeatable rendering noise, then tune with a documented reason. A generous threshold can mask small but meaningful changes.
3. cURL, Python, and Node.js for rendering the source page
PDF visual comparison starts with two PDF files. If your PDF is generated from a web page, a screenshot can help inspect the source page’s appearance, but an image of one page is not a substitute for rendering and comparing the PDF pages themselves. The following examples capture the web page that may produce the document.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.com/invoice \
-o invoice-source.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com/invoice"},
timeout=90,
)
r.raise_for_status()
with open("invoice-source.webp", "wb") as f:
f.write(r.content)
Node.js
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://example.com/invoice',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('invoice-source.webp', bytes));
See the ScreenshotNeo API documentation for request options and response headers. Keep API keys in environment variables or a secret manager in real applications, rather than committing them into source control.
4. Choose a comparison tool for your constraints
| Approach | Useful when | Points to evaluate |
|---|---|---|
| pdf-visual-compare | You want JavaScript/TypeScript integration and CI-oriented, page-level results. | Threshold behavior, page selection, excluded regions, diff artifacts, pinned version, and safe handling of input paths. |
| DiffPDF 6 | You need desktop or command-line comparison, including visual or text-oriented modes. | Its manual describes appearance comparison as pixel-by-pixel and offers word and character comparison modes. Confirm platform and current licensing details from its documentation. |
| Antenna House Regression Testing System | You need a commercial regression suite for PDFs or directories and page exclusions. | The vendor describes change highlighting and Windows, Linux, and macOS support. Treat claimed efficiency or completeness figures as vendor claims; validate fit and current terms directly. |
| Applitools ImageTester | You need hosted visual testing and baseline workflows that include PDF documents. | Its documentation describes page selection, rendering quality, and print-margin trimming. Pages are uploaded to Eyes, so check document confidentiality and applicable data-handling terms first. |
Compare tools on local versus hosted processing, supported CI environment and language, page-count and dimension handling, rendering controls, thresholds, exclusions, quality of diff reports, text comparison, platform support, licensing, current pricing, and data handling. Verify current version and commercial terms with each provider; this article does not independently test the products.
5. Handle dynamic content, exclusions, and text assertions
- Dynamic timestamps or generated IDs: Prefer making test output deterministic. If that is not possible, exclude only the smallest justified region and document why it varies.
- Fonts and glyph changes: Install the intended fonts in the rendering environment and pin renderer versions. A fallback font can shift line wrapping across many pages.
- Page selection: Select representative pages when runtime is a concern, but retain coverage for pages with distinct templates, tables, charts, or conditional sections. Missing pages should be checked independently.
- Print marks and margins: Apply the same trimming rules to both versions. Inconsistent cropping changes dimensions and may shift all pixels.
- Text correctness: Add assertions for required phrases, totals, identifiers, or extracted text when those matter. DiffPDF documents character- and word-comparison modes in addition to visual comparison.
- Exclusion areas: Keep them narrow and review them periodically. A broad exclusion can hide a genuine regression in nearby content.
6. Put PDF visual checks in CI
- Generate the candidate PDF in the same pinned environment used for the baseline.
- Run a page-count and page-dimension check before interpreting pixel differences.
- Run the visual comparator with a strict starting threshold and produce page-level results and diff artifacts.
- On failure, publish the actual PDF, baseline PDF, and per-page diff report as CI artifacts.
- Require a reviewer to classify differences as intended or unintended.
- For an intended change, update the baseline in a reviewed commit and retain the diff in the change record.
Hosted visual testing can simplify baseline review, but may transmit document pages to a vendor service. Confirm that this is permitted for the documents under test before sending real customer or confidential data.
7. Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| Nearly every pixel differs | Renderer version, resolution, fonts, color handling, or page crop changed. | Pin and align render settings and font packages; verify page dimensions and trimming first. |
| Only one page fails, or pages appear offset | Page selection or page correspondence is wrong, or an insertion shifted subsequent pages. | Compare page counts and inspect page-level results; verify selected page indexes and identify inserted or missing pages. |
| Small expected changes fail the gate | Threshold is strict or dynamic content is unstable. | Stabilize the generated PDF first. Then inspect repeated diffs and tune the relevant threshold or narrowly exclude a known dynamic region. |
| A real small defect passes | The threshold is too permissive, or the configured either-threshold rule allows a page to pass. | Review the comparator’s exact semantics and lower the permissive limit; add targeted text or geometry assertions for critical content. |
| Diff image is blank or missing | Diff output was not enabled, output directory is wrong, or the comparison stopped on an input error. | Check the installed CLI help, confirm the artifact path, and inspect the process exit status and logs. |
| CI cannot open a PDF path | Relative path differs in CI, artifact was not generated, or path input is restricted. | Print the resolved working directory and verify artifact existence; pass bytes or constrain paths to an approved root. |
| Hosted comparison is blocked by policy | PDF pages cannot be uploaded under the document’s confidentiality rules. | Use an approved local or self-managed workflow, or obtain the required data-handling approval before sending pages to a hosted service. |
8. Performance, reliability, and cost
Rendering and comparing more pages, at higher resolution, generally means more work and larger artifacts. Keep rendering settings stable and choose page coverage deliberately; do not reduce resolution or loosen thresholds until investigating why the comparison is noisy. Cache or reuse generated PDFs only when their inputs and renderer configuration are identical, so a stale artifact cannot create a false result.
For reliability, pin tool and renderer versions, retain failing inputs and diffs, and treat baseline updates as reviewed changes. A comparator is a signal for review, not a proof that every defect is found. Cost and operational effort depend on local compute, CI artifact retention, commercial licensing, or hosted service pricing and document upload requirements; confirm current vendor terms directly.
Or skip the browser setup
When the PDF is produced from a web page, ScreenshotNeo can capture that source page through one API request. It does not compare PDF files; use your PDF comparator for page-to-page regression checks. The capture API and options are documented at ScreenshotNeo docs.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
ScreenshotNeo accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Response headers identify the page verdict and billing status. Its MCP server lets AI agents use the take_screenshot, get_page_info, and capture_pdf tools. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots.
Sign up for 1,000 free screenshots a month, with no card required.
FAQ
Does a visual PDF test verify that the wording is correct?
No. Add text extraction or targeted content assertions for required wording and values.
Should every PDF page be compared?
Compare every page when feasible. If selecting pages, cover each distinct layout and separately check page count and dimensions.
Can I automatically accept a new baseline after a test fails?
Keep baseline updates under review. Automatically replacing a reference can normalize an unintended regression.
Can a screenshot API replace a PDF comparator?
No. A screenshot API captures a web page; PDF visual regression renders and compares document pages. Use a screenshot to inspect the source page and a PDF comparator to validate the generated PDF.


