ScreenshotNeo

BlogHTML to image & PDF

How to Automate PDF Testing

Build a layered PDF testing workflow for content, visual changes, standards conformance, and accessibility—with practical CI guidance.

By the ScreenshotNeo team4 October 20268 min read

How to automate PDF testing: generate representative PDFs from stable fixtures, assert their content and metadata, compare rendered pages with reviewed baselines, validate PDF/A or PDF/UA when required, and combine accessibility automation with human review. These checks find different defects; passing one does not prove a PDF is correct or accessible.

A reliable workflow separates four questions: Is the expected information present? Does the page look right? Does it meet a specified conformance profile? Can people, including assistive-technology users, use its structure and content? Match a check to each risk in your documents.

1. Make document generation repeatable

Test the PDF your application actually produces, not only a manually prepared sample. Create fixtures that represent real variations: short and long content, optional fields, multiple languages, tables that span pages, images, forms, and boundary values such as empty or unusually long labels.

  1. Use deterministic fixture data and a known generation path.
  2. Store expected PDFs or rendered-page baselines under version control, where practical.
  3. Pin or record the PDF renderer, fonts, and relevant runtime versions in CI. Environment changes can alter rendered pixels even when application data has not changed.
  4. When an intentional change updates output, review the new PDF and its diffs before accepting the baseline.

Keep fixtures small enough to diagnose failures, but include representative documents with many pages, complex tables, or embedded images. A single short sample rarely exercises pagination and layout boundaries.

2. Assert content and metadata

Use your application’s test framework and a PDF text extraction or parsing library selected for your language stack. The research available for this guide does not establish one extraction library as best. Treat extraction as an assertion aid: text order and content may behave differently in scanned PDFs, forms, tables, and PDFs with complex reading order.

Useful checks include:

  • Required text is present and forbidden or stale text is absent.
  • Page count stays within expected bounds.
  • Metadata fields such as title, author, or creation date match policy, when those fields matter.
  • Expected headings, totals, identifiers, and form values survive generation.
  • Expected fonts or embedded resources are present if your application depends on them.

Prefer assertions on meaningful values over a byte-for-byte comparison of the entire PDF. Generated files can contain changing metadata or internal object ordering even when the visible result is unchanged. For scanned documents, OCR output is a separate, error-prone signal; inspect representative pages visually as well.

3. Add PDF visual regression testing

Render each fixture to pages and compare those pages with approved baselines. A visual diff can expose clipping, shifted elements, missing images, changed line breaks, and pagination changes that text extraction misses.

  1. Choose a stable renderer and keep its version and fonts consistent between baseline creation and CI.
  2. Compare pages with a tool appropriate to the workflow. The pdf-visual-compare project documents a CLI workflow that can fail on differences and emit JUnit output. Check its current maintenance, dependencies, supported environments, and output stability before adopting it.
  3. Set a comparison tolerance based on reviewed output and the tool’s behavior. Do not copy a threshold from another project without examining what it hides.
  4. Save diff images and reports as CI artifacts so a failure can be reviewed without rerunning the job.
  5. For intended changes, inspect the changed pages and update only the affected baselines.

Choose comparison behavior for the document. Acrobat Compare Files can compare text, line art, and images; its settings distinguish reflowable reports and spreadsheets, presentation-like pages, and scanned documents compared as image captures. Review the resulting report and settings before treating a difference as a defect. See Adobe’s Compare Files guidance.

4. Validate PDF/A or PDF/UA conformance when required

When a contract, archive policy, or publishing requirement calls for PDF/A or PDF/UA, run a validator against the intended profile. veraPDF formalizes applicable requirements as profiles and reports failed checks with details such as object type, condition, specification, and conformance level.

Install veraPDF using its current official distribution, then validate explicitly against the required profile. For example, the CLI supports selecting a flavour with -f / --flavour; the exact profile identifier should be taken from the installed version’s CLI documentation.

# Check the installed CLI and its supported options first
verapdf --help

# Validate one PDF using the profile required by your policy
verapdf -f PDF/A-2B report.pdf

# Validate a batch; use the installed version's documented file/directory options
verapdf -f PDF/A-2B ./pdf-fixtures/

The profile shown is an example, not a recommendation for every document. Select the exact PDF/A part or PDF/UA profile that applies. Do not rely on automatic detection unless its metadata-based behavior matches your policy: veraPDF documents that detection can depend on XMP conformance declarations and that defaults may apply when declarations are absent or invalid.

Archive machine-readable reports and make failures understandable in CI. A validator result only describes checks in the selected profile and the claims it evaluates. veraPDF states that for PDF/UA it performs machine-verifiable checks only; that is useful evidence, not proof of full accessibility.

5. Check accessibility and route manual findings to people

Automated accessibility checks identify issues that can be evaluated mechanically. They cannot fully judge whether reading order makes sense or whether alternative text conveys the meaning of an image.

  • Desktop review: Acrobat Pro provides an accessibility checker and report, along with reading-order tools. Its documentation advises reviewing all reported issues because the checker does not determine whether content is essential. See Adobe’s Acrobat accessibility guidance.
  • API workflow: Adobe PDF Services offers an accessibility checker API for machine-verifiable PDF/UA and WCAG requirements, with a report. Adobe notes that human remediation may still be needed for reading order and meaningful alternative text. See the PDF Services accessibility checker documentation.
  • Standards validation: veraPDF’s PDF/UA profile can add machine-verifiable conformance checks. It does not replace inspection with assistive technology or review by people familiar with the document and its audience.

In a test report, separate automatically confirmed failures from items requiring manual review. Record who reviewed the latter and what representative pages or assistive-technology checks were examined.

6. Put the checks together in CI

A practical pipeline can run fast content checks on every change and retain the heavier review artifacts for failures. Adapt this sequence to your build system:

  1. Generate PDFs from committed fixtures.
  2. Run content, metadata, and page-count assertions.
  3. Render pages in a stable environment and compare them with approved baselines.
  4. Run the explicit veraPDF profile if the document has a conformance target.
  5. Run automated accessibility checks where they fit the workflow.
  6. Publish extracted text, validator output, visual diffs, and accessibility reports as CI artifacts.
  7. Require a person to review intended visual changes and accessibility findings that need judgment.

Keep failure categories separate in CI: content assertion, visual difference, standards violation, automated accessibility finding, or required human review. This makes triage clearer and prevents a green result in one category from masking a failure in another.

7. Choose tools by the evidence you need

Need Example What to check
PDF/A or PDF/UA conformance veraPDF CLI Required profile, explicit selection, report format, version, and policy
Desktop accessibility review Acrobat Pro Manual-check findings, reading order, and remediation workflow
Programmatic accessibility checks Adobe PDF Services API Machine-verifiable coverage, report format, page range, authentication, and service requirements
Visual regression pdf-visual-compare or Acrobat Compare Files CI versus interactive review, rendering and page matching, tolerance, ignored regions, and report output
Content and metadata Your test framework plus a PDF parser/extractor for your stack Text order, fields, fonts, tables, forms, and native versus scanned PDFs

For every tool, confirm current versions and supported environments before making its output a release gate. The tools answer different questions; combine the checks that match the document’s risks.

8. Troubleshooting common failures

Symptom Likely cause What to do
Text assertion fails although the page looks correct Extraction order, line wrapping, ligatures, or table reading order differs from the assertion. Inspect extracted text; assert stable semantic values and allow appropriate whitespace or ordering variation. Use a visual check for appearance.
Visual diff reports many changed pixels Renderer, fonts, operating system, or fixture content changed. Compare environment versions and inspect the diff artifact. Regenerate a baseline only after confirming the output change is intended.
Only later pages differ Pagination changed due to content length, font metrics, or a layout boundary. Check the first page where output diverges, then inspect preceding content and page-break behavior.
veraPDF fails an unexpected profile Automatic profile detection or a default may not match the required conformance target. Set the policy’s intended profile explicitly and verify the profile name against the installed CLI documentation.
PDF/UA validation passes but assistive use is poor The validator checks machine-verifiable requirements, not all semantic or usability judgments. Review reading order, alternative text, headings, and representative output with assistive technology and a knowledgeable person.
Accessibility report has “Needs Manual Check” items The issue requires human judgment or inspection. Review the flagged content and document the decision; do not treat the status alone as either a confirmed defect or a pass.
CI output is hard to diagnose Only a pass/fail status is retained, without page diffs or machine-readable reports. Publish rendered diffs, validator reports, and extracted text as artifacts; include the failed stage in the job summary.

Or skip the browser setup

If your PDF workflow also needs screenshots of the web pages that produce or accompany a document, ScreenshotNeo is a website screenshot API and MCP server. A single GET request captures a URL as PNG, JPEG, WebP, or PDF. It can provide a visual artifact for a web-to-PDF workflow; it does not replace PDF content assertions, conformance validation, or accessibility review.

For a complete PDF QA pipeline, combine browser output capture with the checks above. ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, and failed loads are never billed. AI agents can take screenshots through its MCP server. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request options, and sign up free for 1,000 screenshots a month, with no card required.

FAQ

Does a passing PDF/A check mean the document looks correct?

No. Conformance validation and visual comparison produce different evidence. Review rendered output for layout and content changes.

Can a PDF/UA validator certify that a PDF is accessible?

No single automated result establishes full accessibility. Automated checks cover machine-verifiable requirements; people still need to evaluate reading order, alternative text, and usability.

Should visual baselines be updated automatically in CI?

Use reviewed baseline updates. A changed baseline should correspond to an intentional, inspected document change.

What should be tested for scanned PDFs?

Include rendered-page comparisons and, when text search or extraction matters, test the OCR pipeline separately. Compare pages as images where appropriate.

References