How to Test PDFs as Images
Render each PDF page with consistent settings, then compare it with an approved image baseline. Catch visual changes and page-structure errors in a repeatable workflow.
To test a PDF as images, render every page with a pinned renderer and consistent settings, then compare each rendered page with its approved baseline. Also check page count and order: a pixel comparison cannot reliably catch a missing or duplicated page if the test pairs the wrong pages.
This guide uses PyMuPDF to render the PDF itself. If your goal is to test a PDF displayed inside a web application, use a browser screenshot test as well; Playwright Test provides toHaveScreenshot() for visual comparisons of browser views. PyMuPDF’s page rendering guide and Playwright’s visual comparison guide document these respective workflows.
1. Choose what you are testing
There are two different targets:
- The PDF file: Rasterize its pages with a PDF renderer, then compare the resulting page images. This catches changes to the document output itself.
- The web application’s PDF viewer: Open the page in a browser and capture the rendered view. This can catch viewer layout or integration problems, but it is a test of the browser experience, not a substitute for checking the PDF’s pages directly.
For a PDF regression test, use one image per page. Record and pin the renderer version and rendering settings so baseline and candidate images are produced under the same conditions.
2. Install PyMuPDF and render every page
Install the library in the Python environment used by your test job:
python -m pip install PyMuPDF
Save this as render_pdf.py. It writes one PNG per page and a JSON manifest containing page count and image dimensions. The explicit DPI makes the output independent of the documented 96 DPI default.
from pathlib import Path
import json
import sys
import pymupdf
source = Path(sys.argv[1])
out_dir = Path(sys.argv[2]) if len(sys.argv) > 2 else Path("rendered")
dpi = int(sys.argv[3]) if len(sys.argv) > 3 else 150
if dpi <= 0:
raise SystemExit("DPI must be a positive integer")
out_dir.mkdir(parents=True, exist_ok=True)
doc = pymupdf.open(source)
pages = []
for index, page in enumerate(doc):
pix = page.get_pixmap(dpi=dpi, colorspace=pymupdf.csRGB, alpha=False, annots=True)
filename = f"page-{index + 1:04}.png"
pix.save(out_dir / filename)
pages.append({
"page": index + 1,
"file": filename,
"width": pix.width,
"height": pix.height,
})
manifest = {
"source": source.name,
"page_count": len(doc),
"dpi": dpi,
"colorspace": "RGB",
"alpha": False,
"annotations": True,
"pages": pages,
}
(out_dir / "manifest.json").write_text(json.dumps(manifest, indent=2) + "\n")
print(f"Rendered {len(doc)} pages to {out_dir}")
Run it with python render_pdf.py report.pdf rendered 150. Change the DPI to suit the defects you need to detect, then keep it unchanged for both baseline generation and later runs. The example renders in RGB without an alpha channel and includes annotations; change those choices only deliberately and record the new settings.
3. Create and review image baselines
Render a known-good PDF and save its per-page images as the approved baseline. Render each candidate PDF into a separate directory using the same script, renderer version, DPI, colorspace, alpha setting, and annotation policy.
Compare corresponding page images with your team’s image-diff tool or visual regression system. PyMuPDF handles rendering; it does not define a universal pass/fail pixel tolerance. A difference identifies output that changed, not whether the change is a defect. Inspect changed pages before accepting a baseline update.
Keep baseline changes reviewable in version control or in the artifact system your team uses. Avoid automatically replacing the approved images whenever a comparison fails: that would turn unintended regressions into accepted output.
4. Check page structure and image geometry
Before comparing pixels, validate the manifest and image dimensions. At minimum, fail the test when:
- The candidate page count differs from the expected count.
- A page image is missing, duplicated, or associated with the wrong page number.
- A page’s width or height changes unexpectedly.
- The rendered bounds indicate clipping or unexpected page-size changes.
Page order matters: compare page 1 with page 1, page 2 with page 2, and so on. Do not pair pages solely by visual similarity or by a filename that could conceal a reordered document.
5. Control the rendering settings
| Setting | What to decide | Why it matters |
|---|---|---|
| Renderer and version | Pin the PyMuPDF and underlying rendering environment used in CI. | Different rendering environments are not guaranteed to produce identical pixels. |
| DPI | Set an explicit value, such as 150 or 300, and use it for every run. | Resolution changes image dimensions and the visibility of fine details. PyMuPDF documents a 96 DPI default and supports explicit DPI rendering. |
| Colorspace | Choose a colorspace, such as RGB, and hold it constant. | Different channel values change pixel comparisons. |
| Transparency | Decide whether to include an alpha channel. | Transparency affects raster output and comparison behavior. |
| Clipping | Use the same page bounds or clip rectangle on every run. | Different bounds can crop content or change output geometry. |
| Annotations | Set whether annotations should appear in the render. | Markup can change the output; PyMuPDF’s page rendering API exposes an annots option. |
See the PyMuPDF Page API reference for rendering options. If you use clipping or other options beyond the example, apply the same values when creating baselines and candidate renders.
6. Test a PDF shown in a browser
If the requirement is that a user sees the right PDF inside your application, test that application with a browser screenshot assertion in addition to rendering the file’s pages. Playwright Test documents toHaveScreenshot() and its visual comparison workflow for browser screenshots:
import { test, expect } from '@playwright/test';
test('PDF viewer page matches its visual baseline', async ({ page }) => {
await page.goto('http://127.0.0.1:3000/reports/sample');
await expect(page).toHaveScreenshot('report-viewer.png');
});
Replace the example address with the route under test. Make the application data, viewport, browser, and loaded state repeatable so screenshots can be compared meaningfully. Playwright’s matcher waits until two consecutive screenshots match before saving the last one. This assertion is for browser screenshots; use a PDF renderer such as PyMuPDF to create page images when the PDF file itself is the subject.
7. cURL, Python, and Node.js with ScreenshotNeo
When the PDF is a public web page that you need to capture as a PDF artifact, ScreenshotNeo can return a PDF from one API request. Its PDF options include paper size, margins, landscape mode, and page ranges. This is useful for capturing a web page as a PDF; for visual regression of an existing PDF file, render that file’s pages locally as described above. See the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
These supplied examples save or request a WebP screenshot of a URL. ScreenshotNeo also supports PDF output and PDF settings; select the PDF response and options described in its documentation when your target is a web page to print. Keep API keys out of source control and check the response status and headers in production code.
Or skip the browser setup
ScreenshotNeo takes a website screenshot or PDF with one GET request. Cookie banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000 shots. Every feature is on every plan. Learn more at ScreenshotNeo and read the API docs.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Sign up free for 1,000 screenshots a month, with no card required.
Performance, reliability, and cost
- Resolution and storage: Higher DPI creates larger images and can increase rendering time and artifact storage. Use the lowest resolution that exposes the defects your check needs to detect.
- Repeatability: Pin the renderer and settings, use stable input PDFs, and keep the baseline creation environment aligned with CI. The cited documentation does not promise pixel identity across renderers or operating systems.
- Review effort: A full-page visual diff can flag many changed pixels for a small meaningful edit. Treat it as a review signal and tune any project-specific tolerance against actual output; the cited sources specify no universal PDF tolerance.
- Cost: PyMuPDF is a local rendering library; the cited sources provide no pricing claim. For ScreenshotNeo, the stated free allowance is 1,000 shots per month, with paid plans from $5 for 3,000. Only clean shots are billed; consult the product documentation and plan details for current configuration.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Every page differs after a dependency update | The renderer version or environment changed. | Pin the version and environment. Regenerate baselines only after reviewing the visual changes. |
| Images have unexpected dimensions | DPI, page bounds, or clipping differs. | Set DPI explicitly and compare the manifest dimensions. Keep page bounds and clip settings stable. |
| Markup appears or disappears between runs | The annotation setting differs. | Set annots explicitly and use the same value for both renders. |
| Colors or transparent areas differ | Colorspace or alpha configuration changed. | Use fixed colorspace and alpha settings, and recreate candidate and baseline images under those settings. |
| Visual test passes despite a missing page | The harness compares only images it found or pairs pages incorrectly. | Assert expected page count, contiguous page numbering, and order before pixel comparison. |
| Browser screenshot is unstable | The application may not have reached a stable state or may vary across runs. | Make test data and viewport repeatable, wait for the relevant viewer state, and inspect the output before updating the baseline. |
| ScreenshotNeo request does not produce the expected PDF | The example request is the screenshot/WebP form, or PDF options are absent. | Use the PDF response format and options documented for the API. Check the response and its headers. |
FAQ
Should I compare PDF files directly or compare page images?
Use page images for visual regression. Also validate document structure when page count, order, or bounds matter.
Is 300 DPI always better for regression tests?
No. Higher DPI produces larger images. Choose a resolution that catches the defects you care about and keep it fixed.
Does Playwright render arbitrary PDF files for this assertion?
The documented screenshot assertion compares browser screenshots. Render PDF pages to images with a PDF renderer when testing the PDF file itself.
Does any changed pixel mean the test should fail?
A difference means the rendered output changed. Review it against the intended document change; the cited sources do not prescribe a universal tolerance.


