How to Archive a Webpage as a PDF with a SHA-256 Checksum
Save a readable webpage snapshot as a PDF, calculate its SHA-256 checksum, and keep the metadata needed to check the file later.
To archive a webpage as a PDF and check later whether that saved file has changed, print the page to PDF, calculate the PDF’s SHA-256 digest, and keep the complete digest with the PDF and capture notes. A matching digest later shows that the compared file bytes match the bytes present when the reference digest was made. It does not prove who published the page, when it existed, or whether its contents were true.
This workflow creates a readable snapshot of a rendered page. It is not a complete website archive: the PDF may omit scripts, linked assets, interactive behavior, or dynamic content. For a broader web capture, use a WARC-capable workflow and preserve its capture metadata.
What the PDF and checksum do—and do not—preserve
Chrome’s print-to-PDF feature saves a rendition of what the browser prints. The SHA-256 digest identifies the bytes of the resulting PDF. NIST describes digests as a way to detect whether data has changed since the digest was generated.
| What you have | What it tells you | What it does not establish |
|---|---|---|
| A portable, readable rendition of printed page content | That every resource, script, interaction, or later-loaded item was captured | |
| SHA-256 digest | Whether the file bytes match a trusted reference digest | Publisher identity, capture time, truth of the page, or legal authenticity |
| URL and capture notes | Context about the source and how the rendition was made | Independent proof that the stated URL or time is accurate |
If you re-save, optimize, annotate, or edit the PDF, its bytes may change and its digest will likely change, even if the page looks the same. Generate the reference digest only after the PDF is in its intended final form.
Save the webpage as a PDF in Chrome
- Open the exact page you want to retain. Wait for the content you intend to capture to appear. Record its full URL and the current date and time, including the timezone.
- Open Print: press Ctrl+P on Windows or Linux, or Command+P on Mac.
- In the print dialog, choose Save as PDF as the destination.
- Review the page range, orientation, scale, margins, headers and footers, and background graphics. Choose settings that preserve the material you need without clipping or adding unwanted browser details.
- Save the file with a clear name, such as
example-com-2026-10-04.pdf. Inspect the saved PDF, including later pages, for missing content, blank pages, clipped columns, or awkward pagination.
Chrome’s interface can vary by operating system and version. See Google’s Chrome printing instructions.
Calculate and record the SHA-256 checksum
Calculate the digest over the final saved PDF using a trusted local checksum utility. This research does not establish platform-specific command syntax, so use the documentation for the checksum tool already available on your system rather than relying on an unverified command.
- Choose the saved PDF and run the tool’s SHA-256 operation on that file.
- Copy the entire digest exactly. A partial digest is not a reliable comparison.
- Save the digest in a text record next to the PDF, along with the full URL, capture date and time with timezone, browser and version if known, and print settings that could affect the result.
- When checking the file later, calculate SHA-256 again over the file and compare the complete new digest with the original reference.
Keep the reference digest somewhere that cannot be silently edited together with the PDF if you need tamper evidence. A checksum stored only beside a file can detect accidental changes, but someone able to replace both can replace the reference too. NIST’s Secure Hash Standard describes the role of message digests in detecting changes since their generation.
Choose PDF or a fuller web archive
Use PDF when a human-readable, fixed rendition is sufficient. Use a WARC-capable capture when you need to retain multiple web resources and related capture information for possible replay. WARC is described by the Library of Congress as a format for combining multiple digital resources into an aggregate archival file with related information. A WARC capture still cannot guarantee perfect replay or capture every kind of content, including some streaming, database-backed, deep-web, or otherwise unsupported material.
The Library of Congress also describes WACZ packages whose manifests can include SHA-256 or MD5 checks for files in the package. That package workflow is broader than calculating one digest for an exported PDF. Learn more from the Library of Congress on web archive formats, WARC, and WACZ. For context about what web archives aim to preserve, see its Web Archiving FAQ.
Or skip the browser setup
ScreenshotNeo can return a page capture as a PDF with one API request. Create an API key, then run this cURL example (replace the target URL as needed):
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -d format=pdf -o page.pdf
See the ScreenshotNeo API documentation for request options. You can calculate a SHA-256 digest of the returned PDF using your local checksum utility and keep it with the same URL and capture-time notes. The digest checks the returned file bytes; it does not authenticate the live webpage or establish capture time.
ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, and failed loads are never billed; an MCP server lets AI agents take screenshots; and 1,000 screenshots per month are free with no card, with paid plans starting at $5 for 3,000. Sign up for the free plan.
Reliability, performance, and cost notes
- Capture reliability: A successful print operation only confirms that a PDF was produced, not that it contains every item you wanted. Inspect the document and record any limits, such as content that loads only after interaction.
- Integrity reliability: The comparison is only as trustworthy as the reference digest and the process used to preserve it. A digest does not act as a digital signature or trusted timestamp.
- Performance: Printing a single page is generally a short manual workflow, but complex pages may take time to finish loading and may paginate unpredictably. Wait for required content and inspect the output.
- Cost: Chrome’s built-in print flow and a local checksum tool avoid a per-capture API charge. ScreenshotNeo’s free tier includes 1,000 shots a month with no card; paid tiers start at $5 for 3,000. Choose based on the number of captures and whether API or MCP automation is useful.
Troubleshooting
| Problem | Likely cause | What to do |
|---|---|---|
| PDF is blank or missing recently loaded content | The page had not finished loading, or the content requires scrolling or interaction | Wait for the desired content, scroll or interact as needed, then print again and inspect the result. |
| Text, columns, or images are clipped | Print scale, orientation, or margins do not fit the page layout | Adjust those settings in the print dialog, save another copy, and review all pages. |
| Background colors or graphics are absent | Background graphics are disabled in print settings or the page does not print them | Check the background graphics option and inspect the new PDF. |
| The digest differs after re-saving | Saving, optimization, annotation, metadata changes, or edits changed the file bytes | Compare against the original untouched PDF. If the new version is the intended record, treat it as a new file and create a new reference digest. |
| The digest tool reports an error | Wrong file path, inaccessible file, or the utility is not set to SHA-256 | Confirm the selected file and permissions, choose SHA-256 explicitly, and copy the complete output. |
| The PDF and digest agree, but provenance is disputed | A checksum proves only a byte comparison against its reference | Preserve the reference independently and follow the applicable records or evidentiary process; a checksum alone does not prove source, author, or time. |
FAQ
Does the checksum verify that the webpage has not changed?
No. It verifies whether the saved PDF bytes match the reference digest. To compare webpage versions, you would need to capture each version and compare those records separately.
Should I hash the page URL or the PDF?
For this workflow, hash the saved PDF. The URL is context to record alongside it, not the input file whose rendered contents are being checked.
Can a matching digest prove when I captured the page?
No. It indicates a match to the reference digest, but does not independently prove when that reference was made. Keep capture time and provenance records, and use a suitable trusted timestamp or records process if your requirements demand one.
Is a PDF enough for long-term website preservation?
Only when a readable rendition meets the need. For resources and replay context, consider a WARC-based web archive while recognizing that no capture method preserves every possible kind of site content.


