ScreenshotNeo

BlogHow-to

How to Archive a Webpage as a PDF Before It Expires

Save a local PDF while the page is available, check that it contains what you need, and optionally preserve a separate Wayback Machine reference.

By the ScreenshotNeo team4 October 20267 min read

To keep a local, readable copy of a webpage, open it while it is still available, use your browser’s print dialog, choose its PDF destination, review the preview, and save the file. Open the PDF afterward and check that the material you need is present and legible. For a second, shareable record, submit the page URL to the Internet Archive’s Save Page Now and keep the resulting archived URL too.

These are different records: a PDF is a static document, while a Wayback capture is an archived web view. Neither should be treated as a guaranteed, complete backup of an entire website. The Internet Archive explains that Save Page Now saves one submitted page, including images and CSS in most cases; it does not save outlinks or start a site-wide crawl.

1. Prepare the page before printing

  1. Open the exact page you need while it is still accessible. If it requires a sign-in or is likely to expire soon, do not wait until you have finished unrelated work.
  2. Scroll through the page once. Expand accordions, tabs, footnotes, or other sections that must appear in the record. Some pages load images or text only as you scroll, so content that has not loaded may not appear in the printout.
  3. Check whether the page has a print view or a reader view that presents the content more clearly. If you use one, confirm it still contains the tables, figures, citations, and context you need.
  4. For pages that change frequently, note the original URL and the date you saved the PDF. The PDF records what your browser rendered at that time; it does not establish that the page was authoritative or unchanged.

2. Save the webpage as a PDF

  1. Open the browser’s print dialog using its menu or print command. Menu names and keyboard shortcuts vary by browser and operating system.
  2. Choose the PDF destination, commonly labeled “Save as PDF” or similar. If the dialog offers paper size, orientation, margins, scale, headers and footers, or background graphics, adjust them to make the content readable.
  3. Inspect the print preview from the first page to the last. Look for clipped columns, blank pages, missing backgrounds, cut-off code, awkward page breaks, and images that did not load.
  4. Save to a known folder with a descriptive filename, such as vendor-policy-2026-10-04.pdf. Use a filename that helps you identify the page and capture date later.
  5. Open the saved PDF in a PDF reader. Confirm that the pages, text, links, images, and tables you care about are present. If something is missing, return to the live page, load or expand that content, and print again.

Print-to-PDF is a practical way to preserve the visible document, not a promise that every browser will reproduce every webpage completely. Dynamic charts, embedded media, lazy-loaded sections, and content requiring interaction may need extra attention.

3. Keep a Wayback Machine reference as well

If you want a public archived URL in addition to your local file, visit Internet Archive Save Page Now, submit the page URL, and wait for the capture result. Copy the resulting archived URL and open it to inspect what was captured.

The Internet Archive says it cannot guarantee that a site has been or will be archived. Treat the archived URL as a useful reference, not as your only copy. See its Wayback Machine guidance for the limits of locating and accessing saved pages.

4. Choose the record that fits your purpose

Method What you keep Useful for Key limitation
Browser print to PDF A local, static document Reading offline, printing, or retaining a specific rendered view Interactive behavior and content that did not load may be absent
Save Page Now A public archived URL for the submitted page Sharing a web reference or checking a captured version One page at a time; not a full-site backup or guaranteed complete capture
Browser extension A convenient way to submit the open page to Save Page Now Saving the current page without copying its URL manually Uses the same one-page capture method and limitations
Archive-It A managed collection and recurring crawl service Organizations preserving collections of web content A separate paid institutional service, rather than a one-off personal PDF workflow

For a one-off page, the browser PDF plus an optional Save Page Now capture is usually the relevant workflow. Organizations that need recurring collection-scale crawls can consider Archive-It, which the Internet Archive describes as a paid service with technical and web-archivist support.

5. Troubleshooting missing or broken content

Symptom Likely cause What to try
A section or image is absent from the PDF It had not loaded, was below the fold, or required interaction Return to the page, scroll to the section, wait for it to load, expand it, then print again.
Text or a table is cut off The page is wider than the selected paper layout or the print scale is too large Review orientation, paper size, margins, and scale in the print dialog. Check the preview before saving.
Colors or background images are missing The print dialog may omit background graphics by default Check for a background-graphics option. If the material remains unreadable, retain another record of the rendered page.
The PDF has extra blank pages or poor page breaks Print layout, margins, or page-break rules split the content awkwardly Adjust scale, orientation, or margins and inspect the full preview again.
Save Page Now fails or returns an incomplete capture The site may block crawling, have SSL-related issues, or depend on content that the archive could not capture Keep the local PDF, inspect the archive result, and retry while the original page remains accessible if appropriate.
The archived page shows an unexpected current image or resource The archived view may have fallen back to a live resource that was not saved Check the timestamp and resource behavior; do not assume the displayed resource was captured.

6. Reliability, performance, and storage

  • Act early: the workflow depends on the source page still being accessible. Save the PDF and submit the archive capture as soon as you know the material matters.
  • Verify both records: a PDF can omit content that did not render, and an archive can be incomplete or later unavailable. Open each result while you can still compare it with the original.
  • Keep useful context: retain the source URL and capture date alongside the PDF. If the page is evidence for a decision or citation, record the archived URL too.
  • Manage storage simply: PDFs are local files; choose a known folder and a descriptive filename. For important records, keep a copy in the storage system you already use for important documents.
  • Do not mistake a screenshot for a PDF: an image can preserve appearance, but it is not the same as a paginated, searchable document. Choose the output format for the way you need to use the record.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. Its API can return a screenshot or PDF; use the API documentation for PDF options such as paper size, margins, landscape, and page ranges. The example below is the documented one-call screenshot request, which saves a WebP image; it is not a PDF request.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Before a capture, ScreenshotNeo accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000 shots.

Sign up free for 1,000 screenshots a month, with no card required.

FAQ

Does a PDF preserve the original webpage exactly?

No. It preserves a static rendering of what the browser printed. It does not preserve the complete website or guarantee that dynamic content, interaction, or every resource will be included.

Does saving a page to the Wayback Machine create a PDF?

No. Save Page Now creates an archived web reference. Use the browser’s print-to-PDF workflow for a local PDF, and keep both if you need both forms.

Can I archive a whole website by submitting its homepage?

Save Page Now is for a single submitted page; it does not start a site-wide crawl. Collection-scale needs are a different use case.

Should I keep the original URL if I have the PDF?

Yes. Keeping the source URL and the date you saved it helps identify what the local document represents. Add the archived URL when you also submit it to Save Page Now.