How to Save a Webpage as a PDF and Verify the File Metadata
Save a webpage as a readable PDF, then inspect its document properties and check that the pages and content you need were captured.
To save a webpage as a PDF, open it in your browser, choose Print, select the browser’s PDF destination, review the preview, and save the file. Then open the PDF’s document properties to inspect recorded fields such as title, author, and creation date. Finally, check the pages and content themselves: metadata describes the file, but it does not prove that the webpage was captured completely or at a particular time.
1. Save a webpage as a PDF in your browser
- Open the webpage and choose Print. In Firefox, this opens print preview. In Edge, use Settings and more > Print, or press Ctrl+P on Windows or Command+P on macOS.
- Choose the PDF destination. Firefox documents Save to PDF; Edge’s available destination depends on the operating system and print setup.
- Review the preview before saving. Check that the intended text, images, and page range appear. Webpages can have a print layout that differs from their on-screen design.
- Adjust settings that affect the result: orientation, paper size, scale, margins, page range, and headers or footers. In Firefox, background printing and an Original or Simplified format may also be available.
- Save with a clear filename and location. Reopen the saved PDF to confirm it opens and contains the pages you expected.
Browser controls vary by version and operating system. For browser-specific steps, see the official guides for printing from Firefox and printing in Microsoft Edge.
2. Save a webpage as a PDF with Chrome Headless
For a repeatable command-line capture, Chrome documents a headless print-to-PDF option. This is useful for automation when you do not need to adjust each page in a visual print preview.
chrome --headless --print-to-pdf https://example.com/
To omit the printed date and time, URL, and page-number header and footer, add --no-pdf-header-footer:
chrome --headless --print-to-pdf --no-pdf-header-footer https://example.com/
You can set a maximum wait before capture with --timeout, in milliseconds. For example, this waits up to five seconds:
chrome --headless --print-to-pdf --timeout=5000 https://example.com/
Choose a wait that suits the page. A fixed delay may be too short for a slow page or unnecessarily long for a fast one. Chrome’s documented command-line options are described in the Chrome Headless documentation.
3. Inspect the PDF’s file metadata
In Acrobat, open the PDF’s document properties. On Windows, Adobe documents Menu > Document properties; on macOS, use File > Document properties. The Description view shows basic properties. To inspect embedded metadata, open Additional Metadata > Advanced.
Depending on how the PDF was created, properties may include:
| Field | What it describes |
|---|---|
| Title | A descriptive document title, if recorded. |
| Author | An author value recorded in the document, if present. |
| Subject and keywords | Descriptive fields that can help identify or categorize the file. |
| Creation and modification dates | Dates recorded in PDF metadata. They are not, by themselves, proof of when a webpage was captured. |
| Security settings, fonts, and initial-view preferences | Additional document properties that may be available in the properties views. |
| Custom or extended metadata | Embedded fields that may be visible in the advanced metadata view. |
PDFs can store document-level metadata in an optional document information dictionary and in a metadata stream. The PDF Reference describes both mechanisms; a given file does not necessarily contain both. For that reason, if a particular metadata field matters, inspect the embedded metadata view as well as the basic properties panel. See Adobe’s guide to viewing document properties and metadata and the PDF Reference.
4. Verify the PDF content separately
Metadata inspection and content review answer different questions. Properties show information recorded about the PDF. They do not confirm that every relevant part of the webpage made it into the saved file.
- Open the PDF and check that it is readable and that all expected pages are present.
- Review the text, images, tables, and page breaks that matter for your purpose.
- Check that the saved page range is correct and that scaling has not made text too small.
- Confirm whether page headers, footers, background colors, or other print settings changed the result.
- If links matter, inspect whether relevant links are present and usable in the PDF.
File-system dates, PDF creation or modification dates, and the date a webpage was captured are different things. Treat PDF dates as recorded metadata, not forensic proof of capture time or authenticity.
5. Choose manual printing or automated capture
| Approach | Useful when | Trade-off |
|---|---|---|
| Browser print preview | You are saving a page once and want to inspect or adjust its print layout. | Requires manual steps for each capture. |
| Chrome Headless | You need a repeatable command-line workflow. | You must choose suitable timing and cannot rely on a per-page visual preview. |
| ScreenshotNeo PDF API | You want a direct API call or need PDF capture in an application or automation. | Requires an API key and a network request. |
These approaches provide different workflows; the cited documentation does not establish comparative speed, fidelity, or quality measurements.
Or skip the browser setup
ScreenshotNeo can return a PDF from one GET request. See the ScreenshotNeo API documentation for the request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -d format=pdf -o page.pdf
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com", "format": "pdf"},
timeout=90,
)
r.raise_for_status()
open("page.pdf", "wb").write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com',
format: 'pdf'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo returned ${res.status}`);
await Bun.write('page.pdf', res);
ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server gives AI agents tools for taking screenshots, getting page information, and capturing PDFs. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Every feature is available on every plan. Learn more at ScreenshotNeo and read the docs, then sign up for 1,000 free screenshots a month, with no card.
Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| PDF is missing text, images, or sections | The page’s print layout differs from its live view, or content had not loaded before printing. | Inspect the preview, wait for the page to finish loading, and check the saved PDF. Try the browser’s available background or simplified-print options if appropriate. |
| Text is clipped or too small | Paper size, orientation, margins, or scale do not fit the content. | Change those settings and review the preview again before saving. |
| Unexpected URL, date, or page numbers appear | Print headers and footers are enabled. | Turn off headers and footers in the print settings, or use Chrome Headless with --no-pdf-header-footer. |
| Chrome Headless output is incomplete | The page needed more time to load than the configured wait allowed. | Increase --timeout and inspect the resulting PDF. A timeout is a maximum wait, not proof that all dynamic content is ready. |
| Expected metadata fields are absent | The PDF may not contain those fields, or the basic properties panel may not expose the embedded data you need. | Check Additional Metadata > Advanced. Some fields may be read-only depending on how the PDF was made. |
| PDF dates do not match the date you expected | Creation and modification metadata record file properties, not a verified webpage capture time. | Do not use those fields alone to establish when the webpage was viewed or captured. |
| Automated capture returns an error or no usable file | The browser command, page load, or network request may have failed. | For local Chrome, confirm the Chrome executable is available and the URL is reachable. For ScreenshotNeo, check the HTTP status, API key, URL, and response before saving the body as a PDF. |
Performance, reliability, and cost notes
- Manual capture: Preview takes extra time but lets you catch layout and page-range problems before saving. It has no API or scripting setup.
- Headless capture: Automation avoids repeated manual printing, but timing and dynamic page content can affect what is rendered. Use a wait suited to the page and verify output when completeness matters.
- Metadata inspection: A properties panel is a quick check; advanced inspection can reveal more embedded fields. Neither view establishes that the webpage content is complete.
- ScreenshotNeo: A request requires an API key and network access. Only clean shots are billed; bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing. Pricing is Free for 1,000 shots/month, Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000; yearly billing gives two months free.
FAQ
Does PDF metadata prove when I saved the webpage?
No. Creation and modification dates are recorded file metadata. They do not independently prove when the webpage was captured.
Can I edit PDF metadata without changing the page content?
Yes. Editing a metadata field does not change the document’s content. Some fields may be read-only depending on how the PDF was created.
Does every PDF contain both kinds of metadata storage?
No. The PDF specification describes an optional document information dictionary and a metadata stream; do not assume every file contains both.
Will a PDF look exactly like the live webpage?
Not necessarily. Pages can be designed to print differently from their screen presentation, so inspect the preview and the saved file.


