ScreenshotNeo

BlogHow-to

How to Archive a Website’s HTML, Images, and PDF for Offline Evidence

Preserve a website's HTML and linked files in a replayable archive, add a readable PDF, document the capture, and check what actually made it offline.

By the ScreenshotNeo team4 October 20269 min read

To archive a website for offline evidence, preserve the HTML and its referenced resources together in a web-archive format such as WARC, then save a PDF or screenshot as a convenient rendered view. Record the source URL, capture time and scope, inspect the archive for missing content, and keep separate copies. A PDF alone is not a complete website archive: it does not preserve the site’s resource relationships or working hyperlinks as a replayable site.

A capture records what the capture process obtained. It does not prove that nothing was omitted, that a dynamic site existed in precisely that state at one instant, or that the files satisfy legal requirements in every jurisdiction. Treat provenance, inspection notes, and preservation of the original capture as part of the record.

1. Define what you need to preserve

“The website” can mean one page, a set of paths, or pages that rely on resources hosted on other domains. Decide and write down the scope before capturing. A single page URL does not imply that the surrounding site or every dependency is included.

  • List the exact URL or URLs and, when relevant, the paths or sections in scope.
  • Record the date and time in UTC, the time zone if the source time was recorded locally, and why the capture was made.
  • Note whether access required a login, special headers, or an interaction. Do not assume a public capture represents content behind authentication.
  • Decide whether you need a replayable package, a readable fixed view, or both. For evidence-oriented preservation, keep the web capture and a PDF or screenshot as separate artifacts.

For a small amount of content, browser “Save As” can be an option, but check whether the result includes linked files and can be reopened correctly. Save descriptive metadata such as the site name and date created. The Library of Congress discusses browser export, linked files, and descriptive metadata in its personal archiving guidance.

2. Capture HTML and the resources it depends on

Saving only HTML can leave a page without its images, stylesheets, scripts, linked PDFs, or other resources. For a useful offline replay, preserve relevant resources and their relationships to the pages that reference them.

WARC is a preservation-oriented container for web resources. The Library of Congress identifies WARC as its preferred web-archive format and explains that a WARC content block can contain resources in different formats, including binary images or audiovisual files embedded or linked in HTML. NARA’s transfer guidance lists WARC 1.0, WARC 1.1, and WACZ among preferred formats for permanent federal web records. These format choices do not guarantee that every resource will be captured. See the Library of Congress WARC description and NARA file-format tables.

Choose a capture method that can save the page and its dependencies in the format you need. The sources listed here describe preservation formats and workflow considerations, not a particular command-line crawler or its syntax, so this guide does not prescribe unverified capture commands. Before relying on a tool, confirm that it exports WARC or another format your intended replay workflow supports, what URL scope it follows, and how it handles external resources, redirects, and errors.

3. Save a PDF or screenshot as a readable companion

A PDF or screenshot is useful for quickly seeing a page as it rendered during capture. Keep it alongside the web archive and label it as a derivative view with its source URL and creation time. It can help a reviewer who does not have a replay tool, but it is not a substitute for the archive: a static view does not preserve the site’s full hyperlink behavior and underlying resource relationships. NARA’s web-record guidance identifies static screenshots as lacking hypertext functionality.

For a browser-based workflow, open the scoped page, wait for the content you need to appear, then use the browser’s print-to-PDF feature or save a screenshot. Record the browser or tool and version if known, and any relevant print settings. Dynamic content, lazy-loaded images, and content revealed only after interaction may not appear in the PDF. Compare the derivative against the live page and the archive replay, and note visible omissions.

4. Create a manifest and inspect the capture

Keep a short README or manifest with the archive. It gives later readers the context needed to understand what the files represent. A practical manifest can include:

  • Site or matter name and the exact source URL or URL list.
  • Capture date and time in UTC, plus the original time zone if applicable.
  • Scope, exclusions, and reason for capture.
  • Capture tool and version, if known, and relevant settings.
  • File inventory, including the web archive, PDF or screenshot, and manifest.
  • Access problems, redirects, login requirements, blocked resources, capture errors, or known gaps.
  • A brief description of the folder layout or site structure where it helps explain page relationships.

Then open the archive with a compatible replay tool and inspect the pages that matter. Check key images, internal links, downloadable documents, and pages at the edge of the requested scope. Record missing assets or content that could not be accessed. This checklist is a practical synthesis of preservation guidance, not a universal evidence protocol.

Completeness has limits. The Library of Congress notes that a site cannot be harvested instantaneously and completely; a dynamic site may change during capture, producing a combination of resources that did not exist together at one exact instant. NARA explains that reconstruction depends on whether referenced static and dynamic files were tracked. See the Library of Congress quality factors and NARA guidance on managing web records.

5. Organize and preserve multiple copies

Use descriptive filenames and keep related artifacts together. For example, group the WARC or WACZ capture, its PDF or screenshot derivative, and the manifest under a dated folder for the site or matter. Preserve original capture files and keep notes about any later conversion or review separately.

Maintain at least two copies in different locations, and periodically verify that the files still open. The Library of Congress suggests checking readability at least once a year and refreshing storage media every five years or when needed. It lists portable hard drives as one possible storage medium; a drive stores a copy but does not perform the website capture or guarantee preservation. These are general preservation recommendations, not a certification against loss. See Library of Congress personal archiving guidance.

Do-it-yourself workflow checklist

  1. Write down the precise pages or paths, access time, time zone, and capture purpose.
  2. Capture HTML together with relevant images, documents, and other referenced resources in WARC or a suitable web-archive format.
  3. Save a PDF or screenshot as a separate, labeled view of the rendered page.
  4. Create a manifest with the source, time, scope, method, file inventory, and known limitations.
  5. Replay the archive and inspect the important pages, assets, links, and downloads.
  6. Record gaps, then place at least two copies in separate locations and schedule readability checks.

Or skip the browser setup

If you need a rendered screenshot to accompany your archive, ScreenshotNeo can return an image or PDF from one request. It is a website screenshot API and MCP server; it creates a rendered view, not a WARC web archive, so keep using a web-archiving workflow for HTML and linked-resource preservation. The call below requests a PDF companion for a page.

cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -d format=pdf -o page.pdf

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com", "format": "pdf"},
    timeout=90,
)
r.raise_for_status()
with open("page.pdf", "wb") as f:
    f.write(r.content)

Node.js:

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com',
  format: 'pdf'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('page.pdf', bytes));

See the ScreenshotNeo API documentation for request options. Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000 screenshots. ScreenshotNeo is a website screenshot API for developers.

Sign up for 1,000 free screenshots a month, with no card required.

Troubleshooting

Symptom Likely cause What to do
HTML opens, but images or styling are missing. The capture saved the document without all referenced resources, or the references point outside the captured scope. Review the capture tool’s scope and resource handling, recapture the needed paths or domains if accessible, and document anything still missing.
A page or link does not replay offline. The linked page was outside scope, required a live server, or was not captured. Check the manifest and archive inventory; expand the stated scope only where needed and recapture accessible content.
The archive differs from the live page. The site changed during capture, content loaded dynamically, or a dependency was unavailable. Record the capture time and observed difference. Inspect the replay and preserve a PDF or screenshot as a companion view; do not claim the capture is complete.
The PDF lacks images or below-the-fold content. Images may be lazy-loaded, the page may not have finished rendering, or print settings may omit content. Wait for the relevant content before printing, inspect the PDF, and use the archive for resource preservation rather than treating the PDF as the complete record.
Replay works on one machine but not another. The second machine may lack a compatible replay tool or the archive copy may be incomplete or unreadable. Preserve the archive format and tool/version information in the manifest, verify the copied file, and test it with a compatible replay tool.
Capture misses login-only or blocked content. The resource may require authentication, permission, or an interaction the capture process did not perform. Record the access restriction and gap. Capture only content you are authorized to access, and describe the method and limitations accurately.

Reliability, performance, and cost considerations

  • Completeness: A one-page capture is narrower and easier to inspect than a site-wide crawl, but it may omit linked pages and cross-domain resources. State the actual scope.
  • Time and change: Crawling is not instantaneous. Pages and resources can change during the capture window, so include timestamps and avoid implying a single perfectly simultaneous snapshot.
  • Verification effort: Replaying key pages and checking files takes time, but it reveals gaps that a successful export alone cannot show.
  • Storage: WARC packages and their companion derivatives require space for the captured resources. Keep redundant copies and check that they remain readable.
  • Tool cost: Capture-tool pricing and limits depend on the selected tool; the preservation sources here do not establish a general price. ScreenshotNeo pricing applies to rendered screenshots, not WARC capture: Free is 1,000 shots/month with no card; Starter is $5 for 3,000; Growth $15 for 15,000; Pro $39 for 60,000; Scale $99 for 250,000; Business $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan.

Frequently asked questions

How do I save a website with all its images?

Use a capture method that includes HTML and referenced resources, then replay the result and check the important images. A saved HTML file by itself may still refer to resources that were not captured.

How can I view a website offline?

Open the saved web archive with a compatible replay tool. A PDF is easier to read in many viewers but presents a fixed rendering rather than a replayable site.

Is a PDF enough to preserve a website?

It is useful as a readable derivative, but on its own it does not preserve the site’s full linked-resource structure or hypertext behavior. Pair it with a web archive when offline replay and resource relationships matter.

Does a WARC prove that a website was complete or unchanged?

No. The format can hold web resources, but completeness depends on what the capture actually collected. Keep provenance and inspection notes, and describe known gaps.

A file format alone cannot establish admissibility in every court or jurisdiction. Preserve the capture circumstances and seek applicable legal guidance when that question matters.