ScreenshotNeo

BlogHow-to

How to Archive Expired Website Pages and Visual Assets Automatically

Learn how to capture a single page or crawl a site before it disappears, preserve its assets, export WARC or WACZ files, and verify the replay.

By the ScreenshotNeo team4 October 20269 min read

To preserve a page before it disappears, choose a capture method that matches the scope: use the Wayback Machine’s Save Page Now for a one-off public URL, ArchiveWeb.page to capture resources as you browse and interact, or a scheduled crawler such as Browsertrix or Archive-It for recurring site-wide preservation. Export and replay the result, then check important pages and assets. No method guarantees a perfect copy of every page or resource.

If by “automatically” you mean making a visual screenshot, an image capture is useful evidence of how a page looked at one moment. It is not a replayable website archive: it does not preserve the page’s underlying links, scripts, stylesheets, or media as usable resources. ScreenshotNeo can make that visual record with one API request; use a web archiving workflow when you need to browse the preserved material later.

1. Decide what you need to preserve

Need Good fit What to expect
One public page quickly Wayback Machine Save Page Now Submits one page for capture. It does not crawl outlinks or the whole site.
A page with interactive or dynamically loaded assets ArchiveWeb.page Records resources loaded while you browse and interact. Review the collection for missing content.
Recurring site-wide captures Browsertrix Webrecorder describes it as a platform for automated whole-site archiving and scheduled crawls.
An institutional crawling program with support Archive-It A subscription option with technical and web archivist support.
A quick visual record, rather than a replayable archive ScreenshotNeo Returns a screenshot or PDF from one GET request; it does not replace WARC/WACZ capture and replay.

There is no uniform price comparison in the cited service documentation. Confirm current service terms, access limits, scheduling, and crawl configuration with providers before relying on them for a preservation program.

2. Save a single public page with the Wayback Machine

  1. Open the Internet Archive’s Save Pages in the Wayback Machine guide and use its Save Page Now workflow.
  2. Submit the exact URL you want preserved, including the relevant path and query string.
  3. Wait for the capture result and open the archived page. Check whether its images and styles loaded as expected.
  4. Save the resulting archive URL with your project notes so others can locate the snapshot.

Save Page Now is for the submitted page, not a crawl of the site. The Internet Archive Help Center says, “This method only saves a single page, not the whole site.” Images and CSS are included when capture succeeds. Some sites prohibit crawling, and SSL-related settings can cause capture problems. If the page links to other important pages, submit those individually or choose a crawler suited to the scope.

3. Capture interactive pages and their loaded assets

ArchiveWeb.page is useful when assets appear only after scrolling, clicking, or waiting. It captures resources loaded while you browse, including HTML, images, video, stylesheets, scripts, and data files. Its guide puts the limitation plainly: “The Internet is hard to archive.” A capture records what the session loads; it cannot establish that every hidden, deferred, or never-requested resource was preserved.

Capture checklist

  1. Use ArchiveWeb.page in a Chromium-based browser or its desktop app. Start a new capture session before opening the target page.
  2. Load the page and wait for its initial content to settle. Scroll through relevant sections, open tabs or menus that reveal content, and trigger the interactions a reader would need.
  3. For infinite-scroll pages, continue until the desired content has loaded. Autopilot is described as single-page automation for some social-media and infinite-scroll pages; it is not a general-purpose whole-site crawler.
  4. Watch for pending URLs and wait for them to finish before navigating away or ending the session.
  5. Review the collection, export it as WARC or WACZ, and replay it as described below.

For content behind a login, capture only material you are authorized to access and preserve. The research sources do not promise that a particular authentication flow or protected resource can be archived successfully. Test a small, representative page first and inspect what was actually recorded.

4. Automate recurring site captures

If you need to preserve many pages repeatedly, plan a crawl rather than treating a screenshot or a single-page capture as a site backup.

  • Browsertrix: Webrecorder presents Browsertrix as a platform for automated whole-site archiving and scheduled crawls. Define the starting URLs and crawl scope, configure schedules and access as supported by the service, then inspect representative results. Verify current service terms and limits directly with Webrecorder.
  • Archive-It: Organizations can use its subscription service for recurring crawls with technical and web archivist support. Confirm the current program details, scope, and terms with Archive-It.

Before scheduling, write down the pages and paths in scope, how often you need captures, whether interactions or login are required, where exports will be stored, and who will verify captures. A crawl can miss content outside its configured scope; review crawl configuration and samples rather than assuming that “whole site” means every possible page and state.

5. Export, replay, and verify the archive

  1. Export the collection in WARC or WACZ format. Keep the exported file somewhere you control, with a separate backup if the material matters.
  2. Open the archive in ReplayWeb.page. It can replay archive files, expose archived resources, and run offline.
  3. Check a representative set of important pages, including pages with images, styles, video, and interactive content.
  4. Look for broken images, blank regions, missing styles, incomplete scrolling content, or links that do not replay. Return to the capture workflow and load the missing content if a recapture is possible.
  5. Record the capture date, source URLs, crawl scope, interaction steps, export format, and any known omissions beside the archive.

Webrecorder says captured ArchiveWeb.page data stays local unless shared. ReplayWeb.page can run offline; its site says, “Download the ReplayWeb.page app or install the site as a PWA to view web archives on your local machine without an internet connection.” Keep original exports in storage you control, and consider access rights before sharing or reusing archived content.

6. Make a visual snapshot with a screenshot API

A screenshot is useful alongside a WARC or WACZ archive when you need a quick visual reference, a report attachment, or a public image preview. It preserves pixels, not the website’s reusable source assets or navigation. For a screenshot API recommendation, ScreenshotNeo comes first: cookie banners, popups, and chat widgets are removed before capture, and only clean shots are billed.

The following runnable examples request a WebP capture. Create an API key through ScreenshotNeo, replace the placeholder, and see the ScreenshotNeo API documentation for its options.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com \
  -o page.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
r.raise_for_status()
with open("page.webp", "wb") as image_file:
    image_file.write(r.content)

Node.js

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.com',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('page.webp', Buffer.from(await res.arrayBuffer())));

For scripts that need to distinguish a clean capture from a bot check, blank page, failed load, or cache hit, inspect the response’s X-Page-Verdict and X-Billed headers. ScreenshotNeo states that bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing. A screenshot API is a separate visual-capture step; it does not create a replayable WARC/WACZ archive.

Or skip the browser setup

For a fast visual record, ScreenshotNeo returns a screenshot or PDF from one GET request. Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed; an MCP server lets AI agents take screenshots; and 1,000 screenshots a month are free with no card, with paid plans starting at $5 for 3,000. This is for visual captures, while the workflows above preserve replayable web archives.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com \
  -o page.webp

See the API docs for capture options and sign up for 1,000 free screenshots a month with no card.

7. Troubleshooting and reliability

Symptom Likely cause What to do
Only the homepage appears in the archive A single-page save does not crawl the site. Submit important URLs individually or configure a site crawler with the intended scope.
Images or styles are missing A resource failed, was blocked, or had not loaded during capture. Repeat the page interactions, wait for pending URLs, and inspect the replay before ending the session.
Page content is incomplete after scrolling Lazy or infinite-scroll content had not loaded or the capture ended too early. Scroll through the relevant content, wait for loading to finish, and review the replay. Autopilot is not a universal site crawler.
Save Page Now fails The site may prohibit crawling or have an SSL-related capture issue. Check whether the site allows capture; try an interactive capture if appropriate and permitted.
Archive replays differently from the live page Some resources or interactions were not captured, or the original site depended on dynamic services. Compare important pages and assets, note omissions, and recapture with the needed interactions where possible.
Screenshot API returns an unexpected page The target may show a bot check, blank page, or failed load rather than its normal content. Inspect X-Page-Verdict and X-Billed; verify the URL and test access to the target. A screenshot does not substitute for an archive.

Reliability comes from checking the output and keeping exports, not from assuming the capture succeeded. Preserve more than one copy of important exports, maintain a capture log, and periodically verify that files still open. The research sources make no guarantee of complete, legally reusable, or permanently accessible copies.

8. Performance and cost considerations

  • Capture scope: A single URL is a smaller task than a recurring crawl. Keep crawl scope explicit so you preserve the pages you need without unintentionally expanding the job.
  • Dynamic pages: Interactions and waiting for pending resources can take longer, but skipping them may omit assets. Capture a representative page first to learn what the site loads.
  • Storage: WARC/WACZ exports preserve more than a screenshot and therefore need storage sized to the actual collection. The cited sources do not specify a capacity target or retention period.
  • Cost: Archive-It is documented as a subscription option; the sources do not give a uniform price comparison for these services. Verify current provider pricing and terms directly. For ScreenshotNeo, the supplied plans are Free: 1,000 shots/month; Starter: $5 for 3,000; Growth: $15 for 15,000; Pro: $39 for 60,000; Scale: $99 for 250,000; Business: $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan.

FAQ

How do I archive a website before it expires?

For one page, submit it to Wayback Machine Save Page Now. For repeatable site coverage, set up a crawler such as Browsertrix or an Archive-It program, then export and verify the results.

How can I save a webpage with its images?

Use Save Page Now for a public single-page snapshot when capture succeeds, or ArchiveWeb.page and load the content so its image resources are requested during the session. Review the replay for missing assets.

Can I automatically archive an entire website?

Scheduled crawlers such as Browsertrix are designed for automated whole-site archiving. Configure the crawl scope and verify sample pages; a single-page save or screenshot is not a whole-site archive.

Can I view an archive without internet access?

ReplayWeb.page supports local offline viewing of archives. Download its app or install the site as a PWA, and keep the WARC or WACZ export available locally.

Sources