ScreenshotNeo

BlogGuides

Why Website Archiving Matters and How to Archive a Site

Learn why websites need preserving and how to save one page, capture a site collection, or keep an organized local copy you can verify.

By the ScreenshotNeo team4 October 20269 min read

Website archiving preserves web content that may change or disappear. The right method depends on what you need: use the Wayback Machine’s Save Page Now for a one-time public capture of one page, a crawl or collection workflow for a defined set of pages, and an owner-controlled export with redundant copies when you need files you can manage yourself. A screenshot records how a page looked; it does not preserve a whole site or its interactive behavior.

Websites document events, organizations, public reactions, government information, and cultural and scholarly material. The Library of Congress preserves selected web content for researchers now and in the future. Preservation is useful for research, records, reference, and recovery, but no capture method guarantees a complete, permanently functional copy.

1. Choose the kind of archive you need

Approach Scope and control Best fit Key limitation
Wayback Machine Save Page Now One specified page in a public archive Quick reference capture of a public page Does not capture an entire site or enroll the URL in future crawls
Site crawl or institutional collection A selected set of URLs or a broader collection; control depends on the operator Organizations and projects preserving a site or topic collection Discovery, access restrictions, and dynamic behavior can limit coverage
Owner-controlled export Files you select and store Personal preservation or site-owner recovery planning You must organize, verify, and maintain copies
Screenshot A visual record of a page at capture time Design review, visual evidence, or a quick reference Does not preserve linked pages, source files, or working interactions

These methods can complement each other. For example, a site owner can export important content, keep redundant local copies, and submit selected public pages to an archive. Decide whether you need a public timestamped reference, a replayable collection, files under your control, a visual record, or several of these.

2. Save one page with the Wayback Machine

  1. Open the Internet Archive’s Wayback Machine.
  2. Use its Save Page Now feature and enter the specific public page URL you want to preserve.
  3. Open the resulting capture and check the page and important linked assets. Do not assume a successful submission means every asset or interaction was saved.
  4. Record the capture URL and date alongside your project notes if you need to cite or revisit it.

Save Page Now is a one-time capture of a specific page. It does not add that URL to future crawls, and it does not save multiple pages, directories, or an entire site. The Internet Archive also says it cannot guarantee that a particular site has been or will be archived, and it no longer offers a service to pack up lost sites as a backup. Treat a public capture as a reference, not as your only preservation copy.

Replay may be incomplete. A crawler can miss pages that are not linked, pages found only through search forms, or content it cannot access. JavaScript-driven behavior, forms, and features that depend on a live server may not work in a replay.

3. Preserve a whole site or collection

A site-wide effort needs a crawler or an export designed to gather multiple pages and resources. Start by defining the collection rather than assuming that a domain name describes the scope.

  1. Set the scope. List the domain, subdomains, paths, documents, or topic pages that matter. The Library of Congress notes that web collections may select a full domain, subdomain, page, or document depending on the intended content.
  2. Choose seed URLs. A crawl typically starts from supplied URLs and discovers content through links. Include important starting points, and identify known pages that may not be linked from them.
  3. Check access and policy. Confirm that you are allowed to collect the content and whether logins or other access controls prevent capture. Do not expect a public crawler to preserve private or restricted areas.
  4. Run the crawl or export. Institutional web archives commonly use crawlers and formats such as WARC (Web ARChive); some older collections use ARC. A format supports storage and replay workflows, but does not guarantee that every site feature will behave as it did live.
  5. Review coverage. Compare collected pages with your scope list. Check representative pages, documents, images, and replay behavior, and note known omissions.
  6. Document the collection. Preserve the crawl date, scope, seed URLs, responsible person or organization, and known limitations with the archive.

For institutional collections, the Internet Archive describes Archive-It as a subscription service for building and preserving born-digital collections. Select an approach based on your collection requirements and operating responsibilities; this guide does not assess service terms or current feature availability.

4. Keep an owner-controlled copy

If you own the site or need files you can manage independently, create an export or local copy in addition to any public archive capture. The Library of Congress’s personal archiving guidance recommends locating relevant sites and services, selecting material with long-term value, exporting it, saving useful metadata, and organizing files with descriptive names and folders.

  1. Inventory what matters. Record the site and services involved, important pages and documents, and any content that is difficult to replace.
  2. Choose a capture method. For a small number of pages, a browser’s Save As command may be enough. For larger sets, use an export or automated process that saves linked files. A screenshot can supplement these copies when the visual appearance matters.
  3. Preserve context. Store the original URL, page or site name, capture or export date, and relevant notes. Keep a list of what was included and excluded.
  4. Organize for retrieval. Use descriptive file names and a consistent folder structure. Keep related pages, assets, metadata, and instructions together.
  5. Make separate copies. Keep at least two copies in different locations where practical. A portable external hard drive can hold one extra copy of files you already exported; the drive does not crawl a website or preserve interactive behavior.
  6. Check the archive. Open representative files to confirm they remain readable. The Library of Congress suggests checking files at least once a year and making new media copies every five years or when necessary.

Redundancy lowers the chance that one lost device or damaged copy destroys your work, but it cannot guarantee against every kind of loss. Keep the copy locations separate enough to reduce shared risks, and make a new copy when the media or files need attention.

5. Capture a visual record with a screenshot

A screenshot is useful when the question is “what did this page look like at this moment?” It is not a substitute for crawling a site or exporting its content. If you need a visual record, choose a screenshot format and page scope that fit the evidence or reference task. For an archive project, store the screenshot with the source URL, capture date, and a note describing what it does and does not represent.

For a local browser-based capture, open the page in your browser and use its screenshot or print-to-PDF controls. Inspect the result for lazy-loaded content, overlays, and content below the fold. Browser output records only what was rendered under that session’s conditions; it does not save the source site as a replayable collection.

Or skip the browser setup

For a quick visual capture, ScreenshotNeo is a website screenshot API and MCP server. Its one-call API returns an image or PDF, and its screenshot features are for visual capture rather than whole-site preservation. See the API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
with open("shot.webp", "wb") as f:
    f.write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo request failed: ${res.status}`);
const image = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', image));

ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Its response headers identify the page verdict and billing status. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. Free includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Yearly billing gives two months free, and every feature is on every plan. This captures a page visually; use an archive or export workflow for site preservation.

Sign up for 1,000 free screenshots a month, with no card required.

6. Common problems and fixes

Problem Likely cause What to do
The archived page is missing The URL was not captured, the crawler could not reach it, or the page was undiscoverable Check the exact URL and capture record. For a collection, add the page as a seed where appropriate and review crawl coverage.
Images, styles, or scripts are missing Assets were hosted separately, blocked, or not collected Inspect asset URLs and collection scope. Record missing assets as a limitation rather than treating the page as complete.
A captured page looks different Dynamic scripts, remote resources, or time-sensitive content did not replay as on the live site Keep a dated visual capture as supplementary evidence and document the replay difference.
Forms or buttons do not work in replay The interaction depends on server state, a live API, or a user session Preserve the resulting documents or data separately if you are authorized to do so; do not assume a static archive can recreate the service.
Pages behind a login are absent The crawler had no authorized session or could not access restricted content Use an owner-controlled export through an authorized account, following applicable rules, and keep it in controlled storage.
A local copy cannot be found later Files lack descriptive names, folder structure, or metadata Add a top-level readme or inventory with URLs, dates, scope, and exclusions; use consistent descriptive file names.
A saved file no longer opens Media degradation, file corruption, or an unsupported format Check files periodically, keep separate copies, and migrate readable files to new media when needed.

7. Improve reliability, effort, and cost

  • Match effort to value. Save a single page when one public reference is enough. Define a bounded collection for a site crawl. Export and maintain owner-controlled copies for recovery needs.
  • Plan for incomplete capture. Dynamic sites, search-only content, unlinked pages, access controls, and changing remote resources all limit what a crawler or replay can preserve. Track known gaps.
  • Preserve metadata. URLs, dates, scope, and notes make an archive easier to interpret after the original site changes.
  • Keep copies in separate places. Redundant copies help with device loss, but require periodic checks and media refreshes.
  • Budget for ongoing work. Public page submissions, crawls, exports, storage, and review each take effort. Institutional collection services may be subscription-based; verify current costs and terms directly. Local storage also has ongoing costs in time and replacement media.
  • Use screenshots for the visual task. ScreenshotNeo bills only clean shots; unsuccessful captures and cache hits cost nothing. Use its usage API and response billing headers to monitor screenshot usage. Screenshot captures do not replace a site archive.

Frequently asked questions

Does saving a page to the Wayback Machine preserve the whole website?

No. Save Page Now captures a specific page once. It does not capture an entire site or schedule future captures.

Can an archived website be fully interactive?

Sometimes parts work, but replay depends on what was collected and whether behavior relies on a live server, scripts, forms, or user state. Do not assume full functionality.

Is a screenshot a website backup?

No. It preserves a visual rendering of a page, not the source files, linked pages, or server-side behavior.

What is WARC?

WARC is a format used in web archiving to store captured web content. A format helps preserve capture data; it does not guarantee that the archived site will work like the live one.

How many local copies should I keep?

The Library of Congress recommends at least two copies stored in different locations where practical, plus periodic checks that files remain readable.

Sources and further reading