What Is Website Archiving? A Practical Guide
Website archiving preserves pages and their resources so they can be revisited later. Learn how captures differ from backups, how to make them, and where they fall short.
Website archiving captures and preserves web pages and associated resources so that a version can be revisited after the live site changes or disappears. The right method depends on your goal: finding an old public page, making a one-time capture, preserving an organization’s records, or maintaining a recovery backup. A capture is not automatically complete, interactive, legally authenticated, or permanently available.
For a historical public version, first look for an existing capture in the Wayback Machine. To preserve one page now, use its Save Page Now service. To preserve an entire site or meet formal recordkeeping requirements, plan a scoped, repeatable collection workflow; a one-page capture is not a site crawl or records program.
1. What website archiving preserves
A web page is more than its visible text. A useful capture may need HTML, images, stylesheets, scripts, linked pages, and information about when and how it was collected. Depending on the archive and capture method, it may also preserve relationships among those resources so that links and navigation can be replayed.
“Archived” describes an attempt to retain a version. It does not promise that every asset was collected or that interactive behavior will still work. Pages can rely on live APIs, authentication, external services, scripts, or media streams that an archive cannot reproduce.
2. Archiving, backups, and screenshots solve different problems
| Method | Main purpose | Typical scope | What it does not establish |
|---|---|---|---|
| Public web archive | Find or preserve historical public versions | One URL or a collected set of URLs, depending on service | Complete coverage, restoration readiness, or legal authenticity |
| Operational backup | Restore a site after data loss or failure | Application, database, files, configuration, and other recovery dependencies | Independent recordkeeping or public historical access |
| Records archive | Retain designated records under a documented schedule | Defined content, metadata, procedures, and preservation copies | That a casual screenshot satisfies retention or evidence requirements |
| Screenshot | Show visual appearance at a moment in time | Rendered viewport or full page image | Links, source structure, searchable page relationships, or interactive behavior |
These purposes can overlap, but one copy may not meet all of them. The U.S. National Archives and Records Administration (NARA) distinguishes keeping current content available for restoration from setting aside recordkeeping copies and tracking revisions. A screenshot can document appearance, but NARA does not accept screenshots as substitutes for transfer formats that retain web functionality.
If you need only a visual reference of a live page, a screenshot can be useful. ScreenshotNeo is a website screenshot API and MCP server; it produces images or PDFs, which can support visual documentation but are not a replacement for a web archive that preserves links and replayable resources.
3. Find a historical page or make a new capture
Look up an existing capture
- Open the Wayback Machine and enter the page URL.
- Choose a date with an available capture.
- Inspect the page and follow important links, noting whether their captured dates match your intended moment.
- Check images, styles, and other important resources. A page appearing in the archive index does not mean all of its contents were captured.
Coverage depends on what crawlers could discover and access. Password-protected pages, pages blocked from crawlers, unlinked pages, and content generated through inaccessible scripts may be absent. The Internet Archive notes that simple HTML is generally easiest to archive.
Save one page now
Use Internet Archive’s Save Page Now for a one-time capture of a specific page. It does not schedule future crawls or save a whole directory or website. After submitting the URL, inspect the resulting capture and verify that the parts you care about are present.
4. Plan an organizational website archive
Organizations should select scope, cadence, and control based on the importance of the content, the consequences of losing it, and the applicable retention requirements. NARA recommends risk profiling: higher-risk portions may need more frequent snapshots. There is no single interval appropriate for every website.
- Set the purpose. Decide whether the goal is public history, operational recovery, formal records preservation, or a combination. Set separate requirements when one workflow cannot serve all goals.
- Define scope. Identify domains, site areas, critical pages, associated assets, and the site structure to retain. Record what is excluded and why.
- Assess risk and cadence. Decide how often to capture each area based on how quickly it changes and the impact of losing a version. Include change tracking where appropriate.
- Check access and dependencies. Identify logins, crawler restrictions, scripts that generate links, external services, and media that may not be collectable. Resolve access issues through authorized, documented methods.
- Keep control information. Store capture dates, scope, site maps, procedures, and relevant harvesting information with the preservation copies.
- Review replay and gaps. Test sample pages and important resources after capture. Record missing pages, broken assets, and behaviors that depend on a live service.
- Apply retention rules. For official records, use the rules and approved schedules for the relevant organization and jurisdiction.
NARA recommends accompanying snapshots with a site map. It also advises organizations to document procedures, protect records from unauthorized alteration or destruction, train staff, and maintain approved retention schedules. These are NARA’s recommendations for its records context; they are not universal legal requirements.
5. Choose a preservation format and collection method
For a preservation copy, prioritize a format and workflow that retain component parts, relationships, useful metadata, and integrity information. NARA’s preferred-format table lists WARC 1.0, WARC 1.1, and WACZ for the specified class of permanent federal web records. Its transfer guidance also addresses component parts, links and functionality, data integrity, dynamic content, internally referenced URLs, and harvesting control information. Apply those requirements only when they govern your records; other organizations should check their own standards.
For institutional collections, Internet Archive describes Archive-It as a subscription service. Check its current scope, terms, and suitability directly with the provider. Compare candidate workflows by:
- Whether they capture one page or crawl multiple pages.
- Whether capture is one-time, scheduled, or triggered by change.
- What control you retain over preservation copies and metadata.
- How they handle dynamic pages, assets, and external dependencies.
- How captures can be searched and replayed.
- Whether the workflow supports organizational retention and evidence requirements.
Do not treat any public archive as a complete backup or a formal legal records system without checking its coverage and the applicable requirements.
6. Why an archived website may be incomplete
- Access restrictions: Password protection or crawler restrictions prevent collection unless access is explicitly supported and authorized.
- Robots rules or exclusions: A site may block crawling or its owner may request exclusion.
- Undiscovered pages: Orphan pages and links created only by JavaScript may never be found by a crawler.
- Missing resources: Images, scripts, styles, or fonts may fail to capture, leaving broken or unstyled replay.
- Live dependencies: APIs, login flows, search, embedded services, and other server-side features may not work in a preserved copy.
- Media limitations: Streaming audio and video can be difficult to collect; guidance from the UK Government Web Archive describes its own crawler workflow and technical recommendations.
- Date mismatch: An archive may use a resource from a nearby capture date when the exact-time version is missing. Inspect timestamps for important material.
When completeness matters, make a list of required URLs and assets before capture, then compare it with what was collected. Treat gaps as findings to resolve or document, not as proof that the archive is complete.
7. Make a visual capture with ScreenshotNeo
A screenshot captures rendered appearance rather than a replayable archival collection. It can be useful for a dated visual record, review, or documentation when that is the intended output. ScreenshotNeo accepts one GET request with a URL and returns PNG, JPEG, WebP, or PDF. Its options include full-page capture with lazy images loaded, CSS-selector element capture, viewport and device presets, retina scale, custom CSS and JavaScript, wait conditions, custom headers and cookies, and caching with a chosen TTL. See the ScreenshotNeo API documentation for request options.
The examples below capture a live page. Store the resulting file with your own capture date, target URL, and other recordkeeping metadata if your workflow requires them. This is visual documentation, not WARC/WACZ preservation.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
with open("shot.webp", "wb") as f:
f.write(r.content)
Node.js
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
Use your API key in server-side code and avoid committing it to source control. The same API supports full-page, element, PDF, wait, and rendering options; consult the linked docs for parameter names and supported values.
8. Or skip the browser setup
With ScreenshotNeo, one GET request returns a screenshot or PDF without setting up a browser automation stack:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed, and response headers identify the page verdict and billing status. An MCP server lets AI agents use screenshot tools. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. These are visual captures, not a substitute for a preservation archive. Sign up for 1,000 free screenshots a month, with no card required.
9. Troubleshooting a capture workflow
| Symptom | Likely cause | What to do |
|---|---|---|
| The URL has no historical result | No crawler captured it, the page was excluded, or access was restricted. | Check the exact URL and relevant variants. If you control the site, make an authorized new capture and plan a recurring collection if future history matters. |
| The page loads but images or styles are missing | Resources were inaccessible, blocked, or not captured at the same time. | Inspect resource URLs and timestamps. Record the gap and adjust the collection scope or access configuration where possible. |
| Navigation works poorly in replay | Links may be generated by scripts or depend on live server behavior. | List required destinations explicitly and verify each one. Consider whether the workflow can capture those URLs and their dependencies. |
| Page looks different from the expected date | The selected page and its resources may come from different capture times. | Check timestamps on individual resources and compare another capture date. |
| Video or interactive content is absent | Streaming media and live interactive services can be difficult to harvest or replay. | Document the limitation and preserve permitted source files or supporting records through an approved workflow. |
| A screenshot API returns an unexpected result | The page may be blank, blocked, slow, or dependent on rendering timing. | Check the response and verdict headers, allow enough time, and use documented wait or authentication options when appropriate. Do not mistake a screenshot result for an archival crawl. |
10. Reliability, performance, and cost
Archive reliability starts with scope and repeatability. A one-off capture is useful for a specific moment, but it cannot establish a history of changes. Scheduled snapshots improve coverage over time, while increasing storage and review work. Set cadence according to risk, preserve capture metadata, and periodically review representative pages and known gaps.
Capture speed and completeness depend on site size, crawl depth, scripts, access controls, media, and external dependencies. A fast capture is not evidence of a complete one. Prioritize critical sections, avoid unnecessary duplication where retention rules permit, and retain an inventory of URLs and exclusions.
Costs depend on the collection provider, scope, frequency, storage, and staff effort. Institutional services may be subscription-based; confirm current terms directly. For ScreenshotNeo visual captures, the free plan is 1,000 shots per month with no card. Paid plans are Starter at $5 for 3,000, Growth at $15 for 15,000, Pro at $39 for 60,000, Scale at $99 for 250,000, and Business at $249 for 1,000,000; yearly billing gives two months free. Every feature is on every plan. Only clean shots are billed; bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing. This pricing applies to ScreenshotNeo visual captures, not long-term website archival storage.
11. Legal, rights, and authenticity considerations
Public availability does not automatically grant permission to republish archived material. Check the relevant archive terms and the rights status of content before reuse. For legal, regulatory, or official recordkeeping, follow the applicable retention schedule and evidence process. The Internet Archive says the Wayback Machine was not expressly designed for legal use, though it receives requests for certified records and provides an affidavit process. A casual capture should not be presented as authenticated evidence.
12. Frequently asked questions
Can I archive only one page?
Yes. Save Page Now is intended for a one-time capture of a specific page. Verify its assets and linked content separately; the capture does not schedule future crawls.
Does an archived page stay available forever?
No preservation method should be assumed permanent without an explicit policy and maintained copies. Check the provider’s terms and keep copies under your own retention controls when required.
Is a screenshot a website archive?
It is a visual record of a rendered page. It does not retain the page’s link structure and interactive functionality in the way web archive formats are intended to.
How often should an organization capture its site?
Use a risk assessment and retention needs. More important or rapidly changing areas may call for more frequent capture; there is no universal interval.
Can an archive prove what a page said?
A historical capture can help document prior content, but it is not automatically legally authenticated. Follow the relevant evidentiary process when proof is required.


