Web Archiving for Research
Learn how to plan, capture, document, review, and preserve web sources as research evidence using WARC, WACZ, and repeatable workflows.
Web archiving for research means preserving a documented representation of selected web content at a particular time so another person can understand what was collected, when it was collected, how it was captured, and which parts did not replay. An archive is not a promise that every page, linked service, database, stream, or interaction has been preserved.
The reliable process is: define the research question and scope, choose seed URLs, decide whether to capture once or repeatedly, use an exportable format such as WARC (or a WACZ package where appropriate), review the replay, record limitations, and preserve a stable reference when one exists.
What web archiving involves
A research capture has four parts:
- Selection: identify the pages, domains, embedded resources, and external dependencies that matter to the question.
- Collection: harvest the selected content at a known date and time with a stated tool and configuration.
- Preservation: retain the captured records, metadata, indexes, and checksums in a format that can be managed over time.
- Interpretation: inspect the replay and describe missing resources, broken interactions, access restrictions, and other limits.
The Library of Congress describes its goal as creating “a reproducible copy of how the site appeared at a particular point in time.” Its Web Archiving FAQ explains why a replay should be treated as evidence of a collection process rather than an exact replacement for the live site.
1. Define the research scope before you capture
Write the research question
State what the archive must let you examine later. Examples include changes to a policy page, how an organization described an event, the evolution of a public dataset landing page, or the availability of a digital exhibit.
Choose seed URLs and boundaries
A seed URL is a starting point, not an automatic definition of everything to collect. Record:
- the exact URL, including path and query parameters that affect content;
- the domain and subdomains included;
- which linked pages, downloads, images, scripts, and third-party resources are in scope;
- which areas are excluded, such as account-only pages or unrelated sections;
- permissions, notices, robots policies, and institutional review requirements that apply.
For a single-page question, a focused capture may be sufficient. For a changing site, define a repeat schedule and revisit it when the site or research need changes. Institutional schedules vary; document the schedule you actually use instead of assuming that one interval fits every project.
Decide whether one snapshot is enough
| Research need | Collection approach | Documentation |
|---|---|---|
| Evidence of one published state | One capture at a recorded date and time | Explain why a single point in time answers the question |
| Change over time | Repeated captures of the same seeds | Keep the schedule, missed runs, and URL changes |
| Broad institutional history | Seed list plus bounded crawl | Define domain, depth, file types, and exclusions |
| Interactive or media-heavy material | Capture plus targeted manual review | List interactions and media that did or did not replay |
2. Select an archive format
WARC for preservation records
The Library of Congress Recommended Formats Statement identifies WARC as its preferred web archive format. WARC stores web transactions as records and supports preservation workflows; the Library describes record-at-a-time GZIP compression in its guidance. Its format description identifies WARC as an international standard. Use WARC when you need a preservation-oriented record that can be managed independently of a particular viewer.
WACZ as a package
Webrecorder specifications describe WACZ as a packaging standard for web archives. A WACZ package can contain archive data together with indexes and supporting metadata. Do not treat a WACZ package as the same thing as an individual WARC record: WARC is the record format, while WACZ packages archive material for distribution or replay workflows.
Other formats and legacy material
The Library lists Internet Archive ARC_IA as an acceptable predecessor format. Prefer open, non-proprietary outputs when selecting a new capture tool so the collection is not tied to one vendor or replay application. If you receive another format, record its provenance and conversion history rather than silently replacing the original.
3. Capture the material
Use a crawler, browser-based recorder, institutional service, or a local workflow that can export the format you selected. Keep a run manifest for every capture:
capture_id: project-2026-001
started_at_utc: 2026-10-01T14:30:00Z
collector: [tool and version]
format: WARC
seeds:
- https://example.org/research-page
scope: example.org; linked PDF files; embedded images
excluded: account pages; unrelated subdomains
schedule: one-time
notes: consent banner accepted; video stream not archived
Capture time should include a timezone, preferably UTC. Preserve the original seed list even if redirects or URL changes occur. If a page requires an interaction, record the steps and whether the interaction was captured.
4. Review replay quality
A completed crawl or downloaded file proves that a collection process ran; it does not prove completeness. The Library of Congress cautions that current tools cannot capture all web content, including some multimedia-rich pages, streaming media, deep web content, and databases.
Use a review checklist
- Open every primary seed in the replay interface.
- Check page text, headings, navigation, and timestamps.
- Inspect embedded images, fonts, stylesheets, scripts, audio, and video.
- Test important links and forms without entering credentials or changing live data.
- Check whether redirects, cookie consent, geolocation, or login walls changed the result.
- Compare a sample of replayed pages with the live page only to identify differences; cite the archive for your evidence.
- Record absent resources, error messages, non-replayable interactions, and any manual substitutions.
State which pages and resources you checked. If a chart, stream, search result, or database query is missing, describe that gap in the collection record and in any publication that relies on the archive.
5. Document the archive so others can use it
At minimum, preserve:
| Field | What to record |
|---|---|
| Archive identity | Collecting person or institution and project name |
| Capture date and time | Timestamp with timezone |
| Seeds and scope | URLs, domains, depth, inclusions, and exclusions |
| Tool and format | Collector and version when known; WARC or WACZ details |
| Frequency | One-time or repeat schedule, including missed runs |
| Replay notes | Known failures, absent media, blocked resources, and altered interactions |
| Persistent reference | Stable URI or archive identifier when available |
The Library of Congress guidance on preservable websites recommends stable website URIs so captures can be viewed along a continuous timeline, and recommends open standards and file formats for preservation. Make clear that an archived replay can behave differently from the original live site.
6. Choose a service model
Researchers generally combine one of three approaches:
- Local capture: You control seeds, settings, storage, and review. This can fit a small, reproducible project but leaves you responsible for preservation and access.
- Hosted institutional service: A provider handles harvesting, scheduling, storage, and replay under an agreement. Archive-It is an example of this service category in a U.S. Government Publishing Office publication; verify current features, terms, and pricing before choosing a service.
- Existing public archive: A public collection may already preserve the page. Check its capture date, scope, replay quality, and permission or citation requirements before relying on it.
Compare options by capture scope, treatment of interactive and external resources, export to non-proprietary formats, access controls and notices, replay and review tools, repeat scheduling, and who is responsible for long-term storage. The right choice depends on your research question and institutional obligations.
7. Use screenshots as a visual research record
A screenshot can document visible layout and presentation at a point in time, but it is not a substitute for a web archive when you need underlying resources, links, scripts, metadata, or replay. Use screenshots alongside WARC or WACZ captures when the visual arrangement itself is evidence. Record viewport, device scale, URL, capture time, and any consent or overlay state.
Or skip the browser setup
For a quick visual record of a research page, ScreenshotNeo provides a GET-based screenshot API. It removes cookie and consent banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, with the result identified by X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
See the ScreenshotNeo API documentation for the complete option list. You can request full-page captures, a CSS-selected element, dark mode, device presets or custom viewports, retina scale, PDF paper sizes and page ranges, custom CSS and JavaScript, clicks, waits, blocked resources, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, caching, signed links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, and usage data.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.org/research-page \
-o research-page.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.org/research-page"},
timeout=90,
)
r.raise_for_status()
open("research-page.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://example.org/research-page'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const buffer = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('research-page.webp', buffer));
Use screenshots as a clearly labeled visual supplement. Keep the archive manifest and replay notes for evidentiary work. ScreenshotNeo offers 1,000 shots per month free with no card; paid plans start at $5 for 3,000 shots, and every feature is available on every plan. Create a free ScreenshotNeo account.
Capture quality, performance, and cost
Performance
- Start with a narrow seed list and expand only when the research question requires it.
- Use repeat captures for pages that change instead of recrawling unrelated sections.
- Separate high-value interactive pages from static assets so review time follows research importance.
- Record redirects and failed requests; retries can improve coverage but should remain visible in the log.
Reliability
- Keep the original files, manifests, and checksums together.
- Store at least one independent copy when the material is important.
- Preserve tool versions and configuration so a later researcher can reproduce the process.
- Use persistent URIs or archive identifiers when available.
Cost
Local tools shift costs to compute, storage, maintenance, and staff review. Hosted services may bundle scheduling, replay, and storage under a subscription; verify current terms. For ScreenshotNeo, only clean shots are billed, while bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing. The Free plan includes 1,000 shots monthly without a card; Starter is $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000. Yearly billing gives two months free.
Troubleshooting common archive problems
| Problem | Likely cause | Fix |
|---|---|---|
| The replay is blank | Scripts, deferred rendering, access controls, or a failed document request | Inspect the capture log, record the failure, and recapture with the required interaction or a permitted alternate representation. |
| Images or styles are missing | Third-party resources were out of scope, blocked, or unavailable at capture time | Check resource URLs and scope; document missing dependencies rather than treating the page as complete. |
| Video does not play | Streaming media or player requests were not captured | Record the limitation and preserve a permitted transcript, poster frame, or separate media record if the research requires it. |
| A database result cannot be reproduced | Deep web content, session state, or a query was not included | Save the query parameters and response when permitted, and explain that the archive does not represent the full database. |
| The page differs from the live site | The site changed, dependencies changed, or the replay rewrites links | Cite the capture date and archive identity; compare only to explain the difference. |
| Repeated captures are inconsistent | Changing content, geolocation, personalization, or timing | Fix timezone and viewport where possible, record configuration, and treat each run as its own observation. |
| ScreenshotNeo returns an unexpected result | Page verdict, timeout, bot check, or blocked resource | Read X-Page-Verdict and X-Billed, then adjust waits, headers, cookies, blocking, or the target URL using the API documentation. |
Short FAQ
Is a screenshot a web archive?
No. It preserves a visual view. A web archive can preserve request records, linked resources, and replay metadata. Use both when visual appearance and underlying evidence matter.
Should every project capture an entire domain?
No. Define the smallest scope that answers the research question, then expand it when dependencies or linked evidence require it.
Does WACZ replace WARC?
No. WARC is a web archive record format; WACZ is a package format that can contain archive data and indexes.
Can an archive prove that a website never changed?
No. It documents the captured state and the collection process. Repeated captures can show observed changes, but no capture proves that every state or resource was preserved.
What should I cite?
Cite the archive or persistent URI, capture date and time, archive identity, seed or page URL, and any known replay limitation. Keep the collection manifest with the project.


