ScreenshotNeo

BlogGuides

Web Archiving with Screenshot APIs: Screenshots, WARC, and Replay

A screenshot records how a page looked; a web archive preserves content for later replay. Learn when to use each and how to capture and document both.

By the ScreenshotNeo team4 October 202610 min read

A screenshot API can save a visual record of a webpage as it appeared when captured. It does not, by itself, create a replayable web archive. If you need to preserve content and its resources for future access, capture an archive in a preservation format such as WARC as well as, or instead of, the screenshot. The Library of Congress recommends WARC with record-at-a-time GZIP compression and lists ARC_IA and WACZ as acceptable alternatives. Library of Congress Recommended Formats for Web Archives.

Use a screenshot for visual reference, review, or evidence of a rendered state. Use a web archive when the goal includes collecting page content and resources for later replay. For many workflows, keep both: the archive for preservation and the screenshot for a quick visual reference.

1. Decide what you need to preserve

Need Use What it gives you
Show how a page looked at a particular moment Screenshot A raster image such as PNG, JPEG, or WebP. It is easy to inspect but does not contain the page’s underlying HTML, scripts, styles, or linked resources.
Collect page content and resources for future replay Web archive A package of captured web resources in an archive format such as WARC. Replay depends on the available archive tools and on which resources were captured.
Document both the appearance and captured content Screenshot plus web archive A convenient visual reference alongside the preservation capture. Record the relationship between the two and their capture context.

The Library of Congress recommends non-proprietary capture output that conforms to standard formats and requirements. Its guidance prefers WARC with record-at-a-time GZIP compression, identifies ARC_IA as a precursor to WARC, and lists WACZ as an acceptable format. CDX is a component file used with WARC content; these formats are not image formats and are not interchangeable with a screenshot. See the Library of Congress format guidance.

2. Capture a rendered screenshot with Playwright

For a reproducible visual capture, run a browser automation script against a page and save the screenshot. Playwright supports viewport screenshots, full-page screenshots, individual-element screenshots, and image buffers. This captures a rendered visual; it does not produce a WARC or another replayable web archive. Playwright screenshot documentation.

Install and run

mkdir page-capture
cd page-capture
npm init -y
npm install playwright
npx playwright install chromium

Save this as capture.mjs, then run node capture.mjs https://example.com. It writes a full-page PNG and prints the capture timestamp and requested URL. Replace the example URL with a page you are permitted to capture.

import { chromium } from 'playwright';

const url = process.argv[2];
if (!url) {
  console.error('Usage: node capture.mjs https://example.com');
  process.exit(2);
}

const capturedAt = new Date().toISOString();
const browser = await chromium.launch({ headless: true });
try {
  const page = await browser.newPage({ viewport: { width: 1440, height: 1000 } });
  const response = await page.goto(url, { waitUntil: 'networkidle', timeout: 60000 });
  await page.screenshot({ path: 'page.png', fullPage: true });
  console.log(JSON.stringify({
    requestedUrl: url,
    finalUrl: page.url(),
    capturedAt,
    httpStatus: response?.status() ?? null,
    screenshot: 'page.png'
  }, null, 2));
} finally {
  await browser.close();
}

networkidle can be unsuitable for pages that keep network connections open or continuously poll. If navigation times out, choose a page-specific readiness condition, such as waiting for a known selector, or use a bounded delay after the page’s main content appears. No single wait condition proves every asynchronous resource has finished loading.

Choose the capture scope

  • Viewport: omit fullPage to capture only the visible browser viewport.
  • Full page: set fullPage: true to capture the full scrollable page as a tall image. Very long pages can produce large images or exceed browser and memory limits.
  • One element: wait for a locator and call locator.screenshot({ path: 'element.png' }). This is useful for a chart, article, or component, but omits context outside the element.
  • Buffer: omit path and use the returned buffer when the next step uploads or processes the image. Store the capture context separately.

For example, to capture a specific element after navigation, replace the screenshot line with:

const article = page.locator('main article');
await article.waitFor({ state: 'visible', timeout: 15000 });
await article.screenshot({ path: 'article.png' });

Selectors are site-specific. A missing or duplicated selector can fail or target the wrong content, so inspect the page structure and use a selector that uniquely identifies the intended region.

3. Save a replayable web archive when preservation matters

A screenshot-only workflow is incomplete when the requirement is future replay or preservation. Use an archiving tool that can capture web resources and produce a standards-based output, then retain its archive output and any associated metadata. The research sources establish the relevant formats and limitations, but do not prescribe one current command-line archiver or a universal capture command; tool-specific installation and commands should follow that tool’s official documentation.

  1. Define scope. Record the seed URL, any included or excluded paths or domains, and whether linked resources are in scope.
  2. Capture the archive. Configure the chosen web archiving tool to produce WARC, preferably with record-at-a-time GZIP compression, or an acceptable format such as WACZ or ARC_IA when appropriate to your workflow.
  3. Preserve the output. Keep the archive files and relevant capture metadata together in storage with access controls and backup appropriate to the value of the material.
  4. Check replay. Open the capture using a compatible replay tool and note missing or altered resources. A successful crawl or archive file does not prove that every page feature will replay.
  5. Capture a visual reference. Optionally produce a screenshot with Playwright and associate it with the archive record.

For public pages, an existing Wayback Machine capture may already be useful. The Internet Archive documents an Availability API for checking whether a URL has an accessible snapshot and a CDX API for more complex capture queries, filtering, and analysis. These lookup capabilities answer questions about existing public captures; they are distinct from creating your own preservation archive. Consult the Internet Archive Wayback APIs documentation for current request details.

4. Record the capture context

A screenshot or archive without context can be difficult to interpret later. The Library of Congress recommends displaying the archiving institution, the date and time of capture, and statements that distinguish archived functionality from the live site. Record at least:

  • The operator or institution responsible for the capture.
  • Capture date and time, including timezone or UTC offset.
  • The requested URL and final URL after redirects.
  • The capture method and output format, such as a Playwright full-page PNG or a WARC archive.
  • Relevant configuration: viewport, browser, wait condition, scope, and any filters.
  • Known gaps, access restrictions, or functionality that differs from the live site.

Keep a small metadata file beside a screenshot or archive. The example Playwright script prints some of these fields, but you should retain them in the same storage and retention workflow as the capture.

5. Understand what captures can miss

Neither a screenshot nor a web archive guarantees a complete record of a site. The Library of Congress specifically identifies multimedia-rich content, streaming media, deep web content, and databases as examples that current tools may not preserve through web capture. It states: “Tools currently available cannot capture all web content, so certain types of web content may not be preservable through web capture at this time.” Library of Congress web archiving FAQ.

  • A screenshot records pixels at capture time; it does not preserve text as searchable page structure or make controls work later.
  • A full-page screenshot can still omit content that only appears after interaction, authentication, scrolling-triggered loading, or a separate request.
  • An archive may miss resources hosted on other domains, content behind access controls, or data that changes dynamically.
  • Streaming media and database-driven content may not be captured or replayed as expected.
  • Replay can differ from the live page because external resources, scripts, services, or browser behavior change.

Describe these limits in the capture record. Do not label a screenshot as a complete website archive.

6. Options that affect screenshot usefulness

Choice When it helps Trade-off
Viewport or full page Viewport preserves a specific screen view; full page provides a long-page overview. Full-page captures can be tall, slow, and memory-intensive.
Element capture Focuses on one region with less irrelevant content. It may lose surrounding layout and page context.
Image format and scale Choose the output and pixel dimensions needed by your review or storage workflow. Higher pixel dimensions and less compressed formats generally require more storage and processing; keep the original when evidentiary needs call for it.
Wait strategy Wait for the content relevant to the record to appear. Long fixed waits waste time; network-idle waits can stall on pages with continuing requests. Page-specific readiness is usually more predictable.
Browser state Set viewport and, where needed, an explicitly documented locale or authenticated state. Cookies and login state can change the captured content and may contain sensitive information. Protect credentials and avoid retaining secrets in metadata.

Playwright’s official documentation describes screenshot behavior and options. The screenshot capabilities of a browser library should not be confused with the behavior of a hosted screenshot API; evaluate a hosted service using its own documentation.

7. Performance, reliability, and cost

Browser rendering consumes time and memory, especially on large pages or when capturing full-page images. Keep captures bounded with navigation and selector timeouts, avoid unnecessary browser launches in high-volume jobs, and record failures rather than silently treating them as successful captures. For repeated captures, browser reuse can reduce startup overhead, but isolate page state so cookies or storage from one target do not leak into another.

Reliability depends on the page as well as the capture process: redirects, transient network errors, bot challenges, authentication, dynamic content, and external resources can change the result. Save the HTTP status and final URL where available, retain a timestamp, and retry only transient failures with a bounded policy. For preservation workflows, verify the archive and replay a sample; for screenshots, verify that the output exists and is non-empty.

Costs include browser compute, storage, and any hosted capture or archive service you choose. Full-page images and repeated captures increase storage needs. This research provides no benchmark or general price comparison, so estimate using your page sizes, capture frequency, retention period, and selected provider’s current pricing. A screenshot API may simplify rendering but does not replace the archive format and preservation workflow.

8. Troubleshooting

Symptom Likely cause Fix
Navigation times out The page is slow, or continuous requests prevent the selected wait condition from completing. Use a bounded timeout and wait for a page-specific selector or content milestone. Check whether the page actually loaded before capturing.
Screenshot is blank or incomplete The page had not rendered relevant content, a navigation failed, or content loads only after interaction. Check the response status and final URL, wait for the target content, and handle required interactions explicitly. Record remaining gaps.
Full-page screenshot is unexpectedly huge or fails The document is very long or the browser runs out of available memory. Capture a viewport or selected elements, or divide the visual review into smaller regions. Keep the chosen scope in metadata.
Element screenshot cannot find its target The selector is wrong, ambiguous, or the element has not appeared. Inspect the page, use a stable unique selector, and wait for it to become visible before capturing.
Archived page does not replay correctly Resources were unavailable, outside the capture scope, access-controlled, or unsupported by the capture/replay workflow. Inspect the archive and its scope, capture relevant permitted resources, and document missing functionality. Some content types cannot be preserved reliably.
Wayback lookup returns no useful snapshot No accessible capture is available for that URL or the query does not match the desired record. Check the exact URL and use the documented Availability API for a closest accessible snapshot or CDX for more complex queries.

9. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server for developers. A GET request returns a PNG, JPEG, WebP, or PDF. It is useful for a rendered visual capture; it does not turn a screenshot into a replayable WARC archive. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await require('node:fs/promises').writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

Cookie and consent banners are accepted like a visitor and more than 60 known consent platforms, newsletter popups, and chat widgets are removed before the shot; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers say which outcome occurred. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000 shots. All features are available on every plan. Keep a separate web archive capture when preservation or replay is required.

Sign up for 1,000 free screenshots a month, no card required.

10. Frequently asked questions

Can a screenshot be used as proof of what a page showed?

It can serve as a visual record, but preserve its timestamp, URL, capture method, and operator context. A screenshot alone does not establish that the page was complete or independently authenticated.

Should I keep the screenshot and archive together?

If you create both, associate them with a shared capture identifier or metadata record so readers can tell which visual reference corresponds to which archive.

Does a public Wayback snapshot replace my own archive?

It may help locate a public historical capture, but it does not automatically satisfy your scope, retention, documentation, or preservation requirements. Check the specific snapshot and its replay limitations.

Which format should I choose?

For web preservation, the Library of Congress prefers WARC with record-at-a-time GZIP compression and lists ARC_IA and WACZ as acceptable alternatives. Choose based on your archive workflow and supported replay tools; use an image format only for the visual record.