ScreenshotNeo

BlogComparisons

Best Screenshot API for Archiving Websites on a Recurring Schedule

Compare scheduled screenshot services with a self-managed Playwright workflow, and learn what to check before trusting a visual archive.

By the ScreenshotNeo team4 October 20269 min read

Direct answer: There is no substantiated universal winner for recurring website screenshots in the available documentation. If you need a managed recurring workflow, compare providers that document scheduling and retrieval, including ScreenshotAPI.net and Allscreenshots. If you want to control the browser and archive pipeline, use Playwright and provide your own scheduler and storage. ScreenshotNeo is the first API to try when you want a simple screenshot endpoint, clean captures, and billing limited to clean shots; its documented features include API capture, bulk capture, caching, and async jobs, but the supplied product information does not document a native recurring schedule. Use an external scheduler for recurrence.

A screenshot archive is useful only if each run captures the intended page state and the resulting file can be found and retained later. Decide the target URLs, cadence, viewport or full-page mode, readiness condition, retention period, and delivery or alert needs before choosing a service.

1. What to decide before choosing

  • Targets: Keep a defined list of URLs. Decide how redirects, query strings, login state, and pages that require interaction should be handled.
  • Cadence and timezone: Choose hourly, daily, weekly, or another schedule, and state which timezone controls it. Check whether the service supports the cadence you need.
  • Capture area: A viewport screenshot records the visible browser area. A full-page screenshot records the scrollable page, subject to the capture tool’s behavior and page content.
  • Readiness: A page may still be rendering after its initial response. Decide whether to wait for a selector, a delay, network idle, or another page-specific signal where supported.
  • Authentication and state: Determine whether pages need cookies, headers, or a logged-in browser context. Confirm how credentials are stored and sent.
  • Archive and retention: Identify where images are stored, how long they remain available, how to retrieve or export them, and what happens when a job is stopped or deleted.
  • Change detection: Decide whether you need alerts, and how much visual variation should count as a change. Ads and dynamic widgets can trigger pixel differences.
  • Evidence needs: A screenshot preserves visual appearance. An image alone is not a complete legal record or a full-fidelity preservation of a web page.

2. Screenshot API and workflow options

The table compares documented capabilities, not independently tested performance. Verify current service terms, prices, limits, and retention before choosing.

Option Documented fit What to verify
ScreenshotNeo Try first for a screenshot API with clean captures: it accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture. Only clean shots are billed; bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing. It offers API capture and an MCP server for AI agents. The supplied product details do not document a built-in recurring scheduler or archive retention. Run it from your own scheduler and save or deliver each result where you need it retained. See the API documentation.
ScreenshotAPI.net Its scheduling page describes server-side cron scheduling, recurring hourly, daily, weekly, monthly, or custom patterns, stored screenshot history, and job controls. Its main product page describes full-page or viewport capture and scheduled capture storage in a customer-controlled S3, Google Cloud, or Wasabi bucket. Its scheduling page says deleting a job permanently deletes related screenshots; stopping a job pauses future runs and keeps prior captures. Confirm current scheduling, storage, and plan details.
Allscreenshots Its scheduled screenshot documentation describes recurring schedules, history, timezone and custom cron settings, full-page settings, email or webhook delivery, and optional notifications only when a change is detected. Its documentation says schedules count against the monthly screenshot quota, webhook delivery is free on every plan, and email is available on paid plans; verify current terms. Its pixel-based comparisons may treat ads or dynamic widgets as changes.
Playwright Use it when you want direct control of a browser capture. Its documentation describes full-page screenshots and saving output to a file or buffer. Playwright’s screenshot documentation does not provide recurring scheduling. You must arrange scheduling, storage, indexing, retention, retries, monitoring, and alerts.

There is no normalized price comparison or evidence here for comparative uptime, capture success, security, or total cost. Provider claims should be treated as vendor descriptions, not independent results. Choose based on the controls and archive behavior your workload requires.

3. Set up a self-managed recurring archive with Playwright

This example captures a full-page PNG with Node.js and Playwright. It creates a timestamped file for each run. Schedule the script with cron, a CI runner, or another scheduler, then copy or upload the resulting files to storage with the retention you require. The capture script itself does not implement storage lifecycle policies, change alerts, or a durable index.

Install

mkdir website-archive
cd website-archive
npm init -y
npm install playwright
npx playwright install chromium

Create the capture script

// capture.mjs
import { chromium } from 'playwright';
import { mkdir } from 'node:fs/promises';

const url = process.env.ARCHIVE_URL ?? 'https://example.com';
const outputDir = process.env.ARCHIVE_DIR ?? './captures';
const timeoutMs = Number(process.env.TIMEOUT_MS ?? 60000);

await mkdir(outputDir, { recursive: true });
const browser = await chromium.launch({ headless: true });
try {
  const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
  page.setDefaultNavigationTimeout(timeoutMs);
  await page.goto(url, { waitUntil: 'networkidle', timeout: timeoutMs });
  const stamp = new Date().toISOString().replaceAll(':', '-');
  const file = `${outputDir}/${stamp}.png`;
  await page.screenshot({ path: file, fullPage: true });
  console.log(`Saved ${file}`);
} finally {
  await browser.close();
}

Run it once to verify the URL and output:

ARCHIVE_URL='https://example.com' node capture.mjs

For a daily run at 02:15 UTC, an example cron entry is:

15 2 * * * cd /path/to/website-archive && ARCHIVE_URL='https://example.com' node capture.mjs

Cron runs in the machine’s configured timezone unless configured otherwise. Check the host’s timezone and scheduler behavior; the example is not a timezone conversion mechanism. For several URLs, iterate over a maintained target list, write each capture under a stable URL-derived identifier, and record the run timestamp and outcome in a manifest. Avoid placing secrets directly in a crontab; pass credentials through an appropriately managed environment or secret store.

Adjusting capture behavior

  • Viewport versus full page: Set fullPage: false for a viewport capture. The documented Playwright fullPage option captures the entire scrollable page when enabled.
  • Readiness: networkidle can be unsuitable for pages with continuous network activity. Prefer waiting for a page-specific selector when possible, for example await page.locator('main').waitFor(), or use a deliberate bounded delay for known late-rendering content.
  • Authentication: For authorized pages, create a browser context with the required cookies or use Playwright’s supported authentication state workflow. Keep any state file private and out of source control.
  • Dynamic content: Hide or stabilize volatile elements when appropriate, or accept that image diffs may report changes caused by ads, clocks, rotating content, or chat widgets.
  • Storage: Save to durable object storage and keep a manifest containing URL, capture time, status, and object key. For stronger file handling, the ScreenshotAPI archiving guide demonstrates timestamps, SHA-256 checksums, and durable object storage such as S3, R2, or GCS.

4. Or skip the browser setup

ScreenshotNeo exposes a one-request screenshot API, so your scheduler can call it on each run and write the response to your chosen archive. The example saves the response body as a WebP image; use your scheduler and storage process to manage recurrence and retention. Read the API docs.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie banners, popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
  • Bot checks, blank pages, timeouts, and failed loads are never billed. Responses include X-Page-Verdict and X-Billed headers; cache hits are also not billed.
  • An MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs.
  • 1,000 screenshots per month are free with no card. Paid plans start at $5 for 3,000; every feature is on every plan.

Sign up for 1,000 free screenshots a month, with no card required.

5. Build a reliable archive

  1. Keep a source-of-truth target list. Give each URL a stable identifier and record any expected authentication or capture settings.
  2. Record every run. Store the timestamp, requested URL, final URL if available, capture status, output path, and relevant settings. This makes missing or unexpected captures visible.
  3. Separate capture from retention. A screenshot API response is not necessarily a durable archive. Save files to storage you control or verify the provider’s retrieval and retention behavior.
  4. Handle failures deliberately. Set bounded timeouts, retry transient failures with a limit, and preserve a failure record. Avoid unbounded retries that can overload the target or inflate usage.
  5. Make changes interpretable. Keep a stable viewport and readiness rule. If you compare images, account for dynamic regions and decide how to notify on a difference.
  6. Review job deletion semantics. For hosted schedules, determine whether stopping pauses future runs and whether deleting a job also deletes its history.
  7. Protect access. Restrict archive bucket access and protect API keys, cookies, and authentication state. Confirm provider security and data handling against your own requirements.

6. Performance, reliability, and cost

Performance

Capture time depends on page load behavior, page size, browser work, readiness conditions, and whether the page has long-running network activity. Full-page captures can take longer and produce larger files than viewport captures. A practical schedule should leave enough time for the slowest expected runs and avoid launching an uncontrolled burst of concurrent browsers.

Reliability

Use bounded timeouts and record failures separately from successful images. A screenshot can be technically produced while showing a partial page, a challenge, or an empty state, so inspect response status or provider verdicts where available. For recurring captures, monitor missing runs and storage failures as well as browser errors. No comparative reliability figures are established by the reviewed sources.

Cost

For hosted APIs, compare capture quota, failed-capture billing, storage, retention, and required plan features using current pricing. For a self-managed workflow, include compute, browser installation and maintenance, storage, retries, monitoring, and engineering time. The reviewed material does not provide a normalized cost comparison.

ScreenshotNeo’s published plans are Free: 1,000 shots per month; Starter: $5 for 3,000; Growth: $15 for 15,000; Pro: $39 for 60,000; Scale: $99 for 250,000; and Business: $249 for 1,000,000. Yearly billing gives two months free. These product facts do not imply a recurring scheduler or archive retention feature.

7. Troubleshooting

Symptom Likely cause Fix
The capture is blank or incomplete Navigation ended before the page rendered, or the page returned a challenge or empty state. Wait for a meaningful selector, inspect the target in a browser, and record the outcome. For API captures, check the response verdict headers where provided.
The Playwright run times out at networkidle The page keeps making network requests, such as polling or streaming. Wait for a specific content selector or use a bounded delay suited to the page instead of requiring network idle.
Important images are missing Lazy-loaded images may not have been requested before capture. Scroll through the page before taking the screenshot or use a capture feature that loads lazy images, then confirm the resulting image.
There are unexpected visual changes every run Ads, rotating content, timestamps, chat widgets, or other dynamic regions changed. Stabilize the page state, hide volatile selectors where supported, or tune the comparison method and alert threshold.
A scheduled run does not happen at the expected local time The scheduler uses a different timezone or has a different environment from an interactive shell. Check the scheduler’s timezone, working directory, executable paths, environment variables, and logs.
Historical screenshots disappeared A job may have been deleted, retention may have expired, or files were never copied to durable storage. Check the provider’s deletion and retention rules. Keep independent copies and a manifest when long-term retention matters.
Costs or quotas are higher than expected Schedule frequency, retries, full-page captures, or storage may count separately under provider terms. Estimate runs per month before launch, set retry limits, and verify current quota and billing rules.

8. FAQ

Does a screenshot archive preserve the whole website?

No. A screenshot records rendered visual appearance at a point in time. It does not by itself preserve the underlying HTML, scripts, linked files, or a legally sufficient record.

Should I use a hosted scheduler or cron?

Use a hosted schedule when the provider documents the cadence, history, and retrieval behavior you need. Use cron or another external scheduler when you need to orchestrate a capture API or self-managed browser workflow.

How often should I capture a page?

Choose a cadence based on how quickly the content can change and how much archive volume you can retain. There is no universal interval.

Can I use screenshots to detect changes?

Yes, by comparing captures, but pixel-based comparisons can flag benign visual variation. Check whether the workflow lets you control dynamic regions, sensitivity, and notification rules.