How to Archive a Website as Screenshots and PDFs Automatically
Choose a local, self-hosted, or hosted workflow to save recurring website screenshots and PDFs, with runnable code and practical checks for reliable archives.
To archive a website automatically as screenshots and PDFs, choose where the capture should run: on a server you control, in a hosted scheduling service, or in a browser that stays available. For a code-driven workflow, run a headless browser on a schedule and save both a full-page image and a PDF. For a managed collection with several preservation formats, ArchiveBox documents scheduled imports and screenshot and PDF outputs. A browser extension can save locally, but its schedule depends on the browser being available.
First decide how often you need a capture, how many pages you need to keep, which formats matter, and whether the capture must run when your computer is off. A screenshot or PDF is a visual record; it does not by itself preserve all page source, interactions, or prove legal context.
1. Choose an automation approach
| Approach | Runs where | Useful when | Key constraint |
|---|---|---|---|
| Self-hosted ArchiveBox | Your server or machine | You want an archive collection with multiple formats and control over storage | You operate the host and its storage |
| Headless-browser script | Your scheduled machine or server | You need custom capture behavior or integration with your own jobs | You own browser setup, retries, and file retention |
| Hosted scheduled capture | Provider infrastructure | You want dated recurring captures without operating a capture host | Cadence, retention, export, and pricing depend on the provider and plan |
| Browser extension | Your local browser | A personal workflow can run while your browser is open | Missed runs may be skipped when the browser is closed |
ArchiveBox describes saving PNG screenshots, PDFs, HTML, and other files, and includes scheduled imports. Its scheduled archiving documentation shows recurring imports. This is a reasonable self-hosted option when retaining a broader archive matters alongside visual output.
Hosted examples include PDF Archive, which describes scheduled dated PDFs and says scheduling is on a paid tier, and Site2PDF, which describes recurring PDF or image captures and dated versions. PeekShot describes recurring screenshots and history, while Snapshot Archive describes schedules, screenshot history, and PDF exports. These are vendor descriptions, not independent comparisons of capture fidelity or reliability. Check current cadence, retention, export, and account limits before choosing.
The Auto Page Capture extension describes scheduled local saves in MHTML, HTML, PDF, and image formats. Its documentation notes scheduled jobs run only while the browser and extension are available; a run missed while the browser is fully closed is skipped.
2. Make a repeatable screenshot and PDF with Python
This example uses Playwright to open each URL, wait for the page to settle, and save a full-page PNG and PDF. Run it once manually before putting it on a schedule. The script creates date-stamped output folders and records failures without discarding successful captures.
- Install Python 3 and Playwright:
python -m pip install playwright python -m playwright install chromium - Save this as
capture.py:
from datetime import datetime, timezone
from pathlib import Path
from urllib.parse import urlparse
import asyncio
import json
from playwright.async_api import async_playwright
URLS = [
"https://example.com/",
]
OUTPUT = Path("archive")
async def main():
stamp = datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%SZ")
async with async_playwright() as p:
browser = await p.chromium.launch()
for url in URLS:
host = urlparse(url).netloc.replace(":", "_") or "page"
folder = OUTPUT / stamp / host
folder.mkdir(parents=True, exist_ok=True)
page = await browser.new_page(
viewport={"width": 1440, "height": 1000},
device_scale_factor=1,
)
try:
response = await page.goto(
url, wait_until="networkidle", timeout=60000
)
# Some pages keep network connections open; in that case,
# use wait_until="domcontentloaded" and a short page.wait_for_timeout().
await page.screenshot(
path=str(folder / "page.png"),
full_page=True,
animations="disabled",
)
await page.pdf(
path=str(folder / "page.pdf"),
format="A4",
print_background=True,
prefer_css_page_size=True,
)
metadata = {
"url": url,
"captured_at": stamp,
"http_status": response.status if response else None,
"title": await page.title(),
}
(folder / "metadata.json").write_text(
json.dumps(metadata, indent=2), encoding="utf-8"
)
print(f"Saved {url} to {folder}")
except Exception as exc:
(folder / "error.txt").write_text(str(exc), encoding="utf-8")
print(f"Failed {url}: {exc}")
finally:
await page.close()
await browser.close()
if __name__ == "__main__":
asyncio.run(main())
Run it with python capture.py. The output is archive/<UTC timestamp>/<host>/, containing a PNG, PDF, and metadata when successful. The HTTP status and title help identify redirects or unexpected pages; a successful navigation does not guarantee that the page content is the expected content.
Schedule it
On a Unix-like machine with cron, use the absolute path to the script and an absolute working directory. This example runs daily at 02:15 UTC; adjust the time and frequency to your needs:
15 2 * * * cd /absolute/path/to/project && /usr/bin/python3 capture.py >> capture.log 2>&1
Keep the machine powered on and the schedule service running. For captures that must happen independently of a personal computer, run the job on an always-available server or use a hosted scheduler. Prevent overlapping jobs if one run can take longer than the interval. Monitor the log and disk usage, and copy completed files to storage that matches your retention needs.
Capture settings to tune
- Wait condition:
networkidlewaits for network activity to settle, which can time out on pages with persistent connections or analytics.domcontentloadedreturns sooner but may capture before client-rendered content appears. When necessary, wait for a page-specific selector or a measured delay. - Viewport and scale: Set viewport dimensions and device scale factor to make repeated images comparable. Responsive pages render differently at different widths.
- Full page versus viewport:
full_page=Truecaptures the full document as a tall image. For a repeatable viewport image, usefull_page=False. Extremely long pages can create large files or exceed browser limits. - PDF paper: Choose a paper format, orientation, margins, and background printing to suit the document. With
prefer_css_page_size=True, page CSS can determine the printed size. Websites may define print-specific styles that differ from their screen layout. - Content readiness: If a known element signals the page is ready, wait for it explicitly before saving. Lazy-loaded images may require scrolling through the page before capture; test the result on the target site.
- Authentication and state: For pages behind login, use a dedicated browser context and securely managed credentials or storage state. Do not place secrets in source code or logs. Check the site’s access rules before automating authenticated pages.
3. Use ArchiveBox for a self-hosted collection
ArchiveBox supports URL imports and documents screenshot and PDF outputs, among other archive files. Its scheduled import feature can be useful when you want recurring collection imports managed alongside the archive. Install and run it using the project’s current installation and usage documentation; deployment commands and configuration can differ by release and installation method.
For a small, stable list of pages, keep the URL list under version control and schedule imports. Before relying on the collection, perform a manual import and inspect the resulting snapshot directory to confirm the screenshot and PDF extractors ran. Decide how the archive directory is backed up, who can access it, and how old captures are pruned. The schedule only helps while its host and archive service are available.
4. Configure a hosted or browser-based schedule
For a hosted service, create a job for each URL or supported page set, choose the cadence, then confirm the first capture and its export path. Review these settings before depending on it:
- Whether it captures a viewport, a full page, a PDF, or more than one format.
- Whether schedules are available on the plan you intend to use, and the current maximum number of jobs or URLs.
- How many historical versions are retained, how to download them, and what happens if a plan ends.
- Whether schedules run in the provider’s cloud independently of your computer.
- How the service reports failed captures, access-denied pages, and changed content.
For an extension, configure the URL, interval, output format, and local download behavior, then leave the browser available at the scheduled time. Treat this as a personal convenience workflow rather than an always-on server schedule when missed runs matter.
5. Keep the archive useful and verifiable
Use a consistent directory pattern such as domain/path/YYYY-MM-DD/, or a timestamp-first pattern as in the script. Keep the source URL, capture time and time zone, output format, and run status next to each capture. Preserve a log of failures and reruns so a missing file is distinguishable from a page that did not change.
Set retention deliberately: screenshots and especially full-page PDFs accumulate. Estimate storage from a sample of captures, then multiply by pages, formats, and retained runs. Compress or move older files only after confirming they remain readable. Keep backups separate from the machine doing scheduled capture if the archive matters.
A hash can help detect whether a stored file changed after capture, but it does not establish that the screenshot faithfully represents a page at a particular time or prove the circumstances of capture. If records have compliance or legal significance, define the required capture process with the responsible reviewer; a timestamp or certificate alone does not establish legal admissibility.
6. Troubleshoot common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| Scheduled run never starts | The machine is off, the scheduler is stopped, the browser is closed for an extension workflow, or the job has the wrong working directory | Check scheduler status and logs; use absolute executable and script paths; keep the host available. For extensions, keep the browser available during scheduled runs. |
| Navigation times out | The site is slow or keeps network requests open | Check reachability and response behavior. Try domcontentloaded, then wait for a page-specific element or a bounded delay. Avoid treating timeout as a valid capture. |
| Image is blank or incomplete | Capture began before client rendering, fonts, or images finished; content is lazy-loaded | Wait for a meaningful selector, scroll to trigger lazy loading, and inspect the saved output. Use a site-specific readiness condition where possible. |
| PDF layout differs from screenshot | Print CSS, paper size, margins, or background settings change the rendered page | Set paper and print-background options intentionally, and review the PDF separately from the screen capture. |
| Login or consent page was archived | The capture context lacks session state, or a banner blocked the intended content | Configure authorized session state, handle consent as appropriate, and verify the captured title and page content. Do not assume a successful HTTP response is the target page. |
| Files are missing or overwritten | Fixed output names reused a directory, concurrent jobs collided, or writes failed | Use unique timestamped paths, prevent overlapping runs, check permissions and available disk space, and log exceptions. |
| Archive storage grows too quickly | Many URLs, formats, or high-frequency runs are retained indefinitely | Reduce cadence or retained history, estimate sizes from real captures, and apply a documented retention and backup policy. |
7. Performance, reliability, and cost
Capture time depends on the page, network, browser startup, readiness condition, and number of URLs; there is no single useful duration for all sites. Reuse a browser process for a batch, but isolate pages or contexts where cookies and state should not cross between URLs. Limit concurrency to avoid overloading your host or the target website. A queue with bounded retries is more robust than launching an unbounded number of simultaneous browser sessions.
For reliability, record one result per URL, distinguish failed navigation from a successful capture, and alert on repeated failures. Retry transient errors with a limit and delay; repeated rapid requests can burden a website and may trigger access controls. Keep an eye on browser and operating-system updates because automated browser environments change. For cost, self-hosting uses your compute and storage; hosted services may charge by plan, volume, or retention. Compare the total number of URLs multiplied by capture frequency and formats, plus historical retention, before selecting a workflow.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. A GET request captures a URL as PNG, JPEG, WebP, or PDF. It can schedule captures through async jobs and signed webhooks, batch up to 100 URLs per call, and lets you choose a cache TTL. See the ScreenshotNeo API documentation for the supported parameters and formats.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Cookie banners are accepted like a visitor and removed before the shot, along with known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. An MCP server lets AI agents use screenshot, page-info, and PDF-capture tools. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots.
Sign up free for 1,000 screenshots a month, with no card required.
FAQ
Can I archive a whole website rather than one URL?
Yes, if your tool can import a URL list, sitemap, or crawl scope. Set boundaries and rate limits first, then check which pages actually completed; a site-wide job can be much larger than a list of selected pages.
Should I save HTML as well as screenshots and PDFs?
If you need a more useful historical record than visual appearance alone, retain an HTML or web archive format when your chosen tool supports it. A screenshot and PDF preserve a rendered view, not the complete original site behavior.
Will a recurring capture show exactly the same page each time?
No. Page content, personalization, experiments, ads, location, and responsive layout can change. Record the capture settings and session context, and compare like-for-like viewports and states.
Does a timestamped PDF prove what a website displayed?
A timestamp is useful metadata, but by itself it does not independently establish capture accuracy or legal sufficiency. Follow the record-handling process required for your use case.


