ScreenshotNeo

BlogHow-to

How to Schedule Daily Website Screenshots and Keep a PDF Archive

Build a daily Playwright capture workflow that saves dated screenshots and PDFs, then check that your archive is complete and readable.

By the ScreenshotNeo team4 October 202612 min read

Use Playwright to render each target page, then have a daily scheduler invoke your capture script and save dated screenshots and PDFs. The browser does the rendering; the scheduler decides when it runs; your storage and checks determine whether the archive stays useful. This guide uses GitHub Actions for scheduling and includes cURL, Python, and Node.js options for managed captures.

1. Decide what each daily record should show

Before writing the workflow, define the URL list, viewport, browser engine, capture scope, and output formats. Keep these conditions stable across runs so that changes in the archive are more likely to reflect the website rather than a changed capture setup.

Capture scope What it records Use it when
Viewport The visible browser area You want a consistent snapshot of the initial view.
Full page The full scrollable document You need a long page in one image. Lazy-loaded content may need scrolling or a page-specific wait first.
Element A selected part of the page You are tracking one chart, price, or content region and can identify it with a stable selector.

Choose a viewport width and height, browser engine, locale, and device scale deliberately. Responsive sites can show different content at different widths. For a visual history, record these settings with each capture. Playwright supports viewport and full-page screenshots, and its screenshot tooling can target an element. Playwright screenshot documentation

2. Build a Playwright capture script

The script below captures each URL as a PNG and PDF, uses a page-specific readiness selector, writes results into a dated directory, and records a JSON manifest. Replace the example URLs and selectors with pages you are allowed to capture. Selectors are optional per target; when a page has no stable selector, use a meaningful readiness condition or a short bounded delay and inspect the result.

Install the project

mkdir daily-site-archive
cd daily-site-archive
npm init -y
npm install playwright
npx playwright install chromium

Save this as capture.mjs:

import { chromium } from 'playwright';
import { mkdir, writeFile } from 'node:fs/promises';
import path from 'node:path';

const targets = [
  { name: 'homepage', url: 'https://example.com', readySelector: 'h1' },
  { name: 'status-page', url: 'https://status.example.com', readySelector: 'main' },
];
const width = 1440;
const height = 1000;
const date = new Date().toISOString().slice(0, 10);
const outputDir = path.join('archive', date);
await mkdir(outputDir, { recursive: true });
const browser = await chromium.launch({ headless: true });
const manifest = [];

try {
  for (const target of targets) {
    const context = await browser.newContext({ viewport: { width, height } });
    const page = await context.newPage();
    const startedAt = new Date().toISOString();
    const safeName = target.name.replace(/[^a-z0-9_-]+/gi, '-').toLowerCase();
    try {
      const response = await page.goto(target.url, {
        waitUntil: 'domcontentloaded',
        timeout: 45000,
      });
      if (target.readySelector) {
        await page.locator(target.readySelector).waitFor({
          state: 'visible',
          timeout: 20000,
        });
      }
      // Scroll to the bottom and back to trigger common lazy-loaded images.
      await page.evaluate(async () => {
        const step = Math.max(400, window.innerHeight);
        for (let y = 0; y < document.body.scrollHeight; y += step) {
          window.scrollTo(0, y);
          await new Promise(resolve => setTimeout(resolve, 120));
        }
        window.scrollTo(0, 0);
      });
      await page.screenshot({
        path: path.join(outputDir, `${safeName}.png`),
        fullPage: true,
        animations: 'disabled',
      });
      // page.pdf() uses print media by default. Emulate screen if the PDF
      // should retain screen CSS; remove this line for print styling.
      await page.emulateMedia({ media: 'screen' });
      await page.pdf({
        path: path.join(outputDir, `${safeName}.pdf`),
        format: 'A4',
        printBackground: true,
        preferCSSPageSize: true,
      });
      manifest.push({
        name: target.name,
        requestedUrl: target.url,
        finalUrl: page.url(),
        httpStatus: response?.status() ?? null,
        startedAt,
        completedAt: new Date().toISOString(),
        viewport: { width, height },
        browser: 'Chromium via Playwright',
        result: 'success',
      });
    } catch (error) {
      manifest.push({
        name: target.name,
        requestedUrl: target.url,
        startedAt,
        completedAt: new Date().toISOString(),
        viewport: { width, height },
        result: 'failed',
        error: String(error),
      });
      console.error(`Capture failed for ${target.name}:`, error);
    } finally {
      await context.close();
    }
  }
} finally {
  await browser.close();
  await writeFile(
    path.join(outputDir, 'manifest.json'),
    JSON.stringify(manifest, null, 2),
  );
}

if (manifest.some(item => item.result === 'failed')) process.exitCode = 1;

The script isolates each target in its own browser context, records failures without preventing later pages from running, and exits unsuccessfully if any target failed. Add authentication only through secrets or a controlled test account; do not commit credentials or cookies to the repository. A capture is a rendering of the accessible page state, not a copy of the site’s underlying application or every interactive behavior.

Choose readiness conditions carefully

domcontentloaded waits for initial document parsing, not every asynchronous widget, image, or API call. A page-specific selector is usually more informative than sleeping for a fixed number of seconds. If a page renders important data after the selector appears, wait for a more specific condition. Network-idle waits can be unreliable on sites with persistent analytics or streaming requests. Set timeouts so a hung page cannot stall the entire job indefinitely.

For full-page screenshots, the example scrolls the document to encourage common lazy-loaded images to load. Some pages use virtual scrolling and only keep a small region in the DOM; a full-page image may not include all virtualized content. For those sites, capture sections separately or use a page-specific export.

3. Export PDFs that match your purpose

Playwright’s page.pdf() uses print CSS media by default. That is often appropriate for a printable document, but a site’s print stylesheet can hide navigation, rearrange content, or change colors. The script calls page.emulateMedia({ media: 'screen' }) before PDF export to request screen styling. Remove that call when you want print styling, and inspect a sample PDF either way. Playwright PDF API reference

The PDF and PNG are separate renderings and may not match pixel for pixel. CSS page size, page breaks, backgrounds, fonts, animations, and content that loads late can all affect the result. Keep the PNG as a visual snapshot when exact screen appearance matters. Use a PDF page format such as A4 or Letter, or let the page’s CSS define the size with preferCSSPageSize. Large full-page pages can produce very tall images and lengthy PDFs; consider element captures or a bounded set of sections if the archive grows too large.

4. Schedule the capture once a day with GitHub Actions

Create .github/workflows/daily-capture.yml in the repository’s default branch:

name: Daily website archive

on:
  schedule:
    - cron: '17 7 * * *'
  workflow_dispatch:

permissions:
  contents: write

jobs:
  capture:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
        with:
          ref: ${{ github.event.repository.default_branch }}
      - uses: actions/setup-node@v4
        with:
          node-version: 22
      - run: npm ci
      - run: npx playwright install --with-deps chromium
      - run: node capture.mjs
      - name: Commit archive
        run: |
          git config user.name 'github-actions[bot]'
          git config user.email '41898282+github-actions[bot]@users.noreply.github.com'
          git add archive
          git diff --cached --quiet || git commit -m "Daily website archive $(date -u +%F)"
          git push

The cron expression has five fields: minute, hour, day of month, month, and day of week. 17 7 * * * means 07:17 every day in the schedule’s timezone. GitHub Actions uses UTC by default and supports an IANA timezone in the schedule configuration. Scheduled workflows run from the latest commit on the default branch. Place the workflow file there before expecting scheduled runs. GitHub schedule event documentation

A scheduled run is not an exact-time guarantee. GitHub says runs may be delayed during periods of high load, and queued scheduled work can be dropped when load is sufficiently high. Picking a minute away from the beginning of an hour can reduce exposure to common scheduling peaks but does not guarantee punctuality. Public repository workflows can also be disabled after 60 days without repository activity. Check the current GitHub documentation and repository settings if a missed capture would matter.

The workflow above commits files to the repository. The job needs write permission, and branch protection or organization policy may prevent the push. For reviewable run outputs without growing Git history, upload artifacts instead; confirm the current artifact action retention and repository limits before relying on it as long-term storage. A workflow run and a successfully archived file are separate outcomes, so inspect both.

Alternative: run it on your own machine

The capture script is independent of GitHub Actions. You can run node capture.mjs from cron, a system scheduler, or another job runner. For a Linux cron entry that starts at 07:17 UTC daily, use the absolute path to the project and log output:

17 7 * * * cd /absolute/path/daily-site-archive && /usr/bin/node capture.mjs >> capture.log 2>&1

Local scheduling avoids a hosted workflow dependency, but the machine must be powered on and connected at run time. Check the system timezone and daylight-saving behavior if the capture must follow local wall-clock time.

5. Store, name, and verify the archive

A date-based path prevents today’s run from overwriting yesterday’s files. For multiple formats and viewport configurations, use names that encode the page, date, viewport, and format, for example:

archive/2026-10-04/homepage-1440x1000.png
archive/2026-10-04/homepage-1440x1000.pdf
archive/2026-10-04/manifest.json

The manifest should make failures and redirects visible. Useful fields include requested URL, final URL, capture time, browser and version, viewport, output formats, HTTP status, and result. Check for missing dates and failed targets instead of assuming a successful workflow run produced every expected file.

For a small history, a repository can make changes easy to review. Frequent large PNG and PDF commits can make repository history cumbersome. Another option is a local preservation tool such as ArchiveBox, which documents saving multiple representations including HTML, PDF, PNG, and WARC. These formats serve different purposes; a rendered PDF or screenshot cannot recreate every interactive feature. ArchiveBox project

For important records, keep a second copy separately from the capture machine or account, and periodically confirm that files open. An external SSD can be a convenient offline copy, but it is optional; a successful write once does not by itself make an archive durable. Choose retention and backup based on the value and volume of the records rather than assuming a particular storage size.

6. Use a managed API instead of maintaining the browser

If you want the capture script to request an image or PDF without installing and updating a browser, ScreenshotNeo is a website screenshot API and MCP server. Its API returns PNG, JPEG, WebP, or PDF from a GET request. It can capture full pages, and its options include viewport and device presets, waiting for a selector or delay, and custom CSS and JavaScript. See the ScreenshotNeo API documentation for parameters and output handling.

One request from cURL, Python, or Node.js

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', new Uint8Array(await res.arrayBuffer()));

Use the PDF output option documented by the API when the archive needs PDFs. Keep the returned bytes under a dated filename and store the response headers and request parameters in your manifest. The API can also cache captures with a chosen TTL, run asynchronous jobs with signed webhooks, and capture up to 100 URLs per bulk call. Cache hits cost nothing. Its response headers report page verdict and billing status, so your job can distinguish a clean capture from a bot check, blank page, timeout, failed load, or cache hit.

Or skip the browser setup

Use this one-call request and schedule it with the same daily job runner:

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.

Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.

7. Reliability, performance, and cost

  • Reliability: A scheduler can miss or delay runs, and a page can fail independently. Keep per-target error records, fail the job when any page fails, and review missed dates. For high-value records, use a second schedule or manual recovery procedure and maintain an independent archive copy.
  • Performance: Each target launches page navigation, waits, image loading, screenshot rendering, and PDF generation. Reuse a browser process for a batch, as the script does, and avoid waiting for all network traffic to stop on pages that keep connections open. Full-page captures and PDFs use more time and storage than viewport-only images.
  • Cost: Playwright itself does not charge per capture, but your runner time, storage, backup, and maintenance have costs. Repository growth can become operationally inconvenient. ScreenshotNeo offers 1,000 shots per month free, then plans of $5 for 3,000, $15 for 15,000, $39 for 60,000, $99 for 250,000, and $249 for 1,000,000; yearly billing gives two months free. Every feature is on every plan.
  • Consistency: Browser version, fonts, timezone, geolocation, user agent, cookies, and authentication can affect the rendered result. Record relevant settings and avoid changing them silently. An archive records what the browser rendered at capture time; it does not prove the full site state or preserve all functionality.

8. Troubleshooting

Symptom Likely cause Fix
No scheduled run appears Workflow is not on the default branch, schedule syntax is invalid, or repository scheduling is disabled. Check the workflow path and default branch, validate the five cron fields, and inspect Actions settings. Run it once with workflow_dispatch.
The job starts late or is missing Scheduled runs can be delayed or dropped during high load. Schedule away from the start of an hour, inspect workflow history, and add monitoring or recovery for important archives.
Browser executable is missing Playwright package is installed but Chromium was not installed in that environment. Run npx playwright install --with-deps chromium in the workflow or npx playwright install chromium locally.
Navigation times out Slow origin, network issue, blocked automation, or page waiting for an unsuitable load condition. Keep a finite timeout, use domcontentloaded plus a meaningful selector, and record the failure. Investigate site access rules before retrying.
Capture is blank or incomplete Content renders asynchronously, a selector appeared too early, or lazy/virtual scrolling delayed content. Wait for a page-specific state, scroll to load ordinary lazy images, and handle virtualized pages with section-specific captures.
PDF differs from the browser view PDF uses print CSS by default, page breaks or colors differ, or fonts/assets were not ready. Choose print or screen media intentionally, enable backgrounds, wait for required fonts/assets, and keep a PNG as a separate visual record.
Git push is rejected Workflow lacks contents write permission, branch protection blocks direct pushes, or repository policy prevents write access. Check workflow permissions and branch rules; use an artifact or a pull request flow if direct commits are disallowed.
Files overwrite or are absent Output names are not date-scoped, capture failed before writing, or the archive step did not run. Use dated paths, write a manifest even on failure, and make archive presence part of a post-run check.
Storage grows too quickly Full-page images and PDFs are committed every day without retention policy. Set a retention period, store only needed scopes, compress or move old records to a separate archive, and keep an independent backup.

9. Frequently asked questions

Can I archive a page that requires login?

Yes, if you are authorized to access it. Use a dedicated account and inject credentials through protected secrets or a secure session setup. Avoid saving session cookies or passwords in the repository, and confirm the site’s rules for automated access.

Should I keep HTML as well as PDFs?

If you need more than a visual record, consider an archive format such as HTML or WARC in addition to screenshots and PDFs. Each format preserves different information, and none guarantees that dynamic services or interactive behavior will work later.

A scheduled capture can help document a rendered page at a recorded time, but this workflow does not establish chain of custody, completeness, or legal admissibility. For evidentiary requirements, define those controls with qualified counsel or the relevant records policy.

How do I handle a daylight-saving time change?

Choose UTC for a stable interval, or use an IANA timezone if the capture should follow local clock time. Verify the scheduler’s current timezone support and record the actual timestamp of every run.