How to Schedule Website Screenshots and Generate a Monthly PDF Report
Build a repeatable monthly screenshot archive with Playwright and GitHub Actions, then assemble, validate, and store a shareable PDF report.
Use a browser automation script to capture each URL, schedule that script to run monthly, and assemble the resulting images into a dated PDF. The screenshot archive and the PDF are separate outputs: the capture step records each page, while a report step arranges those captures with URLs and timestamps. This guide uses Playwright, Node.js, and GitHub Actions.
1. Decide what each monthly capture should contain
Before writing code, record the pages to capture and the conditions that make the result comparable from month to month. Keep the browser, viewport, capture scope, and page state consistent.
- Viewport: the visible browser area, useful for monitoring a page’s first screen.
- Full page: the full scrollable document, useful for broad visual records. Lazy-loaded content may require scrolling or another page-specific preparation step first.
- Element: a selected component such as a pricing table, useful when the whole page is unnecessary.
Playwright supports viewport, full-page, and locator screenshots. See the Playwright screenshot guide. Choose a stable viewport and browser. Decide whether to dismiss consent banners, sign in, set a locale, or wait for a selector before the shot. Only automate account access on systems you are authorized to use, and protect credentials as secrets.
2. Create a reproducible Playwright capture script
The following small project captures a list of URLs into a directory named for the reporting month. It also writes a manifest with each URL, timestamp, and screenshot filename, which the PDF step uses for captions and completeness checks.
npm init -y
npm install --save-exact playwright pdf-lib
npx playwright install chromium
Create capture.mjs:
import { chromium } from 'playwright';
import { mkdir, writeFile } from 'node:fs/promises';
const urls = (process.env.TARGET_URLS ?? 'https://example.com')
.split(',').map(value => value.trim()).filter(Boolean);
const month = process.env.REPORT_MONTH ?? new Date().toISOString().slice(0, 7);
const outDir = `artifacts/${month}`;
const width = Number(process.env.VIEWPORT_WIDTH ?? 1440);
const height = Number(process.env.VIEWPORT_HEIGHT ?? 1000);
const scope = process.env.CAPTURE_SCOPE ?? 'full'; // full or viewport
if (!/^\d{4}-\d{2}$/.test(month)) throw new Error('REPORT_MONTH must be YYYY-MM');
if (!urls.length) throw new Error('Set TARGET_URLS to one or more comma-separated URLs');
if (!['full', 'viewport'].includes(scope)) throw new Error('CAPTURE_SCOPE must be full or viewport');
await mkdir(outDir, { recursive: true });
const browser = await chromium.launch({ headless: true });
const results = [];
try {
const context = await browser.newContext({ viewport: { width, height } });
for (let i = 0; i < urls.length; i++) {
const url = urls[i];
const page = await context.newPage();
const capturedAt = new Date().toISOString();
const filename = `${String(i + 1).padStart(2, '0')}.png`;
try {
const response = await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 60000 });
if (!response) throw new Error('Navigation returned no main-document response');
if (!response.ok()) throw new Error(`HTTP ${response.status()} ${response.statusText()}`);
// Optional per-page readiness check, for example: await page.locator('main').waitFor();
await page.screenshot({ path: `${outDir}/${filename}`, fullPage: scope === 'full', animations: 'disabled' });
results.push({ url, capturedAt, filename, status: 'ok' });
} catch (error) {
results.push({ url, capturedAt, filename: null, status: 'error', error: String(error) });
} finally {
await page.close();
}
}
await writeFile(`${outDir}/manifest.json`, JSON.stringify({ month, viewport: { width, height }, scope, results }, null, 2));
} finally {
await browser.close();
}
if (results.some(result => result.status !== 'ok')) process.exitCode = 1;
Set TARGET_URLS to comma-separated addresses, for example https://example.com,https://example.org. This baseline waits for DOM content, not every network request; that avoids waiting indefinitely on analytics or long-lived connections. For pages with delayed content, add a meaningful selector wait after navigation (for example, await page.locator('main article').waitFor()) or a page-specific readiness condition. A single fixed sleep is less reliable because load times vary.
Change fullPage to false for viewport capture. For an element, use a locator screenshot, such as await page.locator('.pricing-table').screenshot({ path: ... }); validate the selector because a missing or ambiguous element should be treated as a failed capture. If the site lazy-loads images below the fold, scroll through the page before taking a full-page image and wait for the relevant images to finish loading.
3. Assemble the month’s images into a PDF
Create make-report.mjs. This produces one landscape PDF page per successful capture, scales each image to fit, and prints its URL and capture time. It fails if any expected URL lacks a successful image, so a partial archive does not silently look complete.
import { PDFDocument, StandardFonts, rgb } from 'pdf-lib';
import { readFile, writeFile } from 'node:fs/promises';
const month = process.env.REPORT_MONTH ?? new Date().toISOString().slice(0, 7);
const dir = `artifacts/${month}`;
const manifest = JSON.parse(await readFile(`${dir}/manifest.json`, 'utf8'));
const failed = manifest.results.filter(item => item.status !== 'ok');
if (failed.length) throw new Error(`Cannot publish incomplete report: ${failed.map(x => x.url).join(', ')}`);
const pdf = await PDFDocument.create();
const font = await pdf.embedFont(StandardFonts.Helvetica);
for (const item of manifest.results) {
const bytes = await readFile(`${dir}/${item.filename}`);
const image = await pdf.embedPng(bytes);
const page = pdf.addPage([792, 612]); // landscape letter, points
page.drawText(item.url, { x: 30, y: 578, size: 10, font, maxWidth: 732 });
page.drawText(`Captured: ${item.capturedAt} | Report month: ${month}`, { x: 30, y: 561, size: 9, font, color: rgb(0.35, 0.35, 0.35) });
const box = { x: 30, y: 24, width: 732, height: 520 };
const scale = Math.min(box.width / image.width, box.height / image.height);
const width = image.width * scale;
const height = image.height * scale;
page.drawImage(image, { x: box.x + (box.width - width) / 2, y: box.y + (box.height - height) / 2, width, height });
}
await writeFile(`${dir}/monthly-report.pdf`, await pdf.save());
console.log(`Wrote ${dir}/monthly-report.pdf`);
Install and pin the package versions in the lockfile, then use npm ci in automation. The page size above is landscape letter; change the PDF dimensions for A4 or portrait. Very tall full-page captures are scaled down to fit one page and can become hard to read. For those, split the image across multiple PDF pages or use viewport captures; check the final PDF at its intended viewing size.
4. Run it on a monthly schedule with GitHub Actions
Create .github/workflows/monthly-report.yml. The example schedules the job at 09:00 UTC on the first day of each month, and also allows a manual run with a selected month. GitHub’s workflow syntax documents scheduled triggers and timezone behavior; review the current schedule syntax for your repository’s needs.
name: Monthly website screenshot report
on:
schedule:
- cron: '0 9 1 * *'
workflow_dispatch:
inputs:
report_month:
description: Reporting month in YYYY-MM (blank means current UTC month)
required: false
type: string
permissions:
contents: read
jobs:
capture-and-report:
runs-on: ubuntu-latest
timeout-minutes: 30
env:
TARGET_URLS: https://example.com,https://example.org
REPORT_MONTH: ${{ inputs.report_month || '' }}
VIEWPORT_WIDTH: '1440'
VIEWPORT_HEIGHT: '1000'
CAPTURE_SCOPE: full
steps:
- uses: actions/checkout@v6
- uses: actions/setup-node@v6
with:
node-version: '22'
cache: npm
- run: npm ci
- run: npx playwright install --with-deps chromium
- name: Set default report month
run: echo "REPORT_MONTH=${REPORT_MONTH:-$(date -u +%Y-%m)}" >> "$GITHUB_ENV"
- run: node capture.mjs
- run: node make-report.mjs
- uses: actions/upload-artifact@v5
if: always()
with:
name: website-report-${{ env.REPORT_MONTH }}
path: artifacts/${{ env.REPORT_MONTH }}/
retention-days: 90
Replace the example URLs with the pages you own or are authorized to capture. Set the reporting month explicitly for manual reruns so the archive lands in the intended folder. Scheduled workflows use the repository’s default branch and can be delayed during busy periods; if the exact delivery time matters, monitor the run and choose a scheduler with suitable guarantees. Retention is configured on the uploaded artifact and should match your archive and access needs. For long-lived records, copy the PDF and manifest into storage with an appropriate retention policy.
Playwright’s CI setup requires installing the browser and system dependencies. Its documentation recommends one worker in CI when prioritizing stability and reproducibility, and demonstrates uploading artifacts. This simple script runs sequentially. Read the Playwright CI guide for CI setup details.
5. Verify the output and keep failures visible
- Confirm the manifest includes every configured URL and has no error statuses.
- Open sample screenshots and check the viewport, page state, consent dialogs, lazy content, and legibility.
- Open the PDF and check the reporting month, URL captions, order, and page count.
- Keep failed URLs in logs and fail the job instead of quietly publishing a partial PDF.
- Use stable filenames and preserve the manifest with the PDF so later reviewers can identify capture time and source page.
- Review artifact access: screenshots may contain personal information or private site content. Store them only where intended readers are authorized.
6. Configuration choices and tradeoffs
| Decision | Options | Practical guidance |
|---|---|---|
| Schedule | CI workflow or server scheduler | Choose a timezone and expected delay tolerance; document manual reruns. |
| Scope | Viewport, full page, element | Use the smallest capture that answers the review question; full pages can shrink too much in a one-page PDF. |
| Readiness | Navigation event, selector, page-specific condition | Wait for meaningful content; avoid assuming one fixed delay works on every site. |
| PDF layout | One image per page, split tall captures, multiple images per page | Choose for legibility and report size; keep URL and timestamp visible. |
| Retention | CI artifact or longer-term object storage | Set expiry, access control, and naming for the archive’s intended lifetime. |
Performance, reliability, and cost
Sequential captures are slower as the URL list grows, but use resources predictably. Parallelize only after accounting for CI memory, target site rate limits, and independent failure reporting. Reusing one browser context can reduce startup overhead; use separate contexts when cookies or state must not cross between sites. Pin Node and package versions, preserve the lockfile, and install the matching browser in CI. Browser and operating system updates can affect rendering, so a stable archive should record its runtime versions. CI minutes, artifact storage, and any external storage have provider-specific limits and costs; check those current terms rather than assuming a fixed price. The supplied sources do not establish a universal cost or reliability comparison.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. One request captures a page, and it can also produce PDFs. For an individual capture, use the documented API parameters in the ScreenshotNeo API docs:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
For a recurring report, trigger the API call from your scheduler for each URL, save the returned files with the reporting month, and assemble them into a report if you need a custom multi-page layout. ScreenshotNeo accepts cookie or consent banners like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, and failed loads are never billed, and response headers say which page verdict and billing result applied. An MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up for 1,000 free screenshots a month with no card.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Browser launch fails in CI | Browser binaries or operating-system dependencies are missing or mismatched. | Run npx playwright install --with-deps chromium after installing locked dependencies; inspect Playwright’s CI guide. |
| Navigation times out | Slow site, blocked automation, or waiting for a load event that never finishes. | Use a bounded timeout and wait for a meaningful selector or DOM content rather than every network connection; investigate access restrictions. |
| Screenshot is blank or incomplete | Capture happened before content was rendered, or the page requires interaction/state. | Wait for a page-specific selector, establish required state, and inspect the screenshot and logs. |
| Full-page image misses lazy content | Images or sections load only after scrolling into view. | Scroll through the page and wait for required assets before capture, or choose viewport/element scope. |
| Element capture errors | Selector no longer matches, matches unexpectedly, or element is outside the expected state. | Verify the selector against the current page and wait for the element to be visible. |
| PDF looks too small | A long full-page image was scaled to fit one page. | Use viewport captures or split the image over multiple PDF pages; select a page size that preserves readability. |
| Job succeeds but report is missing URLs | Failures were ignored or the PDF step consumed incomplete inputs. | Keep manifest status checks and nonzero exit codes; upload logs and partial artifacts for diagnosis. |
| Scheduled run appears late or absent | Scheduled workflow delay, workflow configuration, or default-branch assumptions. | Check workflow run history and current GitHub schedule rules; use manual dispatch to recover a missed report. |
FAQ
Should the PDF include the screenshots themselves or links to them?
Embed the images when readers need a portable report. Include stable links or archive identifiers as well when the screenshots are large or need separate access controls.
Can the report compare this month with last month?
Yes. Keep the same capture configuration and retain each month’s images. A visual comparison is a separate report feature; this workflow creates an archive and monthly PDF, not an automatic change analysis.
Can I use another scheduler?
Yes. A server scheduler or another CI platform can run the same Node scripts. The scheduler starts the work; Playwright captures, and the PDF library assembles the report.
Is this a turnkey monthly report generator?
No. The capture script and PDF assembly are explicit pieces that you own and can adapt to your layout, archive rules, and review process.


