ScreenshotNeo

BlogHow-to

How to Run a Scheduled Playwright Screenshot Job in Docker

Run Playwright screenshots on a schedule in Docker with a version-matched browser image, durable output, and a GitHub Actions or self-managed scheduler.

By the ScreenshotNeo team4 October 20268 min read

To run a scheduled Playwright screenshot job in Docker, keep a capture script and locked dependencies in your repository, run them in a Playwright container whose release matches the Playwright package, and use a scheduler to start the container. Write screenshots to a known directory, then upload that directory as a workflow artifact or mount it to durable storage. This guide uses Node.js and Chromium with GitHub Actions; it also covers a self-managed Docker host.

1. Create a small Playwright screenshot project

The official Playwright Docker image includes browser binaries and their operating-system dependencies, but it does not include your project’s Playwright package. Install that package in the job from a lockfile, and keep its release aligned with the image release so Playwright can locate the expected browser executable. See the Playwright Docker documentation and Playwright CI documentation.

Start with this project layout:

scheduled-shot/
  package.json
  package-lock.json
  screenshot.mjs
  .github/
    workflows/
      screenshot.yml

Create package.json with an exact Playwright version. Replace 1.x.y with a real release and use the corresponding version in the Docker image tag below; do not leave the placeholder in a runnable project.

{
  "name": "scheduled-shot",
  "private": true,
  "type": "module",
  "scripts": {
    "screenshot": "node screenshot.mjs"
  },
  "dependencies": {
    "playwright": "1.x.y"
  }
}

Generate and commit the lockfile with npm install --package-lock-only after substituting the version. The scheduled job will use npm ci, which installs the exact dependency tree described by the committed lockfile.

2. Write the capture script

This script accepts the target URL and output directory from environment variables, creates the directory, navigates to the page, and saves a full-page PNG. It closes Chromium even when navigation or capture fails.

// screenshot.mjs
import { chromium } from 'playwright';
import { mkdir } from 'node:fs/promises';
import { resolve } from 'node:path';

const url = process.env.TARGET_URL;
if (!url) throw new Error('Set TARGET_URL to the page to capture');

const outputDir = resolve(process.env.OUTPUT_DIR ?? 'screenshots');
await mkdir(outputDir, { recursive: true });

const browser = await chromium.launch();
try {
  const page = await browser.newPage({
    viewport: { width: 1440, height: 1000 },
    deviceScaleFactor: 1,
  });
  await page.goto(url, { waitUntil: 'networkidle', timeout: 60_000 });
  await page.screenshot({
    path: `${outputDir}/page.png`,
    fullPage: true,
    animations: 'disabled',
  });
} finally {
  await browser.close();
}

The screenshot API supports output paths, full-page capture, format and quality choices, and clipped regions. PNG is suitable when pixel fidelity matters. For JPEG or WebP, set the corresponding type and a quality value; quality applies to lossy formats. For a stable history, incorporate a UTC timestamp into the filename rather than overwriting page.png. See the Playwright screenshot API.

3. Schedule it with GitHub Actions

Add .github/workflows/screenshot.yml. The sample runs every day at 17:23 UTC and can also be started manually from the Actions tab. Replace the Playwright tag with the same release pinned in package.json. Check that the action major versions and Node version remain supported in your repository.

name: Scheduled screenshot

on:
  schedule:
    - cron: '23 17 * * *'
  workflow_dispatch:

jobs:
  capture:
    runs-on: ubuntu-latest
    container:
      image: mcr.microsoft.com/playwright:v1.x.y-noble
    steps:
      - uses: actions/checkout@v6
      - uses: actions/setup-node@v6
        with:
          node-version: '22'
      - run: npm ci
      - run: npm run screenshot
        env:
          TARGET_URL: https://example.com
          OUTPUT_DIR: screenshots
      - uses: actions/upload-artifact@v5
        if: ${{ !cancelled() }}
        with:
          name: scheduled-screenshots
          path: screenshots/
          retention-days: 14

GitHub scheduled workflows use five-field POSIX cron, run from the latest commit on the default branch, and support IANA timezones as well as the default UTC behavior. The documented shortest interval is five minutes. Starts can be delayed during high load, and queued runs can be dropped; a cron entry is not an exact-time guarantee. GitHub also documents higher load at the start of an hour, so selecting a minute such as 23 can avoid that peak. The workflow file must exist on the default branch. Consult GitHub’s schedule event documentation for current syntax and behavior.

The upload step stores the generated directory as a workflow artifact, which you can retrieve after the job ends. Artifacts are for job output; dependency caches are for reusable inputs. Choose a retention period that fits how long you need each capture. See GitHub’s artifact documentation.

4. Run the same job on a self-managed Docker host

For a Linux machine with Docker, a host scheduler such as cron or a systemd timer can run the same project in a short-lived container. Mount the project directory so dependencies and output are available outside the container, and ensure the output path is writable.

docker run --rm --init --ipc=host \
  -e TARGET_URL=https://example.com \
  -e OUTPUT_DIR=/work/screenshots \
  -v "$PWD:/work" \
  -w /work \
  mcr.microsoft.com/playwright:v1.x.y-noble \
  sh -lc 'npm ci && npm run screenshot'

Substitute the same real release used by the package. --init helps handle processes correctly in the container, while --ipc=host gives Chromium adequate shared memory and can prevent memory-related browser crashes. A container’s local files disappear when it is removed unless they are mounted or sent to durable storage.

Put the Docker command in a script and have the host scheduler invoke it. Keep the URL and other configuration outside the image; use a secrets store for credentials. Send container output to the host’s logging system and configure alerting for failed jobs. The right scheduler and storage depend on the host’s operations requirements.

Before adding the schedule, run the command once manually and confirm that the image is produced at the expected mounted path. For long-lived archives, configure durable storage rather than relying on a short-lived workflow artifact or container filesystem.

5. Make captures repeatable and reliable

  • Pin the inputs: lock the Playwright package, match the browser image release, and choose a consistent Node version.
  • Keep the rendering environment stable: operating system, browser release, settings, hardware, power source, and headless mode can all affect pixels. Compare captures from like-for-like environments. See Playwright’s visual comparison guidance.
  • Choose a page readiness condition deliberately: networkidle waits for network activity to settle, but analytics, polling, and other persistent connections can make it unsuitable for some sites. If that happens, wait for a meaningful selector or a known page state, or use a bounded delay when the page has no better signal.
  • Control concurrency: a single capture script launches one browser. If you expand the job into parallel captures or a test suite, keep concurrency modest; Playwright recommends one worker in CI when stability and reproducibility matter.
  • Keep browser caching purposeful: Playwright notes that browser caching in CI often is not worthwhile because cache restoration can take as long as downloading browsers, and Linux system dependencies cannot be cached.
  • Preserve useful evidence: use timestamps or URL-specific filenames if multiple captures share an output directory. Upload or copy logs and screenshots on failure when diagnosing intermittent page issues.

6. Handle security when targets are not fully trusted

Playwright describes its Docker image as intended for testing and development and does not recommend it for visiting untrusted websites. The image runs as root by default, which disables Chromium’s sandbox. For scraping or crawling scenarios, Playwright documents using a separate user and a seccomp profile. Treat target URLs as a security boundary, avoid exposing secrets to pages, and apply container isolation appropriate to your threat model. Read the Docker guidance for crawling and scraping.

7. Hosted workflow or self-managed scheduler?

Decision GitHub Actions Self-managed container
Scheduler Workflow cron; convenient, but starts may be delayed or dropped at high load. Host scheduler or orchestrator; depends on host availability and operations.
Output Upload workflow artifacts with a chosen retention period. Mount a host directory or configure durable storage.
Operations Maintain workflow configuration, dependencies, and artifact retention. Maintain Docker host, scheduler, logs, alerts, and storage.
Reproducibility Pin image, package, runtime, and capture settings. The same pinning helps, with more control over the host environment.

Or skip the browser setup

For a scheduled capture where managing Chromium, its dependencies, and storage is unnecessary, ScreenshotNeo returns a screenshot from one GET request. It is a website screenshot API and MCP server from ScreenshotNeo. See the ScreenshotNeo API documentation for parameters.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

Cookie banners are accepted like a visitor and removed along with known consent platforms, newsletter popups, and chat widgets before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are never billed, and response headers say which page verdict and billing result applied. An MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs. The free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000. Other options include full-page and element capture, device presets, custom CSS and JavaScript, wait conditions, request blocking, PDF output, caching, signed links, async jobs, and bulk capture.

Sign up for 1,000 free screenshots a month, with no card required.

Troubleshooting

Symptom Likely cause Fix
Playwright cannot find the browser executable The installed package and Docker image releases do not match, or the package was not installed. Pin the same release in the manifest and image tag; run npm ci in the job.
Chromium crashes or exits unexpectedly Insufficient shared memory or container process handling. Use --ipc=host and --init for the Docker invocation; inspect available memory.
Navigation times out The page is slow, unreachable from the runner, or never becomes idle because of ongoing traffic. Check the URL and network access, then wait for a specific selector or another bounded readiness condition rather than requiring network idle.
Screenshot is blank or incomplete The capture ran before the relevant content rendered, or the page uses lazy loading. Wait for a page-specific selector or state. For long pages, verify that the content is loaded before full-page capture.
The screenshot file is missing after the job The file remained in the container or workspace and was not persisted. Upload the output as a workflow artifact or mount/copy it to durable storage before the container exits.
Scheduled run did not start at the expected minute GitHub scheduled starts are subject to load, and runs can be delayed or dropped. Do not use workflow cron as an exact-time trigger. Choose a minute away from the top of the hour and monitor for missed runs if timing matters.
Captures differ between runs Browser or OS versions, rendering settings, content, or environment changed. Pin versions and viewport/device scale, keep the environment consistent, and account for dynamic page content.

For browser launch diagnostics, set DEBUG=pw:browser in the job environment to enable Playwright browser debug output. Confirm the image/package version pairing and inspect the runner or host logs.

FAQ

Does the Playwright image install the npm package for me?

No. It provides browsers and operating-system browser dependencies. Install the project package from its lockfile during the job.

Can I use this for a one-off screenshot as well as a schedule?

Yes. Keep workflow_dispatch for manual GitHub runs, or invoke the Docker command directly on a self-managed host.

Will GitHub Actions run at the exact cron time?

No. Scheduled starts can be delayed or dropped during periods of high load.

Where do the screenshots live after a GitHub job?

They are available through the uploaded workflow artifact until its configured retention period expires.