ScreenshotNeo

BlogHow-to

How to Schedule Website Screenshots Using a GitLab CI Pipeline

Use GitLab pipeline schedules and Playwright to capture websites on a recurring basis, save screenshots as artifacts, and troubleshoot missed or failed runs.

By the ScreenshotNeo team4 October 20269 min read

To schedule website screenshots with GitLab CI, add a job that runs a browser script and saves screenshots in a directory, publish that directory as a job artifact, then create a pipeline schedule for the branch or tag containing your .gitlab-ci.yml. The schedule uses cron and runs independently of commits. This guide uses Playwright with Node.js; it also covers schedule permissions, readiness, artifacts, troubleshooting, and an API alternative.

1. Add Playwright to the repository

Commit your dependency manifest and lockfile so the scheduled job installs a repeatable package version. For example, add @playwright/test as a development dependency and commit the resulting package-lock.json. Playwright recommends its public Docker image for GitLab CI. Pin the image tag and the package version to compatible releases; do not use an unpinned floating tag for recurring visual captures.

npm install --save-dev @playwright/test

The GitLab job below uses the Playwright image shape documented by Playwright. Replace <pinned-version> with a real release tag compatible with the version in your lockfile. The placeholder is intentional; choose and pin the matching versions for your repository.

2. Write the capture script

Create scripts/capture-screenshots.mjs. This runnable script reads newline-separated URLs from the SCREENSHOT_URLS CI variable, creates an output directory, captures each URL, and exits with an error if any navigation fails. The variable should contain URLs, not credentials. Store login secrets separately as protected or masked CI variables and adapt the script to use them only when the target site requires authentication.

import { chromium } from '@playwright/test';
import { mkdir } from 'node:fs/promises';

const rawUrls = process.env.SCREENSHOT_URLS ?? 'https://example.com';
const urls = rawUrls
  .split(/\r?\n/)
  .map((value) => value.trim())
  .filter(Boolean);

if (urls.length === 0) {
  throw new Error('SCREENSHOT_URLS must contain at least one URL');
}

await mkdir('screenshots', { recursive: true });
const browser = await chromium.launch();
let failed = false;

try {
  for (const [index, url] of urls.entries()) {
    const page = await browser.newPage({ viewport: { width: 1440, height: 1000 } });
    try {
      const response = await page.goto(url, {
        waitUntil: 'domcontentloaded',
        timeout: 45000,
      });
      if (!response || !response.ok()) {
        throw new Error(`Navigation returned ${response?.status() ?? 'no response'}`);
      }
      // Replace this with a site-specific readiness check if content loads later.
      await page.screenshot({
        path: `screenshots/page-${String(index + 1).padStart(2, '0')}.png`,
        fullPage: true,
        animations: 'disabled',
      });
      console.log(`Captured ${url}`);
    } catch (error) {
      failed = true;
      console.error(`Failed to capture ${url}:`, error);
    } finally {
      await page.close();
    }
  }
} finally {
  await browser.close();
}

if (failed) {
  process.exitCode = 1;
}

domcontentloaded is a useful starting wait condition when pages keep connections open, but it does not guarantee that client-rendered content or images are ready. For a page with a known ready marker, wait for it explicitly after navigation, for example:

await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 45000 });
await page.locator('[data-page-ready="true"]').waitFor({ state: 'visible', timeout: 20000 });

Use a selector that actually exists on your site. If the page has no ready marker, use a short, deliberate delay or wait for a specific element that indicates the content you want has appeared. A fixed delay is simple but may waste time on fast pages and still be too short on slow ones.

3. Configure the GitLab CI job

Add a capture job to .gitlab-ci.yml. Artifacts are archived job outputs; artifacts:paths tells GitLab which files to retain. when: always also preserves screenshots produced before a later URL fails, which helps diagnose partial runs.

stages:
  - capture

capture_screenshots:
  stage: capture
  image: mcr.microsoft.com/playwright:<pinned-version>-noble
  variables:
    # One browser worker favors stable, repeatable CI captures.
    CI: "true"
  script:
    - npm ci
    - node scripts/capture-screenshots.mjs
  artifacts:
    when: always
    paths:
      - screenshots/
    expire_in: 7 days

The image tag above is a placeholder: use an actual Playwright image release compatible with the Playwright package in your lockfile. Playwright also documents installing browser dependencies with its CLI in a custom job environment. That gives more control over system packages but requires maintaining those dependencies yourself. Use the image when its environment fits your job; use a custom image when your project needs additional system packages or a managed runtime.

Set expire_in to the review window your team needs. Artifact retention and file-size limits depend on your GitLab configuration and offering; check the target instance before relying on a particular maximum. Large full-page screenshots across many URLs can consume artifact storage quickly.

4. Create the pipeline schedule

  1. Open the project in GitLab and go to Build > Pipeline schedules (the exact navigation label can vary by GitLab version).
  2. Create a schedule, choose the branch or tag containing the CI configuration, and enter a cron expression for the cadence.
  3. Choose a timezone and set any non-secret schedule inputs or variables your pipeline needs. Put credentials in the project’s protected or secret-managed variables, not in a plain schedule input or committed YAML.
  4. Save the schedule, then run it manually once. Confirm the intended job ran, the pages loaded, screenshot files were produced, and the artifacts can be downloaded.

A cron schedule has five fields: minute, hour, day of month, month, and day of week. For example, 0 7 * * * means 07:00 every day in the schedule’s selected timezone. GitLab’s documented 0 * * * * example runs at the start of every hour. The instance can enforce a maximum schedule frequency, so a desired cadence may not be accepted or run as often as expected. Check your instance’s limit. Avoid choosing a heavily shared minute mark when many schedules are created together; distributing start times can reduce synchronized load.

Scheduled pipelines run independently of commits and are associated with a target ref. They run with the schedule owner’s permissions. The owner needs suitable project access, and protected refs may require permission to merge. If the owner is blocked or removed, the schedule can stop being active. Arrange ownership so someone with the needed access remains responsible for it.

5. Choose capture and output settings

Choice Use it when Trade-off
Viewport screenshot You monitor a fixed visible area, such as a dashboard header or landing-page fold. Content below the viewport is not included.
Full-page screenshot You need to inspect a long page in one image. Very long pages can create large images and take longer to capture.
Playwright Docker image You want a documented browser environment with less dependency setup. The package and image versions must stay compatible.
Custom job image plus browser install You need specific system packages or a custom runtime. You own more environment and dependency maintenance.
Short artifact expiry Captures are only needed for near-term review. Older runs become unavailable sooner.
Longer artifact expiry People need a longer investigation or comparison window. Retained files use more artifact storage.

For a viewport image, remove fullPage: true. For a full-page image, keep it enabled. Use explicit viewport dimensions and the same browser image for scheduled captures and any comparison baseline. Playwright notes that rendering can vary with host operating system, browser version, settings, hardware, power source, and headless mode, so scheduled captures are not guaranteed to be pixel-identical across environments.

6. Get the screenshots from the pipeline

After the job finishes, open its job page and download the screenshots/ artifact. With when: always, artifacts can remain available even when the script exits unsuccessfully, provided the runner and artifact upload step complete. If the pipeline fails before the job starts or the runner cannot upload artifacts, there may be no archive to download.

For recurring review, choose a cadence that matches how quickly you need to notice changes and an artifact expiry that covers the review window. GitLab’s instance schedule limit, runner availability, page load time, screenshot size, and artifact storage all constrain how frequent and broad a capture job should be.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. A single GET request can return a PNG, JPEG, WebP, or PDF. For a scheduled capture, call the API from your GitLab job and save the response as an artifact; see the ScreenshotNeo API documentation for parameters and setup.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key="$SCREENSHOTNEO_API_KEY" \
  --data-urlencode url="https://example.com" \
  -o screenshots/example.webp

Keep SCREENSHOTNEO_API_KEY in a protected or masked GitLab CI variable, and create the screenshots/ directory before saving. Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up for 1,000 free screenshots a month, with no card required.

Troubleshooting

Symptom Likely cause What to check or change
Browser launch fails Playwright package and image versions are incompatible, or browser dependencies are missing. Pin a matching image and package version. If using a custom image, install the browser dependencies using Playwright’s documented setup for that environment.
Scheduled pipeline does not start The schedule is inactive, targets the wrong ref, exceeds an instance frequency limit, or its owner lost access. Review the schedule status, cron cadence, target branch or tag, and owner permissions. Reassign responsibility to an eligible project member if needed.
Pipeline starts late or is missed Runner capacity, schedule limits, timezone assumptions, or a synchronized start time may be involved. Confirm the selected timezone and cron fields, check runner availability, verify the instance’s maximum frequency, and distribute start times where practical.
Screenshot is blank or incomplete The page has not rendered its content by the chosen readiness condition, or the destination returned an error. Check the navigation response and job log. Wait for a site-specific ready selector or use a different wait condition; verify the page is accessible from the runner.
Navigation times out The target is slow, unreachable from the runner, or keeps network connections open. Check runner network access and the URL. Consider domcontentloaded instead of networkidle for pages with persistent requests, and set a deliberate timeout. Add a readiness check for content that loads after navigation.
Artifacts are missing The output path differs from the declared path, the script did not create files, or upload failed. Make sure the script writes into screenshots/, the path is listed under artifacts:paths, and the job log includes the artifact upload step. Inspect job failure details.
Images vary between runs Browser or host rendering environment changed, or the site itself has dynamic content. Use the same pinned image and package version for baseline and scheduled capture, fix viewport and readiness behavior, and account for genuinely dynamic page content.
Some URLs succeed and others fail One destination may require authentication, block the runner, or return an unsuccessful response. Inspect the logged URL and response status. Configure required access with protected secrets, and decide whether one failure should fail the entire scheduled job.

Performance, reliability, and cost

  • Keep CI predictable: Playwright recommends one worker in CI for stability and reproducibility. Increase parallelism only when runner resources and destination behavior support it.
  • Do not cache browser binaries by default: Playwright notes that restoring cached browser binaries can take as long as downloading them, and Linux system dependencies cannot be cached.
  • Limit capture scope: Capture only the URLs and page area you need. Full-page captures and a large URL list increase job duration and artifact size.
  • Set realistic timeouts: A timeout that is too short creates avoidable failures; one that is too long can hold a runner for a stalled destination. Pair navigation timeouts with a content readiness check.
  • Keep secrets out of logs: Do not place credentials in committed YAML or plain schedule inputs. Use the project’s protected or secret-management features, and avoid printing secret values.
  • Budget storage and compute: GitLab retains the files according to artifact policy, while each scheduled browser run uses runner time. Set schedule cadence and artifact expiry to match the operational review need.

FAQ

Can GitLab take a screenshot every day?

Yes. Create a pipeline schedule with a daily cron expression, select the correct timezone and ref, and make sure the instance’s schedule frequency policy permits it.

Do scheduled pipelines need a commit to run?

No. A pipeline schedule runs independently of new commits, using the ref selected in its schedule configuration.

Can I keep screenshots after the job ends?

Yes. Declare the output directory in artifacts:paths; GitLab makes the archived files available according to the configured retention policy.

Will repeat captures be pixel-identical?

Not necessarily. Browser, host, rendering settings, hardware, and dynamic site content can change the output even when the URL is unchanged.

Can I use a custom Docker image?

Yes. Install a compatible Playwright package, browser, and system dependencies in the job environment. The Playwright image reduces that setup, while a custom image gives more control over the runtime.