ScreenshotNeo

BlogHow-to

How to Take Bulk Website Screenshots with GitLab CI

Capture many pages with Playwright in GitLab CI, shard large test suites across jobs, and keep screenshot files as controlled job artifacts.

By the ScreenshotNeo team4 October 202611 min read

To capture many website pages in GitLab CI, run a Playwright script or Playwright Test suite in a CI job, save image files under a directory in the checkout, and archive that directory with artifacts:paths. For a sufficiently large Playwright Test suite, set GitLab parallel and pass --shard=$CI_NODE_INDEX/$CI_NODE_TOTAL so each job runs its assigned share. A plain list of URLs does not get split by that Playwright Test option: a URL-list script needs its own partitioning logic.

This guide covers both designs, artifact retention and access, scaling limits, troubleshooting, and a hosted API option. The examples are adapted from official documentation and are not claims of an executed pipeline.

1. Choose how to distribute the capture work

“Bulk screenshots” can describe two different workloads:

  • A URL list: a script reads URLs and visits each one. A single job can process the list sequentially or concurrently. If using multiple jobs, the script must divide the list among them.
  • A Playwright Test suite: tests capture pages, and Playwright Test’s shard mechanism divides test work across GitLab’s parallel job instances.

Use the URL-list approach when the input is simply a set of pages and you want direct control over the work queue. Use test sharding when captures are already represented as Playwright tests and the suite can be divided according to Playwright’s shard semantics. For either design, more jobs help only when runners can execute them concurrently and the destination site can handle the requests.

2. Create a Playwright project that saves screenshots

The examples below use Playwright Test. Install the project dependencies locally and commit the lockfile so CI can use npm ci. A minimal test can capture a page into a stable output directory:

// tests/capture.spec.js
const { test } = require('@playwright/test');
const fs = require('node:fs/promises');

const pages = [
  { name: 'home', url: 'https://example.com/' },
  { name: 'docs', url: 'https://example.com/docs' },
];

test('capture configured pages', async ({ page }) => {
  await fs.mkdir('screenshots', { recursive: true });

  for (const item of pages) {
    const response = await page.goto(item.url, {
      waitUntil: 'domcontentloaded',
      timeout: 45_000,
    });

    if (!response || !response.ok()) {
      throw new Error(`Page did not load successfully: ${item.url}`);
    }

    await page.screenshot({
      path: `screenshots/${item.name}.png`,
      fullPage: true,
      animations: 'disabled',
    });
  }
});

This intentionally small example captures pages sequentially inside one test. For a larger suite, prefer one test per page or logical page group so Playwright Test can distribute work. Ensure output filenames are unique across tests and any job matrix. If the target requires authentication, provide credentials through protected CI variables and configure the browser context or saved storage state without committing secrets.

Playwright’s official [screenshot guide](https://playwright.dev/docs/screenshots) describes screenshot options and its [CI guide](https://playwright.dev/docs/ci) documents GitLab CI patterns. Set the Playwright package version and CI image version to matching releases; update both deliberately when upgrading.

3. Run captures in a GitLab CI job and retain the files

A basic pipeline runs the suite in a Playwright Docker image and archives the output directory. The image tag below follows the version shown in the research dossier; verify the current official CI guide and match it to the Playwright dependency in your project before using it.

stages:
  - capture

screenshots:
  stage: capture
  image: mcr.microsoft.com/playwright:v1.63.0-noble
  script:
    - npm ci
    - npx playwright test
  artifacts:
    when: always
    paths:
      - screenshots/
    expire_in: 1 week

The test code must create screenshots/. Artifact paths are relative to the job’s checkout. The example’s when: always retains outputs even when a capture fails, which is useful for diagnosis; choose an expiry that fits your retention needs. GitLab otherwise uploads artifacts from successful jobs by default. See [GitLab job artifacts](https://docs.gitlab.com/ci/jobs/job_artifacts/) for artifact configuration and access settings.

4. Shard a Playwright Test suite across GitLab jobs

For a suite with enough independent tests, add GitLab parallel and pass each job’s index and total to Playwright:

stages:
  - capture

screenshots:
  stage: capture
  image: mcr.microsoft.com/playwright:v1.63.0-noble
  parallel: 4
  script:
    - npm ci
    - npx playwright test --shard=$CI_NODE_INDEX/$CI_NODE_TOTAL
  artifacts:
    when: always
    paths:
      - screenshots/
    expire_in: 1 week

GitLab creates parallel job instances and provides CI_NODE_INDEX and CI_NODE_TOTAL; Playwright uses the shard argument to select a portion of the test suite. Each parallel job uploads its own artifacts. Job artifacts remain associated with their producing job, so inspect or download the outputs from all shard jobs when you need the complete set.

GitLab’s YAML reference documents parallel values from 1 through 200, but that is a configuration range, not a throughput promise. Jobs can queue if runner capacity is limited. Playwright recommends one worker in CI as a stable default; sharding across jobs is the documented way to widen parallelism. Tune the job count and any Playwright worker setting to available CPU and memory, browser process load, target-site request limits, and runner concurrency.

Use a matrix for browser and shard combinations

If captures must run against multiple browser projects, GitLab’s parallel:matrix can create combinations of variables. Configure Playwright projects to consume the browser variable and shard variable as appropriate. Keep the matrix small enough that runner capacity and artifact volume remain manageable. Each browser and shard combination adds work and potentially another set of images.

screenshots:
  stage: capture
  image: mcr.microsoft.com/playwright:v1.63.0-noble
  parallel:
    matrix:
      - BROWSER: [chromium, firefox]
        SHARD: ['1/2', '2/2']
  script:
    - npm ci
    - npx playwright test --project=$BROWSER --shard=$SHARD
  artifacts:
    paths:
      - screenshots/
    expire_in: 1 week

Adapt the project names and output naming to your configuration. Avoid assuming that two jobs can safely write to the same shared filesystem: GitLab jobs normally have separate checkouts. Make filenames deterministic and descriptive, such as a normalized page identifier plus browser name, to make downloaded outputs easy to reconcile.

5. Distribute a raw URL list yourself

The Playwright --shard argument partitions Playwright Test work, not an arbitrary array in a Node.js script. For a URL-list design, divide the list using a stable index rule based on GitLab’s job index and total, or generate a separate input list for each job. A minimal partition helper looks like this:

function urlsForShard(urls, index, total) {
  return urls.filter((_, position) => position % total === index - 1);
}

const index = Number(process.env.CI_NODE_INDEX || '1');
const total = Number(process.env.CI_NODE_TOTAL || '1');
const assignedUrls = urlsForShard(pages, index, total);

Use the helper in a script that launches Playwright and writes each result under screenshots/. Validate that index and total are positive integers and that index is no greater than total. The modulo scheme is simple and deterministic, though it may leave some jobs with more work if page load times vary. If capture durations are uneven, use a queue or prepare balanced partitions instead. Make sure output names cannot collide, especially when URLs normalize to the same filename.

6. Capture choices that affect output

  • Full page or viewport: fullPage: true captures the page’s full scrollable extent. Viewport screenshots are smaller and useful for consistent visual checks.
  • Readiness: choose a navigation wait condition that matches the site. domcontentloaded can return before client-rendered content is ready; wait for a selector or a site-specific readiness signal when needed.
  • Lazy content: full-page capture does not guarantee that every lazy-loaded image or infinite-scroll section has appeared. Scroll or trigger the site’s loading behavior before capture when required.
  • Animations: disabling animations can reduce visual variation in screenshots. Dynamic data, rotating banners, timestamps, and personalization can still cause differences.
  • Viewport and device: set viewport dimensions explicitly for reproducible captures. Use separate projects or configurations for desktop and mobile layouts.
  • Authentication and consent: supply a valid session securely and decide whether the capture should include consent overlays. Website behavior varies by region, session, and policy.

7. Artifacts, access, and size

Use artifacts:paths to choose which output files GitLab archives. Set expire_in when captures should not be retained indefinitely; if omitted, the instance default applies. Later stages fetch earlier artifacts by default, while dependencies and needs:artifacts can control which outputs are fetched. Configure artifacts:access according to who should view or download potentially sensitive screenshots.

GitLab’s job artifacts documentation states that the default maximum final artifact archive size is 100 MB. This is a limit on the archive, not an individual screenshot. If a capture archive reaches the effective project or instance limit, reduce image dimensions or page count per job, split output into multiple archives, or ask an administrator about the configured limit. Check the actual project setting before scaling a large run.

Screenshots can reveal private pages or account data visible to the runner. Keep secrets out of filenames and logs, restrict artifact access, and check access configuration before publishing captured files to a public Pages site. GitLab notes that UI and API artifact controls do not necessarily prevent job-token access through runner APIs.

8. Performance, reliability, and cost

  • Parallelism: more GitLab jobs can reduce elapsed time only when concurrent runner capacity exists. Otherwise jobs wait in a queue. Browser startup, dependency installation, and artifact upload also consume time.
  • Resource use: each active browser consumes CPU and memory. Start with a conservative job count and monitor runner pressure, then adjust based on observed behavior rather than assuming a universal optimum.
  • Target-site limits: a large burst can trigger rate limiting or bot defenses. Limit concurrency, use an authorized test environment where possible, and honor the site’s access rules.
  • Repeatability: pin the package lockfile and browser image version together. Use explicit viewports, stable page readiness checks, and unique filenames. Dynamic content can still vary between runs.
  • Failure isolation: one job per shard limits the scope of a runner or shard failure, but a failed shard means its pages may be missing. Retain failure artifacts when they aid debugging and make pipeline status visible.
  • Storage: full-page images, multiple viewport sizes, and browser matrices increase artifact volume. Choose expiry and dimensions with review needs and the archive limit in mind.
  • Cost: GitLab runner consumption and artifact storage depend on your GitLab plan, runner setup, and retention. The cited documentation provides configuration limits, not a per-screenshot cost or throughput benchmark.

9. Troubleshooting

Symptom Likely cause Fix
Browser executable is missing The installed Playwright package and container image are mismatched, or the project uses an image without the required browsers. Match the Playwright dependency and Docker image versions. Follow Playwright’s current CI guide for the image appropriate to the project.
All jobs appear to run the same tests The command does not pass the shard expression, or the workload is a raw URL list rather than a Playwright Test suite. Pass --shard=$CI_NODE_INDEX/$CI_NODE_TOTAL for Playwright Test. Partition URL lists explicitly in the script.
Parallel jobs wait for a long time There are not enough concurrent runners, or an active-job limit has been reached. Check runner capacity and pipeline limits. Reduce parallel or provide more runner concurrency.
Screenshot directory is missing The test did not create the configured output directory, or wrote files somewhere else. Create the directory before capture and make artifacts:paths match the output location.
Artifacts are absent after a failed job Artifact upload defaults to successful jobs. Set artifacts:when: always or on_failure when failure output is needed, and confirm files existed before job exit.
Only some pages appear in the downloaded results A shard failed, some URLs were excluded, or outputs are spread across separate job artifacts. Inspect every shard’s status and artifact. Confirm partition boundaries and download or fetch artifacts from all producing jobs.
Artifact upload fails due to size The final archive exceeds the effective maximum. Check the configured archive limit; reduce dimensions or pages per archive, split jobs, or adjust the project/instance limit.
Page is blank or incomplete Navigation completed before rendering, the site needs authentication, or content loads after scrolling. Wait for a page-specific selector or readiness condition, provide authorized authentication state securely, and trigger lazy loading as needed.
Different runs produce different images Dynamic content, animation, responsive layout, locale, or timing changed. Fix viewport and locale settings, disable animations where suitable, and wait for stable page content. Mask or avoid inherently dynamic regions if using visual comparisons.
Site returns a challenge or rate limit Request frequency, network reputation, or access policy triggered site defenses. Lower concurrency, use an approved environment or allowlisted runner, and follow the target site’s rules. Do not attempt to bypass access controls.

10. Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server. A single GET request can return a PNG, JPEG, WebP, or PDF. For bulk work, its API also supports up to 100 URLs per call. See the [ScreenshotNeo documentation](https://screenshotneo.com/docs/) for request options and configuration.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));

ScreenshotNeo accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up for ScreenshotNeo’s free 1,000 screenshots per month, with no card required.

11. Frequently asked questions

Does GitLab’s parallel keyword split a file of URLs automatically?

No. It creates parallel job instances. Playwright Test can select a suite shard with the shard argument; a custom URL-list script must partition its own input.

Do I need one screenshot test per page?

No, but separate tests or clearly divided page groups make sharding, retries, and missing-output diagnosis easier. Keep each output filename unique.

Can later GitLab stages use the screenshots?

Yes. GitLab can fetch artifacts from earlier stages by default. Use dependencies or needs:artifacts to control which job outputs a later job downloads.

What should I set for the parallel job count?

There is no universally correct count. Choose a value your runners can execute concurrently and your target site can tolerate, then account for memory, job limits, and artifact size.

Can I keep screenshots after the artifact expires?

Artifacts follow their configured expiry and instance behavior. If longer retention is required, arrange an authorized storage destination and access policy that fits the sensitivity of the captured pages.