ScreenshotNeo

BlogHow-to

How to Schedule Website Screenshots with a Custom User Agent

Set a custom user agent in Playwright, save scheduled website screenshots with GitHub Actions, and understand timing, consistency, and failure modes.

By the ScreenshotNeo team4 October 20269 min read

To schedule website screenshots with a custom user agent, set the user agent when creating a Playwright browser context, navigate to the page, save a screenshot, and run that script on a schedule. GitHub Actions can run it from a cron schedule. Its scheduled runs can be delayed under load, so it is suitable for periodic monitoring but not an exact-time guarantee.

This guide uses Node.js, Playwright, and GitHub Actions. The same separation applies with another scheduler: configure the browser identity before navigation, capture the page, then invoke the script on the interval you need.

1. Create a screenshot script with Playwright

Install Playwright and its Chromium browser in a project:

npm init -y
npm install --save-dev playwright
npx playwright install chromium

Save the following as capture.mjs. It reads the target URL and user agent from environment variables, creates the context with that user agent, navigates, and writes a full-page PNG. The explicit checks make configuration mistakes fail early.

import { chromium } from 'playwright';
import { mkdir } from 'node:fs/promises';

const targetUrl = process.env.TARGET_URL;
const userAgent = process.env.SCREENSHOT_USER_AGENT;

if (!targetUrl) throw new Error('Set TARGET_URL');
if (!userAgent) throw new Error('Set SCREENSHOT_USER_AGENT');

let parsedUrl;
try {
  parsedUrl = new URL(targetUrl);
} catch {
  throw new Error('TARGET_URL must be an absolute URL');
}
if (!['http:', 'https:'].includes(parsedUrl.protocol)) {
  throw new Error('TARGET_URL must use http or https');
}

await mkdir('artifacts', { recursive: true });
const browser = await chromium.launch({ headless: true });
try {
  const context = await browser.newContext({
    userAgent,
    viewport: { width: 1440, height: 1000 },
    deviceScaleFactor: 1
  });
  try {
    const page = await context.newPage();
    page.setDefaultNavigationTimeout(60_000);
    await page.goto(targetUrl, { waitUntil: 'networkidle' });
    await page.screenshot({ path: 'artifacts/page.png', fullPage: true });
  } finally {
    await context.close();
  }
} finally {
  await browser.close();
}

Playwright documents userAgent as the specific user agent to use in a browser context. Setting it at context creation means it applies to pages opened in that context, including the initial navigation. The capture sequence follows Playwright’s documented context, page, navigation, and screenshot APIs: Browser API and Page API.

Run it locally by supplying a real URL and the exact user agent string you intend to use:

TARGET_URL='https://example.com' \
SCREENSHOT_USER_AGENT='Mozilla/5.0 (compatible; ScreenshotMonitor/1.0)' \
node capture.mjs

The sample string is illustrative. If the target site expects a particular browser identity, use a valid string appropriate to your use case. A custom user agent changes the browser’s reported identity; it does not make a browser identical to another device or guarantee access to a site.

Choose the page-ready condition deliberately

The example waits for networkidle, which is useful for pages that settle after loading but may never occur on pages with persistent requests. No single wait condition fits every site. Alternatives include:

  • domcontentloaded when the initial document is enough and you will wait separately for a specific element.
  • load when the page’s load event is a suitable signal, though client-rendered content may continue afterward.
  • networkidle when the page becomes quiet after its requests complete; analytics, polling, or streaming can prevent it.
  • A targeted selector wait such as await page.locator('main article').waitFor() when a known content element marks readiness.

For pages with delayed content, wait for a meaningful selector or a short, measured delay after navigation. For animated pages, disable animations in the screenshot call if a stable visual state matters. Do not treat a successful navigation event as proof that every image, font, or application widget has finished rendering.

2. Schedule the script with GitHub Actions

Commit capture.mjs, package.json, and the lockfile to a repository. Add a workflow such as .github/workflows/screenshot.yml:

name: Scheduled website screenshot

on:
  schedule:
    - cron: '17 */6 * * *'
  workflow_dispatch:

jobs:
  capture:
    runs-on: ubuntu-latest
    timeout-minutes: 10
    env:
      TARGET_URL: ${{ vars.TARGET_URL }}
      SCREENSHOT_USER_AGENT: ${{ vars.SCREENSHOT_USER_AGENT }}
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: 22
          cache: npm
      - run: npm ci
      - run: npx playwright install --with-deps chromium
      - run: node capture.mjs
      - uses: actions/upload-artifact@v4
        with:
          name: website-screenshot
          path: artifacts/page.png
          retention-days: 14

Set TARGET_URL and SCREENSHOT_USER_AGENT as repository variables under the repository’s Actions settings. If the target needs credentials, store them as Actions secrets instead of putting them in the workflow or source code, and read them through the secrets context. The sample target is treated as public; authentication needs an intentional, secure setup.

The cron expression above means minute 17 every six hours. GitHub Actions interprets scheduled times in UTC by default. GitHub also supports an IANA timezone in a schedule. The shortest supported interval is five minutes. Scheduled workflows run from the default branch, and public-repository scheduled workflows are automatically disabled after 60 days without repository activity. See GitHub’s schedule event documentation.

GitHub cautions that scheduled runs can be delayed during high load and queued jobs can be dropped. The start of an hour can be especially busy. Choosing a minute away from 00 can avoid that peak, but does not create an exact-time guarantee. If a capture must happen at a precise moment, choose a scheduler with an appropriate service guarantee and verify its behavior.

Adjusting the schedule

Goal Example cron Meaning (UTC unless timezone configured)
Every day at 08:17 17 8 * * * Once daily
Every hour at minute 17 17 * * * * Hourly
Every six hours at minute 17 17 */6 * * * Four times daily
Every five minutes */5 * * * * Shortest documented interval

Check the cron expression in the workflow before relying on it. Consider daylight saving changes if you use a timezone other than UTC, and decide whether the schedule should follow local clock time or a fixed UTC interval.

3. Save, retain, and compare captures

The workflow uploads the image as a GitHub Actions artifact for 14 days. Change retention-days to fit your review window, subject to the limits of your repository plan. Artifacts are convenient for inspection but are not a permanent image archive. For long-term retention, upload the file to storage you control, apply an explicit retention policy, and avoid exposing screenshots that contain private page data.

If you are comparing screenshots over time, keep the capture environment consistent. Playwright notes that output can vary with operating system, browser version, settings, hardware, power source, and headless mode. Pin the Node and browser versions where practical, use the same runner image, viewport, and device scale factor, and avoid silently changing dependencies. Dynamic content such as timestamps, rotating banners, ads, personalized data, and animations may still differ even on an identical runner. See Playwright’s visual comparison guidance.

4. Useful capture options and edge cases

Need Playwright approach Consideration
Fixed viewport Set viewport in browser.newContext() Use the same dimensions each run for comparison.
Higher pixel density Set deviceScaleFactor, for example 2 Increases output dimensions and file size.
One element only page.locator('selector').screenshot({ path: 'artifacts/element.png' }) Wait for the element and handle cases where it is absent or hidden.
Full page page.screenshot({ fullPage: true }) Long pages can produce large images; lazy-loaded sections may need scrolling or a full-page capture strategy.
Hide volatile content Use screenshot options such as style or add a controlled stylesheet Keep the same masking rules for every comparison.
Page-specific identity Set userAgent in a context for that capture Use separate contexts for different identities so settings do not leak between jobs.

For multiple URLs, capture them sequentially or with a deliberate concurrency limit. Unbounded parallel browsers can exhaust runner memory and cause navigation failures. Give each output a unique name, such as a sanitized hostname and UTC timestamp, so scheduled runs do not overwrite files. Also decide how to handle redirects, authentication expiry, consent overlays, and pages that return an error document with an otherwise successful HTTP navigation.

5. Troubleshooting

Symptom Likely cause Fix
The site still shows the default browser variant The user agent was applied after navigation, or the site uses other signals too. Set userAgent in newContext() before creating or navigating the page. Remember that user agent alone does not emulate all device properties.
Navigation times out at networkidle The page keeps connections open or sends background requests. Use domcontentloaded or load, then wait for a specific content selector. Increase the timeout only if the page genuinely needs more time.
Screenshot is blank or missing content Capture happened before client rendering, fonts, images, or lazy content were ready. Wait for a meaningful locator, scroll relevant content into view, and inspect the page at the same viewport. Avoid assuming that a fixed delay solves every page.
GitHub does not run at the expected minute Scheduled Actions may be delayed by load; cron is UTC unless configured otherwise. Check the workflow’s schedule and timezone, choose a minute away from the top of the hour, and inspect the Actions run history.
No scheduled runs appear in a public repository The repository may have had no activity for 60 days, which disables scheduled workflows. Restore repository activity and re-enable or verify the workflow as needed in Actions.
Executable doesn't exist or browser launch fails Chromium was not installed in the runner environment or required system packages are missing. Run npx playwright install --with-deps chromium after installing dependencies.
Upload step says no files were found The capture failed, wrote elsewhere, or the artifact path differs. Check the preceding step logs and ensure the screenshot path matches artifacts/page.png.
Visual diffs change between runs Browser, OS, fonts, timing, dynamic content, or viewport changed. Pin the environment and capture settings; wait for stable content and mask intentionally variable regions.

6. Performance, reliability, and cost

A run takes longer when it installs browser dependencies, opens more pages, waits for long network activity, or captures very tall pages. Caching package downloads can reduce setup time, but the browser version still needs to match the installed Playwright package. For frequent schedules or many URLs, estimate work as runs per day multiplied by URLs per run, then account for retries and manual runs. Set a job timeout and use bounded concurrency so a slow target does not consume the entire runner indefinitely.

Reliability depends on both the scheduler and the page. A successful schedule trigger does not ensure a successful capture: DNS, TLS, rate limits, bot checks, authentication, and site changes can all interfere. Preserve logs, make failed runs visible through Actions notifications, and decide whether to retry transient failures. Retries should be bounded to avoid making a site outage or rate limit worse. GitHub’s documented schedule behavior does not promise exact start times.

GitHub Actions has plan-specific usage and storage limits; consult the current Actions billing documentation for the repository’s applicable rates and included minutes. Browser setup and screenshot work consume runner time, while retaining many artifacts consumes storage. Costs therefore depend on frequency, duration, concurrency, and retention. The self-managed route also requires maintaining the workflow and browser environment.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. You can keep the schedule in GitHub Actions or another scheduler and call the API on each run. See the ScreenshotNeo API documentation for request options, including custom user agent settings.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));

Keep the access key in a scheduler secret, not in a checked-in workflow file. ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, timeouts, and failed loads are not billed; and an MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000.

Sign up free for ScreenshotNeo to get 1,000 screenshots per month with no card.

FAQ

Does a custom user agent make the screenshot look like a phone?

No. It changes the reported user agent string. Set a mobile viewport and device scale factor as needed; a complete mobile emulation can also involve touch and other browser context settings.

Can GitHub Actions run every minute?

No. GitHub documents five minutes as the shortest interval for scheduled workflows.

Will the same scheduled capture always produce identical pixels?

No. Environment differences and changing page content can alter the image. Keep the browser environment and capture settings consistent, and control or mask content that changes by design.

Can I use the same script for authenticated pages?

Yes, if you provide credentials securely and the site permits automated access. Avoid saving sensitive screenshots as broadly accessible artifacts, and use short-lived credentials where possible.