ScreenshotNeo

BlogHow-to

How to Take Recurring Screenshots of a Website After JavaScript Finishes Loading

Use Playwright to wait for the page content you need, capture it, and schedule repeat runs. Includes GitHub Actions, failure handling, and a one-call API option.

By the ScreenshotNeo team4 October 202610 min read

To take recurring screenshots after JavaScript finishes loading, run a browser automation script on a schedule and make the script wait for a page-specific signal that the content you care about is ready. The scheduler decides when a capture starts; the browser wait decides when to capture during each visit. For most pages, Playwright plus a visible target element is a practical starting point.

1. Install Playwright and choose a readiness signal

A browser’s load or DOMContentLoaded event does not necessarily mean a JavaScript application has finished updating the content you want. Choose a condition tied to that content: a chart container becoming visible, a status element appearing, or expected text being rendered.

The example below waits for an element marked data-ready="true". Replace that locator with one that reflects the target page. The script uses a finite timeout, creates a unique filename for each run, and closes the browser even if navigation or waiting fails.

const { chromium } = require('playwright');
const { mkdir } = require('node:fs/promises');

(async () => {
  const url = process.env.TARGET_URL || 'https://example.com';
  const outputDir = process.env.OUTPUT_DIR || 'captures';
  const readySelector = process.env.READY_SELECTOR || '[data-ready="true"]';

  await mkdir(outputDir, { recursive: true });
  const browser = await chromium.launch({ headless: true });

  try {
    const page = await browser.newPage({
      viewport: { width: 1440, height: 900 },
      deviceScaleFactor: 1
    });
    page.setDefaultTimeout(15_000);

    await page.goto(url, {
      waitUntil: 'domcontentloaded',
      timeout: 30_000
    });
    await page.locator(readySelector).waitFor({ state: 'visible' });

    const stamp = new Date().toISOString().replaceAll(':', '-');
    await page.screenshot({
      path: `${outputDir}/capture-${stamp}.png`,
      fullPage: true
    });
  } finally {
    await browser.close();
  }
})().catch(error => {
  console.error('Screenshot run failed:', error);
  process.exitCode = 1;
});

Install the Playwright package and browser in the environment where this script runs. For example, with npm, run npm install playwright and npx playwright install chromium, then save the script as capture.js and run it with node capture.js. Pass TARGET_URL, READY_SELECTOR, and OUTPUT_DIR as environment variables to configure a run without editing the script.

Other useful readiness conditions

  • Visible element: page.locator('.report-ready').waitFor({ state: 'visible' }) waits until the selected element is visible.
  • Expected text: page.getByText('Updated').waitFor({ state: 'visible' }) can work when the page exposes reliable text that appears after rendering.
  • Page-specific predicate: await page.waitForFunction(() => window.myApp?.reportLoaded === true) waits for a condition in the page. Use this only when the page exposes a suitable state.

Use a selector or predicate that distinguishes completed content from a placeholder, spinner, or stale value. If the site has no single ready marker, wait for the specific content elements needed in the screenshot, or add a short, bounded delay after a stronger readiness signal.

2. Understand load states and timeouts

The sample navigates at domcontentloaded and then waits for the target element. This keeps navigation and application readiness as separate steps. You can use another documented navigation load state when it suits the page, but do not assume a generic browser event proves that asynchronously rendered content is ready.

Playwright defines networkidle as no network connections for at least 500 ms and marks it as discouraged for readiness checks. Pages can maintain open connections or fetch data later, while network silence does not prove that the exact content you want is present. Prefer a page-specific locator or condition. Use Playwright’s Page API documentation for navigation, waiting, timeout, and screenshot options.

Give both navigation and readiness waits finite timeouts. When the expected signal does not arrive, treat the run as failed: log the URL and error, exit with a failure status, and notify or retry according to your workflow. Do not silently save a blank or partially rendered capture as if it were valid.

3. Choose the screenshot shape and output

page.screenshot({ path: 'capture.png' }) captures the viewport. Add fullPage: true to capture the full scrollable page. Full-page output can be very tall and may take longer or use more memory; use viewport capture if the question you are monitoring only concerns above-the-fold content.

  • Unique filenames: Include a timestamp or run identifier if you need a history. Otherwise, repeated runs can overwrite the same path.
  • Format: Playwright supports screenshot output types including PNG, JPEG, and WebP; select a supported type and matching filename for your storage and review workflow.
  • Consistent rendering: Keep viewport, device scale factor, browser version, locale, timezone, and other relevant context settings consistent across runs when comparing images. Different environments or page content can still produce visual differences.
  • Retention: Decide where captures should live after the process exits. A local directory on a disposable runner is not long-term storage; upload or copy files to storage your workflow retains.

4. Schedule the script with GitHub Actions

A scheduled workflow can launch the same script repeatedly. This example runs daily at 08:17 UTC. GitHub Actions schedule expressions use POSIX cron syntax, default to UTC, and have a documented shortest interval of five minutes. Scheduled events can be delayed during periods of high load, so the cron time is not an exact-time guarantee. Choosing a minute away from the top of the hour can help avoid some congestion.

name: Website screenshot

on:
  schedule:
    - cron: '17 8 * * *'
  workflow_dispatch:

jobs:
  capture:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: 22
      - run: npm install playwright
      - run: npx playwright install --with-deps chromium
      - run: node capture.js
        env:
          TARGET_URL: https://example.com
          READY_SELECTOR: '[data-ready="true"]'
      - uses: actions/upload-artifact@v4
        if: always()
        with:
          name: website-capture
          path: captures/
          if-no-files-found: ignore

Commit capture.js and the workflow file under .github/workflows/. The workflow installs the browser, runs the script, and uploads any output as a workflow artifact. Set the scheduled workflow up on the repository’s default branch, since scheduled runs use the latest commit on that branch. For a private or authenticated page, store credentials in repository secrets and pass them as environment variables; do not put passwords, tokens, or cookies in committed code.

If you need a different cadence, edit the cron expression. For example, 0 */6 * * * requests a run every six hours. Convert the intended local time to UTC unless you have configured a supported timezone for the workflow, and account for daylight-saving changes if the capture must follow local clock time. Use GitHub’s schedule event documentation for the current syntax and platform behavior.

5. Run on a machine you manage instead

You can use the same script with a scheduler on a machine where Node.js and Chromium are installed. This can be useful when the target is only reachable from a private network or when you already manage a persistent capture host. Configure the scheduler to run from the project directory, provide the needed environment variables and secrets, and direct output to retained storage. Add logging and failure alerts in the surrounding job runner; the Playwright script returns a nonzero exit status on failure.

Choose the scheduler based on where the browser can reach the site, how credentials are supplied, how precisely the timing matters, where images should be retained, and how you will learn about failures. A hosted workflow is convenient for repository-based jobs; a managed machine gives you control over its network and environment. Neither choice makes the browser wait condition unnecessary.

6. Troubleshoot common failures

Symptom Likely cause Fix
Screenshot is blank or shows a spinner The script waits for a generic load event, or the chosen selector appears before the real content is ready. Wait for the rendered result or a page-specific ready state. Check that the selector identifies final content, not a shell or placeholder.
Timeout waiting for selector The selector is wrong, the element is hidden, the page failed to load data, or the application changed. Inspect the page and update the selector. Check navigation and console errors, and keep a finite timeout so failure is visible.
Navigation timeout The site is slow, unreachable from the runner, or keeps loading resources while the selected navigation condition is too strict. Choose an appropriate navigation state such as domcontentloaded, set a sensible finite navigation timeout, then wait separately for the target content.
It hangs on networkidle Long polling, streaming, analytics, or other requests keep network activity present; network quiet is also not a semantic readiness signal. Replace the network-idle wait with a locator, expected text, or page predicate for the content you need.
Every run replaces the prior image The script writes to a fixed path. Include a timestamp or unique run ID in the output filename, then apply a retention policy to control storage growth.
No image remains after a CI run The runner’s filesystem is temporary or the workflow did not upload the file. Upload the output as an artifact or copy it to storage with the retention you need. Confirm the output directory and artifact path match.
Different-looking images across runs Viewport, browser, device scale, locale, timezone, fonts, dynamic content, or site state changed. Keep the rendering environment and context stable, and account for legitimate page changes such as timestamps or rotating content.
Scheduled run starts late or does not appear exact Scheduled workflow events can be delayed under load; cron is a trigger, not an exact-time capture guarantee. Choose a minute away from the hour boundary where practical. If strict timing is required, evaluate a scheduler and environment appropriate to that requirement.
Works locally but fails in automation The runner lacks the browser dependencies, environment variables, network access, or credentials used locally. Install Chromium and its dependencies in the job, configure secrets explicitly, and verify the runner can reach the URL.

7. Performance, reliability, and cost

Capture frequency and capture duration are different. The scheduler controls how often a browser starts; navigation, readiness waits, rendering, and image encoding determine how long each run takes. Full-page screenshots and pages with heavy assets can increase runtime and output size. Use the smallest viewport or capture region that meets the monitoring need, and avoid an unnecessarily long fixed sleep when a specific readiness condition is available.

For reliability, make failures observable, use finite timeouts, and preserve enough run context to diagnose them: target URL, timestamp, error, and, when appropriate, a log or failure screenshot. Decide whether a transient error should trigger a bounded retry. Retries should not turn a broken readiness condition into an apparently successful capture. Keep credentials in secrets, and consider that capturing a page may expose personal or private data in stored images.

Costs depend on where the browser runs, execution frequency, storage and retention, and any hosted infrastructure you choose. The research sources do not provide a cost comparison among schedulers or browser hosts. Estimate monthly runs from the cadence, then include expected retries and the storage for retained images.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. Make one request for a URL to receive an image or PDF, then schedule that request at the cadence you need. Its capture flow accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each cleanup step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and the response headers say the page verdict and billing status. AI agents can use its MCP server tools: take_screenshot, get_page_info, and capture_pdf.

For example, this cURL request saves a WebP screenshot. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

Python equivalent:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

Node.js equivalent:

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));

To make captures recurring, run one of these calls from the scheduler you already use. ScreenshotNeo includes full-page capture with lazy images loaded, CSS selector element capture, dark mode, device presets and custom viewport, retina scale, custom CSS and JavaScript, click-before-capture, hide selectors, wait conditions, request blocking, custom headers and cookies, timezone and geolocation, image resizing, chosen cache TTL, async jobs with signed webhooks, bulk capture for up to 100 URLs per call, usage API, and an OpenAPI spec. It also supports PDF settings, HTML/CSS input, transparent backgrounds, signed links for public image tags, and parameter names used by other screenshot APIs.

One thousand screenshots per month are free with no card. Paid plans start at $5 for 3,000 screenshots; higher plans are $15 for 15,000, $39 for 60,000, $99 for 250,000, and $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan. Sign up for 1,000 free screenshots a month with no card.

FAQ

Can I schedule screenshots more frequently than once a day?

Yes. GitHub Actions documents a five-minute shortest scheduled interval. A locally managed scheduler may offer different timing controls, depending on its configuration.

Should I wait for the whole page to stop making requests?

Usually not. Network silence is not proof that the content of interest is ready, and some pages keep requests open. Wait for the content itself.

Will the screenshot be identical every time?

Not necessarily. The site may change, and rendering conditions or dynamic content can vary. Keep the browser context stable and account for expected changes.

Can a scheduled capture access a page behind a login?

It can if the browser run is given valid authentication and can reach the page. Supply secrets securely and avoid storing sensitive screenshots where unintended users can access them.