ScreenshotNeo

BlogHow-to

Run Scheduled Website Screenshots with Node.js and cron

Use Playwright to capture a website with Node.js, then schedule the script with cron. Includes runnable code, deployment guidance, and troubleshooting.

By the ScreenshotNeo team4 October 202610 min read

Use Playwright to open the page and save a screenshot, then let cron launch that Node.js script at the times you choose. The script should use stable paths, wait for the page state you need, record failures, and close the browser so each scheduled run exits cleanly.

This guide uses Playwright with Chromium on a Linux host. Cron syntax and installation commands can vary by operating system, but the same design applies wherever cron and Node.js are available.

1. Create the Node.js screenshot project

Install Node.js, create a project directory, and install Playwright. Playwright’s browser binaries must also be available to the user account that will run the cron job.

mkdir -p "$HOME/website-screenshots"
cd "$HOME/website-screenshots"
npm init -y
npm install playwright
npx playwright install chromium

On a Linux server, Playwright may also need operating-system libraries for Chromium. If launch errors mention missing shared libraries, use Playwright’s documented Linux dependency setup for your distribution. Run the install as the same account that will execute the scheduled job, or ensure that account can access the installed browser.

2. Write a screenshot script that exits cleanly

Save this as capture.js. It accepts the target URL and optional output path as command-line arguments. It creates the output directory, sets a predictable viewport, waits for navigation, captures either the viewport or full page, and closes Chromium in a finally block. A nonzero exit code lets cron or a surrounding monitor detect failure.

const { chromium } = require('playwright');
const fs = require('node:fs/promises');
const path = require('node:path');

async function main() {
  const targetUrl = process.argv[2] || 'https://example.com';
  const outputPath = process.argv[3] || path.join(__dirname, 'screenshots', 'latest.png');
  const outputDir = path.dirname(outputPath);

  // Use a fixed viewport for repeatable captures.
  const width = Number(process.env.VIEWPORT_WIDTH || 1440);
  const height = Number(process.env.VIEWPORT_HEIGHT || 900);
  const fullPage = process.env.FULL_PAGE !== 'false';

  let browser;
  try {
    await fs.mkdir(outputDir, { recursive: true });
    browser = await chromium.launch({ headless: true });
    const context = await browser.newContext({
      viewport: { width, height },
      deviceScaleFactor: 1,
    });
    const page = await context.newPage();

    page.setDefaultNavigationTimeout(45_000);
    page.setDefaultTimeout(15_000);

    const response = await page.goto(targetUrl, { waitUntil: 'domcontentloaded' });
    if (response && response.status() >= 400) {
      throw new Error(`Navigation returned HTTP ${response.status()}`);
    }

    // Add a site-specific readiness wait here when DOMContentLoaded is too early.
    await page.screenshot({ path: outputPath, fullPage, type: 'png' });
    console.log(`Saved ${outputPath} (${new Date().toISOString()})`);
  } finally {
    if (browser) await browser.close();
  }
}

main().catch((error) => {
  console.error(`[${new Date().toISOString()}] Screenshot failed:`, error);
  process.exitCode = 1;
});

Run it once from the project directory before scheduling:

node capture.js https://example.com "$PWD/screenshots/example.png"

For a viewport-only capture, set FULL_PAGE=false. The defaults capture the full scrollable page at a 1440 by 900 CSS-pixel viewport. A full-page screenshot is useful for a page record; viewport mode is usually a better fit for checking a specific above-the-fold layout.

Wait for the content you need

DOMContentLoaded means the initial document has been parsed; it does not guarantee that every image, client-rendered component, font, or API-backed widget is ready. If the page has a stable landmark, wait for it explicitly:

await page.goto(targetUrl, { waitUntil: 'domcontentloaded' });
await page.locator('main .report-ready').waitFor({ state: 'visible' });
await page.screenshot({ path: outputPath, fullPage: true });

For a simple page where all network activity settles promptly, waitUntil: 'networkidle' is another option. Analytics, polling, ads, and chat connections can keep a page active, so a specific selector or a bounded delay may be more reliable. Avoid unbounded waits: a single page that never reaches the expected state can occupy the scheduled process indefinitely.

3. Choose capture settings for the artifact

Choice Use it when Trade-off
Viewport screenshot You need a consistent visible screen or visual check of a fixed region. Content below the viewport is omitted.
fullPage: true You need the full scrollable document. Very long pages can use more memory and create large files; sticky elements may render differently in a full-page capture.
PNG You need lossless output for archiving or pixel comparison. Files can be larger than lossy image formats.
JPEG or WebP Smaller image files matter more than lossless pixels. Compression can alter pixels and affect visual comparisons.
CSS pixel scale You want output dimensions tied to CSS viewport dimensions. Less dense output on high-density displays.
Device pixel scale You need a higher-density raster. More pixels increase memory use and file size.

Playwright supports screenshot options including output path, format, quality for lossy formats, clipping, and full-page capture. Use a fixed viewport and device scale for comparisons. Browser version, operating system, hardware, and headless configuration can still change rendering, so screenshots from different environments may not be pixel-identical.

Capture a region or element

For a fixed region, pass a clip rectangle in page coordinates. For a single element, use the locator screenshot API:

await page.screenshot({
  path: outputPath,
  clip: { x: 0, y: 0, width: 800, height: 500 },
});

await page.locator('[data-report-card]').screenshot({
  path: outputPath,
});

Ensure the target element exists and is visible before capturing it. If the element is below the fold, Playwright’s locator screenshot scrolls it into view. For pages that lazy-load images as they enter the viewport, a full-page screenshot is generally preferable to a manual viewport capture; if specific lazy content is essential, scroll it into view and wait for its images before taking the screenshot.

4. Schedule the script with cron

Cron starts a process according to a crontab schedule. It does not keep the Node.js script alive between runs, unlike a timer inside a long-running Node process. Use absolute paths because cron usually has a smaller environment and may start in a different working directory than your interactive shell.

Find the Node executable and project directory:

command -v node
pwd

Edit the current user’s crontab:

crontab -e

For example, run the capture every day at 06:30 according to the cron host’s configured timezone. Replace the example paths with the actual absolute paths from your machine:

30 6 * * * cd /home/alex/website-screenshots && /usr/bin/node /home/alex/website-screenshots/capture.js https://example.com /home/alex/website-screenshots/screenshots/latest.png >> /home/alex/website-screenshots/cron.log 2>&1

The five schedule fields are minute, hour, day of month, month, and day of week. This example runs at 06:30 every day. Check the host’s timezone and cron implementation when the local run time matters; do not assume the scheduler uses the timezone of your laptop or browser.

Schedule Crontab expression
Every 15 minutes */15 * * * *
At 08:00 every day 0 8 * * *
At 09:00 on weekdays 0 9 * * 1-5
At 02:15 on the first day of each month 15 2 1 * *

Some cron implementations treat daylight-saving transitions and timezone settings differently. Verify the behavior on the target host, especially for jobs that must run at a precise local time. If runs must not overlap, add a lock mechanism or use a scheduler that provides concurrency control; a slow page or stalled browser can otherwise still be running when the next interval begins.

5. Keep output, logs, and runs manageable

Keep a history instead of overwriting

The example writes latest.png, so each successful run replaces the prior capture. To retain history, include a timestamp in the output filename. Use an ISO-like UTC timestamp so names sort naturally:

const stamp = new Date().toISOString().replaceAll(':', '-');
const outputPath = path.join(__dirname, 'screenshots', `${stamp}.png`);

Decide how long to retain captures and logs, and ensure the destination has enough disk space. A full-page image, high device scale, frequent schedule, or many target URLs can grow storage quickly. Use a cleanup policy appropriate for the value of the archive; keep the latest image separately if another process expects a stable filename.

Limit concurrency and resource use

  • Close the browser in a finally block, and close contexts or pages if you keep a browser process open for several captures in one run.
  • Use bounded navigation and selector timeouts so a failing page does not consume a worker forever.
  • For many URLs, capture sequentially or use a small concurrency limit; launching a browser for every URL at once can exhaust memory and CPU.
  • Use viewport captures and CSS-pixel scale when a full-resolution full-page record is unnecessary.
  • Keep browser, Node.js, and operating-system versions stable when comparing captures. Changes to fonts or rendering libraries can alter pixels.

For deterministic visual checks, control dynamic content where possible, use the same browser and operating-system environment as the baseline, and inspect diffs rather than treating every changed pixel as a product defect. Playwright Test can capture repeatedly until consecutive screenshots match before comparing with an expectation; that feature is useful for visual tests, while a scheduled archival capture may simply need a timestamp and a record of the run.

6. Troubleshoot common cron and browser failures

Symptom Likely cause Fix
It works in a terminal but not from cron Cron has a limited PATH, a different working directory, or a different environment. Use absolute paths for Node, the script, and output; set the project working directory in the crontab command; log stdout and stderr.
node: command not found The Node installation is not on cron’s PATH, common with version managers. Use the absolute path from command -v node in the same account’s shell, or set an explicit PATH in the crontab.
Chromium executable missing Playwright’s browser was not installed, or it was installed for another account or browser cache directory. Run npx playwright install chromium as the scheduled user and confirm that user’s browser cache is accessible.
Browser fails to launch with missing library errors Required operating-system dependencies are absent. Install the Linux dependencies documented by Playwright for the host distribution.
Navigation times out The site is slow, unreachable, or keeps network activity open; the selected wait condition may be too strict. Check the URL from the host, set a suitable finite navigation timeout, and wait for a relevant selector instead of global network idle where appropriate.
Screenshot is blank or missing client content The page was captured before its app rendered, or the selector wait did not match the page. Wait for a visible application landmark or specific content, and inspect the page URL and console/log output.
Images are absent in a full-page shot Lazy-loaded images may not have been requested before the capture, or the source image failed. Scroll relevant sections into view, wait for image loading, and verify source availability; for important pages, add a page-specific readiness routine.
Output file cannot be written The path is relative to an unexpected directory, the directory does not exist, or the cron user lacks write access. Use an absolute destination, create its parent directory, and grant the scheduled user permission to write there.
Two captures run at once A prior run took longer than the schedule interval. Increase the interval, make the capture faster, or add a lock/concurrency control around the job.
Images differ between runs Content changed, animation or time-dependent data rendered, or the browser/OS environment differs. Stabilize the runtime and page state, disable or wait out relevant animations where appropriate, and review meaningful visual regions rather than assuming all pixel changes are regressions.

Retain a timestamp with each image and include errors in the log. If the screenshots are used for monitoring, periodically check that new files are being produced; a configured schedule alone does not prove that the page was captured successfully.

7. Alternatives to operating a browser host

A self-managed cron job gives you control over the browser environment, private network access, and where files are stored, but you maintain Node.js, Playwright, browser dependencies, disk capacity, logs, and scheduling. A hosted scheduled workflow can avoid maintaining a persistent machine, but its schedule semantics, runtime limits, storage, timezone behavior, and private-site access depend on the chosen platform and should be checked in that platform’s documentation. For pixel comparisons, use the same environment that produced the baseline either way.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. Instead of installing Chromium and maintaining a browser runtime, make one GET request for the URL. Its cookie/consent handling accepts the banner like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture.

For scheduled use, put the request in a Node.js script and run that script from cron as shown above. See the ScreenshotNeo API documentation for parameters and response details.

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo returned HTTP ${res.status}`);
await require('node:fs/promises').writeFile('screenshot.webp', Buffer.from(await res.arrayBuffer()));

ScreenshotNeo also supports full-page capture, element selectors, dark mode, device and viewport settings, PDF output, custom CSS and JavaScript, wait conditions, request blocking, headers and cookies, caching, async jobs, bulk capture, signed public image links, and a usage API. It accepts parameter names used by other screenshot APIs to make switching easier. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. The plans are Free (1,000), Starter ($5 for 3,000), Growth ($15 for 15,000), Pro ($39 for 60,000), Scale ($99 for 250,000), and Business ($249 for 1,000,000) per month; yearly billing gives two months free, and every feature is on every plan.

Start free with 1,000 screenshots a month and no card.

FAQ

Should I use Node.js timers instead of cron?

Use cron when the task can start, capture, and exit on each run. A Node.js timer is appropriate when a process is already meant to stay running; it will not run if that process stops.

Will a scheduled screenshot be identical every time?

Not necessarily. Dynamic page content and differences in browser, operating system, hardware, or headless mode can affect rendering. Keep the capture environment consistent and review visual differences in context.

Can cron capture a page that requires login?

Yes, if the browser session is configured with the required authentication, such as a controlled storage state or login flow. Protect credentials and session files with restrictive permissions, and avoid logging secrets.

Where do scheduled screenshots go?

They are written to the path passed to the script. Choose a persistent location if the job runs in a container or hosted workflow whose local filesystem may be temporary.