ScreenshotNeo

BlogHow-to

How to Run Recurring Website Screenshot Captures on a Low-Cost Indian VPS

Schedule reliable Playwright screenshots on an Indian VPS with Node.js, cron, logging, retention, and practical plan-selection guidance.

By the ScreenshotNeo team4 October 202611 min read

To run recurring website screenshots on a low-cost Indian VPS, install Node.js and Playwright on a supported Debian or Ubuntu server, schedule a small capture script with cron or a systemd timer, and keep logs, timeouts, and screenshot retention under control. This guide uses Node.js and Chromium. The cheapest advertised VPS plan is not automatically a reliable fit: browser resource use depends on the pages and workload, and the available provider listings are not independent benchmarks.

Use this for sites you own or are authorized to capture. Keep credentials out of source control, avoid sending excessive traffic to target sites, and restrict or rotate screenshots if they may contain personal or confidential information.

1. Choose a VPS based on the workload

Playwright documents Linux support for Debian 12 and 13 and Ubuntu 22.04, 24.04, and 26.04 on x86-64 or arm64. Choose a current supported distribution and check that the provider offers the architecture you intend to use. A vendor’s advertised memory and CPU do not prove that a particular capture cadence or page complexity will work reliably.

Compare plans using the full recurring cost and workload constraints, not the introductory monthly figure alone:

  • Memory and CPU: Find out how much memory is available to Chromium and whether CPU cores are shared. Begin with one browser and one page at a time.
  • Disk: Allow space for the OS, browser, logs, and the screenshots you retain. Full-page images can be large.
  • Transfer: Check monthly bandwidth limits, especially if images are downloaded or copied off-server.
  • Location and network: Pick a data center that suits your target websites and users. Location can affect network latency and page content.
  • Billing terms: Check currency, GST, promotional duration, renewal price, backups, restore terms, support, and whether the provider permits your intended workload.

Examples surfaced in provider listings include VPSWala advertising ₹298/month for 1 vCPU, 2 GB memory, 20 GB NVMe, and 100 GB transfer; EricHost advertising VPS Lite at ₹449/month for 1 vCPU, 2 GB RAM, 20 GB NVMe, a 100 Mbps port, and 300 GB traffic; and Hoststack advertising Linux VPS starting at ₹699/month on an overview page. These are vendor-published offers, not independently verified performance or availability. Promotions and renewal terms can differ, so confirm current specifications and total price directly with the provider. No plan here has been proven to be the cheapest reliable choice for your particular pages and schedule.

For a website screenshot API alternative that avoids maintaining a browser on a VPS, ScreenshotNeo returns a screenshot or PDF from one GET request. It removes known consent banners, newsletter popups, and chat widgets before capture; only clean shots are billed, with response headers indicating the page verdict and billing status. Its paid plans start at $5 for 3,000 shots/month, and 1,000 shots/month are free without a card. The browser setup below remains useful when you need to control the runtime and store captures directly on your server.

2. Install Node.js and Playwright

Connect to the server over SSH and install a current Node.js version using your distribution’s supported package source or a maintained Node.js installation method. Verify that both Node.js and npm are available:

node --version
npm --version

Create a project directory, initialize npm, and install Playwright. Commit the generated package lockfile so deployments use a pinned dependency graph:

mkdir -p ~/site-captures
cd ~/site-captures
npm init -y
npm install playwright

Install Chromium and the Linux packages Playwright requires:

npx playwright install --with-deps chromium

Run this installation step as a user with the required privileges. If you update the Playwright package later, check whether its supported browser binary also needs to be reinstalled. Keep the package and browser installation aligned.

3. Write a capture script

This runnable example captures a list of authorized URLs sequentially, waits for a page-specific selector, applies timeouts, saves timestamped PNGs, and reports failures with a nonzero process exit. Replace the sample URLs and ready selectors with values for your sites. Save as capture.mjs:

import { chromium } from 'playwright';
import { mkdir } from 'node:fs/promises';
import path from 'node:path';

const targets = [
  { name: 'home', url: 'https://example.com', readySelector: 'body' },
  // Add authorized targets, for example:
  // { name: 'status', url: 'https://status.example.org', readySelector: 'main' },
];

const outputDir = path.resolve(process.env.OUTPUT_DIR ?? './screenshots');
const navigationTimeoutMs = Number(process.env.NAVIGATION_TIMEOUT_MS ?? 45000);
const readyTimeoutMs = Number(process.env.READY_TIMEOUT_MS ?? 15000);
const timestamp = new Date().toISOString().replaceAll(':', '-');

await mkdir(outputDir, { recursive: true });
const browser = await chromium.launch({ headless: true });
let failures = 0;

try {
  for (const target of targets) {
    const page = await browser.newPage({
      viewport: { width: 1440, height: 900 },
      deviceScaleFactor: 1,
      timezoneId: 'UTC',
    });
    page.setDefaultNavigationTimeout(navigationTimeoutMs);
    page.setDefaultTimeout(readyTimeoutMs);

    try {
      const response = await page.goto(target.url, { waitUntil: 'domcontentloaded' });
      if (!response) throw new Error('Navigation returned no main resource response');
      if (!response.ok()) throw new Error(`Main response status ${response.status()}`);
      await page.locator(target.readySelector).waitFor({ state: 'visible' });
      const filename = `${target.name}-${timestamp}.png`;
      await page.screenshot({ path: path.join(outputDir, filename), fullPage: false });
      console.log(`${new Date().toISOString()} OK ${target.url} -> ${filename}`);
    } catch (error) {
      failures += 1;
      console.error(`${new Date().toISOString()} FAIL ${target.url}: ${error.message}`);
    } finally {
      await page.close();
    }
  }
} finally {
  await browser.close();
}

if (failures > 0) process.exitCode = 1;

Run it once by hand before scheduling:

cd ~/site-captures
node capture.mjs

The script uses domcontentloaded so navigation does not wait for every image, analytics request, or long-lived connection. It then waits for a selector that represents useful page content. For an app that renders after an API call, choose a selector that appears after rendering. If the page is genuinely ready at DOM content load, body is a simple check, though it does not prove that client-side content has finished loading.

4. Configure what the screenshot captures

Playwright’s page.screenshot() saves a viewport image by default. Set fullPage: true to capture the full scrollable page. The official documentation describes that mode as a screenshot of the full scrollable page as if it could fit on a very tall screen. Use full-page capture only when needed; it can produce tall files and trigger additional content or lazy loading.

Need Playwright setting Tradeoff
Viewport image fullPage: false (default) Predictable size; content below the fold is omitted.
Whole scrollable page fullPage: true Captures beyond the viewport; images may be tall and page content can change during capture.
PNG type: 'png' (default) Lossless output, often larger than JPEG.
JPEG type: 'jpeg', quality: 80 Smaller lossy output; quality is a number from 0 to 100.
WebP type: 'webp', quality: 80 WebP image output; confirm downstream viewers accept it.
CSS-pixel scale scale: 'css' Output dimensions follow CSS pixels.
Device-pixel scale scale: 'device' (default) Output dimensions reflect device scale and may use more storage.

For example, replace the screenshot line with:

await page.screenshot({
  path: path.join(outputDir, filename.replace('.png', '.webp')),
  type: 'webp',
  quality: 82,
  fullPage: true,
  scale: 'css',
});

Other capture decisions affect consistency:

  • Viewport and device scale: Keep width, height, and deviceScaleFactor stable across runs.
  • Time zone and locale: Set them explicitly if the page formats dates or regional content. Keep them constant for comparisons.
  • Animations and changing data: If a page’s own motion or live content matters, capture it as-is. For visual comparisons, consider Playwright screenshot options such as animation disabling and caret hiding, or use a page stylesheet override to mask known changing regions.
  • Lazy-loaded content: A full-page screenshot does not guarantee every lazy image has loaded. Decide what readiness means for that site; scroll through the page and wait for the relevant images if they must appear.
  • Authentication: Use a protected storage-state file or environment-provided credentials where necessary. Restrict file permissions and never commit secrets or session state to source control.
  • One element: Use a locator’s screenshot method when only a component is needed, for example await page.locator('#report').screenshot({ path: 'report.png' }).

Playwright supports PNG, JPEG, and WebP screenshot formats and CSS-pixel or device-pixel scale. The exact options are documented in its screenshot guide and Page API.

5. Schedule daily captures

For a simple daily run, cron is sufficient. Edit the crontab for the same Unix account that installed the project and browser:

crontab -e

Add a schedule for 02:15 UTC each day, using absolute paths and logging both output streams:

15 2 * * * cd /home/ubuntu/site-captures && /usr/bin/node /home/ubuntu/site-captures/capture.mjs >> /home/ubuntu/site-captures/capture.log 2>&1

Replace /home/ubuntu and the Node.js path with the actual account and the result of command -v node. Cron uses the server’s time zone; check it with timedatectl, or schedule in the intended local time after confirming how that system handles the time zone.

For clearer service management, a systemd timer is another standard Linux option. Create ~/.config/systemd/user/site-captures.service:

[Unit]
Description=Capture authorized website screenshots

[Service]
Type=oneshot
WorkingDirectory=%h/site-captures
ExecStart=/usr/bin/node %h/site-captures/capture.mjs

Then create ~/.config/systemd/user/site-captures.timer:

[Unit]
Description=Run website screenshot captures daily

[Timer]
OnCalendar=*-*-* 02:15:00 UTC
Persistent=true

[Install]
WantedBy=timers.target

Enable and inspect the user timer:

systemctl --user daemon-reload
systemctl --user enable --now site-captures.timer
systemctl --user list-timers
journalctl --user -u site-captures.service

User services may stop when the user logs out unless lingering is enabled by the system administrator. If the timer must run without an interactive login, confirm the host’s systemd user-service policy or use a system service configured by an administrator.

6. Retain files and make failures visible

Timestamped filenames avoid silently overwriting yesterday’s capture. Pick a retention period that matches your storage budget and delete only files your job owns. For example, to remove PNG files older than 30 days from this dedicated output directory, schedule or run:

find /home/ubuntu/site-captures/screenshots -type f -name '*.png' -mtime +30 -delete

Before enabling deletion, verify the directory and pattern carefully. Check free disk space with df -h. Rotate or truncate logs periodically so the log file does not grow without limit.

The sample script exits nonzero if any target fails. Cron records no built-in alert, so inspect logs or connect the exit status to an alerting process you already operate. For larger jobs, write a separate status file or send a notification through an authorized monitoring system. Do not put credentials in command-line arguments that may be visible to other users; use restricted environment files or a secret store supported by the host.

7. Keep captures consistent and efficient

Start sequentially with one browser and one page at a time. Browser startup and rendering consume resources, and no verified benchmark here says how many pages a specific VPS can handle. Measure your own representative pages and watch memory, CPU, disk, network transfer, and run duration before increasing concurrency.

  • Set navigation and selector timeouts; avoid unbounded waits.
  • Use a stable browser version, operating system, viewport, locale, time zone, and capture settings when comparing images.
  • Choose a page-specific ready selector instead of relying on a fixed sleep. A fixed delay can waste time or still capture too early.
  • Use viewport captures unless the whole page is necessary; choose JPEG or WebP when smaller files are useful and acceptable.
  • Do not run overlapping jobs. If a capture can last longer than its interval, add a lock or use a scheduler policy that prevents concurrent runs.
  • Set a sensible cadence and respect the target site’s capacity. Repeated captures cause real page requests and may trigger site controls.

Pixel-identical results are not guaranteed across environments. Playwright identifies operating system, browser version, settings, hardware, and headless mode as sources of variation. Keep the environment stable for visual comparisons, and treat unexpected diffs as a signal to inspect the capture conditions as well as the site.

8. Troubleshooting

Symptom Likely cause Fix
Browser executable is missing The Playwright package is installed but its matching Chromium binary is not. From the project directory, run npx playwright install chromium; if system libraries are missing, use npx playwright install --with-deps chromium.
Shared library or browser launch error Linux dependencies are absent, or the browser and Playwright versions do not match. Install dependencies with Playwright’s CLI and reinstall the supported browser after package updates.
Works over SSH but fails in cron Cron has a smaller environment, different working directory, PATH, or user account. Use absolute paths, set the project working directory, and run the exact command as the scheduled user. Capture stdout and stderr to a log.
Navigation timeout The site is slow, a request hangs, or the chosen wait condition is too strict. Use a realistic timeout, navigate with domcontentloaded, then wait for a meaningful selector. Check server connectivity and target behavior.
Screenshot is blank or incomplete The app renders after navigation, the selector is too broad, or content is lazy-loaded. Wait for a visible page-specific element; for lazy content, scroll and wait for the needed images before capturing.
Runs sometimes overlap A prior capture has not finished before the next scheduled run. Reduce the target list or cadence, or add a lock so only one job runs at a time.
Disk fills up Images or logs accumulate without retention. Use a retention rule for the dedicated screenshot directory, rotate logs, and check disk capacity.
Images differ between days without a site change Dynamic content, fonts, browser or OS changes, hardware, or headless rendering can alter pixels. Pin the environment and capture settings, wait for stable content, and mask known dynamic regions for comparison.

Or skip the browser setup

ScreenshotNeo’s screenshot API takes a URL and returns an image or PDF without you installing and maintaining Chromium on the VPS. See the API documentation. One GET request can be called from cron as well; save the API key in a protected environment or secret file rather than in source control.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits cost nothing, and response headers identify the page verdict and billing status. An MCP server gives AI agents such as Claude and Cursor the tools take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.

Sign up free for 1,000 screenshots a month, with no card required.

FAQ

Can I capture pages that require login?

Yes, if you are authorized. Handle authentication through a protected browser storage state or a secure credential source, restrict file access, and keep session data out of version control. Confirm that the login flow is stable for unattended runs.

Should I use cron or a systemd timer?

Either can schedule a daily job. Cron is compact and widely available; a systemd timer integrates with service logs and can catch up a missed run with Persistent=true. Choose the mechanism you can inspect and maintain on your VPS.

Will the lowest advertised plan be enough?

There is no verified minimum server size for every page and capture cadence. Begin with a small sequential workload, measure resource use and completion times on representative pages, and adjust the plan or schedule based on those observations.

Can I compare screenshots pixel by pixel?

You can, but first stabilize the operating system, browser, viewport, device scale, fonts, time zone, and page state. Dynamic content and rendering differences can produce image changes unrelated to a meaningful site change.