ScreenshotNeo

BlogHow-to

How to schedule bulk screenshots of Indian websites with cron

Use cron and Playwright to capture a list of Indian websites on a reliable schedule, with clear timezone, setup, logging, and troubleshooting guidance.

By the ScreenshotNeo team4 October 20269 min read

Use cron to start a browser automation script on a schedule; use Playwright to visit each URL and save its screenshot. For reliable recurring captures, make the timezone, runtime, browser dependencies, output paths, and page readiness explicit. This guide uses Playwright with Node.js on Linux and shows how to run one URL per line from a file.

1. Decide what each screenshot should capture

Make a list of authorized URLs and decide whether each image should show a fixed viewport or the full scrollable page. A viewport screenshot gives a consistent frame and usually a smaller image. A full-page screenshot includes content below the fold but can be very tall. Playwright supports both modes through its screenshot API. See the Playwright screenshot guide and Page API.

For a repeatable run, also decide:

  • How often to capture and which timezone defines the schedule.
  • Where to save output and logs, and how long to retain them.
  • What counts as ready for each site: a particular selector, a delay, or a navigation state.
  • Whether a failed URL should stop the batch or be recorded while the remaining URLs continue.

2. Install Playwright and its browser

Use a dedicated project directory and install Playwright there. Its browser binaries are tied to the Playwright release, and Linux may require system dependencies. After changing the Playwright version, run the browser installation workflow again. Consult the official Playwright browser installation documentation for the commands and dependencies for your operating system.

mkdir -p /opt/site-captures
cd /opt/site-captures
npm init -y
npm install playwright
npx playwright install chromium

On Linux, use the documented dependency-install command if Chromium reports missing shared libraries. Run installation as the same account that will own the scheduled job, or ensure that account can read the installed browser and dependencies.

3. Create the URL inventory

Put one URL on each line. Keep the file readable by the cron account. The script below ignores blank lines and lines beginning with #.

mkdir -p /opt/site-captures/urls /opt/site-captures/output /opt/site-captures/logs
cat > /opt/site-captures/urls/sites.txt <<'EOF'
https://example.com/
https://www.example.org/
EOF

Replace the sample entries with sites you are authorized to access. Avoid putting credentials or private tokens in a world-readable inventory.

4. Write the bulk capture script

This script launches Chromium once, then processes URLs sequentially to keep resource use predictable. It saves deterministic files under a date-and-run directory, continues after an individual URL fails, and exits nonzero if any capture failed. It defaults to viewport screenshots; set FULL_PAGE=1 in the environment for full-page images.

cat > /opt/site-captures/capture.mjs <<'EOF'
import { chromium } from 'playwright';
import { mkdir, readFile } from 'node:fs/promises';
import path from 'node:path';

const listPath = process.env.URL_LIST ?? '/opt/site-captures/urls/sites.txt';
const outputRoot = process.env.OUTPUT_DIR ?? '/opt/site-captures/output';
const fullPage = process.env.FULL_PAGE === '1';
const timeoutMs = Number(process.env.PAGE_TIMEOUT_MS ?? 45000);
const runId = new Date().toISOString().replaceAll(':', '-');
const outputDir = path.join(outputRoot, runId);

const lines = (await readFile(listPath, 'utf8')).split(/\r?\n/);
const urls = lines.map(line => line.trim()).filter(line => line && !line.startsWith('#'));
if (urls.length === 0) {
  console.error(`No URLs found in ${listPath}`);
  process.exit(2);
}
await mkdir(outputDir, { recursive: true });

const safeName = value => {
  const url = new URL(value);
  return `${url.hostname.replace(/[^a-zA-Z0-9.-]/g, '_')}${url.pathname
    .replace(/[^a-zA-Z0-9.-]+/g, '_').replace(/^_+|_+$/g, '')}`.slice(0, 150) || 'page';
};

const browser = await chromium.launch({ headless: true });
let failures = 0;
try {
  const page = await browser.newPage({ viewport: { width: 1440, height: 1000 } });
  page.setDefaultNavigationTimeout(timeoutMs);
  for (const url of urls) {
    try {
      // domcontentloaded avoids waiting indefinitely for analytics or long-lived requests.
      await page.goto(url, { waitUntil: 'domcontentloaded', timeout: timeoutMs });
      // Add a site-specific locator wait here when the page renders key content later.
      const filename = `${safeName(url)}.png`;
      const destination = path.join(outputDir, filename);
      await page.screenshot({ path: destination, fullPage });
      console.log(`OK ${url} -> ${destination}`);
    } catch (error) {
      failures += 1;
      console.error(`FAIL ${url}: ${error instanceof Error ? error.message : String(error)}`);
    }
  }
} finally {
  await browser.close();
}
console.log(`Finished ${urls.length} URL(s); ${failures} failed; output=${outputDir}`);
if (failures > 0) process.exitCode = 1;
EOF

Run it manually from its project directory before scheduling:

cd /opt/site-captures
node /opt/site-captures/capture.mjs

To capture full pages for a run, use FULL_PAGE=1 node /opt/site-captures/capture.mjs. To choose a different inventory or output directory, set URL_LIST or OUTPUT_DIR. The script names files by hostname and path; if your inventory has distinct query-string variants of the same path, adapt the filename function to include a sanitized query or unique inventory identifier to avoid collisions.

Choose a page readiness condition

domcontentloaded means the document has been parsed, but client-rendered text, charts, and lazy images may not yet be ready. Use a condition tied to what the screenshot must show. For a site with a stable content selector, after navigation add:

await page.locator('main article').waitFor({ state: 'visible', timeout: timeoutMs });

For a brief animation or delayed widget, a bounded delay can help, but it is less reliable than waiting for a meaningful selector. Full-page capture does not guarantee that every lazy-loaded asset has completed; scroll or wait for the specific content the capture requires. Readiness is site-specific, so record failures instead of silently saving an image that looks complete but is missing content.

5. Test the exact scheduled environment

Cron runs commands as the owner of that crontab and uses the configured shell context. An interactive terminal may inherit a different working directory, PATH, environment variables, and permissions. Use absolute paths and test as the intended account:

sudo -u CAPTURE_USER /usr/bin/node /opt/site-captures/capture.mjs

Replace CAPTURE_USER with the account name. If that account owns the crontab, run the command directly as that user instead. Confirm that it can read the URL list and project dependencies and write to the output and log directories.

6. Add the cron schedule

A crontab entry has five time and date fields followed by the command. For a daily run at 06:30 on a host whose cron daemon uses Indian Standard Time, add:

30 6 * * * /usr/bin/node /opt/site-captures/capture.mjs >> /opt/site-captures/logs/capture.log 2>&1

Edit the crontab for the account that should run the job:

crontab -e

Examples of the five schedule fields:

Schedule Fields
Every day at 06:30 30 6 * * *
Every Monday at 06:30 30 6 * * 1
At minute 0 every 6 hours 0 */6 * * *
At 06:30 on the first day of each month 30 6 1 * *

Verify the host’s actual timezone and cron implementation. The Cronie manual documents CRON_TZ, but this is not a portable guarantee across all cron variants. Where supported, a Cronie crontab can set CRON_TZ=Asia/Kolkata before the entries. Check the installed manual rather than assuming that variable works everywhere. The Cronie manual also notes that when both day-of-month and day-of-week are restricted, either matching field can trigger the job. See crontab(5); it documents crontab -T for syntax checks on implementations that support it.

Redirecting both output streams gives you a log to inspect. Rotate or periodically remove logs and old screenshot directories so they do not fill the disk. Cron may also deliver output by mail if configured, but a file log is easier to inspect on many servers.

7. Make recurring runs reliable

  • Bound execution time: use a navigation timeout and consider wrapping the cron command with the host’s timeout utility so a hung process cannot overlap future runs.
  • Keep concurrency modest: sequential capture is a safe starting point. Add limited parallelism only after measuring CPU, memory, site load, and failure rates on the actual host. Playwright’s CI guidance favors one worker for stability and reproducibility in CI, while noting that more parallel work can suit capable self-hosted systems; that is guidance rather than a universal limit for screenshot scripts. See Playwright CI guidance.
  • Keep versions controlled: pin the Playwright dependency in the project lockfile and install the matching browser binary. Browser version, operating system, settings, hardware, and headless mode can change rendering. Keep the capture environment consistent when comparing images. See Playwright visual comparison guidance.
  • Track outcomes: monitor the process exit code, log, expected file count, and output directory. The script returns a failure status if any URL fails.
  • Plan storage: estimate output volume from the number of URLs and schedule frequency, then set a retention policy. Full-page images can be much larger in pixel dimensions than viewport captures.
  • Respect site access: capture only pages you are authorized to access and follow applicable site terms, access controls, and published policies. Do not evade technical restrictions.

8. Troubleshoot common failures

Symptom Likely cause Fix
Works in a terminal but cron produces no files Different user, working directory, environment, or relative paths Use absolute paths, log stdout and stderr, and reproduce the command as the crontab owner.
Executable doesn't exist or Chromium launch fails Playwright browser was not installed for this version or is unavailable to the cron user Run the documented browser install for the installed Playwright release as the job account; check browser directory permissions.
Missing shared library / browser process exits Linux system dependencies are missing Install the dependencies listed by Playwright for the host distribution and retry as the job user.
Permission denied writing screenshots or logs Output directory ownership or permissions do not include the cron account Create the directories and assign write access to the intended account; verify with a manual run as that account.
Screenshot is blank or missing application content Capture happens before client rendering or content loading completes Wait for a site-specific visible selector or a bounded delay that matches the page; record timeout errors.
Some URLs fail while others save Individual navigation errors, timeouts, redirects, or site availability Inspect the URL-specific log line, check reachability from the host, and adjust a bounded timeout where appropriate. The script continues through the list and returns nonzero if any failed.
Job runs at the wrong local time Cron daemon timezone differs from expectation, or the implementation handles timezone settings differently Check system timezone and installed cron documentation; use supported timezone configuration and verify an actual scheduled run.
Different images from one run to another Site content changes or browser/host rendering differs Keep the browser and host environment consistent and wait for the same readiness condition; dynamic content can still vary.
Later run overwrites an earlier image Output names are not unique Keep the run ID or timestamp in the path, and include a unique URL identifier when paths or query variants collide.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. A single request captures a URL, and bulk capture accepts up to 100 URLs per call. For recurring bulk work, your scheduler can call the API from a small script instead of maintaining a browser installation. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

Cookie banners are accepted like a visitor would accept them, and more than 60 known consent platforms, newsletter popups, and chat widgets are removed before capture; each step can be turned off. Bot checks, blank pages, and failed loads are never billed, and response headers say the page verdict and whether it was billed. An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up free for 1,000 screenshots a month, with no card required.

FAQ

Does cron itself take the screenshots?

No. Cron starts the command on schedule; the Playwright script performs navigation and capture.

Can I schedule screenshots at a time in India?

Yes, if the cron daemon’s timezone or supported per-crontab timezone setting matches the intended local time. Verify the implementation on the host.

Should I use full-page screenshots for every URL?

Only when below-the-fold content is part of the output you need. For fixed-frame monitoring or visual comparison, viewport capture may be a better fit.

Can I use this for pages behind a login?

The example does not implement login. Add an authorized authentication flow using securely managed credentials, and avoid storing secrets in a broadly readable crontab or URL list.

How many URLs can I process in one run?

The local script reads the whole inventory and processes it sequentially; practical limits depend on runtime, host resources, and the sites. Split very large inventories into manageable batches and monitor duration.