ScreenshotNeo

BlogHow-to

How to Capture Weekly Screenshots of Competitor Websites in Bulk

Build a repeatable weekly screenshot archive for competitor pages with Playwright, stable capture settings, and a scheduled bulk workflow.

By the ScreenshotNeo team4 October 202611 min read

To capture weekly screenshots of competitor websites in bulk, keep a list of exact page URLs, run a browser capture script over that list, save each run under a timestamped archive path, and schedule the script weekly. Playwright can capture either the visible viewport or the full scrollable page. The scheduler and storage are separate parts of a self-managed workflow.

This guide uses Node.js and Playwright. It also covers consistent rendering, safe retries, archive design, visual review, common failures, and a hosted option. Playwright’s screenshot and visual comparison documentation is available at Playwright screenshots and visual comparisons.

1. Decide what to capture

Start with the smallest useful watch list. Each row should represent one page whose visual changes matter; a competitor homepage and pricing page are separate entries. Keep exact URLs rather than just domain names, because pages can have distinct paths, query parameters, regions, or product states.

Field Purpose Example
competitor Groups pages belonging to one company Acme
page Stable human-readable page name Pricing
url Exact address to capture https://example.com/pricing
scope viewport or full full
notes Region, account state, or review reminder Public US page

Use viewport captures when the first screen is the subject, such as the hero area or navigation. Choose full-page captures when content below the fold matters, such as pricing tiers or feature lists. Full-page images can be much taller and larger, so avoid capturing full pages by default when a viewport is enough.

2. Create the URL list

Save the watch list as competitors.csv in the project directory. The example below contains public placeholder URLs; replace them with pages you are permitted to access and monitor.

competitor,page,url,scope,notes
Acme,Home,https://example.com/,viewport,Public landing page
Acme,Pricing,https://example.com/pricing,full,Public pricing page
Beta,Product,https://example.org/product,full,Public product overview

Keep URL edits reviewable in version control or another change-tracked location. Avoid putting passwords, session cookies, or private URLs into a CSV that could be exposed to other users of the job.

3. Install Playwright and prepare a capture script

Use a pinned project dependency and install its browser. Run the initial commands from a new project directory:

npm init -y
npm install playwright
npx playwright install chromium

Create capture.mjs. It reads the CSV, captures each page sequentially, and preserves each weekly run under a UTC timestamp. It writes a JSON manifest containing the URL, capture settings, status, and any error. Sequential capture limits load on the runner and makes failures easier to inspect; increase concurrency only after measuring the effect on your own environment.

import { chromium } from 'playwright';
import { readFile, mkdir, writeFile } from 'node:fs/promises';
import path from 'node:path';

const csvPath = process.env.WATCHLIST ?? 'competitors.csv';
const archiveRoot = process.env.ARCHIVE_DIR ?? 'archive';
const width = Number(process.env.VIEWPORT_WIDTH ?? 1440);
const height = Number(process.env.VIEWPORT_HEIGHT ?? 1000);
const timeoutMs = Number(process.env.PAGE_TIMEOUT_MS ?? 45000);
const settleMs = Number(process.env.SETTLE_MS ?? 1200);

function parseCsvLine(line) {
  // This deliberately supports simple CSV fields without embedded commas or quotes.
  // Use a full CSV parser if your names or notes need CSV escaping.
  return line.split(',').map(value => value.trim());
}

function safeName(value) {
  return value.toLowerCase().replace(/[^a-z0-9]+/g, '-').replace(/^-|-$/g, '') || 'page';
}

const lines = (await readFile(csvPath, 'utf8')).split(/\r?\n/).filter(line => line.trim());
if (lines.length < 2) throw new Error(`No URL rows found in ${csvPath}`);
const headers = parseCsvLine(lines[0]);
const rows = lines.slice(1).map(line => {
  const values = parseCsvLine(line);
  return Object.fromEntries(headers.map((header, i) => [header, values[i] ?? '']));
});
for (const [index, row] of rows.entries()) {
  if (!row.url || !/^https?:\/\//i.test(row.url)) {
    throw new Error(`Row ${index + 2}: url must be an absolute http(s) URL`);
  }
  if (row.scope && !['viewport', 'full'].includes(row.scope)) {
    throw new Error(`Row ${index + 2}: scope must be viewport or full`);
  }
}

const now = new Date();
const runId = now.toISOString().replaceAll(':', '-').replace(/\.\d{3}Z$/, 'Z');
const runDir = path.join(archiveRoot, runId);
await mkdir(runDir, { recursive: true });
const browser = await chromium.launch({ headless: true });
const manifest = [];

try {
  for (const row of rows) {
    const id = `${safeName(row.competitor)}-${safeName(row.page)}`;
    const page = await browser.newPage({
      viewport: { width, height },
      deviceScaleFactor: 1,
      locale: 'en-US',
      timezoneId: 'UTC'
    });
    const output = path.join(runDir, `${id}.png`);
    const record = {
      id, competitor: row.competitor, page: row.page, url: row.url,
      scope: row.scope || 'viewport', capturedAt: now.toISOString(),
      viewport: { width, height }, deviceScaleFactor: 1,
      locale: 'en-US', timezoneId: 'UTC', path: output
    };
    try {
      const response = await page.goto(row.url, {
        waitUntil: 'domcontentloaded', timeout: timeoutMs
      });
      // A short fixed settle interval allows ordinary client rendering to finish.
      // Use a longer site-specific wait or a selector when the page requires it.
      await page.waitForTimeout(settleMs);
      await page.screenshot({ path: output, fullPage: row.scope === 'full', animations: 'disabled' });
      record.httpStatus = response?.status() ?? null;
      record.status = response && response.status() >= 400 ? 'http-error-captured' : 'captured';
    } catch (error) {
      record.status = 'failed';
      record.error = String(error?.message ?? error);
    } finally {
      await page.close();
    }
    manifest.push(record);
    console.log(`${record.status}: ${row.url}`);
  }
} finally {
  await browser.close();
  await writeFile(path.join(runDir, 'manifest.json'), JSON.stringify(manifest, null, 2));
}

if (manifest.some(item => item.status === 'failed')) process.exitCode = 1;

The small CSV reader is intentionally limited: it expects plain fields without quoted commas or embedded newlines. For more complex watch lists, use a maintained CSV parser and preserve the same validation and manifest behavior.

4. Run it once and inspect the archive

Run the script manually before scheduling it:

node capture.mjs

Expected output resembles archive/2026-10-04T09-00-00Z/acme-pricing.png and a manifest.json beside the images. Confirm that the URL, visible content, viewport, and full-page choice are right. The script records HTTP error responses if a page returns an error status, because an error page can still be visually useful. A navigation timeout or browser failure is recorded as failed and makes the process exit with a nonzero status after the manifest is saved.

5. Schedule the capture weekly

Choose a stable runner that remains available at the scheduled time. Install the same Node.js version, project dependencies, and Playwright Chromium build there. A weekly schedule can be configured with cron, a CI scheduler, or the scheduler provided by your operating environment.

For a Unix-like machine with cron, find the absolute paths to Node.js and the project, then add a line like this with crontab -e to run every Monday at 09:00 UTC:

0 9 * * 1 cd /absolute/path/to/project && /absolute/path/to/node capture.mjs >> /absolute/path/to/project/capture.log 2>&1

Set the machine or scheduler timezone deliberately. Cron uses the host’s timezone unless configured otherwise; confirm the actual scheduled timezone in your environment. In a CI runner, use its weekly schedule syntax and make the archive persist to durable storage, since many hosted runners discard local files when a job ends.

  1. Run the job manually from the same environment used by the schedule.
  2. Check that the job has permission to launch Chromium and write to the archive location.
  3. Choose an archive retention policy that matches how long you need historical comparisons.
  4. Send job failures to the monitoring or notification mechanism you already use.
  5. Review at least one scheduled run to verify timezone, destination, and expected URLs.

6. Keep captures comparable over time

A visual history is useful only if capture conditions are reasonably consistent. Playwright notes that browser rendering can vary with the host operating system, browser version, settings, hardware, power source, and headless mode. Run captures in the same environment and keep the browser version, viewport, device scale, locale, timezone, and wait strategy stable. See the Playwright visual comparison guidance.

  • Browser and host: use the same Playwright and Chromium versions and execution image where practical.
  • Viewport and scale: use fixed dimensions and device scale factor. A responsive breakpoint can change the entire layout.
  • Locale and timezone: set these when dates, currency, or localized content could change.
  • Wait behavior: keep the same navigation and settling conditions. Prefer a meaningful ready selector for a site with delayed content.
  • Consent and personalization: cookie banners, A/B variants, logged-in state, and region can change screenshots. Record the conditions or mask volatile regions during review.
  • Animations: disabling animations reduces transient frames, but other dynamic content such as rotating promotions may remain.

Full-page capture records the scrollable page as a tall image. It does not guarantee that every lazy-loaded item appeared: some sites load content only as it approaches the viewport or after interaction. If below-the-fold content is missing, scroll through the page before capture or use a site-specific readiness procedure, then keep that behavior consistent between runs.

7. Organize history and review changes

Never overwrite the previous run if the purpose is to see change over time. Store immutable weekly run directories or use a date-based object storage prefix. Keep the manifest with each run so a reviewer can tell what was captured, when, with which URL and settings, and whether navigation returned an error.

Visual differences are a review aid, not proof of a meaningful competitor change. Ads, personalization, timestamps, rotating content, and browser rendering can all create noise. Compare like-for-like captures, investigate surprising differences in the page itself, and consider excluding unstable regions. Playwright’s visual comparison tooling supports applying a stylesheet to filter dynamic content when generating comparison screenshots; the same idea can be used in a custom review workflow.

For a very large archive, set retention deliberately and keep a separate backup if the images are business records. Image size depends on page length, dimensions, and format; full-page PNGs can consume substantially more storage than viewport captures. Avoid silently replacing old files as a space-saving measure, since that defeats the historical record.

8. Scale the weekly run safely

Sequential capture is simple and predictable. If a watch list grows, limited concurrency can reduce wall-clock time, but too many open pages can consume memory, overload the runner, or trigger site defenses. Add concurrency gradually, preserve per-URL timeouts, and keep a result record for every page even when other captures fail.

  • Use bounded concurrency rather than launching every URL at once.
  • Retry only transient failures, with a small retry count and delay; do not retry indefinitely.
  • Keep failures visible in the manifest and job exit status.
  • Separate a browser crash or network outage from a page that consistently fails.
  • Respect the target sites’ access rules and avoid capture rates that disrupt them.
  • Use stable identifiers for names; if two pages share a competitor and page name, add a unique key to the watch list.

Cost for a self-managed setup comes from the runner, storage, and any monitoring or orchestration you choose. The dossier does not establish comparable service pricing or performance figures. Estimate your own storage from a representative run and apply an explicit retention policy.

9. Hosted option: compare operational fit

A hosted capture service can package bulk URL intake, recurring schedules, an archive, exports, or alerts, but verify that its current behavior and limits fit your list. The research surfaced vendor pages for PeekShot and Site-Shot that advertise recurring or bulk workflows; these are vendor descriptions, not independent benchmarks. Check current retention, export, schedule, and pricing terms directly before relying on a service.

For a developer-oriented screenshot API, ScreenshotNeo offers a bulk capture endpoint workflow alongside its screenshot API and MCP server. Compare any hosted option on setup burden, schedule flexibility, archive and export behavior, alerting, capture controls, and current plan limits. This guide does not claim independent performance or reliability measurements.

Or skip the browser setup

For one URL, ScreenshotNeo returns an image or PDF from a GET request. Its API also supports bulk capture of up to 100 URLs per call; this call is the single-URL form. See the ScreenshotNeo API documentation for request parameters and bulk usage.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot; each cleanup step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots.

Sign up for ScreenshotNeo free: 1,000 screenshots a month, no card required.

Troubleshooting

Symptom Likely cause Fix
Executable doesn't exist or browser launch error Chromium was not installed in the runtime used by the scheduled job. Run npx playwright install chromium in that environment and ensure the scheduled job uses the project installation.
Navigation timeout The site is slow, stalled, or never reaches the chosen wait condition. Use domcontentloaded as in the example, adjust the timeout for your environment, or wait for a specific ready selector. Record timeouts rather than dropping the URL silently.
Screenshot is blank or incomplete Client rendering or lazy content had not completed when the image was captured. Wait for a stable page-specific selector; scroll through lazy sections before taking the image if needed.
Cookie dialog or popup dominates the image The site displays consent or promotional overlays for this browser context. Record the state consistently, handle consent in a permitted and repeatable way, or mask the overlay for visual review. A hosted clean-capture service may remove known consent and widget overlays.
Unexpected differences every week Viewport, browser, locale, timezone, dynamic content, or execution host changed. Compare manifest settings, pin the runtime, standardize rendering conditions, and filter volatile page regions.
Some pages fail but others save One URL returned an error, redirected, or was unreachable. Inspect that row’s manifest entry and HTTP status. Fix changed URLs or transient connectivity, then rerun failed entries without overwriting the original run.
Scheduled job succeeds but files disappear The runner’s local filesystem is temporary or the archive path differs from the manual run. Write to persistent storage or upload the run after capture; use absolute paths and check job permissions.
Names collide or overwrite Two rows produce the same sanitized competitor/page identifier. Add a stable unique key column and include it in each filename.

FAQ

Should I capture the viewport or the whole page?

Capture the viewport for a specific first-screen design or message. Use full-page when the complete page structure matters, and account for longer capture files and lazy-loaded content.

Can screenshots prove that a competitor changed its offer?

No. They preserve visual evidence for review, but rendering noise, personalization, and dynamic content can produce apparent differences. Confirm important findings on the live page and retain the capture conditions.

How much history should I keep?

Choose retention based on the period you need to compare and available storage. Keep dated runs and manifests; do not overwrite the prior week if historical review matters.

Can I run this from a laptop?

Yes, if it is available and awake at the scheduled time, has the browser installed, and writes to storage that persists. For dependable recurring history, use a runner and archive that remain available when the laptop is off.