ScreenshotNeo

BlogHow-to

How to schedule screenshots of a competitor’s website without overloading it

Schedule a few browser screenshots responsibly: check access guidance, pace requests, watch for overload signals, and back off when needed.

By the ScreenshotNeo team4 October 20268 min read

Use a browser automation script to capture only the public pages you need, schedule it at a modest cadence, and limit work to one page at a time per hostname. Check the site’s published access guidance and terms first. Watch status codes, retries, and response time; pause or back off if you see repeated 429 or 5xx responses or rising latency. No universal screenshot interval is safe for every site: the site’s instructions, behavior, and capacity matter.

This workflow uses Playwright for rendered screenshots. It is designed for a small, recurring watchlist, not broad crawling.

1. Check whether automated access is appropriate

  1. Limit the watchlist to the few public pages whose visual changes matter.
  2. Review the site’s robots.txt, published terms, and any official API, feed, or export. Google’s guidance says robots.txt communicates which URLs its crawlers can access; it is not a security mechanism or permission to access restricted material. Google also does not support the crawl-delay field, so do not assume every crawler interprets it the same way. Google’s robots.txt introduction and its robots.txt specification explain the scope and supported directives.
  3. Prefer structured access when it supplies the information you need. Scrapy recommends an API, bulk export, or search endpoint where suitable because it is faster for the user and cheaper for the website than crawling pages. Scrapy optimization guidance.
  4. If guidance disallows your planned access, or the page requires authentication or other restricted access you do not have, do not automate that access.

A robots.txt directive is crawler guidance, not a substitute for the site’s terms or access controls. This article does not determine legal permission for a particular site or jurisdiction.

2. Install Playwright and create a screenshot script

Use a stable browser, viewport, and execution environment so that differences between captures are easier to interpret. This example processes URLs sequentially, uses a finite navigation timeout, takes one screenshot per page, and spaces visits to the same host. Set WATCH_URLS to a comma-separated list of public pages you are permitted to monitor.

mkdir competitor-shots && cd competitor-shots
npm init -y
npm install playwright
npx playwright install chromium

Save as capture.mjs:

import { chromium } from 'playwright';
import { mkdir } from 'node:fs/promises';

const urls = (process.env.WATCH_URLS ?? 'https://example.com/')
  .split(',')
  .map((value) => value.trim())
  .filter(Boolean);
const intervalMs = Number(process.env.INTERVAL_MS ?? 60_000);
const outputDir = process.env.OUTPUT_DIR ?? 'shots';
const timeoutMs = Number(process.env.NAVIGATION_TIMEOUT_MS ?? 30_000);

if (!Number.isFinite(intervalMs) || intervalMs < 0) {
  throw new Error('INTERVAL_MS must be a non-negative number');
}
await mkdir(outputDir, { recursive: true });
const browser = await chromium.launch({ headless: true });

try {
  const context = await browser.newContext({
    viewport: { width: 1440, height: 900 },
    deviceScaleFactor: 1,
  });
  const page = await context.newPage();
  page.setDefaultNavigationTimeout(timeoutMs);

  for (let i = 0; i < urls.length; i += 1) {
    const url = urls[i];
    let response;
    const startedAt = Date.now();
    try {
      response = await page.goto(url, { waitUntil: 'domcontentloaded' });
      const status = response?.status() ?? 'no main-document response';
      const elapsedMs = Date.now() - startedAt;
      console.log(JSON.stringify({ url, status, elapsedMs }));

      if (response && (response.status() === 429 || response.status() >= 500)) {
        console.warn(`Backing off after HTTP ${response.status()} from ${url}`);
        break;
      }
      if (!response || !response.ok()) {
        console.warn(`Skipping screenshot for unsuccessful navigation: ${url}`);
        continue;
      }

      // A timestamp in the filename preserves a simple capture history.
      const safeHost = new URL(url).hostname.replace(/[^a-z0-9.-]/gi, '_');
      const timestamp = new Date().toISOString().replace(/[:.]/g, '-');
      await page.screenshot({
        path: `${outputDir}/${safeHost}-${timestamp}.png`,
        fullPage: true,
        animations: 'disabled',
      });
    } catch (error) {
      console.error(JSON.stringify({ url, error: String(error) }));
      // A timeout or navigation failure is a reason to stop this run and
      // investigate, rather than immediately retrying the same target.
      break;
    }

    if (i < urls.length - 1) {
      await new Promise((resolve) => setTimeout(resolve, intervalMs));
    }
  }
  await context.close();
} finally {
  await browser.close();
}

Run it locally:

WATCH_URLS='https://example.com/,https://example.org/' \
INTERVAL_MS=60000 \
node capture.mjs

The 60-second value here is only an example delay between pages in this one run. It is not a safe or recommended universal interval. Choose a cadence and spacing based on the site’s published guidance, how often the pages change, and observed behavior. For a recurring schedule, run the script from a scheduler such as a CI workflow or a system scheduler, with a single active run per hostname.

3. Choose readiness and cadence deliberately

Playwright’s page.goto() supports navigation completion choices such as commit, domcontentloaded, load, and networkidle. For a screenshot, choose the least waiting that produces the content you need, then wait for a specific selector or a short, bounded delay if a known element renders afterward. Playwright discourages using networkidle as a general readiness signal because pages may keep network connections open. See the Playwright Page API.

For example, after page.goto, a site-specific readiness check can be:

await page.locator('main article').waitFor({ state: 'visible', timeout: 10_000 });
await page.screenshot({ path: 'article.png', fullPage: true });

Use selectors that are stable and specific to the page. A missing selector should fail with a bounded timeout and a useful log entry rather than stall the scheduled job indefinitely.

Scheduling and request pacing

  • Schedule only as often as the page’s change rate makes useful. There is no source-backed interval that is safe for all sites.
  • Keep concurrency low per hostname; this example is sequential. Avoid overlapping scheduled runs.
  • Spread page visits over time instead of launching a burst. If you monitor several hosts, track limits and responses separately for each host.
  • Record timestamp, URL, main-document status, duration, and errors. Retain screenshots only as long as your comparison workflow needs them.
  • Do not retry an error immediately in a tight loop. Treat repeated failures as feedback to pause, reduce frequency, or investigate.

Scrapy describes per-domain concurrency and download delay controls and identifies increasing 429/503 responses, retries, and download latency as signs a crawler may have exceeded a site’s tolerance. AWS guidance also recommends delays, smaller batches, and pausing when a crawler encounters 429 responses: AWS ethical crawler practices. Google’s own crawlers reduce their rate in response to certain error responses, but that is Google’s behavior and not a universal rate formula; see Google’s crawl-rate guidance.

4. Schedule the job and keep its output

Use a scheduler that fits your environment, and prevent concurrent instances from running against the same hostname. In CI, store screenshots and logs as run artifacts with a retention period that matches the review process. Playwright documents browser setup and artifact handling in its continuous integration guide.

For a system scheduler, configure a single recurring invocation at a cadence you have chosen for the particular pages. Confirm the scheduler’s timezone and missed-run behavior; do not queue several missed invocations to run at once. If your scheduler supports concurrency groups or locks, use them to ensure there is at most one active capture job for the target hostname.

5. Troubleshoot common failures

Symptom Likely cause Response
HTTP 429 The target is rate limiting or refusing the current request volume. Stop the run, do not retry in a loop, reduce cadence and concurrency, and resume cautiously only after responses return to normal.
Repeated 5xx responses or rising latency The site may be overloaded or experiencing an issue; the job may also be adding load. Pause captures, inspect logs, and try again later at a lower cadence. Do not treat a transient error as a reason to send more requests.
Navigation timeout The page is slow, waiting for ongoing activity, or unreachable from the runner. Use a finite timeout, log the failure, and investigate. Try a more suitable readiness condition such as domcontentloaded; avoid extending timeouts without limit.
Screenshot is blank or missing content The capture happened before the relevant content rendered, a selector changed, or the page requires an interaction or access the script does not have. Check the saved status and logs, wait for a stable page-specific selector with a bounded timeout, and respect access controls. Do not attempt to bypass a bot check or restriction.
Many duplicate or inconsistent captures Overlapping scheduled jobs, varying viewport/browser settings, or changing page content. Use a single-run lock, hold browser and viewport settings constant, and record capture times.
Images or below-the-fold content are absent Lazy-loaded elements may not have loaded before the full-page screenshot. Use a site-appropriate bounded scroll or wait strategy, and compare output before increasing capture frequency. A full-page screenshot alone does not guarantee every lazy element loaded.

6. Reliability, performance, and cost

A browser screenshot costs more local compute and usually more page work than reading a structured endpoint, because the browser loads and renders resources. Keep the URL list short, reuse one browser process for a small batch, and close contexts cleanly. An API or export is usually preferable when it answers the monitoring question without rendering the page.

Reliability comes from bounded waits, sequential work, stable capture settings, a single active job per host, and logs that make failures visible. Build in a stop condition for 429, repeated 5xx responses, abnormal latency, and navigation errors. Backoff protects the site and keeps a failing job from compounding the problem. Retention and storage cost depend on your scheduler and artifact storage; set an explicit retention policy rather than assuming old images are free to keep.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. A single request can return a screenshot or PDF, with options including full-page capture, viewport and device settings, waits, and custom headers. See the ScreenshotNeo API documentation for parameters. It can simplify capture infrastructure, but a hosted API does not remove your responsibility to check the target site’s access guidance, choose a low cadence, and back off on overload signals.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
  • Cookie and consent banners are accepted and removed before the shot; more than 60 known consent platforms, newsletter popups, and chat widgets are supported, and each step can be turned off.
  • Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing; response headers report the page verdict and billing status.
  • An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
  • The Free plan includes 1,000 screenshots a month with no card. Paid plans start at $5 for 3,000 screenshots; yearly billing gives two months free, and every feature is on every plan.

Use the free ScreenshotNeo sign-up to get 1,000 screenshots a month with no card.

FAQ

Is robots.txt permission to take screenshots?

No. It communicates crawler guidance. Check the site’s terms and access controls as well, and do not treat a robots.txt file as authorization for restricted access.

What is a safe interval between screenshots?

There is no universal safe interval in the cited guidance. Choose a modest cadence based on the site’s instructions, how often the page changes, and the responses and latency you observe.

Should I use screenshots if an API exists?

If an official API or export provides the information you need, it is generally a lower-work alternative for the site. Use screenshots when the visual rendering itself is what you need to monitor.

Can this monitor pages behind a login or bot check?

This workflow is for publicly accessible pages. Do not bypass access restrictions or bot checks; seek an authorized access method from the site.