ScreenshotNeo

BlogHow-to

How to Bulk Screenshot a List of URLs and Detect Blank Images

Capture URL lists with Playwright, log failures separately, and flag blank-looking screenshots using a calibrated image-variation heuristic.

By the ScreenshotNeo team4 October 202611 min read

To bulk screenshot URLs and find blank-looking results, run one Playwright browser process, visit each URL, save the screenshot, and measure pixel variation in the resulting image. Keep navigation failures and low-variation screenshots as separate outcomes: a flat image is only a review candidate, not proof that the browser failed to render the page.

The runnable Node.js script below reads one URL per line from urls.txt, captures each page, writes a JSON Lines report, and flags candidate blank images using a simple luminance-variation score. The score and cutoff are examples to calibrate on your own pages; Playwright provides screenshot capture and bytes for processing, but does not define a universal blank-image detector. Playwright Page API · Playwright screenshot documentation.

1. Install and prepare a URL list

Use Node.js and Playwright. Create a project, install the package and browser, then place one absolute HTTP or HTTPS URL on each line. Empty lines and lines beginning with # are ignored.

npm init -y
npm install playwright
npx playwright install chromium
# urls.txt
https://example.com
https://www.wikipedia.org
https://www.iana.org/domains/reserved

Give every input a stable index (the script does this) and preserve the original URL in the report. That lets you identify a failed or suspicious result even if redirects change the final page address.

2. Capture the pages and flag low-variation images

Save this as bulk-screenshot.mjs. It uses one Chromium process and one page at a time, so it avoids launching a browser for every URL. It writes screenshots to screenshots/ and one JSON object per line to results.jsonl.

import { chromium } from 'playwright';
import { readFile, mkdir, writeFile, appendFile } from 'node:fs/promises';
import path from 'node:path';

const INPUT = 'urls.txt';
const OUT_DIR = 'screenshots';
const REPORT = 'results.jsonl';
const TIMEOUT_MS = 30_000;
const FULL_PAGE = false; // Set true when the whole scrollable document matters.
const VIEWPORT = { width: 1365, height: 900 };

// Example starting point, not a universal cutoff. Calibrate with your own set.
const LOW_VARIATION_STDDEV = 7;

function luminanceStdDev(buffer, width, height) {
  // PNG is decoded by Playwright's bundled/browser dependency only in the browser;
  // use a small image package for reliable Node-side pixel decoding.
  // This placeholder is replaced below by sharp, which decodes PNG bytes.
  throw new Error('Install sharp and use the implementation below.');
}

For a fully runnable pixel detector, add sharp and replace the placeholder function and setup with this complete version (the following is the complete script):

import { chromium } from 'playwright';
import sharp from 'sharp';
import { readFile, mkdir, writeFile, appendFile } from 'node:fs/promises';
import path from 'node:path';

const INPUT = 'urls.txt';
const OUT_DIR = 'screenshots';
const REPORT = 'results.jsonl';
const TIMEOUT_MS = 30_000;
const FULL_PAGE = false;
const VIEWPORT = { width: 1365, height: 900 };
const LOW_VARIATION_STDDEV = 7;

async function imageStats(png) {
  // Sample a reduced-size grayscale image to keep analysis memory bounded.
  const { data, info } = await sharp(png)
    .resize({ width: 160, height: 160, fit: 'inside' })
    .greyscale()
    .raw()
    .toBuffer({ resolveWithObject: true });
  let sum = 0;
  let sumSquares = 0;
  for (const value of data) {
    sum += value;
    sumSquares += value * value;
  }
  const count = data.length;
  const mean = sum / count;
  const variance = Math.max(0, sumSquares / count - mean * mean);
  return { width: info.width, height: info.height, luminanceStdDev: Math.sqrt(variance) };
}

const lines = (await readFile(INPUT, 'utf8')).split(/\r?\n/);
const urls = lines.map((line, index) => ({ line: index + 1, url: line.trim() }))
  .filter(({ url }) => url && !url.startsWith('#'));

await mkdir(OUT_DIR, { recursive: true });
await writeFile(REPORT, '');
const browser = await chromium.launch({ headless: true });
const context = await browser.newContext({ viewport: VIEWPORT });
const page = await context.newPage();

try {
  for (let index = 0; index < urls.length; index++) {
    const { line, url } = urls[index];
    const id = String(index + 1).padStart(5, '0');
    const screenshotPath = path.join(OUT_DIR, `${id}.png`);
    const record = {
      id, inputLine: line, url, capturedAt: new Date().toISOString(),
      finalUrl: null, navigationStatus: null, screenshotPath: null,
      image: null, blankCandidate: null, error: null
    };

    try {
      const response = await page.goto(url, {
        waitUntil: 'domcontentloaded',
        timeout: TIMEOUT_MS
      });
      record.finalUrl = page.url();
      record.navigationStatus = response?.status() ?? null;

      // domcontentloaded avoids waiting forever on long-lived analytics or sockets.
      // Add a site-specific selector wait or short delay here if content renders later.
      const png = await page.screenshot({ path: screenshotPath, fullPage: FULL_PAGE });
      record.screenshotPath = screenshotPath;
      record.image = await imageStats(png);
      record.blankCandidate = record.image.luminanceStdDev < LOW_VARIATION_STDDEV;
    } catch (error) {
      record.error = String(error?.message ?? error);
      // Keep any screenshot path only when a screenshot was actually saved.
    }

    await appendFile(REPORT, `${JSON.stringify(record)}\n`);
    console.log(`${id} ${record.error ? 'ERROR' : record.blankCandidate ? 'REVIEW' : 'OK'} ${url}`);
  }
} finally {
  await context.close();
  await browser.close();
}

Install the additional image decoder used by this script, then run it:

npm install sharp
node bulk-screenshot.mjs

Each successful image receives its own PNG. For each URL the report records its stable ID, original input line, timestamp, final URL, HTTP response status when available, image dimensions, score, and candidate flag. Navigation or capture errors are written to the same report with an error string and are not mislabeled as blank screenshots.

What the detector measures

The example resizes the screenshot to at most 160 pixels on its longest side, converts it to grayscale, and calculates the standard deviation of luminance values. A near-uniform image tends toward a low score; text, edges, photos, and layout tend to increase variation. LOW_VARIATION_STDDEV is a starting value only. There is no universal cutoff: a solid-color page may be legitimate, and a largely white page with a small error message can score low.

Before using this flag operationally, inspect representative examples: a normal content page, an intentionally sparse page, a consent or bot-check screen, a known error page, a dark page, and a page with delayed rendering. Adjust the cutoff to minimize the kind of false alarm that matters for your workflow. Treat the result as review, not as a pass/fail diagnosis.

3. Choose capture and wait behavior

Playwright’s page.screenshot() saves a file when given path, returns image bytes for processing, and supports full-page capture. See the Page API screenshot options.

Choice Use it when Trade-off
Viewport (fullPage: false) You need a consistent first-screen image or want lower memory and file sizes. Content below the fold is not included; lazy images farther down may not load.
Full page (fullPage: true) You need the full scrollable page for review or archival. Images are larger and long pages affect the variation score. Very long documents consume more memory.
domcontentloaded Pages render content after document parsing and you want to avoid waiting for every resource. Client-rendered content or images may still be incomplete.
load The page’s load event is a reasonable signal for your target sites. Some pages wait on slow or unnecessary resources; this event does not guarantee app-specific readiness.
networkidle A page becomes quiet after its requests finish and that reflects readiness. Analytics, polling, and persistent connections can make network quietness an unreliable criterion.

For an application that renders after navigation, wait for a stable selector or a bounded delay before capture. Replace the screenshot line with, for example:

await page.locator('main article').waitFor({ state: 'visible', timeout: 10_000 });
await page.screenshot({ path: screenshotPath, fullPage: FULL_PAGE });

Use a selector that indicates the content you actually need; waiting for body only confirms that a document exists. Avoid arbitrary long sleeps across a large list because they increase total runtime. If you use a short delay, document why the target site needs it.

4. Scale the batch without losing useful records

The example processes sequentially, which is the simplest option for modest lists and keeps browser memory predictable. For faster batches, use a small fixed number of pages or workers, not one browser per URL. Each worker should own a page, and each URL should still get an independent result record. Increase concurrency gradually while watching memory, target-site rate limits, and failure rates.

  • Reuse the browser: browser startup is shared across all URLs. Close the context and browser in a finally block so errors do not leave browser processes behind.
  • Bound time: keep navigation timeouts finite. A timeout should create a failed-navigation record, not stop the entire batch.
  • Retry selectively: retry transient network errors or selected server responses with a small retry limit and backoff. Do not retry every result indefinitely; a persistent bot check or invalid URL will not be fixed by repeated requests.
  • Keep outcomes separate: distinguish navigation error, non-success HTTP status, screenshot error, and blank candidate. An HTTP 404 may still produce a real page image, so decide whether to retain it based on your purpose.
  • Make runs resumable: use stable IDs or a normalized URL key and persist each record as it finishes. On restart, skip only records that meet your chosen completion policy.
  • Control storage: decide how long to retain PNGs and reports. Full-page images can use substantial disk space; archive or delete them according to your review and audit needs.

For visual regression, where every URL has a trusted expected reference, use Playwright Test’s toHaveScreenshot() assertion instead of inventing a blankness threshold. It waits for two consecutive screenshots to stabilize before comparing and supports image-difference options. This detects deviation from a baseline; it is a different task from discovering blank-looking images without a reference. Rendering can vary across operating systems, browser versions, settings, hardware, power source, and headless mode, so keep the capture environment consistent. See visual comparisons and PageAssertions options.

5. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. One GET request captures a URL and returns an image or PDF; the API documentation describes the available parameters. Its bulk capture option accepts up to 100 URLs per call, and its parameter names also work with those used by other screenshot APIs.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie banners, newsletter popups, and chat widgets are removed before the screenshot; each cleanup step can be turned off.
  • Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed. Response headers identify the page verdict and billing outcome.
  • An MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.
  • The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Every feature is available on every plan.

Sign up for 1,000 free screenshots a month, with no card required.

6. Troubleshooting

Symptom Likely cause Fix
Navigation times out The server is slow, the connection stalls, or the chosen readiness event never arrives. Keep a finite timeout, capture the error in the report, and consider domcontentloaded plus a page-specific selector wait. Retry only likely transient failures.
Screenshot is nearly blank but navigation succeeded The app rendered late, content is below the viewport, or the site returned a consent, bot-check, or error page. Inspect the screenshot and final URL, choose a more meaningful wait condition, try full-page capture when appropriate, and keep the result as a candidate for review.
Many valid pages are flagged The threshold is too high, the pages have sparse layouts, or viewport capture contains large flat backgrounds. Review a sample of flagged and unflagged outputs, lower or recalibrate the cutoff, and consider measuring regions or edges instead of a single whole-image score.
A blank page is missed A browser error or small visible element adds enough pixel variation to cross the threshold. Combine the heuristic with checks relevant to your site, such as expected text or a required selector. No global pixel score can establish semantic correctness.
Images are absent in full-page captures Images may load lazily only after scrolling, or may be blocked or delayed. Scroll in steps before capture when the site requires it, or wait for known image selectors. Record this capture behavior because it changes runtime and output.
HTTP error page is recorded as a successful image Navigation can return an HTTP response such as 404 while the browser still renders a page. Keep response status separate from screenshot success; set your own policy for whether non-success statuses should be reviewed or excluded.
Results differ between machines or runs Fonts, browser version, operating system, rendering settings, headless mode, or dynamic page content changed. Pin the browser and runtime environment where repeatability matters, standardize viewport and scale, and suppress or wait for known dynamic content.
sharp fails to install or decode The native image package may not match the platform, or the screenshot buffer may not be a supported/complete image. Use a supported Node environment, inspect the original PNG, and record the analysis error separately from the navigation result. Keep the screenshot so it can be reprocessed.
Batch stops at one bad URL An exception escaped the per-URL handler or cleanup was not protected. Keep each URL’s navigation, capture, and analysis inside its own try/catch, append a result in every case, and close browser resources in finally.

7. Performance, reliability, and cost

For this self-hosted workflow, the main costs are the machine time to render pages, memory for browser pages and image buffers, and disk space for outputs. Full-page capture and high concurrency raise resource use; sequential capture is slower but easier to reason about. The script reduces detector memory by analyzing a downscaled copy while saving the original screenshot.

Use a bounded queue if you add concurrency, and cap retries. If each page takes roughly t seconds and you process n URLs sequentially, total runtime is roughly the sum of page times plus browser and analysis overhead; parallel workers can reduce elapsed time but may increase resource contention and target-site errors. This is a planning relationship, not a benchmark.

For reliability, persist each record as it completes, include timestamps and both input and final URLs, and keep a copy of suspicious images until review. A screenshot alone cannot say whether the page was semantically correct. The score is useful for triage; use known selectors, text assertions, or visual baselines where the task calls for stronger validation.

Self-hosted Playwright has no per-shot API charge, but requires you to operate the browser environment and maintain storage, retries, and detection logic. ScreenshotNeo offers usage-based plans: Free 1,000 shots per month, Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000; yearly billing gives two months free. Features are available on every plan. Select based on expected volume and whether the managed capture and billing classifications fit your workflow.

8. FAQ

Does a low-variation score prove the page is broken?

No. It identifies an image worth checking. Intentional solid backgrounds, sparse pages, consent screens, and viewport choices can all produce low variation.

Should I use blank detection or screenshot comparison?

Use a calibrated heuristic when you have no reference image and want candidate blanks. Use visual comparison when you have a trusted baseline and need to detect changes against it.

Should I capture the viewport or the whole page?

Choose the area your task evaluates. Viewport capture is smaller and consistent; full-page capture includes scrollable content but can be larger and changes how the detector behaves.

Can I safely raise batch concurrency?

Only as far as the machine and target sites tolerate. Start with a small worker count, track timeouts and memory, and reduce concurrency when errors or resource pressure rise.

Sources