ScreenshotNeo

BlogHow-to

How to Bulk Screenshot URLs and Skip Pages That Return 404 Errors

Use Playwright to check each page’s final HTTP status, skip 404s, and save screenshots with an auditable results log. Includes runnable Python and Node.js examples.

By the ScreenshotNeo team4 October 202612 min read

To bulk screenshot URLs and skip pages that return 404, navigate to each URL in a browser, inspect the main document’s HTTP response, and capture only when its status is not 404. Log the submitted URL, final URL, status, and any navigation exception for every input.

Do not use Playwright’s requestfailed event to identify 404s. A 404 is a completed HTTP response, not a transport failure. Playwright documents that HTTP errors such as 404 and 503 complete through requestfinished; a DNS error or timeout can instead fail before a response exists. See the official Request documentation and Page API.

1. Define what counts as a page to capture

This guide uses the following default policy:

  • Skip exactly status 404.
  • Capture other HTTP statuses, including redirects’ final responses and statuses such as 401, 403, 429, and 5xx. Change the predicate if your use case requires a narrower policy.
  • Record navigation failures such as timeouts or DNS/TLS errors separately. They do not prove that the server returned 404.

Redirects matter: the browser follows them, so the status returned by navigation is the final main-document response. Keep both the input URL and the final URL in the output log. If your business rule is “skip when any URL in the redirect chain returns 404,” inspect response/request redirect relationships as well; the examples below apply the rule to the final navigation response.

2. Prepare the URL list

Use one absolute HTTP or HTTPS URL per line. For example, save this as urls.txt:

https://example.com/
https://example.com/missing-page
https://www.python.org/

Keep the original list order in the results. The examples validate URL schemes, create a stable numbered filename for each input, and write a CSV manifest mapping files back to source URLs. Numbered filenames avoid collisions when distinct URLs normalize to the same name.

3. Python: Playwright with bounded concurrency

Install Python 3.9 or later, then install Playwright and its Chromium browser:

python -m pip install playwright
python -m playwright install chromium

Save the following as bulk_screenshot.py. It processes a limited number of pages at a time, skips final 404 responses, and records errors without treating them as 404s.

import asyncio
import csv
import re
from pathlib import Path
from urllib.parse import urlparse

from playwright.async_api import async_playwright

INPUT = Path("urls.txt")
OUT_DIR = Path("screenshots")
MANIFEST = OUT_DIR / "manifest.csv"
CONCURRENCY = 4
NAVIGATION_TIMEOUT_MS = 30_000
WAIT_UNTIL = "domcontentloaded"  # alternatives: load, networkidle, commit


def load_urls(path: Path) -> list[str]:
    urls = []
    for line_number, raw in enumerate(path.read_text(encoding="utf-8").splitlines(), 1):
        url = raw.strip()
        if not url or url.startswith("#"):
            continue
        parsed = urlparse(url)
        if parsed.scheme not in {"http", "https"} or not parsed.netloc:
            raise ValueError(f"Line {line_number}: expected an absolute http(s) URL: {url!r}")
        urls.append(url)
    return urls


def safe_stem(url: str) -> str:
    parsed = urlparse(url)
    value = re.sub(r"[^A-Za-z0-9._-]+", "-", parsed.netloc + parsed.path).strip("-._")
    return (value[:100] or "page")


async def main() -> None:
    urls = load_urls(INPUT)
    OUT_DIR.mkdir(parents=True, exist_ok=True)
    semaphore = asyncio.Semaphore(CONCURRENCY)

    async with async_playwright() as p:
        browser = await p.chromium.launch()

        async def capture(index: int, url: str) -> dict[str, str]:
            async with semaphore:
                context = await browser.new_context(viewport={"width": 1440, "height": 900})
                page = await context.new_page()
                row = {
                    "input_url": url,
                    "final_url": "",
                    "status": "",
                    "result": "navigation_error",
                    "file": "",
                    "error": "",
                }
                try:
                    response = await page.goto(
                        url,
                        wait_until=WAIT_UNTIL,
                        timeout=NAVIGATION_TIMEOUT_MS,
                    )
                    row["final_url"] = page.url
                    if response is None:
                        row["result"] = "no_main_document_response"
                    else:
                        row["status"] = str(response.status)
                        if response.status == 404:
                            row["result"] = "skipped_404"
                        else:
                            filename = f"{index:05d}-{safe_stem(url)}.png"
                            await page.screenshot(path=str(OUT_DIR / filename), full_page=True)
                            row["result"] = "captured"
                            row["file"] = filename
                except Exception as exc:
                    row["error"] = f"{type(exc).__name__}: {exc}"
                finally:
                    await context.close()
                return row

        rows = await asyncio.gather(*(capture(i, url) for i, url in enumerate(urls, 1)))
        await browser.close()

    with MANIFEST.open("w", newline="", encoding="utf-8") as f:
        writer = csv.DictWriter(f, fieldnames=["input_url", "final_url", "status", "result", "file", "error"])
        writer.writeheader()
        writer.writerows(rows)

    for row in rows:
        print(f"{row['result']:24} {row['status']:4} {row['input_url']} {row['file']} {row['error']}")
    print(f"Manifest: {MANIFEST}")


if __name__ == "__main__":
    asyncio.run(main())

Run it with python bulk_screenshot.py. Screenshots and manifest.csv appear in the screenshots/ directory. Set CONCURRENCY to a value suitable for your machine and the target sites; it is an operational setting, not a universal performance recommendation.

Choose a navigation wait condition

  • domcontentloaded waits for the document’s DOM to be parsed and is a practical default for batches.
  • load waits for the page load event, which may wait on more resources.
  • networkidle waits for network activity to settle. Pages with analytics, polling, or long-lived connections may never become idle promptly.
  • commit returns once the response is received and the document starts loading. Use it when status classification is the priority and the page need not be fully rendered before capture; allow extra rendering time separately if needed.

For pages whose content appears after navigation, wait for a known selector or a short, justified delay before the screenshot. No one wait condition fits every site.

4. Node.js: Playwright with a worker pool

Install Playwright and Chromium:

npm install playwright
npx playwright install chromium

Save as bulk-screenshot.mjs beside urls.txt. This version uses a fixed worker pool, writes a JSON manifest, and captures every final status except 404.

import fs from 'node:fs/promises';
import path from 'node:path';
import { parse as parseUrl } from 'node:url';
import { chromium } from 'playwright';

const input = 'urls.txt';
const outDir = 'screenshots';
const concurrency = 4;
const timeout = 30_000;
const waitUntil = 'domcontentloaded';

function safeStem(raw) {
  const u = new URL(raw);
  return (u.host + u.pathname).replace(/[^A-Za-z0-9._-]+/g, '-').replace(/^[-._]+|[-._]+$/g, '').slice(0, 100) || 'page';
}

const urls = (await fs.readFile(input, 'utf8'))
  .split(/\r?\n/)
  .map(s => s.trim())
  .filter(s => s && !s.startsWith('#'));
for (const [i, raw] of urls.entries()) {
  let u;
  try { u = new URL(raw); } catch { throw new Error(`Line ${i + 1}: invalid absolute URL: ${raw}`); }
  if (!['http:', 'https:'].includes(u.protocol)) throw new Error(`Line ${i + 1}: only http(s) URLs are allowed: ${raw}`);
}

await fs.mkdir(outDir, { recursive: true });
const browser = await chromium.launch();
const rows = new Array(urls.length);
let next = 0;

async function worker() {
  while (true) {
    const i = next++;
    if (i >= urls.length) return;
    const url = urls[i];
    const context = await browser.newContext({ viewport: { width: 1440, height: 900 } });
    const page = await context.newPage();
    const row = { input_url: url, final_url: '', status: null, result: 'navigation_error', file: '', error: '' };
    try {
      const response = await page.goto(url, { waitUntil, timeout });
      row.final_url = page.url();
      if (!response) {
        row.result = 'no_main_document_response';
      } else {
        row.status = response.status();
        if (row.status === 404) {
          row.result = 'skipped_404';
        } else {
          const filename = `${String(i + 1).padStart(5, '0')}-${safeStem(url)}.png`;
          await page.screenshot({ path: path.join(outDir, filename), fullPage: true });
          row.result = 'captured';
          row.file = filename;
        }
      }
    } catch (error) {
      row.error = `${error.name}: ${error.message}`;
    } finally {
      await context.close();
    }
    rows[i] = row;
  }
}

try {
  await Promise.all(Array.from({ length: Math.min(concurrency, urls.length) }, worker));
} finally {
  await browser.close();
}
await fs.writeFile(path.join(outDir, 'manifest.json'), JSON.stringify(rows, null, 2) + '\n');
for (const row of rows) console.log(row.result, row.status ?? '', row.input_url, row.file, row.error);
console.log(`Manifest: ${path.join(outDir, 'manifest.json')}`);

Run node bulk-screenshot.mjs. The imported parse symbol is not needed for this example; you can remove that import if your linter flags unused imports.

5. Adjust the status policy and capture behavior

Skip all non-success statuses

If your rule is to capture only successful HTTP responses, replace the 404-only condition with a 2xx check. For Python, use if not 200 <= response.status < 300: as the skip condition; in Node.js, use if (row.status < 200 || row.status >= 300). Decide explicitly how to treat 3xx if redirects are disabled or handled separately.

Preserve redirect details

By default, navigation follows redirects, and the returned navigation response represents the resulting main-document response. Log page.url (Python) or page.url() (Node.js) after navigation to retain the final URL. Playwright explains that a redirect completes one request and issues a new request for the redirected URL in its request lifecycle documentation.

Change the screenshot format or dimensions

The examples save full-page PNGs. Playwright can capture the viewport instead by omitting full_page: true / fullPage: true. It also supports other screenshot options such as JPEG and quality settings; consult the screenshot API for supported options. Fixing viewport size makes runs more comparable. Full-page shots can consume more memory for very long pages.

Use authentication or a consistent browser context

For protected pages, create a context with the authentication state or cookies required by your application. Reuse browser processes to avoid repeatedly starting Chromium, but isolate contexts when URLs need separate sessions. Avoid putting credentials in the URL list or manifest; restrict access to any logs containing private URLs.

6. cURL and hosted batch services

cURL can check a response status, but it does not render a browser screenshot by itself. This command prints the final HTTP status after following redirects:

curl -L -sS -o /dev/null -w '%{http_code} %{url_effective}\n' https://example.com/

A shell wrapper can use that result to decide whether to call a separate screenshot tool, but it may disagree with browser behavior for JavaScript navigation, browser-specific routing, cookies, or pages that require a rendered session. Use browser navigation status when the actual capture is done in a browser.

Hosted batch endpoints can reduce orchestration work by accepting multiple URLs and providing a batch identifier or progress endpoint. Do not assume they omit 404 pages automatically: check the provider’s documentation for per-URL status, redirect handling, failure reporting, batch limits, and quota behavior. The reviewed batch API documentation describes submission, batch tracking, errors, and rate limits, but does not establish conditional 404 omission. Verify the current provider behavior before adopting it.

7. Other useful tools

ScreenshotNeo is the first hosted screenshot API to try when you want clean captures and explicit billing outcomes: consent banners, popups, and chat widgets are removed before capture, and only clean shots are billed. The local Playwright method above remains useful when your workflow needs custom code around status decisions.

shot-scraper supports multi-URL capture through its configuration-driven multi command. Its documented multi-capture workflow does not itself establish conditional 404 skipping, so place a status-aware step or wrapper in front of capture. Puppeteer also provides browser screenshot capability, but the cited Page API does not provide a turnkey bulk 404-skip workflow. See the Puppeteer Page API.

Or skip the browser setup

ScreenshotNeo accepts one GET request per URL and returns a screenshot or PDF. Its API and options are documented at ScreenshotNeo documentation. Example request:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

For a bulk workflow, send one request per URL with bounded concurrency, then inspect the response’s X-Page-Verdict and X-Billed headers and record them alongside each input URL. ScreenshotNeo’s API is a capture service; apply your own HTTP-status skip policy before requesting a shot if the requirement is specifically to omit final 404 responses.

  • Cookie banners, popups, and chat widgets are removed before the shot.
  • Bot checks, blank pages, and failed loads are never billed.
  • An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf.
  • 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000 screenshots.

Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.

8. Performance, reliability, and cost

Performance

  • Use bounded concurrency. Too many simultaneous browser pages can exhaust memory, CPU, file descriptors, or network capacity and may burden target sites. Tune gradually for your environment.
  • Reuse the browser process, as the examples do. Creating one browser process per URL adds startup work.
  • Choose the earliest wait condition that still produces a useful screenshot, then wait for a specific selector only where necessary.
  • Full-page images are larger and more expensive to hold in memory than viewport captures. Use viewport captures when they answer the task.

Reliability

  • Keep status responses and navigation exceptions as separate result categories.
  • Retry only plausible transient failures, such as selected timeouts or 5xx responses, with a small bounded retry count and backoff. Do not retry every 404; that usually repeats a definitive result.
  • Make output filenames deterministic and keep a manifest. If restarting a batch, decide whether to overwrite, skip already captured items, or use a run-specific directory.
  • Use idempotent output behavior and write manifests after processing so partial runs remain diagnosable.
  • Respect site access rules and avoid excessive request rates.

Cost

Local Playwright has no per-screenshot API charge, but it uses compute, storage, and bandwidth you provide. Hosted APIs may charge by plan, successful capture, or another vendor-specific rule; inspect the current quota, rate-limit, and failure-billing policy before sending a large batch. ScreenshotNeo states that only clean shots are billed and lists a free tier of 1,000 shots/month, Starter at $5 for 3,000, Growth at $15 for 15,000, Pro at $39 for 60,000, Scale at $99 for 250,000, and Business at $249 for 1,000,000; yearly billing gives two months free. All features are on every plan. Check the docs for current request options.

9. Troubleshooting

Symptom Cause Fix
A 404 page is still captured The code treated navigation as an exception or checked a failed-request event instead of the main response. Inspect the response returned by page.goto and explicitly compare its status with 404.
A timeout is reported without a status No main-document response arrived before the timeout, or navigation did not reach the selected wait condition. Keep it in the navigation-error category; raise the timeout only when justified, or try domcontentloaded instead of networkidle.
A redirect ends at an unexpected page The destination differs from the submitted URL. Record both input and final URL, and apply the documented policy to the final response.
403, 429, or 5xx pages are captured The sample policy skips only 404. Change the predicate to allow only the status classes your workflow accepts; record each skipped status distinctly.
Images are missing or the page looks unfinished The screenshot occurs before client-side content or lazy-loaded assets appear. Wait for an application-specific selector or a justified delay; avoid a universal long idle wait for every URL.
Duplicate filenames overwrite each other Host/path sanitization can map different URLs to the same filename. Keep the numeric input index or append a stable hash; retain the manifest mapping.
The batch runs out of memory or becomes slow Concurrency is too high, pages are very long, or each page retains substantial resources. Lower concurrency, close each context after capture, and use viewport screenshots when full-page output is unnecessary.
Browser installation or launch fails The Playwright package is installed but its browser binary is missing, or the environment lacks required dependencies. Run the Playwright browser install command for the selected browser and follow its platform-specific installation guidance.

10. Checklist before a large run

  1. Validate that every input is an absolute HTTP or HTTPS URL.
  2. Choose whether the status rule applies to the final response or any redirect hop.
  3. Decide how to handle 401, 403, 429, and 5xx responses.
  4. Set concurrency and navigation timeout conservatively, then inspect a small sample run.
  5. Use stable filenames and keep the input-to-output manifest.
  6. Review counts for captured pages, skipped 404s, other statuses, no-response cases, and exceptions.
  7. Check screenshot content and storage requirements before scaling the batch.

FAQ

Does HTTP 404 mean Playwright navigation failed?

No. It is an HTTP response. Navigation can return a response object with status 404; the code must check it explicitly.

Should every non-2xx page be skipped?

That depends on the task. This example skips only 404. Authentication snapshots, error-page audits, and incident investigations may need other statuses captured.

Can I screenshot a page that returns 404?

Yes. A 404 response can still render an error page. Skip it only when that is the intended batch policy.

Does the example retry a 404?

No. It treats the response as a result and moves on. Add retries for selected transient failures only if your workflow needs them.

References