ScreenshotNeo

BlogHow-to

How to Fix Navigation Errors When Capturing Screenshots of Many Websites

Diagnose navigation exceptions, HTTP errors, timeouts, and incomplete pages in multi-site screenshot jobs with a reliable Playwright workflow.

By the ScreenshotNeo team4 October 20269 min read

When a screenshot job visits many websites, first distinguish a navigation exception from an HTTP error response and from a page that loaded but is not ready to capture. Catch navigation exceptions, record the response status, wait for the page condition your screenshot needs, and keep failures visible in your job results. A longer timeout helps only when a site needs more time; it will not fix an invalid URL, an SSL error, or an unreachable server.

This guide uses Playwright with Node.js for the do-it-yourself workflow. It shows how to capture a batch of URLs, classify outcomes, tune readiness and timeouts, and investigate common failures. The same diagnostic principles apply to other browser automation tools.

1. Separate navigation failures from HTTP errors

A navigation call can fail in two different ways:

  • It throws: the browser could not complete navigation, for example because the URL is invalid, a certificate check failed, the server was unreachable, the main resource failed, or the navigation timed out.
  • It returns an HTTP response with an error status: the page loaded with a status such as 404 or 500. A resolved navigation does not mean the page returned a successful status.

Keep these outcomes separate in your logs and output. A thrown exception does not always have a response to inspect; a resolved response should have its status checked. Also record the final URL when available, since redirects can change the destination.

With Puppeteer, page.goto() returns the main-resource response; a redirect resolves with the final response. Navigations to about:blank or a hash-only URL change can return null. In headless shell, HTTP 404 and 500 responses do not by themselves make page.goto() throw, so inspect the response status explicitly. See the Puppeteer Page.goto() reference.

2. Choose a page readiness condition that matches the screenshot

Playwright offers four navigation milestones:

Condition What it means When it may fit
commit The response was received and document loading started. When you want to begin checking the page as early as possible.
domcontentloaded The initial HTML document has been parsed. When you plan to wait for a specific element after the document becomes available.
load The page’s load event fired. This is Playwright’s default. When the page’s load event is a useful baseline for the capture.
networkidle The network has reached Playwright’s network-idle condition. Use cautiously; background analytics, streaming, and persistent connections can make network idleness a poor proxy for visible readiness.

Playwright discourages using networkidle for tests and recommends checking whether the page is ready with an assertion. For screenshots, navigate to a suitable milestone, then wait for an observable condition tied to the content you need, such as the main heading becoming visible. No one milestone is right for every site. See Playwright navigation options.

3. Build a batch capture that records each outcome

Install Playwright and its browser in your project, then save the following as capture.mjs. Run it with node capture.mjs. The example sets the viewport before navigation, checks HTTP status, waits for a page-specific selector, writes successful screenshots, and records failures without stopping the rest of the batch.

import { chromium } from 'playwright';
import { mkdir, writeFile } from 'node:fs/promises';

const urls = [
  'https://example.com',
  'https://www.wikipedia.org',
  'https://httpbin.org/status/404',
];

const timeoutMs = Number(process.env.NAVIGATION_TIMEOUT_MS ?? 30000);
const outputDir = 'screenshots';
const results = [];

await mkdir(outputDir, { recursive: true });
const browser = await chromium.launch({ headless: true });

try {
  for (const [index, inputUrl] of urls.entries()) {
    const startedAt = Date.now();
    let page;
    const result = {
      inputUrl,
      index,
      status: 'failed',
      responseStatus: null,
      finalUrl: null,
      elapsedMs: null,
      error: null,
    };

    try {
      const normalizedUrl = new URL(inputUrl).toString();
      result.normalizedUrl = normalizedUrl;
      page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
      page.setDefaultNavigationTimeout(timeoutMs);

      const response = await page.goto(normalizedUrl, {
        waitUntil: 'domcontentloaded',
        timeout: timeoutMs,
      });

      result.finalUrl = page.url();
      result.responseStatus = response?.status() ?? null;

      // Replace this with a selector that identifies the content you need.
      // Do not make every site depend on a selector that some sites lack.
      const readySelector = process.env.READY_SELECTOR;
      if (readySelector) {
        await page.locator(readySelector).waitFor({
          state: 'visible',
          timeout: timeoutMs,
        });
      }

      const filename = `${outputDir}/page-${index + 1}.png`;
      await page.screenshot({ path: filename, fullPage: true });
      result.screenshot = filename;
      result.status = response && response.status() >= 400 ? 'http-error-captured' : 'captured';
    } catch (error) {
      result.finalUrl = page?.url() ?? null;
      result.error = error instanceof Error ? error.message : String(error);
      result.status = result.error.toLowerCase().includes('timeout') ? 'timeout' : 'navigation-or-capture-error';
    } finally {
      result.elapsedMs = Date.now() - startedAt;
      await page?.close();
    }

    results.push(result);
    console.log(JSON.stringify(result));
  }
} finally {
  await browser.close();
  await writeFile('capture-results.json', JSON.stringify(results, null, 2));
}

The sample uses a finite 30-second default as a configurable example, not a universal recommendation. Set NAVIGATION_TIMEOUT_MS to fit your job’s latency budget. For sites with different page structures, use per-site readiness selectors in your input data rather than one global selector.

In production, consider validating schemes and allowed destinations before navigation, limiting concurrency, and storing one result record for every input URL. If an HTTP error page is still useful to your workflow, label it clearly as captured with an error status. If only successful responses should produce an artifact, check the status before writing the screenshot and record the rejected response instead.

4. Set navigation timeouts deliberately

A timeout is the maximum time an operation will wait. Increasing it can help a slow but reachable site; it cannot repair an invalid URL, an SSL failure, an unreachable host, or a failed main resource. Use a finite timeout and report timeout separately from other errors so you can tell whether the job is waiting too long or failing for another reason.

In Playwright, you can pass timeout to page.goto() or configure a default navigation timeout with page.setDefaultNavigationTimeout(). The navigation-specific default takes priority over page.setDefaultTimeout() and context-wide general timeout settings for navigation methods. Make the limit configurable for batch jobs, and tune it from observed job needs rather than assuming one value fits every site. See the Playwright timeout API.

5. Set the viewport before opening each site

Set viewport dimensions before navigation so responsive sites render at the intended size from the start. In Playwright, the example creates each page with its viewport before calling goto(). If pages share a browser context, a context viewport setting can apply to all of them. See Playwright viewport configuration.

Pick dimensions that match the screenshot you need. A desktop capture and a mobile capture can trigger different layouts and different navigation flows, so make the viewport part of the capture configuration and record it with the result.

6. Diagnose failures by category

Observed result What to check Practical next step
Invalid URL error Scheme, spelling, whitespace, and URL parsing. Normalize and validate each input before launching navigation. Include https:// or http:// where required.
SSL or certificate error Whether the site has a valid certificate and whether the browser trusts its certificate chain. Use the correct HTTPS URL and resolve the site’s certificate problem. Do not treat certificate bypasses as a general fix.
Unreachable server or timeout DNS, connectivity, server response, and whether the timeout is appropriate for the job. Retry selectively where appropriate, keep the timeout finite, and preserve the exact error and elapsed time.
HTTP 404 or 500 with no exception The main-resource response status. Check response.status(); classify the result as an HTTP error even if you choose to save its error page.
Navigation completes but screenshot is incomplete Whether the content is rendered after the chosen navigation milestone. Wait for a page-specific visible selector or another observable condition before capture.
net::ERR_BLOCKED_BY_CLIENT on a remote HTTP URL in Puppeteer Chrome for Testing’s HTTPS-First behavior and whether the browser showed a warning page. Prefer HTTPS if the site supports it. Puppeteer’s troubleshooting guide documents a launch argument to disable the behavior for a controlled environment when this specific cause is confirmed.

Playwright documents navigation throws such as SSL errors, invalid URLs, timeouts, unreachable servers, and main-resource failures. Puppeteer documents the HTTP-status distinction and a Chrome for Testing HTTPS-First case that can block remote HTTP navigation with net::ERR_BLOCKED_BY_CLIENT. Use browser-specific remedies only after confirming the matching condition. See Puppeteer troubleshooting.

7. Make large capture jobs more reliable

  • Keep one record per input. Include the original and normalized URL, browser/framework version, elapsed time, final URL, response status, error text, viewport, readiness condition, and artifact path when available.
  • Do not collapse distinct failures. Preserve categories for thrown navigation errors, HTTP error responses, timeouts, and captures that appear incomplete.
  • Keep failed evidence. Retain logs and, where practical, an error-state screenshot or page details for later review instead of silently marking the URL successful.
  • Retry selectively. A retry may help a transient network or server issue, but repeating an invalid URL or certificate failure unchanged is unlikely to help. Record attempts so retries do not hide flaky sites.
  • Control concurrency. Start with a workload your browser and network can handle, then adjust based on memory, CPU, and target-site behavior. Excessive parallel pages can make slowdowns harder to diagnose.
  • Close pages and browsers. The example closes each page after its result and always closes the browser, which limits resource leakage in long jobs.
  • Use per-site readiness rules when needed. A selector that identifies the main content is more meaningful than waiting for a global network condition that some sites may never reach.

These are operational practices, not guarantees of a particular success rate. A remote site can remain unavailable or change its markup; the goal is to make each outcome understandable and recoverable.

8. Or skip the browser setup

If you only need screenshots and do not want to manage a browser for each destination, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns an image or PDF. Its clean-capture steps accept cookie or consent banners like a visitor and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and whether it was billed.

See the ScreenshotNeo API documentation for request options. This cURL example saves a WebP screenshot:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Equivalent Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

Equivalent Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));

ScreenshotNeo also offers an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Create a free account and get 1,000 screenshots a month with no card.

9. Performance, reliability, and cost considerations

For a self-hosted browser workflow, total job time depends on site response and rendering, readiness waits, screenshot work, and how many pages run at once. Raising concurrency can shorten a batch but also increases local browser and network load. Track elapsed time by URL and keep timeouts bounded so one slow page does not hold the entire job indefinitely.

Retries consume additional browser time and resources. Retry only categories that might be transient, and keep the first failure in the result record. A response with status 404 or 500 may still be worth capturing for diagnostics; decide whether it should count as an artifact according to the job’s purpose.

With ScreenshotNeo, only clean shots are billed; bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed. The available plans are Free with 1,000 shots per month and no card, Starter at $5 for 3,000, Growth at $15 for 15,000, Pro at $39 for 60,000, Scale at $99 for 250,000, and Business at $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan. Choose based on your expected volume; do not count failed pages as billable clean shots.

10. FAQ

Should I treat every 4xx or 5xx page as a navigation failure?

No. A browser may complete navigation and return an HTTP response with an error status. Record the status separately from thrown navigation errors, then decide whether the error page is useful to save.

Is networkidle the best condition before every screenshot?

No single condition fits every site. Background requests may keep a page active, so prefer a page-specific observable readiness check when you know which content must appear.

Will a longer timeout solve a browser navigation error?

Only if the page needs more time and can otherwise load. A longer wait does not fix invalid URLs, SSL failures, or unreachable servers.

Why might an HTTP site show a browser warning in Puppeteer?

Puppeteer’s troubleshooting documentation describes Chrome for Testing HTTPS-First behavior that can block a remote HTTP navigation and produce net::ERR_BLOCKED_BY_CLIENT. Prefer HTTPS where available, and investigate the documented targeted workaround only when this specific behavior is the cause.