ScreenshotNeo

BlogHow-to

Best Way to Bulk Capture Screenshots of SEO Landing Pages

Capture SEO landing pages in bulk with Playwright, consistent viewports, stable filenames, retries, and a clear workflow for comparing results.

By the ScreenshotNeo team4 October 202610 min read

The most flexible way to bulk capture SEO landing pages is to automate a real browser with Playwright: feed it a canonical URL list, set one viewport and capture mode, wait for the page state you need, then save each image under a stable filename. For a small or self-managed job, a script gives you control of retries, logs, and storage. Use viewport captures for consistent above-the-fold comparisons and full-page captures when reviewers need the entire page.

Playwright can navigate pages, save screenshots to files, return image bytes, capture full pages, and capture selected elements. The official screenshots guide and Page screenshot API document these operations; a batch importer, scheduler, retry policy, and comparison dashboard are application logic you build around them.

1. Decide what the screenshots should show

Before capturing, define what a useful comparison means. Use the same browser engine, viewport, device scale factor, wait condition, and capture mode for every URL in a run. Record the date and those settings alongside the output so a later reviewer can interpret differences.

Goal Capture mode Why
Compare above-the-fold layout across pages Viewport screenshot Every image has the same dimensions and visual focus.
Review all scrollable content Full page Captures the page as one tall image.
Inspect a particular module, such as a hero or pricing table Element screenshot Focuses the output on a selected CSS element.

Choose a wait condition that matches the page. Navigation completion alone may not mean a client-rendered page or its images are ready. A fixed delay is easy to understand but can waste time on fast pages and still be too short on slow pages. For predictable content, wait for a selector that indicates the section of interest is present.

2. Prepare a canonical URL list

Keep one URL per line in a UTF-8 text file named urls.txt. Prefer the canonical landing-page URL you want to inspect, and remove duplicates or tracking variants unless those variants are part of the test. The script below rejects non-HTTP URLs and duplicate entries. It also derives filenames from a short hash, avoiding collisions between URLs that share a path.

https://example.com/
https://example.com/pricing/
https://example.org/product/

Do not put secrets in URLs if the list or screenshot directory will be shared. For pages that need authentication, use an appropriate controlled browser context and protect both credentials and resulting images.

3. Install Playwright

Use Node.js and Playwright for the runnable example below. In a new project directory:

npm init -y
npm install playwright
npx playwright install chromium

Playwright runs a browser it controls, so the machine or container needs the browser dependencies and enough memory for the chosen concurrency. Start with one page at a time for stability; raise concurrency only after checking resource use and target-site behavior.

4. Run a bulk capture script

Save this as capture.mjs next to urls.txt. It captures a desktop viewport by default, supports optional full-page mode, waits for a selector or a delay, retries navigation failures, and writes a JSONL manifest with each result. Set FULL_PAGE=1, WAIT_FOR_SELECTOR, or WAIT_MS as needed.

import { chromium } from 'playwright';
import { createHash } from 'node:crypto';
import { mkdir, readFile, appendFile } from 'node:fs/promises';
import { URL } from 'node:url';

const input = await readFile('urls.txt', 'utf8');
const urls = [...new Set(input.split(/\r?\n/).map(s => s.trim()).filter(Boolean))];
const outDir = 'screenshots';
const manifestPath = `${outDir}/manifest.jsonl`;
const width = Number(process.env.VIEWPORT_WIDTH || 1440);
const height = Number(process.env.VIEWPORT_HEIGHT || 1000);
const fullPage = process.env.FULL_PAGE === '1';
const waitSelector = process.env.WAIT_FOR_SELECTOR;
const waitMs = Number(process.env.WAIT_MS || 0);
const maxAttempts = Number(process.env.RETRIES || 2) + 1;

if (!Number.isInteger(width) || width < 1 || !Number.isInteger(height) || height < 1) {
  throw new Error('Viewport width and height must be positive integers.');
}
await mkdir(outDir, { recursive: true });
const browser = await chromium.launch({ headless: true });
const context = await browser.newContext({ viewport: { width, height }, deviceScaleFactor: 1 });
const page = await context.newPage();

try {
  for (const url of urls) {
    let parsed;
    try {
      parsed = new URL(url);
      if (!['http:', 'https:'].includes(parsed.protocol)) throw new Error('Only HTTP and HTTPS URLs are supported.');
    } catch (error) {
      await appendFile(manifestPath, JSON.stringify({ url, ok: false, error: error.message }) + '\n');
      continue;
    }

    const id = createHash('sha256').update(url).digest('hex').slice(0, 12);
    const path = `${outDir}/${parsed.hostname.replace(/[^a-z0-9.-]/gi, '_')}-${id}.png`;
    let result;
    for (let attempt = 1; attempt <= maxAttempts; attempt++) {
      try {
        const response = await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30000 });
        if (waitSelector) await page.locator(waitSelector).waitFor({ state: 'visible', timeout: 15000 });
        if (waitMs > 0) await page.waitForTimeout(waitMs);
        await page.screenshot({ path, fullPage });
        result = {
          url, finalUrl: page.url(), status: response?.status() ?? null,
          ok: true, path, width, height, fullPage, attempt,
          capturedAt: new Date().toISOString()
        };
        break;
      } catch (error) {
        result = { url, ok: false, attempt, error: error.message, capturedAt: new Date().toISOString() };
        if (attempt < maxAttempts) await page.waitForTimeout(500 * attempt);
      }
    }
    await appendFile(manifestPath, JSON.stringify(result) + '\n');
    console.log(result.ok ? `Saved ${result.path}` : `Failed ${url}: ${result.error}`);
  }
} finally {
  await context.close();
  await browser.close();
}

Run the default viewport capture:

node capture.mjs

Capture the full scrollable page, wait for a page-specific element, or add a short settling delay:

FULL_PAGE=1 node capture.mjs
WAIT_FOR_SELECTOR='main h1' node capture.mjs
WAIT_MS=1500 node capture.mjs

On Windows PowerShell, set the values for the current process before running the script:

$env:FULL_PAGE='1'; node capture.mjs
$env:WAIT_FOR_SELECTOR='main h1'; node capture.mjs

The manifest records the requested and final URL, HTTP status when available, file path, capture mode, dimensions, attempt number, timestamp, and errors. Treat an HTTP error status separately from a navigation exception: a page can render an error response and still produce an image. Review the manifest before using the set as a complete comparison.

5. Add concurrency carefully

The example deliberately processes URLs sequentially. For larger lists, a small worker pool can reduce wall-clock time, but each browser page consumes resources, and rapid concurrent requests can burden target sites or trigger rate limits. Use a bounded concurrency limit, per-origin pacing where appropriate, and retries with backoff for transient navigation failures. Avoid retrying every failure blindly: invalid URLs and persistent 404 responses are not fixed by repeated attempts.

For repeatable reviews, keep each run in its own dated directory and store a manifest with the viewport, browser version, capture mode, wait strategy, and requested URLs. Do not overwrite a previous run until you have decided whether you need historical comparisons.

6. Capture mobile, full-page, or a selected element

To capture a mobile layout, create the context with a mobile viewport and device settings. For example, change the context creation in the script to:

const context = await browser.newContext({
  viewport: { width: 390, height: 844 },
  deviceScaleFactor: 1,
  isMobile: true,
  hasTouch: true
});

A viewport screenshot uses the current viewport unless fullPage: true is set. To capture one element rather than the page, wait for it and pass a path to its screenshot method:

const hero = page.locator('main .hero');
await hero.waitFor({ state: 'visible' });
await hero.screenshot({ path: 'hero.png' });

Element screenshots can fail when a selector matches nothing, is hidden, or is covered by a changing layout. Make selectors specific and wait for the target element to become visible. A full-page capture can be very tall; huge pages may require more memory and produce large image files.

7. Use bytes when you need downstream processing

Playwright can return screenshot bytes instead of writing a file directly. That lets a workflow send the image to object storage, an image-difference tool, or another processing step. The screenshot operation itself does not determine whether an SEO page is indexed or ranks well.

const bytes = await page.screenshot({ fullPage: true, type: 'png' });
// Pass bytes to your chosen storage or image-processing client.

Google Search Console is for selected diagnostics

Google Search Console URL Inspection can show how Google’s inspection crawler rendered a page, but it is not a general bulk screenshot export workflow. Google says the rendered screenshot is available only for a successful live test; the page must be reachable, and screenshots are not available for indexed results or unsuccessful fetches. Live testing also does not check every possible indexing issue. See Google’s URL Inspection documentation.

Use URL Inspection to investigate selected URLs from Google’s perspective. For requesting indexing of many new or updated pages, Google recommends submitting a sitemap; indexing requests also have a daily limit. Sitemap submission is indexing guidance, separate from taking screenshots.

Or skip the browser setup

ScreenshotNeo is a screenshot API and MCP server. For a one-off capture, make one GET request; use the API’s documented bulk option for batches of up to 100 URLs per call. See the API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie and consent banners are accepted where possible, and 60+ known consent platforms, newsletter popups, and chat widgets can be removed before capture; each cleanup step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing. Response headers say the page verdict and whether the capture was billed. An MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000.

Create a free account for 1,000 screenshots a month, with no card required.

Reliability, performance, and cost considerations

  • Throughput: Total runtime depends on navigation time, page rendering, selected waits, image size, and concurrency. Measure a representative subset before estimating a whole batch.
  • Stability: Use bounded retries for transient errors, log failures, and preserve a manifest. A screenshot alone cannot tell you whether a page was a successful SEO landing page; inspect response status and final URL too.
  • Consistency: Fix the viewport, device scale factor, browser, and wait rule across a run. Ads, personalization, geolocation, A/B tests, and changing content can still make captures differ.
  • Storage: Full-page PNGs may be large. Choose a format and retention period based on whether you need visual fidelity, compact storage, or pixel comparison.
  • Cost: Self-managed Playwright has no per-screenshot API fee, but it uses compute, storage, maintenance, and engineering time. A hosted API trades browser setup and maintenance for service pricing; compare batch support, retries, output controls, and total workload against your needs.

Troubleshooting

Symptom Likely cause What to do
Navigation timeout Slow server, long-running requests, or a page that never reaches the chosen event. Use a suitable navigation event such as domcontentloaded, wait for the content you need, set an intentional timeout, and retry transient failures with backoff.
Screenshot is blank or incomplete Client rendering or lazy content has not appeared yet. Wait for a meaningful selector or a reasonable delay. For long pages, scroll or otherwise trigger lazy loading before the full-page capture.
Cookie banner, popup, or chat obscures content The page presents overlays to a new visitor. For a self-managed browser, handle the consent flow or close the overlay in site-specific automation when permitted. Hosted cleanup behavior depends on the service and its settings.
File overwritten or wrong image matched Distinct URLs produced the same simplistic filename. Use a stable hash of the full URL, as in the example, and keep per-run directories.
Element screenshot fails Selector is absent, hidden, or ambiguous. Wait for a specific visible locator and verify the selector against the rendered page.
HTTP error page saved as an image Navigation succeeded at the network level, but returned an error status. Check and record the response status; treat the image as diagnostic output, not a successful landing-page capture.
Browser fails to launch in a container Browser binaries or system dependencies are missing, or memory is constrained. Install Playwright’s browser and required dependencies, reduce concurrency, and check the container’s resource limits.
Run is slow or target returns rate limits Too many simultaneous navigations or expensive waits. Reduce concurrency, pace requests per origin, and avoid fixed delays when a selector can express readiness.

Frequently asked questions

Can I use screenshots to prove that a page has an SEO problem?

No. A screenshot helps inspect rendered appearance. It does not establish indexing status, ranking, or the cause of a search issue; use appropriate search diagnostics for those questions.

Should every capture be full-page?

No. Use viewport captures for consistent above-the-fold review and full-page captures when the whole page needs inspection. Full-page output is larger and can be harder to compare at a glance.

Can Search Console capture all my landing pages at once?

The documented URL Inspection screenshot is part of a qualifying live test for an inspected URL, with availability limits. It is not a general bulk screenshot export tool.

What should I keep with each image?

At minimum, preserve the source URL, final URL, capture timestamp, viewport, capture mode, and outcome. This makes the screenshot useful as a review artifact later.