ScreenshotNeo

BlogHow-to

How to Capture Screenshots of Multiple Web Pages

Learn how to capture viewport or full-page screenshots from a URL list with browser tools, scripts, retries, filenames, and ScreenshotNeo.

By the ScreenshotNeo team1 October 20268 min read

To capture screenshots of multiple web pages, put the URLs in a list, visit each page with a browser automation tool, wait for the right page state, capture either the viewport or the full scrollable page, and save every result with a unique filename. For a few pages, a browser’s built-in screenshot command is fastest. For repeated work, use Playwright, Puppeteer, shot-scraper, webshot2, or a screenshot API.

Choose the right capture method

Need Good starting point Decisions to make
A few pages once Built-in browser screenshot Whether your browser supports full-page capture and where it saves or copies the image
A repeatable URL list Playwright script or CLI wrapper Browser engine, viewport, full-page mode, waits, filenames, retries
JavaScript browser automation Puppeteer Chrome setup, navigation timing, output handling
R batch jobs webshot2 Parallelism, output names, available CPU and memory
An API or server workflow ScreenshotNeo Authentication, capture options, caching, billing behavior

Firefox documents visible-area and full-page screenshots in its built-in tools. See Mozilla’s screenshot instructions. Playwright and Puppeteer are better when the same operation must run against many URLs.

Viewport screenshots versus full-page screenshots

A viewport screenshot contains the browser area currently visible at the chosen width and height. A full-page screenshot includes the page’s scrollable content. In Playwright, use fullPage: true for the latter; its screenshot API also supports format, clipping, and quality options. Playwright screenshot documentation

  • Use viewport capture for responsive layout checks, hero sections, or a consistent card-sized image.
  • Use full-page capture for audits, documentation, archives, and long landing pages.
  • Use a clip or element capture when only one region matters; this keeps files smaller and makes comparisons easier.

Prepare a URL list

Keep one URL per line and validate it before launching a browser. Assign names yourself when page identity matters; otherwise derive a sanitized name from the hostname and path.

https://example.com/
https://example.com/pricing
https://www.mozilla.org/en-US/firefox/

Use HTTPS where available, preserve query strings when they change the page, and remove duplicate URLs if repeated captures are not intentional.

Capture multiple pages with Playwright (Node.js)

The following script is a complete illustrative workflow. It visits each URL, waits for navigation to settle, captures a viewport or full page, writes distinct filenames, and continues after failures.

import { chromium } from 'playwright';
import { readFile } from 'node:fs/promises';
import path from 'node:path';
import fs from 'node:fs';

const urls = (await readFile('urls.txt', 'utf8'))
  .split(/\\r?\\n/)
  .map(line => line.trim())
  .filter(Boolean);

fs.mkdirSync('screenshots', { recursive: true });

function filenameFor(url, index) {
  const u = new URL(url);
  const identity = `${u.hostname}${u.pathname}`
    .replace(/[^a-z0-9]+/gi, '-')
    .replace(/^-|-$/g, '')
    .toLowerCase() || 'page';
  return path.join('screenshots', `${String(index + 1).padStart(3, '0')}-${identity}.png`);
}

const browser = await chromium.launch();
const context = await browser.newContext({
  viewport: { width: 1440, height: 900 },
  deviceScaleFactor: 1
});

for (let i = 0; i < urls.length; i++) {
  const url = urls[i];
  const page = await context.newPage();
  const output = filenameFor(url, i);
  try {
    await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 45_000 });
    await page.waitForLoadState('networkidle', { timeout: 15_000 }).catch(() => {});
    await page.screenshot({
      path: output,
      fullPage: true,
      type: 'png'
    });
    console.log(`saved ${url} -> ${output}`);
  } catch (error) {
    console.error(`failed ${url}: ${error.message}`);
  } finally {
    await page.close();
  }
}

await browser.close();

Install and run it with:

npm install playwright
npx playwright install chromium
node capture.mjs

Change fullPage: true to false for viewport captures. Set type to jpeg or webp when those formats suit your pipeline. The Page API documents navigation and screenshot controls.

Capture pages with Playwright (Python)

from pathlib import Path
from urllib.parse import urlparse
import re
from playwright.sync_api import sync_playwright

urls = [line.strip() for line in Path("urls.txt").read_text().splitlines() if line.strip()]
Path("screenshots").mkdir(exist_ok=True)

def filename_for(url, index):
    parsed = urlparse(url)
    raw = f"{parsed.netloc}{parsed.path}"
    name = re.sub(r"[^a-zA-Z0-9]+", "-", raw).strip("-").lower() or "page"
    return Path("screenshots") / f"{index:03d}-{name}.png"

with sync_playwright() as p:
    browser = p.chromium.launch()
    context = browser.new_context(viewport={"width": 1440, "height": 900})
    for index, url in enumerate(urls, start=1):
        page = context.new_page()
        output = filename_for(url, index)
        try:
            page.goto(url, wait_until="domcontentloaded", timeout=45_000)
            try:
                page.wait_for_load_state("networkidle", timeout=15_000)
            except Exception:
                pass
            page.screenshot(path=str(output), full_page=True)
            print(f"saved {url} -> {output}")
        except Exception as exc:
            print(f"failed {url}: {exc}")
        finally:
            page.close()
    browser.close()
python -m pip install playwright
python -m playwright install chromium
python capture.py

Use a command-line wrapper

shot-scraper is a Playwright-based command-line tool. It can be useful when you want shell scripts, scheduled jobs, or simple URL-driven capture without maintaining browser code. Check its current installation and browser setup instructions before deploying it.

Use Puppeteer in JavaScript

Puppeteer is a JavaScript browser automation library; Chrome’s documentation lists screenshots and PDF generation among its uses. Puppeteer documentation

import puppeteer from 'puppeteer';
import { readFile } from 'node:fs/promises';

const urls = (await readFile('urls.txt', 'utf8')).split(/\\r?\\n/).filter(Boolean);
const browser = await puppeteer.launch();
const page = await browser.newPage();
await page.setViewport({ width: 1440, height: 900 });

for (let i = 0; i < urls.length; i++) {
  try {
    await page.goto(urls[i], { waitUntil: 'networkidle2', timeout: 45_000 });
    await page.screenshot({ path: `screenshots/page-${i + 1}.png`, fullPage: true });
  } catch (error) {
    console.error(`failed ${urls[i]}: ${error.message}`);
  }
}
await browser.close();

Use webshot2 for an R workflow

The webshot2 documentation describes loading and capturing multiple URLs in parallel. Parallel capture can reduce elapsed time, but increase it gradually so the machine has enough memory and browser capacity.

Make captures repeatable

  • Fix the viewport: use the same width, height, and device scale factor for every URL.
  • Choose a wait condition: domcontentloaded is quick; a later network or selector wait is safer for dynamic pages.
  • Control animations: inject CSS to pause transitions when pixel comparisons require stable frames.
  • Handle lazy content: full-page capture may trigger loading, but pages differ; scroll or wait for a known selector when needed.
  • Keep URL identity: include an index plus sanitized host/path so two pages do not overwrite each other.
  • Record outcomes: write URL, timestamp, output path, status, and error text to a log.

Authentication, cookies, and private pages

Public pages can be opened directly. For private pages, create an authenticated browser context, load saved storage state, or set cookies and headers before navigation. Never place credentials in a URL or commit storage-state files to source control. If a page depends on a short-lived session, refresh the login state before the batch begins and fail clearly when authentication redirects to a sign-in page.

Performance and reliability

  • Reuse one browser process and create pages or contexts per job; launching a new browser for every URL adds overhead.
  • Limit concurrency to what the host can handle. Too many simultaneous pages can cause memory pressure, throttling, or timeouts.
  • Retry transient navigation failures with a small bounded retry count and a delay. Do not retry invalid URLs indefinitely.
  • Save each image as soon as it succeeds so a later failure does not discard completed work.
  • Use deterministic settings for visual regression: browser engine, viewport, fonts, timezone, locale, and color scheme.
  • Expect third-party ads, consent dialogs, bot checks, and personalized content to change results between runs.

File size and cost considerations

PNG preserves detail but can be large. JPEG is smaller for photographic pages and supports quality settings. WebP can reduce transfer size when your consumers support it. Full-page images consume more memory than viewport captures, especially on very long pages. Browser automation costs are your machine, runtime, and any hosted browser resources; an API adds request charges and removes much of the browser maintenance.

Troubleshooting

Symptom Likely cause Fix
Screenshot is only the top section Viewport capture was selected Enable full-page capture or capture the required element.
Images or fonts are missing Capture happened before resources finished loading Wait for a selector, a longer delay, or an appropriate network state.
Page is blank Navigation failed, JavaScript crashed, or a bot check blocked the browser Log the final URL and response, increase diagnostics, and inspect the page manually.
Timeout errors Slow origin, blocked request, or overly short timeout Increase the timeout for that site, retry once, and continue recording failures.
Files overwrite each other Names were based only on a repeated slug Add an index and sanitized host/path, or assign explicit names.
Different runs do not match Responsive layout, personalization, animation, or changing content Fix viewport and locale, pause animations, authenticate consistently, and capture at a defined time.
Browser executable is missing Playwright or Puppeteer package is installed without its browser Run the tool’s browser installation command for the chosen engine.

Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF, and you can call it once per URL or use its bulk capture option for up to 100 URLs per call.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo API documentation for request parameters. Its options include full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets or custom viewports, retina scale, PDF settings, custom CSS and JavaScript, clicks, waits, ad and tracker blocking, custom headers/cookies/user agents, timezone and geolocation, transparent backgrounds, resizing, TTL caching, signed image links, asynchronous jobs with signed webhooks, and a usage API.

Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot, with each cleanup step configurable. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers report the page verdict and whether it was billed. An MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots, and every feature is available on every plan.

Create a free ScreenshotNeo account and start with 1,000 screenshots a month at no charge.

FAQ

Can I capture a list without opening each page manually?

Yes. A Playwright, Puppeteer, shot-scraper, webshot2, or ScreenshotNeo workflow can iterate over the list and save each result.

Should I capture the viewport or the full page?

Use viewport mode for a fixed screen view. Use full-page mode when the entire scrollable document matters.

How do I prevent one failed URL from stopping the batch?

Wrap each navigation and screenshot in its own error handler, log the failure, and continue to the next URL.

Why do screenshots of the same URL differ?

Dynamic content, cookies, authentication, responsive dimensions, animations, ads, and page changes can all alter pixels. Fix those inputs when repeatability matters.

Can a screenshot workflow produce PDFs too?

Browser tools such as Puppeteer support PDF generation, and ScreenshotNeo provides PDF capture with paper size, margins, orientation, and page-range options.