ScreenshotNeo

BlogHow-to

How to Bulk Screenshot Pages and Preserve the URL in Each Image Filename

Capture a list of pages with Playwright, save URL-based filenames safely, and keep a manifest so every image maps back to its source.

By the ScreenshotNeo team4 October 20269 min read

To bulk screenshot pages and preserve each URL in its image filename, read URLs from a list, capture each page with browser automation, and derive a safe filename from the URL. Add a short hash to avoid collisions, and keep a manifest containing the exact original URL and saved filename. The runnable Playwright example below captures a list of URLs, supports viewport or full-page screenshots, and records failures without stopping the batch.

1. Choose a capture workflow

For a recurring or customized batch, Playwright gives you direct control over URL-derived names, navigation timeouts, readiness checks, and failure handling. For a known, mostly static URL set, shot-scraper can take multiple screenshots from a YAML file with a URL and output name per item. For a one-off capture, the Playwright CLI accepts a custom filename and a full-page option. These approaches differ in how much custom code and configuration they need; choose based on how often the job runs and how much control it needs.

Approach Good fit URL naming
Playwright API Custom naming, readiness conditions, and error handling Build any safe path in code
shot-scraper multi A defined set of pages maintained in YAML Set an explicit output for each entry
Playwright CLI Occasional manual captures Pass a filename; otherwise the CLI uses a timestamp-based default

2. Install Playwright

This example uses Node.js and Chromium. In a new directory, install Playwright and its browser:

npm init -y
npm install playwright
npx playwright install chromium

Create urls.txt with one absolute HTTP or HTTPS URL per line:

https://example.com/
https://example.com/docs/getting-started
https://example.org/products?category=books

3. Bulk capture with URL-derived filenames

Save this as capture.mjs. It writes screenshots into screenshots/ and a manifest.jsonl file with the requested URL, final URL, output path, capture time, and status. A short SHA-256 suffix distinguishes URLs that normalize to the same readable filename, such as URLs with different queries or trailing slash forms.

import { chromium } from 'playwright';
import { createHash } from 'node:crypto';
import { mkdir, readFile, appendFile } from 'node:fs/promises';
import path from 'node:path';

const inputFile = process.argv[2] ?? 'urls.txt';
const outputDir = process.argv[3] ?? 'screenshots';
const fullPage = process.env.FULL_PAGE === '1';
const width = Number(process.env.VIEWPORT_WIDTH ?? 1440);
const height = Number(process.env.VIEWPORT_HEIGHT ?? 1000);
const navigationTimeoutMs = Number(process.env.NAVIGATION_TIMEOUT_MS ?? 45000);

function filenameFor(rawUrl) {
  const u = new URL(rawUrl);
  // Keep host and path readable; query and fragment are represented in the hash.
  const readable = `${u.hostname}${u.pathname}`
    .replace(/[^a-zA-Z0-9._-]+/g, '-')
    .replace(/-{2,}/g, '-')
    .replace(/^[.-]+|[.-]+$/g, '')
    .slice(0, 100) || 'page';
  const hash = createHash('sha256').update(rawUrl).digest('hex').slice(0, 10);
  return `${readable}-${hash}.png`;
}

const urls = (await readFile(inputFile, 'utf8'))
  .split(/\r?\n/)
  .map(line => line.trim())
  .filter(line => line && !line.startsWith('#'));

await mkdir(outputDir, { recursive: true });
const manifestPath = path.join(outputDir, 'manifest.jsonl');
const browser = await chromium.launch({ headless: true });

try {
  const context = await browser.newContext({ viewport: { width, height } });
  const page = await context.newPage();
  for (const requestedUrl of urls) {
    const output = path.join(outputDir, filenameFor(requestedUrl));
    const record = {
      requested_url: requestedUrl,
      final_url: null,
      filename: output,
      captured_at: new Date().toISOString(),
      status: 'error'
    };
    try {
      const response = await page.goto(requestedUrl, {
        waitUntil: 'domcontentloaded',
        timeout: navigationTimeoutMs
      });
      // Replace this with a page-specific selector when the content has a known ready marker.
      await page.waitForLoadState('load', { timeout: 10000 }).catch(() => {});
      record.final_url = page.url();
      record.http_status = response?.status() ?? null;
      await page.screenshot({ path: output, fullPage, animations: 'disabled' });
      record.status = response && response.status() >= 400 ? 'http_error_captured' : 'ok';
    } catch (error) {
      record.error = error instanceof Error ? error.message : String(error);
    }
    await appendFile(manifestPath, `${JSON.stringify(record)}\n`);
    console.log(`${record.status}: ${requestedUrl} -> ${output}`);
  }
  await context.close();
} finally {
  await browser.close();
}

Run the batch in viewport mode:

node capture.mjs urls.txt screenshots

To capture the full scrollable document instead, set FULL_PAGE=1:

FULL_PAGE=1 node capture.mjs urls.txt screenshots

The script uses one browser page sequentially, which limits resource use and makes failures easy to associate with their URLs. The initial navigation waits for DOM content, then gives the page up to ten seconds to reach the load state. If the page has a reliable ready marker, add await page.locator('[data-ready="true"]').waitFor() before the screenshot. Use a selector that actually exists on the target site.

4. Build safe, traceable filenames

A URL is not automatically a safe filename. It can contain slashes, query parameters, fragments, reserved punctuation, Unicode, or enough characters to exceed filesystem limits. A practical name contains a readable host and path plus a short hash of the full URL.

  • Sanitize characters: replace characters outside letters, digits, dot, underscore, and hyphen.
  • Limit length: keep the readable portion bounded and leave room for the hash and extension.
  • Prevent collisions: hash the original URL, including its query and fragment, so similar paths remain distinguishable.
  • Keep provenance: preserve the exact URL in a manifest. A sanitized filename is useful to humans, but it cannot reliably reconstruct every original URL.

The example hashes the input URL exactly as provided. If equivalent URL spellings should share one image name, normalize them first according to your own rules; for auditability, retaining the requested and redirected final URL separately is safer.

5. Viewport versus full-page capture

A viewport capture records what is visible at the selected browser dimensions. Use it for layout comparisons at a consistent screen size. A full-page capture records the entire scrollable page in one tall image; it can become unwieldy for very long or infinite-scroll pages. Playwright documents the fullPage screenshot option and supports CSS-pixel or device-pixel output scale. CSS scale produces one output pixel per CSS pixel; device scale can create larger images on high-DPI configurations. See the Playwright Page API.

For the same site comparison across runs, keep the browser version, operating system, viewport, device scale, and capture timing consistent. Playwright notes that rendering may vary with the operating system, browser version, settings, hardware, power source, and headless mode; see its visual comparisons documentation.

6. Other batch options

shot-scraper YAML batch

shot-scraper’s multi command reads a YAML list whose entries specify a URL and output path. Its documentation also describes per-image dimensions, waits, and selectors. Use explicit output names for predictable URL-based naming; do not depend on a default name when the batch needs a stable naming convention.

- url: https://example.com/
  output: screenshots/example-com-home.png
- url: https://example.org/products?category=books
  output: screenshots/example-org-products-books.png
  width: 1440
  height: 1000

Use the syntax documented for your installed shot-scraper release, especially for wait and selector fields. The source documentation shows the multi-URL YAML workflow and output names: shot-scraper documentation.

Playwright CLI

For a quick single-page capture, the Playwright CLI supports a custom filename and full-page screenshots. Its screenshot command documentation describes --filename and --full-page: Screenshots & PDF. For a large URL list with URL-derived names and a manifest, a script or configured multi-page job is easier to maintain than repeated manual CLI commands.

7. Reliability, speed, and cost

  • Readiness: fixed sleeps are simple but waste time on fast pages and can still be too short on slow pages. Prefer a relevant selector or explicit page state when available.
  • Concurrency: the sample runs sequentially to keep memory use predictable. For faster batches, use a small worker pool with separate pages, and tune it against the sites’ rate limits and the machine’s available memory. Excess concurrency can cause timeouts or make pages behave differently.
  • Retries: retry transient navigation failures a limited number of times, with a delay between attempts. Record every final status in the manifest; do not silently replace a failed capture with a stale file.
  • Dynamic content: animations, personalization, consent prompts, delayed assets, and live data can change the result. Disable animations where appropriate and choose a capture state suited to the content.
  • Long pages: full-page images consume more memory and disk and may be difficult to review. Use viewport captures when the visible layout is the actual comparison target.
  • Direct costs: Playwright and shot-scraper are browser automation approaches; the workflow itself does not require a paid screenshot API. Budget for the machine, storage, and operational time used to run and retain captures.
  • Audit trail: save requested URL, final URL, output path, timestamp, HTTP status, and error. This helps distinguish redirects, HTTP error pages, and failed navigations from successful images.

8. Troubleshooting

Symptom Likely cause Fix
Two pages overwrite the same file The filename uses only a sanitized host or path Add a hash of the complete URL or a unique batch index; keep the original URL in the manifest.
Screenshot is blank or incomplete The page had not rendered the relevant content when capture began Wait for a page-specific selector or state. Increase navigation timeout only if the site legitimately needs more time.
Images are missing Lazy-loaded images have not entered the viewport, or external resources are delayed Use full-page capture where appropriate and add a site-specific scroll-and-wait routine if lazy assets must load before capture.
Capture timed out Slow navigation, a hanging resource, or an overly strict readiness condition Use domcontentloaded as the navigation milestone, set a bounded timeout, and wait separately for the content you need.
Filename is too long or rejected URL path was copied directly to the filesystem Sanitize it, cap the readable segment, and retain a short hash.
Images differ between runs Browser environment or page content changed Pin browser/runtime settings where possible and keep viewport, scale, timing, and page state consistent. Dynamic pages may still vary.
Some URLs stop the whole job An uncaught navigation or screenshot exception terminates the loop Catch errors per URL, write an error record to the manifest, and continue with the remaining entries, as the example does.
Manifest contains old entries The script appends to an existing JSONL file Move or remove the previous manifest before a fresh run, or change the script to create a run-specific manifest name.

9. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. One GET request captures a URL as an image; its API also supports bulk capture of up to 100 URLs per call. Each capture request can return an image, and your application can save the response bytes using a filename derived from the requested URL. See the ScreenshotNeo documentation.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests

url = "https://stripe.com"
r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": url},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await (await import('node:fs/promises')).writeFile('shot.webp', bytes);

To preserve names in your own batch, apply the filename function from the Playwright example to each input URL, then save each response body to that path. ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are never billed; its MCP server lets AI agents take screenshots; and 1,000 screenshots a month are free with no card, with paid plans starting at $5 for 3,000. Sign up for 1,000 free screenshots a month with no card.

10. Frequently asked questions

How do I save a screenshot with the page URL in the filename?

Sanitize the URL’s host and path for readability, add a short hash of the full URL to prevent collisions, and save the exact URL in a manifest.

Should query parameters be included in the visible filename?

Usually keep them out of the readable portion because they can make names long or expose sensitive values. Hash the full URL so query-distinct URLs still get distinct names, and keep the source URL in a protected manifest if it contains private parameters.

How can I capture full pages instead of the viewport?

With Playwright, use fullPage: true in page.screenshot. Long pages produce tall files, so use viewport mode when only a consistent visible area is needed.

Can I recover the exact URL from the image filename?

No. Sanitization and hashing are intentionally not reversible. Store the original URL alongside the filename in a manifest.

What should a batch manifest contain?

At minimum, retain the requested URL, output filename, capture timestamp, and status. For redirect and failure review, also record the final URL, HTTP status, and error details.