ScreenshotNeo

BlogHow-to

How to Take Screenshots of a URL List and Create an HTML Gallery

Capture a list of URLs with Playwright, save consistent screenshots, and build a responsive HTML gallery with source links and clear failure states.

By the ScreenshotNeo team4 October 202615 min read

To take screenshots of a URL list and create an HTML gallery, use a browser automation library to visit each URL, save an image with a stable filename, and write an index.html that displays those images with captions and links back to their source pages. This guide uses Playwright with Node.js, then shows how to run the same capture idea with Python, cURL, and ScreenshotNeo.

1. Choose the capture approach

Playwright or Puppeteer can navigate a real browser page and save a screenshot. Use the library that fits your existing runtime and browser setup. The documentation reviewed here supports the navigation and screenshot workflows; it does not establish a performance winner.

Choice Use it when
Playwright You want a browser automation script with explicit control over navigation, viewport, screenshot type, and full-page capture.
Puppeteer Your project already uses Puppeteer and you want to keep its existing browser dependency.
Screenshot API You want to make HTTP requests instead of installing and operating a browser. See the ScreenshotNeo option below.

For a gallery, keep each original URL in a record alongside its image filename. That makes it possible to show the source link and status even if two pages have the same title, share a hostname, or redirect.

2. Prepare the URL list and project

Use a plain text file with one absolute URL per line. Include the scheme (https:// or http://); a bare hostname is not a complete navigation URL.

https://example.com/
https://developer.mozilla.org/
https://www.wikipedia.org/

Create a project and install Playwright. The following commands use npm:

mkdir url-gallery
cd url-gallery
npm init -y
npm install playwright
npx playwright install chromium

Save the URL list as urls.txt in this directory. The script below creates a gallery/ folder with an image, a JSON manifest, and the gallery page.

Save this as gallery.mjs. It uses one browser page in sequence to keep the script simple and the output deterministic. It records navigation failures per URL and continues, while making sure the browser closes even if an unexpected error occurs.

import { chromium } from 'playwright';
import { mkdir, readFile, writeFile } from 'node:fs/promises';
import path from 'node:path';

const inputPath = process.argv[2] ?? 'urls.txt';
const outputDir = process.argv[3] ?? 'gallery';
const fullPage = process.env.FULL_PAGE === '1';
const viewport = { width: 1440, height: 900 };
const timeoutMs = Number(process.env.NAVIGATION_TIMEOUT_MS ?? 30000);

function escapeHtml(value) {
  return String(value).replace(/[<>"'&]/g, (char) => ({
    '&': '&amp;', '<': '&lt;', '>': '&gt;',
    '"': '&quot;', "'": '''
  })[char]);
}

function parseUrls(text) {
  return text.split(/\r?\n/)
    .map((line) => line.trim())
    .filter((line) => line && !line.startsWith('#'));
}

const urls = parseUrls(await readFile(inputPath, 'utf8'));
if (urls.length === 0) throw new Error(`No URLs found in ${inputPath}`);
for (const url of urls) {
  let parsed;
  try { parsed = new URL(url); } catch { throw new Error(`Invalid absolute URL: ${url}`); }
  if (!['http:', 'https:'].includes(parsed.protocol)) {
    throw new Error(`Only http and https URLs are supported: ${url}`);
  }
}

await mkdir(outputDir, { recursive: true });
const browser = await chromium.launch({ headless: true });
const records = [];

try {
  const page = await browser.newPage({ viewport });
  page.setDefaultNavigationTimeout(timeoutMs);

  for (const [index, url] of urls.entries()) {
    const filename = `page-${String(index + 1).padStart(3, '0')}.png`;
    const record = { index: index + 1, url, filename, status: null, finalUrl: null, title: '', error: null };

    try {
      // domcontentloaded is a starting point, not a guarantee that every
      // lazy image or site-specific widget has finished rendering.
      const response = await page.goto(url, { waitUntil: 'domcontentloaded', timeout: timeoutMs });
      record.status = response?.status() ?? null;
      record.finalUrl = page.url();
      record.title = await page.title().catch(() => '');

      // Add a task-specific wait here when the target requires it, e.g.:
      // await page.locator('main article').waitFor({ state: 'visible', timeout: 10000 });
      await page.screenshot({
        path: path.join(outputDir, filename),
        fullPage,
        type: 'png'
      });
    } catch (error) {
      record.error = error instanceof Error ? error.message : String(error);
    }
    records.push(record);
  }
} finally {
  await browser.close();
}

await writeFile(path.join(outputDir, 'manifest.json'), JSON.stringify(records, null, 2));

const cards = records.map((record) => {
  const title = record.title || record.finalUrl || record.url;
  const source = escapeHtml(record.url);
  const caption = escapeHtml(title);
  const status = record.status === null ? 'No HTTP response recorded' : `HTTP ${record.status}`;
  const details = record.error
    ? `<p class="error">Capture failed: ${escapeHtml(record.error)}</p>`
    : `<img src="${escapeHtml(record.filename)}" alt="Screenshot of ${caption}" loading="lazy">`;
  return `
    <article class="card">
      <a class="preview" href="${source}" target="_blank" rel="noopener noreferrer">${details}</a>
      <div class="meta">
        <h2>${caption}</h2>
        <p>${escapeHtml(status)}</p>
        <a class="source" href="${source}" target="_blank" rel="noopener noreferrer">Open source page</a>
      </div>
    </article>`;
}).join('\n');

const html = `<!doctype html>
<html lang="en">
<head>
  <meta charset="utf-8">
  <meta name="viewport" content="width=device-width, initial-scale=1">
  <title>URL Screenshot Gallery</title>
  <style>
    * { box-sizing: border-box; }
    body { margin: 0; padding: 2rem; color: #172033; background: #f4f6fa; font: 16px/1.5 system-ui, sans-serif; }
    main { max-width: 1500px; margin: 0 auto; }
    header { margin-bottom: 1.5rem; }
    h1 { margin: 0 0 .25rem; font-size: clamp(1.5rem, 3vw, 2.25rem); }
    .grid { display: grid; grid-template-columns: repeat(auto-fit, minmax(min(100%, 320px), 1fr)); gap: 1rem; }
    .card { overflow: hidden; border: 1px solid #dce2ec; border-radius: 12px; background: white; box-shadow: 0 2px 8px #1720330d; }
    .preview { display: grid; min-height: 180px; place-items: center; background: #e9edf4; color: #8b2530; text-decoration: none; }
    .preview img { display: block; width: 100%; height: 240px; object-fit: contain; background: #e9edf4; }
    .meta { padding: 1rem; }
    .meta h2 { overflow-wrap: anywhere; margin: 0 0 .35rem; font-size: 1rem; }
    .meta p { margin: 0 0 .5rem; color: #536078; font-size: .875rem; }
    .source { overflow-wrap: anywhere; }
    .error { padding: 1rem; margin: 0; }
  </style>
</head>
<body>
  <main>
    <header><h1>URL Screenshot Gallery</h1><p>${records.length} URL(s) processed. Select a card to open its source page.</p></header>
    <section class="grid" aria-label="Page screenshots">${cards}
    </section>
  </main>
</body>
</html>`;
await writeFile(path.join(outputDir, 'index.html'), html);
console.log(`Wrote ${records.length} records to ${outputDir}/index.html`);
console.log(`${records.filter((record) => record.error).length} capture(s) failed; see manifest.json.`);

Note: The source above is HTML-escaped for display in this article. In your gallery.mjs file, replace entities such as &lt;, &gt;, and &amp; with their literal characters (<, >, and &). This keeps the displayed code readable in HTML while ensuring the file itself contains valid JavaScript and markup.

Run it with:

node gallery.mjs

To capture each page’s full scrollable content instead of only the viewport:

FULL_PAGE=1 node gallery.mjs

Open gallery/index.html in a browser. Keep the HTML, manifest, and screenshot files together when you share the gallery; its image paths are relative to the HTML file.

4. Pick the right screenshot and readiness settings

Viewport or full page

The default script captures the 1440 by 900 CSS-pixel viewport. This is useful for comparing many sites because the thumbnails show a consistent above-the-fold area. Set FULL_PAGE=1 to capture the entire scrollable document. Full-page images can be very tall and slower to review in a grid. Playwright describes full-page capture as making the whole scrollable page fit on a very tall screen.

Wait for the content your task needs

domcontentloaded means the initial document has been parsed; it does not guarantee that every image, font, animation, or delayed component is visually ready. If a page has important content that appears later, wait for a meaningful selector or add a deliberate short delay. For example, after navigation:

await page.locator('main article').waitFor({ state: 'visible', timeout: 10000 });

Playwright also offers load, commit, and networkidle navigation conditions. Network idle represents a period with no network connections for 500 ms, but Playwright discourages using it as a general testing readiness signal. Pages with long polling or continuously loaded resources may never become idle; delayed content can also appear after an idle interval.

Image format and scale

PNG is the default in the example and preserves detail without lossy compression. Playwright also supports JPEG and WebP screenshots; JPEG and WebP quality settings can reduce file size when a smaller gallery matters more than lossless pixels. Device scale factor can increase image dimensions and file size. Use a consistent viewport and scale for comparisons; choose higher scale only when the extra detail is useful.

Save to a file or use a buffer

A path is convenient for this folder-based gallery. If another processing step needs image bytes, capture a buffer instead and write or transform it yourself:

const image = await page.screenshot({ type: 'png', fullPage: false });
// image is a Buffer; pass it to an image-processing step or write it to storage.
  • Stable filenames: The row number avoids collisions when multiple URLs use the same host or title. The manifest preserves the input URL, final URL, status, and filename together.
  • HTTP errors: A 404 or 500 response is still a response. Playwright navigation does not throw just because the server returned a valid HTTP error status; check and record the status yourself.
  • Escaping: Escape titles and URLs before inserting them into HTML. A title is page content and should not be treated as trusted markup. The example does this for text and attribute values.
  • Failed captures: Keep a visible failure card and record the error rather than silently omitting a URL. This makes a best-effort batch easy to audit.
  • Private pages: A screenshot can expose account or other private information. Only capture pages you are permitted to access, and avoid sharing a gallery containing private content publicly.
  • Local sharing: The gallery references local image files, so share the complete output directory or host the directory contents together.

6. Python alternative with Playwright

If your project is Python-based, install Playwright and its Chromium browser:

python -m pip install playwright
python -m playwright install chromium

Save this as gallery.py. It reads one URL per line and captures a viewport screenshot for each valid HTTP or HTTPS URL.

from pathlib import Path
from html import escape
from urllib.parse import urlparse
import json
from playwright.sync_api import sync_playwright

urls = [line.strip() for line in Path('urls.txt').read_text().splitlines()
        if line.strip() and not line.lstrip().startswith('#')]
out = Path('gallery')
out.mkdir(exist_ok=True)
records = []

for url in urls:
    parsed = urlparse(url)
    if parsed.scheme not in ('http', 'https') or not parsed.netloc:
        raise ValueError(f'Expected an absolute HTTP(S) URL: {url}')

with sync_playwright() as p:
    browser = p.chromium.launch()
    try:
        page = browser.new_page(viewport={'width': 1440, 'height': 900})
        page.set_default_navigation_timeout(30000)
        for number, url in enumerate(urls, 1):
            filename = f'page-{number:03}.png'
            record = {'index': number, 'url': url, 'filename': filename,
                      'status': None, 'final_url': None, 'title': '', 'error': None}
            try:
                response = page.goto(url, wait_until='domcontentloaded')
                record['status'] = response.status if response else None
                record['final_url'] = page.url
                record['title'] = page.title()
                page.screenshot(path=str(out / filename), full_page=False)
            except Exception as exc:
                record['error'] = str(exc)
            records.append(record)
    finally:
        browser.close()

(out / 'manifest.json').write_text(json.dumps(records, indent=2))
cards = []
for record in records:
    source = escape(record['url'], quote=True)
    title = escape(record['title'] or record['final_url'] or record['url'])
    status = 'No HTTP response recorded' if record['status'] is None else f"HTTP {record['status']}"
    if record['error']:
        preview = f"<p class='error'>Capture failed: {escape(record['error'])}</p>"
    else:
        preview = f"<img src='{escape(record['filename'], quote=True)}' alt='Screenshot of {title}' loading='lazy'>"
    cards.append(f"<article class='card'><a class='preview' href='{source}' target='_blank' rel='noopener noreferrer'>{preview}</a>"
                 f"<div class='meta'><h2>{title}</h2><p>{status}</p>"
                 f"<a href='{source}' target='_blank' rel='noopener noreferrer'>Open source page</a></div></article>")
html = """<!doctype html><html lang='en'><head><meta charset='utf-8'>
<meta name='viewport' content='width=device-width, initial-scale=1'><title>URL Screenshot Gallery</title>
<style>body{margin:0;padding:2rem;font:16px/1.5 system-ui,sans-serif;background:#f4f6fa;color:#172033}
.grid{display:grid;grid-template-columns:repeat(auto-fit,minmax(min(100%,320px),1fr));gap:1rem}
.card{overflow:hidden;border:1px solid #dce2ec;border-radius:12px;background:white}.preview{display:grid;min-height:180px;place-items:center;background:#e9edf4}
.preview img{width:100%;height:240px;object-fit:contain}.meta{padding:1rem}.error{padding:1rem;color:#8b2530}</style></head>
<body><main><h1>URL Screenshot Gallery</h1><section class='grid'>""" + ''.join(cards) + """</section></main></body></html>"""
(out / 'index.html').write_text(html, encoding='utf-8')
print(f'Wrote {len(records)} records to {out / "index.html"}')
print(f"{sum(bool(r['error']) for r in records)} capture(s) failed; see manifest.json")

As with the Node.js listing, the HTML entities shown in this article are escaped for display. In the actual gallery.py file, use literal markup characters in the HTML strings. If you prefer generating the HTML in a separate template file, write the template as normal HTML and substitute only escaped record values.

7. cURL example for one URL

cURL can request a hosted screenshot service for an individual URL, but it does not itself launch a browser or create a gallery from a list. To batch with an API, make one request per input URL, save each returned image under a stable row-based name, record the mapping in a manifest, then use the gallery-generation step above.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com/ \
  -o page-001.webp

For multiple URLs, repeat the request with a distinct output filename for each row. Keep secrets out of shared scripts and galleries.

8. Troubleshooting common problems

Symptom Likely cause Fix
Invalid URL or navigation rejects the address The input is missing a scheme, contains whitespace, or is malformed. Use a complete URL such as https://example.com/; trim each line and validate before launching the browser.
Navigation timeout The site is slow, unreachable, or waiting on a condition that does not occur. Use a task-appropriate wait condition, raise the per-navigation timeout for slow sites, or record the failure and continue. Avoid relying on network idle for pages with ongoing requests.
Screenshot is blank or incomplete The screenshot was taken before the relevant client-rendered content appeared, or the page failed to load. Check the recorded status and final URL, then wait for a visible content selector or a known site-specific readiness signal before capture.
Lazy images are missing Images may load only after scrolling or when brought near the viewport. For pages where those images matter, scroll through the page before capturing or use an explicit full-page strategy and verify the resulting image. A navigation event alone is not proof that all lazy content is ready.
Some cards show HTTP 404 or 500 The server returned an HTTP error status; browser navigation can still complete normally. Use the manifest status to distinguish an HTTP error page from a navigation failure. Decide whether to keep the error-page screenshot or mark the URL as unsuccessful for your use case.
Images do not appear when the gallery is shared The HTML was copied without its relative screenshot files. Share the entire gallery/ directory or host its contents together.
Duplicate files overwrite each other Names were derived only from hostnames or titles. Use the input row number or another unique identifier and retain the original URL in the manifest.
Markup appears in a title or URL Untrusted text was inserted into HTML without escaping. Escape text and attribute values, and use safe URL schemes. Do not interpolate page markup into the gallery.

9. Performance, reliability, and cost

The example captures sequentially in one browser page. That is straightforward and limits the number of browser pages open at once, but total run time grows with the number of URLs and how long each site takes. If you add parallel workers, keep concurrency modest and monitor memory, browser stability, and the target sites’ capacity; the reviewed documentation does not establish a universally safe concurrency level or performance winner.

Viewport capture generally produces more manageable gallery thumbnails than full-page capture. Full-page and high device-scale screenshots can create large image files. JPEG or WebP may reduce storage and transfer size at the cost of lossy output. For a repeatable review, use the same viewport, scale, wait condition, and image format for every URL.

Failures are expected in a multi-site batch: a URL can be invalid, a server can be unreachable, TLS can fail, or navigation can time out. Recording each error and continuing preserves useful results from the other rows. A valid HTTP status such as 404 or 500 is different: Playwright returns the response rather than throwing solely due to that status.

With local Playwright, the direct costs are your compute, storage, and any browser infrastructure you operate; no per-screenshot service price is implied here. With an API, account for the provider’s plan, request limits, and any network transfer charges that apply to your environment. ScreenshotNeo’s stated plans are listed in its product block below.

10. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server for developers. Send one GET request per URL, save the returned image, and generate the same manifest and HTML gallery as above. Its parameter names used by other screenshot APIs also work, which can make switching easier. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/ -o page-001.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com/"},
    timeout=90,
)
r.raise_for_status()
open("page-001.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: HTTP ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('page-001.webp', Buffer.from(await res.arrayBuffer())));

Repeat with a unique output filename for each URL, then point the gallery manifest at those files. ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before the shot; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs.

The free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; higher plans are $15 for 15,000, $39 for 60,000, $99 for 250,000, and $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan. Sign up for 1,000 free screenshots a month, with no card.

11. FAQ

No. The generated page and its relative image files can be opened locally. A web server is useful when you want to share the gallery over a network or deploy it, but keep the files together.

Should I store the final URL or the input URL?

Keep both when possible. The input URL is the user’s original list entry; the final URL shows where redirects ended. The example records both in the manifest and links the card to the original input.

Can I capture pages that need authentication?

Browser automation can be configured for an authorized session, but this example does not log in or manage credentials. Treat resulting screenshots and galleries as sensitive if they contain private page content.

The manifest is machine-readable and preserves status, errors, original URL, final URL, title, and filename for later review or regeneration.