ScreenshotNeo

BlogHow-to

How to Take Screenshots of URLs from a Text File with Bash

Capture every URL in a text file with Bash and Playwright, with reliable filenames, error handling, and practical batch options.

By the ScreenshotNeo team4 October 202610 min read

Direct answer: Use Bash to read one URL per line, and let a browser tool such as Playwright render and capture each page. For a small batch, the Playwright CLI can open a URL and save a screenshot. For repeatable filenames, per-URL error handling, and larger batches, call Playwright’s API from a short Node.js script launched by Bash.

This guide assumes urls.txt contains one URL per line. Blank lines and lines whose first non-space character is # are skipped in the examples below.

1. Prepare the URL list and install Playwright

Create a plain UTF-8 text file with one absolute URL per line:

https://example.com/
https://developer.chrome.com/
# lines beginning with # are comments
https://playwright.dev/

Install the Playwright CLI and its browser using the current instructions in the Playwright installation documentation. The exact package and browser installation steps depend on your environment and the Playwright version. Confirm the CLI is on your PATH and that the browser it uses is installed before running a batch.

Keep the input file separate from generated screenshots. Use absolute URLs, including the scheme (https:// or http://), so navigation does not depend on browser URL guessing.

2. Capture a small list with the Playwright CLI

For a quick run, use the CLI’s documented open-then-screenshot workflow. The filename option provides a distinct output path per URL:

#!/usr/bin/env bash
set -u

input=${1:-urls.txt}
outdir=${2:-screenshots}
mkdir -p -- "$outdir"

line_no=0
while IFS= read -r url || [[ -n $url ]]; do
  [[ -z ${url//[[:space:]]/} ]] && continue
  [[ $url =~ ^[[:space:]]*# ]] && continue

  line_no=$((line_no + 1))
  name=$(printf 'page-%04d.png' "$line_no")
  printf 'Capturing %s -> %s/%s\n' "$url" "$outdir" "$name"
  playwright-cli open "$url"
  playwright-cli screenshot --filename="$outdir/$name"
done < "$input"

Save it as capture-cli.sh, make it executable with chmod +x capture-cli.sh, then run ./capture-cli.sh urls.txt screenshots. The CLI documentation includes the open URL and screenshot commands, explicit filenames, full-page capture, and high-resolution options. Check the CLI’s current help and documentation for exact flags supported by your installed version.

CLI session caveat: The simple loop follows the documented open-then-screenshot sequence, but CLI session behavior can depend on the installed version and environment. If each iteration reuses a session unexpectedly, the browser remains open, or errors need isolated recovery, use the API approach below, which controls browser and page lifetime directly.

3. Use the Playwright API for controlled batches

For a repeatable batch, let Bash start one Node.js script. It launches a browser once, reads and filters the file, visits each URL, assigns stable numbered filenames, records failures, and closes the browser in a finally block.

Install Playwright in a project using the official installation instructions, then save this as capture.mjs:

import { chromium } from 'playwright';
import { readFile, mkdir, writeFile } from 'node:fs/promises';
import path from 'node:path';

const input = process.argv[2] ?? 'urls.txt';
const outdir = process.argv[3] ?? 'screenshots';
const timeoutMs = Number(process.env.NAVIGATION_TIMEOUT_MS ?? 30000);

const lines = (await readFile(input, 'utf8')).split(/\r?\n/);
const urls = lines
  .map((line) => line.trim())
  .filter((line) => line.length > 0 && !line.startsWith('#'));

await mkdir(outdir, { recursive: true });
const browser = await chromium.launch({ headless: true });
const failures = [];

try {
  for (let i = 0; i < urls.length; i += 1) {
    const url = urls[i];
    const filename = path.join(outdir, `page-${String(i + 1).padStart(4, '0')}.png`);
    const page = await browser.newPage({ viewport: { width: 1280, height: 800 } });

    try {
      const response = await page.goto(url, {
        waitUntil: 'load',
        timeout: timeoutMs,
      });
      if (!response) {
        console.warn(`No main-document response for ${url}; saving rendered page anyway`);
      } else if (!response.ok()) {
        console.warn(`HTTP ${response.status()} for ${url}; saving rendered page anyway`);
      }
      await page.screenshot({ path: filename, fullPage: false });
      console.log(`${url} -> ${filename}`);
    } catch (error) {
      const message = error instanceof Error ? error.message : String(error);
      failures.push({ url, message });
      console.error(`Failed ${url}: ${message}`);
    } finally {
      await page.close();
    }
  }
} finally {
  await browser.close();
}

if (failures.length > 0) {
  await writeFile(
    path.join(outdir, 'failures.json'),
    JSON.stringify(failures, null, 2) + '\n',
  );
  process.exitCode = 1;
}

Run it with node capture.mjs urls.txt screenshots. To increase the navigation timeout for slow sites, set NAVIGATION_TIMEOUT_MS=60000 node capture.mjs. This script deliberately continues after an individual navigation or screenshot failure and writes those failures to failures.json. A non-2xx main-document response is reported but still captured because an error page can be useful evidence.

Playwright’s page API provides navigation and screenshot methods; see the official Page API reference. The code uses Chromium, a 1280 × 800 viewport, PNG output, and viewport-only capture as explicit defaults. Change them to fit your comparison or archiving task.

4. Choose capture size, format, and timing

Choice Use it when How to change it
Viewport screenshot You want the initial visible screen at a known window size. Keep fullPage: false; set the viewport when creating the page.
Full page You need the full scrollable document in one image. Set fullPage: true in page.screenshot(), or use the CLI’s --full-page option where supported.
PNG You need lossless output or expect to inspect fine details. Use a .png filename.
JPEG or WebP Smaller files matter more than lossless pixel preservation. Use the API’s supported screenshot type and quality options, or the CLI format option documented for your version. JPEG and WebP are lossy.
High resolution You need denser output for visual inspection or display. Use the CLI’s documented --hires flag or set the page’s device scale factor in the API.
Wait for page load The main page and its load event are sufficient. The example uses waitUntil: 'load'.
Wait for app content A client-rendered page fills in after the load event. After navigation, wait for a stable, page-specific selector with page.waitForSelector(), or use a short explicit delay only when necessary.

For example, change the API capture call to await page.screenshot({ path: filename, fullPage: true }) for a full-page image. Full-page captures can be very tall and use more memory. A device scale factor above 1 increases pixel dimensions and output size; keep it consistent across runs.

Do not treat “network idle” as a universal readiness signal: analytics, streaming, and long polling can keep requests active. Prefer a selector that represents the content you need. If pages animate or load personalized content, screenshots can differ even when the script is unchanged.

5. Run Chrome Headless directly

Chrome Headless is a simpler command-line option for one-off captures. Chrome’s official command-line reference demonstrates this form:

chrome --headless --screenshot --window-size=412,892 https://developer.chrome.com/

Chrome saves this screenshot as screenshot.png in the current working directory. In a loop, that default name can be overwritten on every iteration. To avoid losing earlier files, run each invocation with a distinct working directory, or use a supported output-path option if the Chrome version you installed documents one. Verify the available flags with the official Chrome Headless command-line reference.

For a batch requiring named outputs, full-page capture, or per-URL recovery, the Playwright API example is easier to manage. Headless Chrome and Playwright can both render pages, but results may vary with browser version, operating system, settings, fonts, and hardware.

6. Handle real-world URL files

  • Windows line endings: The API script splits on \r?\n and trims whitespace, so CRLF files are handled. The Bash loop uses read -r; a trailing carriage return can remain with some files. If navigation fails on every line, normalize the file with a tool such as sed or convert it to Unix line endings.
  • Duplicate URLs: Numbered filenames preserve one output per non-comment input line, including duplicates.
  • Unsafe URL characters: Keep each URL on one line and use properly percent-encoded URLs. Lines are passed as arguments, not evaluated as shell code.
  • Authentication: A plain URL list does not carry login state. For authenticated pages, configure a browser context with the required storage state or login flow; protect any saved credentials and state files.
  • Redirects and errors: Navigation may redirect. The API code captures the final rendered page and logs non-2xx main-document responses; failed navigation is recorded in failures.json.
  • Lazy-loaded content: Full-page capture does not guarantee every lazy image has loaded. If the page requires scrolling to trigger images, scroll through it or wait for the relevant images/selectors before capturing.
  • Large lists: The example uses one page at a time. This is slower but limits simultaneous browser memory use and makes failures easier to trace. Add bounded concurrency only after measuring resource use and site rate limits.
  • Reproducibility: Pin the Playwright dependency and keep the browser version, operating system, fonts, viewport, device scale, and headless settings consistent when comparing images. Playwright notes that screenshots can vary across browsers, platforms, and hardware; see its visual comparison guidance.

7. Troubleshooting

Symptom Likely cause Fix
playwright-cli: command not found The CLI is not installed or its executable is not on PATH. Follow the current Playwright CLI installation instructions and confirm the executable location.
Browser executable missing The Playwright package is installed but the matching browser was not installed. Install the browser required by the installed Playwright version using its documented browser installation step.
Every screenshot has the same page The CLI reused a session or the next navigation did not complete before capture. Use the API script with an explicit page per URL, or consult the installed CLI’s session controls and wait behavior.
Earlier screenshots disappear Repeated Chrome Headless invocations overwrote the default screenshot.png. Give each run a unique output location or use Playwright’s explicit filename support.
Navigation times out The site is slow, unreachable, blocked, or never reaches the chosen readiness event. Check the URL from the same machine, increase the timeout for genuinely slow pages, and wait for a specific content selector when appropriate.
Screenshot is blank or incomplete Capture began before client-rendered content appeared, or the page blocked automation. Wait for a page-specific selector; inspect the page manually. Some sites require authentication or may challenge automated browsers.
Images are missing Lazy loading, blocked resources, or a capture before image requests finished. Scroll to trigger lazy loading and wait for important images or a page-specific readiness condition.
Filename contains odd characters or URLs fail Input has whitespace, CRLF residue, comments not filtered, or malformed URLs. Trim and normalize entries, skip blank/comment lines, and use absolute, percent-encoded URLs.
Different runs look different Browser, platform, fonts, viewport, scale, animation, or dynamic content changed. Hold the environment and settings constant; disable or wait for animations if pixel comparison requires stable rendering.
Process runs out of memory Very tall pages, high device scale, or too many parallel pages consume memory. Capture the viewport, lower scale, process URLs sequentially, and close each page after saving.

8. Performance, reliability, and cost

Local browser automation has no per-screenshot API charge, but it uses your machine’s CPU, memory, disk, and network. Browser installation and maintenance are part of the cost. A single browser reused across a sequential list avoids repeated launch overhead; closing each page after capture helps bound memory use. Very long full-page images take more time and storage than viewport captures.

Reliability comes from making each URL independent: assign stable filenames, catch each failure, preserve a failure log, and close browser resources even when an entry fails. For large batches, keep concurrency bounded and respect target sites’ access policies and rate limits. For reproducible comparisons, keep browser and system settings fixed and expect dynamic pages to change.

9. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. A GET request returns a PNG, JPEG, WebP, or PDF. One request can replace local browser installation for straightforward URL capture:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));

See the ScreenshotNeo API documentation for request options. Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000.

Sign up for 1,000 free screenshots a month, with no card required.

10. FAQ

Can Bash render a webpage by itself?

No. Bash can process the file and run commands, but a browser such as Chromium must render the page and capture its pixels.

Should I use one screenshot per URL or combine them?

Use one file per URL for easier review and retry. A contact sheet can be generated afterward if you need a visual overview.

Can I use a text file with hundreds of URLs?

Yes. Process them sequentially or with carefully bounded parallelism, keep stable names, and log failures so individual URLs can be retried.

Why does a screenshot differ from what I see in my regular browser?

The browser environment, viewport, login state, fonts, timing, and dynamic page content may differ. Match those conditions as closely as practical.

References