ScreenshotNeo

BlogHow-to

How to capture screenshots for URLs in a text file with Chrome headless

Read URLs from a text file and capture each with Chrome headless, saving screenshots under unique names. Includes shell, PowerShell, and troubleshooting guidance.

By the ScreenshotNeo team4 October 20268 min read

To capture screenshots for URLs in a text file with Chrome headless, put one URL on each line, then run Chrome once per non-empty line with --headless and --screenshot. Give each invocation its own output path so every page gets a separate PNG. Chrome writes screenshot.png in the current directory by default, so relying on that default in a batch will overwrite earlier captures.

The examples below capture each URL at a 1280 × 900 viewport. The basic command-line screenshot is a viewport capture; it is not documented as a guaranteed full-page screenshot. Chrome’s headless command-line reference documents --screenshot, --window-size, --timeout, and --virtual-time-budget.

1. Prepare the URL list and output directory

Create a UTF-8 plain text file named urls.txt, with one complete URL per line. For example:

https://developer.chrome.com/
https://example.com/
https://www.wikipedia.org/

Lines should contain the URL only. Blank lines are skipped by the shell loop below. Create a directory for the images before running the batch:

mkdir -p screenshots

Check that Chrome is available from your shell. The executable name depends on the operating system and installation; common names include chrome and google-chrome. If neither works, find the installed executable and use its full path in the examples.

2. Capture every URL on macOS or Linux

Save this as capture.sh, or paste it into a POSIX-compatible shell from the directory containing urls.txt:

#!/usr/bin/env bash
set -u

chrome_bin="${CHROME_BIN:-chrome}"
input_file="urls.txt"
output_dir="screenshots"

mkdir -p "$output_dir"
if ! command -v "$chrome_bin" >/dev/null 2>&1; then
  printf 'Chrome executable not found: %s\n' "$chrome_bin" >&2
  printf 'Set CHROME_BIN to its path, for example: CHROME_BIN=/path/to/chrome\n' >&2
  exit 1
fi
if [ ! -f "$input_file" ]; then
  printf 'URL list not found: %s\n' "$input_file" >&2
  exit 1
fi

n=0
while IFS= read -r url || [ -n "$url" ]; do
  # Ignore empty lines and lines containing only spaces or tabs.
  case "$url" in
    *[!$' \t']*) ;;
    *) continue ;;
  esac

  n=$((n + 1))
  output="$output_dir/page-$n.png"
  printf 'Capturing %s -> %s\n' "$url" "$output"
  if ! "$chrome_bin" --headless --screenshot="$output" \
      --window-size=1280,900 --timeout=10000 "$url"; then
    printf 'Chrome returned an error for line %s: %s\n' "$n" "$url" >&2
  fi
done < "$input_file"

printf 'Processed %s non-empty URL lines.\n' "$n"

Run it with bash capture.sh. If Chrome is installed at a nonstandard path, set CHROME_BIN, for example CHROME_BIN="/path/to/chrome" bash capture.sh. The numbered output names correspond to the order of non-empty lines in the file. The script reports a failed invocation and continues to later lines; check that each expected PNG exists and is non-empty before treating the batch as complete.

Use host-based filenames when useful

Sequential names are safe and predictable. If you need filenames that identify their source, derive a name from the URL hostname and add a counter to handle repeated hosts. Do not use the raw URL as a filename: paths and query strings can contain characters that have special meaning to filesystems and shells. Always sanitize any URL-derived value and retain a unique suffix.

3. Capture every URL on Windows with PowerShell

In PowerShell, set $chrome to the executable name or full path for your installation. This loop skips blank lines, creates unique numbered paths, and continues after a failed Chrome process:

$chrome = "chrome"
$inputFile = "urls.txt"
$outputDir = "screenshots"

New-Item -ItemType Directory -Force -Path $outputDir | Out-Null
if (-not (Test-Path -LiteralPath $inputFile)) {
    throw "URL list not found: $inputFile"
}

$index = 0
Get-Content -LiteralPath $inputFile | ForEach-Object {
    $url = $_.Trim()
    if ([string]::IsNullOrWhiteSpace($url)) { return }

    $index++
    $output = Join-Path $outputDir ("page-{0}.png" -f $index)
    Write-Host "Capturing $url -> $output"
    & $chrome --headless "--screenshot=$output" --window-size=1280,900 --timeout=10000 $url
    if ($LASTEXITCODE -ne 0) {
        Write-Warning "Chrome returned exit code $LASTEXITCODE for $url"
    }
}

Write-Host "Processed $index non-empty URL lines."

Run this in the directory containing urls.txt. If the command name is not recognized, use the full path to chrome.exe and keep it quoted, for example $chrome = 'C:\Path With Spaces\chrome.exe'. The call operator & runs that quoted path.

4. Tune the capture for your pages

Option What it controls When to use it
--headless Runs Chrome without its regular visible browser window. Batch work from a shell or automation job.
--screenshot Writes a screenshot. With no output path, Chrome uses screenshot.png in the current working directory. Supply a unique path for each URL in a batch.
--window-size=WIDTH,HEIGHT Sets the viewport dimensions used by the documented screenshot example. Choose dimensions that match the screen size you need to inspect.
--timeout=MILLISECONDS Sets a maximum wait before capture, including while a page is still loading. Bound the wait for slow or unsettled pages. A timeout does not prove the page finished rendering.
--virtual-time-budget=MILLISECONDS Fast-forwards time-dependent page code in the headless browser. Try it when page behavior depends on timers. Results depend on the page’s code.

For example, replace the Chrome invocation with this to allow a longer maximum wait and provide a virtual-time budget:

chrome --headless --screenshot="screenshots/page-1.png" \
  --window-size=1440,1000 --timeout=20000 \
  --virtual-time-budget=5000 "https://developer.chrome.com/"

These flags do not provide a general readiness guarantee. A site may continue loading after a timeout, defer content until user interaction, or require authentication. If the page needs a specific element to appear before capture, a browser automation framework with an explicit wait condition may be a better fit than a fixed command-line delay.

5. Know what the screenshot contains

  • Viewport versus full page: --window-size sets the viewport dimensions in the documented example. Do not assume this basic command captures the entire scrollable page. The cited CLI reference does not document a dedicated full-page option for this flag.
  • Lazy-loaded content: Images and sections that load only after scrolling may not be present. A fixed timeout may not trigger them.
  • Authentication and access: Pages behind a login, bot check, CAPTCHA, or other gate may show a challenge or an access-denied screen instead of the intended content.
  • Redirects and browser differences: A URL can redirect, and a site can render differently in headless mode. The command-line options do not guarantee a particular result for every site.
  • Output files: The default screenshot path repeats on every run. Set a unique output name and check the resulting files to avoid silently replacing earlier captures.

6. Troubleshoot common problems

Symptom Likely cause What to do
chrome: command not found or command not recognized Chrome is not on the shell’s executable path, or its command has a different name. Find the installed executable and set CHROME_BIN or the PowerShell $chrome variable to its full path.
Every run leaves only one image Each run used the default screenshot.png path. Pass a distinct value to --screenshot, such as screenshots/page-1.png.
No output image appears The output directory is missing, the path is invalid, Chrome could not launch, or a URL argument was malformed. Create the directory first, quote the output path and URL, and inspect Chrome’s error output and exit status.
The image is blank or shows a loading state The page may need more time, block headless traffic, depend on interaction, or have failed to load. Try a larger --timeout, inspect the page in a regular browser, and confirm the URL is publicly accessible. A longer timeout cannot resolve a bot check or authentication requirement.
The bottom of a long page is missing The basic screenshot command captured the viewport rather than a guaranteed full page. Use a browser automation workflow that scrolls or invokes a full-page capture capability, or capture at a viewport matching the specific area you need.
Some pages fail while later pages succeed Individual sites may redirect, reject automation, time out, or fail for unrelated reasons. Keep per-URL status in the batch output, retry only the affected URLs when appropriate, and verify each file rather than assuming the whole batch succeeded.
URLs with spaces or shell punctuation behave strangely The URL was not passed as a single quoted argument. Keep the URL in a variable and quote it, as in the shell loop. Do not interpolate it into a command string.

7. Performance, reliability, and cost

The simple loops run one Chrome command at a time. This is easy to inspect and keeps the workflow straightforward, but total elapsed time depends on the pages and their load behavior; the source documents no benchmark. Start with a small list, inspect the images and error output, and scale only after confirming your shell, Chrome path, URLs, and output naming are correct.

For repeatable results, keep the input list stable, record which line maps to each output, use explicit viewport and timeout values, and verify output files. Re-running the same numbered batch overwrites those numbered files, so copy prior results elsewhere if you need to retain multiple runs. Avoid running many Chrome processes in parallel without accounting for available memory and process limits.

Chrome headless is included as a browser workflow, so the examples do not call a paid screenshot service. Your operational costs are the machine, browser installation, and time spent handling retries and page-specific behavior. A fixed timeout bounds a wait but may trade a faster result for a capture taken before a page is ready.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. One GET request returns an image or PDF; the API accepts common screenshot parameter names used by other screenshot APIs. Cookie banners are accepted and removed before the shot, along with 60+ known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in headers.

Use this cURL call to capture one URL as WebP. See the ScreenshotNeo API documentation for options and formats.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://developer.chrome.com/ \
  -o shot.webp

The same GET request in Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://developer.chrome.com/"},
    timeout=90,
)
r.raise_for_status()
with open("shot.webp", "wb") as f:
    f.write(r.content)

And in Node.js with a runtime that provides fetch:

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://developer.chrome.com/'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));

The service also provides an MCP server for AI agents, with take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Other listed plans are Growth at $15 for 15,000, Pro at $39 for 60,000, Scale at $99 for 250,000, and Business at $249 for 1,000,000; yearly billing gives two months free. Every feature is on every plan. Sign up for 1,000 free screenshots a month with no card.

FAQ

Does the loop capture one screenshot per line?

It captures one per non-empty line, in order. Each invocation writes to a numbered path, so repeated URLs still get separate files.

Can I safely use URLs with query parameters?

Yes, as long as the URL is passed as one quoted argument. The examples quote the value; do not build a shell command by concatenating untrusted URL text.

Does a successful Chrome command mean the page is correct?

No. Check the image itself. Chrome can produce an image of a challenge, error page, or incomplete rendering even when the invocation created an output file.