How to Take Bulk Website Screenshots with wkhtmltoimage
Use a Bash loop to capture many URLs with wkhtmltoimage, name files safely, record failures, and choose rendering options that fit your pages.
To take bulk website screenshots with wkhtmltoimage, put one URL per line in a text file and run the single-page command once for each line. The tool accepts one input and one output per invocation; the loop below supplies the bulk orchestration, names each output, and records failures.
wkhtmltoimage is a command-line HTML-to-image tool based on Qt WebKit. The project describes it as running headlessly, without requiring a display service. Its upstream GitHub repository is archived and read-only, so check your installed version against your target sites before relying on it for a new workflow. Project documentation · Upstream repository
1. Install and check the command
Install wkhtmltoimage using a package or binary appropriate for your operating system, or build it from source using the project’s instructions. Installation steps vary by platform and package source. Confirm that the executable is available:
wkhtmltoimage --version
wkhtmltoimage --help
The project is archived, and compatibility can depend on the binary, operating system, libraries, network access, and behavior of the pages you capture. Test several representative URLs before scaling up.
2. Create a URL list
Save one URL per line in urls.txt:
https://example.com/
https://example.org/docs/
https://example.net/pricing/
Keep credentials out of this file where possible. If a URL contains private query parameters, avoid printing it into shared logs. Use a safe label or redact sensitive values when recording failures.
3. Run a serial batch
This Bash script skips blank lines, derives a readable filename, captures each page, and logs failed commands. It is intentionally serial, which makes it easier to trace the URL-to-file mapping and diagnose problems:
#!/usr/bin/env bash
set -u
input_file="${1:-urls.txt}"
output_dir="${2:-screenshots}"
mkdir -p "$output_dir"
while IFS= read -r url || [ -n "$url" ]; do
# Ignore empty lines and lines whose first non-space character is #.
[[ -z "${url//[[:space:]]/}" ]] && continue
[[ "$url" =~ ^[[:space:]]*# ]] && continue
# Make a readable filename from the URL. Sanitization can cause collisions;
# use the indexed manifest version below for production batches.
name=$(printf '%s' "$url" | sed 's#^[[:alpha:]][[:alnum:]+.-]*://##; s#[^A-Za-z0-9._-]#_#g')
output="$output_dir/$name.png"
if wkhtmltoimage "$url" "$output"; then
printf 'OK: %s -> %s\n' "$url" "$output"
else
status=$?
printf 'FAILED (exit %s): %s\n' "$status" "$url" >&2
fi
done < "$input_file"
Save it as capture.sh, make it executable, and run it:
chmod +x capture.sh
./capture.sh urls.txt screenshots
The documented command shape is wkhtmltoimage [OPTIONS]... <input file> <output file>. The loop, output naming, and logging are workflow code built around that one-input/one-output command. See the wkhtmltoimage manual for its command options.
Use collision-safe names and keep a manifest
Different URLs can become the same filename after punctuation is replaced. Query strings may also make names unwieldy or expose private values. For repeatable batches, use a stable index in the filename and save the URL-to-file mapping in a manifest:
#!/usr/bin/env bash
set -u
input_file="${1:-urls.txt}"
output_dir="${2:-screenshots}"
mkdir -p "$output_dir"
manifest="$output_dir/manifest.tsv"
failures="$output_dir/failures.tsv"
: > "$manifest"
: > "$failures"
index=0
while IFS= read -r url || [ -n "$url" ]; do
[[ -z "${url//[[:space:]]/}" ]] && continue
[[ "$url" =~ ^[[:space:]]*# ]] && continue
index=$((index + 1))
output=$(printf '%s/%06d.png' "$output_dir" "$index")
printf '%s\t%s\n' "$output" "$url" >> "$manifest"
if ! wkhtmltoimage "$url" "$output"; then
# This file may contain sensitive URL data. Restrict access or redact it.
printf '%s\t%s\n' "$output" "$url" >> "$failures"
fi
done < "$input_file"
printf 'Manifest: %s\nFailures: %s\n' "$manifest" "$failures"
Restrict access to manifests and failure logs if their URLs contain tokens, identifiers, or private query values. For a public or shared pipeline, log a stable input number or a redacted URL instead.
4. Choose rendering options
Pass options before the input and output. The manual documents output format and quality, viewport dimensions, JavaScript timing and status waiting, cookies, authentication, and load-error behavior. Start with a small sample and change one setting at a time.
| Need | Option or approach | What to consider |
|---|---|---|
| Consistent viewport | --width and --height |
Set dimensions to match the layout you need. The manual describes width as a guide unless smart width is disabled. |
| Wait for JavaScript | --javascript-delay <msec> |
A fixed delay may be too short for a slow page or waste time on a fast one. It does not guarantee that all modern site code will finish. |
| Wait for page signal | --window-status <value> |
Useful when the page sets a known status value after rendering. The page must cooperate by setting it. |
| Choose image type | --format |
Pick a supported output format for your downstream use. Confirm that the output extension matches the selected format. |
| Reduce lossy image size | --quality <0-100> |
The documented quality range is 0 to 100. Quality applies to lossy output formats; compare visual results before choosing a lower value. |
| Access a protected page | Cookie, username/password, header, or client-certificate options | Use only the access material needed. Avoid hard-coding secrets into a reusable script or exposing them in logs and process listings. |
| Handle resource or page failures | Load-error options | Choose whether an error should abort, be ignored, or be skipped based on whether the batch must be complete. Log failures for review and retry. |
Option names and accepted values can vary by version. Consult the installed command’s help and the manual before adding flags to a production script.
Example: set a viewport and wait for JavaScript
This illustrates where common options go; choose the delay and dimensions for the pages you are capturing:
wkhtmltoimage --width 1365 --height 900 --javascript-delay 1500 \
"https://example.com/" "screenshots/example.png"
A fixed delay is a timing guess. If the page can set a known completion status, the manual also documents --window-status. Neither approach guarantees that every site will render as intended.
5. Run controlled batches and retry failures
Begin with a few pages that cover the layouts, scripts, authentication, and network conditions in your list. Inspect the output images and failure log. Once the results are acceptable, process the full list.
The serial script is simple to debug but its total elapsed time grows with the number of pages and each page’s load time. A worker pool can reduce wall time, but it needs an explicit concurrency limit so it does not overload your machine or the target sites. Concurrency is external shell orchestration, not a documented built-in bulk mode in wkhtmltoimage. Keep failed URLs in a separate retry list, and avoid retrying permanent errors indefinitely.
6. Troubleshoot common problems
| Symptom | Likely cause | Fix |
|---|---|---|
wkhtmltoimage: command not found |
The executable is not installed or is not on PATH. |
Install a platform-appropriate build, check its location, and run wkhtmltoimage --version. |
| Blank or incomplete page | Rendering may have occurred before page scripts or content finished, or the page may not work with the installed renderer. | Try a measured JavaScript delay or a page-provided window status, verify the URL is reachable, and test the exact page manually. |
| Some images or page sections are missing | Resources may load late, fail over the network, or depend on browser behavior the installed Qt WebKit renderer does not support. | Check the page in the same environment, review network and load errors, and test whether a longer wait changes the result. Compatibility is not guaranteed. |
| Output dimensions vary | Viewport options may be unset or the page may respond to viewport width. | Set width and height explicitly, then inspect pages with responsive layouts. Width is documented as a guide unless smart width is disabled. |
| Output files overwrite each other | Sanitized URL names collided or multiple inputs mapped to one name. | Use a stable index or hash and preserve a manifest mapping each output to its source URL. |
| The batch stops or outputs are missing | A command failed, the output directory is unwritable, or error handling is too strict for the chosen completeness policy. | Capture the exit status, check directory permissions and disk space, select the appropriate documented load-error behavior, and retain a retry list. |
| Authentication fails | The required cookie, header, credentials, or client certificate may be absent, expired, or incorrectly passed. | Use the relevant documented access option, confirm the page works with those credentials, and keep secrets out of source control and logs. |
| Behavior differs across machines | Installed binaries, dependencies, network access, and page behavior can differ; the upstream repository is archived. | Record the binary version and environment, test on the deployment host, and verify representative outputs after environment changes. |
7. Performance, reliability, and cost
- Throughput: A serial loop runs one command at a time. Page load time and JavaScript waits add up across the batch. Limited parallel workers can shorten elapsed time, but increase host and network load.
- Reliability: Make outputs uniquely identifiable, preserve a manifest, capture nonzero exits, and retry only the failures you intend to retry. Test the installed binary and representative pages in the same environment as the batch.
- Output size: PNG is suitable when fidelity matters. For a lossy format, set and evaluate the documented quality option. The right choice depends on the image’s use and storage constraints.
- Cost:
wkhtmltoimageis a local command-line workflow, so this method does not add a per-shot API charge. You still use machine time, storage, and network capacity; the supplied project sources do not establish a current support or maintenance commitment.
Or skip the browser setup
If you need a managed capture call, ScreenshotNeo is a website screenshot API and MCP server. Its API accepts one GET request with a URL and can return PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation for parameters and configuration.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const image = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', image));
ScreenshotNeo accepts cookie and consent banners like a visitor, then removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Every feature is on every plan.
Sign up for 1,000 free screenshots a month, with no card.
FAQ
Does wkhtmltoimage have a built-in bulk command?
The documented command takes one input and one output. A shell loop or another external job runner applies it to many URLs.
Can it capture pages that require login?
The manual documents cookies and several authentication-related options. Whether a particular login flow works depends on the page and your installed version; test it with a non-sensitive sample first.
Will a JavaScript delay make every page complete?
No. A delay can help with content that appears shortly after initial load, but it cannot guarantee that scripts, network requests, or modern page behavior will finish successfully.
Is wkhtmltoimage actively maintained?
The upstream GitHub repository notice says it was archived and made read-only on January 2, 2023. Check the repository and your installed build for current compatibility information.
Sources
- Debian Manpages: wkhtmltoimage(1) — command syntax and documented options.
- wkhtmltopdf project documentation — project description of the headless Qt WebKit tool.
- wkhtmltopdf GitHub repository — upstream archive notice.


