ScreenshotNeo

BlogHow-to

How to Bulk Capture Website Screenshots with Urlbox

Discover URLs from a sitemap, then capture them with Urlbox using its API, CLI, or CaptureDeck. Learn how to batch jobs, save results, and handle failures.

By the ScreenshotNeo team4 October 20268 min read

Urlbox’s documented bulk workflow is to discover the pages first, then submit one render job per URL. It does not describe an endpoint that captures every page on a site in one request. You can automate the work with the Urlbox API, loop over a CSV with its CLI, or paste a URL list into CaptureDeck. For a large site, use batches, bounded concurrency, and a plan for storing the output.

1. Find the URLs to capture

Start with the site’s XML sitemap, commonly available at https://example.com/sitemap.xml. A sitemap may be an index that points to several child sitemaps, so use a parser that supports nested indexes rather than assuming every entry is a page.

If you cannot find a sitemap, check /robots.txt for a sitemap declaration, then look at the site footer, help pages, or an existing URL inventory. Some sites publish no sitemap; in that case, use another source you are authorized to access or provide a manual list. A sitemap is a discovery aid, not a guarantee that every URL is public, current, or suitable for capture.

Before rendering, normalize and review the list. Remove duplicates, exclude URLs that should not be captured, and decide whether query strings represent distinct pages for your use case. Retain the original URL alongside each result so files can be traced back later.

2. Choose a bulk capture workflow

Route Best fit Output and control Scale notes
ScreenshotNeo API Developers who want a one-call screenshot API or an MCP tool for AI agents PNG, JPEG, WebP, or PDF, with capture options and verdict/billing headers Bulk capture is available for up to 100 URLs per call; only clean shots are billed
Urlbox API script An application or processing pipeline that needs per-URL render control Submit render options to the documented synchronous or asynchronous endpoints Use batches and bounded concurrency; asynchronous jobs suit long-running work
Urlbox CLI with CSV A developer with a CSV who wants local files A shell loop can save a full-page PNG for each row Sequential loops are simple; Urlbox recommends asynchronous rendering for very long lists
CaptureDeck A no-code URL-list workflow Paste a URL list, run captures, then view or download organized screenshots as a ZIP; presets include social, full-page, and mobile captures For very large lists, follow Urlbox’s batching and rate-awareness guidance

For the Urlbox routes below, follow the authentication and endpoint instructions for the specific API flow you choose. Its POST render API accepts a URL or HTML input and render options as JSON or form data. The POST API documentation describes HTTP Basic authentication with the secret key as the username; the current API reference and quickstart also document Bearer authentication for the /v1/render/sync and /v1/render/async workflows. Keep secret keys on a server or in a protected shell environment, never in browser code.

3. Capture a URL list with the Urlbox API

The following JavaScript pattern sends one synchronous render request per URL using the endpoint and options shown in Urlbox’s bulk guide. It is suitable for a small list. For a large list, replace unbounded Promise.all with a bounded worker queue or use the documented asynchronous workflow so a single process does not create a burst of requests.

const urls = [
  'https://example.com/',
  'https://example.com/pricing',
  'https://example.com/docs'
];

const secretKey = process.env.URLBOX_SECRET_KEY;
if (!secretKey) throw new Error('Set URLBOX_SECRET_KEY first');

async function render(url) {
  const response = await fetch('https://api.urlbox.io/v1/render/sync', {
    method: 'POST',
    headers: {
      'Authorization': `Bearer ${secretKey}`,
      'Content-Type': 'application/json'
    },
    body: JSON.stringify({ url, full_page: true, format: 'png' })
  });
  if (!response.ok) {
    throw new Error(`Render failed for ${url}: HTTP ${response.status} ${await response.text()}`);
  }
  return { url, image: Buffer.from(await response.arrayBuffer()) };
}

const results = await Promise.all(urls.map(render));
for (let i = 0; i < results.length; i++) {
  const { writeFile } = await import('node:fs/promises');
  await writeFile(`capture-${i + 1}.png`, results[i].image);
  console.log(`Saved ${results[i].url} as capture-${i + 1}.png`);
}

Set URLBOX_SECRET_KEY in your shell or secret manager before running the script. The code intentionally uses a small sample. For thousands of URLs, add a bounded queue, retries for transient failures, per-URL result logging, and asynchronous jobs as appropriate. Confirm the exact request schema and authentication method against the Urlbox documentation for your account and API version.

4. Capture a CSV locally with the Urlbox CLI

Urlbox documents a shell loop that reads CSV rows and writes one full-page PNG per URL. This example assumes a simple one-column CSV with a header named url; adapt the field selection if your CSV contains commas or quoted values by using a real CSV parser rather than shell splitting.

#!/usr/bin/env bash
set -euo pipefail

mkdir -p captures
# Requires the Urlbox CLI to be installed and authenticated as described in its docs.
# Input: urls.csv with a header row named "url".
tail -n +2 urls.csv | while IFS=, read -r url; do
  [ -n "$url" ] || continue
  filename=$(printf '%s' "$url" | sed 's#https\?://##; s#[^A-Za-z0-9._-]#_#g')
  urlbox render "$url" --full-page --format png --output "captures/${filename}.png"
done

Check the current Urlbox CLI documentation for the installed version’s exact command and option spelling. The simple loop is useful for a modest CSV. A very long list should use asynchronous rendering rather than waiting for every capture sequentially.

5. Choose full-page behavior and image format

Urlbox’s full_page option captures the scrollable page rather than only the initial viewport. Its default stitch mode scrolls the page to trigger lazy-loaded content and animations, freezes fixed or sticky elements, captures sections, and stitches them together. The documentation describes this mode as favoring accuracy. The native mode uses the browser’s native full-page screenshot facility; it is faster, but can be less reliable on some pages.

  • skip_scroll disables the initial scroll. It may reduce render time, but can leave lazy-loaded images or content out of the screenshot.
  • full_width includes the full horizontal scroll width for pages that scroll sideways.
  • PNG is a sensible choice for full-page captures when image dimension limits matter. Urlbox documents maximum dimensions of 65,535 × 65,535 pixels for JPEG and 16,383 × 16,383 pixels for WebP.

Choose a format based on what you need to preserve and how you will store or process the files. Very tall pages can create large images and longer render times; consider whether a viewport capture or a page-specific capture is more useful than a single stitched image.

6. Run large jobs safely and keep the results

  1. Split the URL set into batches. Record the batch and URL for each result so a failed subset can be rerun without repeating successful work.
  2. Set a concurrency limit. Avoid sending an uncontrolled burst to the capture service or the target website. Urlbox’s guide advises spacing requests for very large sites, including lists of 1,000 or more pages.
  3. Use asynchronous rendering for long jobs. Synchronous calls keep a worker waiting for each result. Urlbox documents an asynchronous endpoint for jobs that should continue without holding the request open.
  4. Log status and retry selectively. Keep a record of URL, attempt, response status, and output location. Retry transient timeouts or server errors with a limit and a delay; do not blindly retry invalid URLs or permanent access failures.
  5. Save the image bytes. Urlbox’s quickstart says render URLs expire after 30 days. Download captures or arrange cloud storage, such as an S3 bucket, if you need them for longer.
  6. Respect the target site. Keep traffic within a responsible rate, and capture only URLs you are allowed to access. Authentication, robots policies, and site terms may affect what can be captured.

Large stitched pages need more time and storage than viewport images. If downstream work only needs a visual check of the first screen, avoid paying the render and storage cost of full-page output. If the full document matters, retain enough metadata to reproduce the capture settings.

Or skip the browser setup

ScreenshotNeo takes a URL in one API request and returns an image or PDF. See the ScreenshotNeo API documentation for options. For example, this cURL request saves a WebP screenshot:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents use screenshot tools. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.

Sign up for ScreenshotNeo free: 1,000 screenshots a month, no card.

Troubleshooting

Problem Likely cause What to do
A sitemap URL returns an error or no page URLs The site may not publish a sitemap at that path, or it may use a sitemap index or another location Check robots.txt, the footer, and help pages. If there is no sitemap, use an authorized URL source or a manual list.
API returns an authentication error The secret is missing, wrong, or sent using a method that does not match the selected API flow Check the endpoint’s current documentation and use the authentication method specified there. Keep the secret on the server.
Some URLs fail while others succeed A URL may be malformed, unavailable, access-controlled, or timing out Log failures individually, inspect the URL and response, and retry only transient errors with a bounded retry policy.
Lazy-loaded images are missing The page was not scrolled before capture, or the chosen mode did not trigger the relevant content Use full-page stitch mode and avoid skip_scroll when lazy content is required.
Fixed headers repeat or overlap in a long image Full-page mode and page behavior interact differently across sites Use the documented stitch mode, which freezes fixed or sticky elements during capture, and compare with native mode if the result is unsuitable.
Very large capture fails or takes too long A tall or wide page can exceed format dimension limits or consume substantial render time Use PNG where dimension limits are a concern, reduce the captured area if possible, or split the work into smaller page-specific captures.
A saved Urlbox render link no longer works Quickstart render URLs expire after 30 days Download the image or save it to durable storage when the job completes.
CLI filenames collide Different URLs can normalize to the same filename Append a stable row number or hash to each filename and preserve a manifest mapping names to URLs.

FAQ

Can Urlbox capture every page on a site in one request?

The described bulk workflow discovers URLs first and submits a render for each URL. The guide does not describe a native all-pages-in-one-request endpoint.

Does a sitemap guarantee a complete page inventory?

No. A sitemap can omit pages or be absent. Check other site references or use an authorized URL list when completeness matters.

Should I use stitch or native mode?

Use stitch when full-page accuracy and lazy content matter. Try native when render speed matters and verify that the page layout is captured correctly.

Where do the screenshots go?

The CLI workflow can save files locally, CaptureDeck offers organized viewing and ZIP download, and a custom API script can save outputs wherever your application has access. Urlbox render links are temporary, so download or store outputs for long-term use.