ScreenshotNeo

BlogHow-to

How to Take Screenshots of Multiple URLs with Urlbox

Capture a list of pages with Urlbox using a script or CLI. Learn how to batch work, handle failures, choose full-page settings, and keep the results.

By the ScreenshotNeo team4 October 202610 min read

Urlbox does not capture an entire website in one bulk request. To screenshot multiple known URLs, send a separate render request for each URL and manage the work with a script, the Urlbox CLI, or CaptureDeck, its no-code option described in Urlbox’s guide. For a site-wide capture, first find the pages—typically from a sitemap—then process the list in paced batches. Urlbox’s guide to capturing all pages describes this approach.

This guide shows a Node.js API workflow, a CSV-oriented CLI workflow, and the operational choices that matter for longer lists: full-page mode, concurrency, failures, and output retention.

1. Prepare the URL list

If you already have the URLs, put one URL per line in a text file or one URL per row in a CSV. For a whole site, check for a sitemap. A site’s robots.txt or footer may point to it. Some sites do not publish a sitemap, so you may need another source for the URLs you are authorized to capture.

Normalize and review the list before rendering:

  • Remove duplicates so you do not pay for or store repeated work unnecessarily.
  • Keep the scheme, host, and path intact; query parameters may identify distinct pages.
  • Decide whether redirects should be captured as their destination pages or kept as separate input records.
  • Use descriptive, filesystem-safe output names. A URL’s path alone can collide across hosts or query strings.
  • Confirm that the target site permits the volume and timing of requests you plan to send.

2. Render each URL with the Urlbox API

The API workflow is one render request per URL. Urlbox’s guide uses the synchronous JSON endpoint and demonstrates full_page and format settings. The example below reads URLs from a file, runs a bounded number of requests at a time, saves successful image responses, and records failures. It uses Node.js 18 or later for the built-in fetch API.

Create urls.txt with one fully qualified URL on each line, then set your secret key in the environment as URLBOX_SECRET. Keep the secret on the server or in a protected environment variable; do not put it in browser-side code. Urlbox documents secret-key authentication with the Authorization header in its API reference.

// save as capture.mjs
import { mkdir, readFile, writeFile, appendFile } from 'node:fs/promises';
import path from 'node:path';

const secret = process.env.URLBOX_SECRET;
if (!secret) throw new Error('Set URLBOX_SECRET before running this script.');

const urls = (await readFile('urls.txt', 'utf8'))
  .split(/\r?\n/)
  .map((line) => line.trim())
  .filter(Boolean);

const outputDir = 'screenshots';
const concurrency = 3; // Tune conservatively for your account and target sites.
await mkdir(outputDir, { recursive: true });
await writeFile('failures.jsonl', '');

function safeName(url, index) {
  const parsed = new URL(url);
  const pathPart = parsed.pathname.replace(/[^a-z0-9]+/gi, '-').replace(/^-|-$/g, '') || 'page';
  return `${String(index + 1).padStart(4, '0')}-${parsed.hostname}-${pathPart}.png`;
}

async function capture(url, index) {
  try {
    const response = await fetch('https://api.urlbox.com/v1/render/sync', {
      method: 'POST',
      headers: {
        'Content-Type': 'application/json',
        'Authorization': `Bearer ${secret}`
      },
      body: JSON.stringify({ url, full_page: true, format: 'png' })
    });

    if (!response.ok) {
      const detail = await response.text();
      throw new Error(`HTTP ${response.status}: ${detail.slice(0, 500)}`);
    }

    const filename = safeName(url, index);
    await writeFile(path.join(outputDir, filename), Buffer.from(await response.arrayBuffer()));
    console.log(`Saved ${url} -> ${filename}`);
  } catch (error) {
    const record = { index, url, error: String(error) };
    await appendFile('failures.jsonl', `${JSON.stringify(record)}\n`);
    console.error(`Failed ${url}: ${error}`);
  }
}

for (let start = 0; start < urls.length; start += concurrency) {
  const batch = urls.slice(start, start + concurrency);
  await Promise.all(batch.map((url, offset) => capture(url, start + offset)));
}

console.log(`Finished ${urls.length} input URLs. Review failures.jsonl for errors.`);

Run it with URLBOX_SECRET set in your shell. For example, on macOS or Linux, use export URLBOX_SECRET='your-secret-key', then node capture.mjs. The output images go into screenshots/; failures are written to failures.jsonl so you can retry only the affected URLs.

Adjust the render request

The example sets full_page: true and format: 'png'. Add or change render options according to the Urlbox API reference and your capture requirements. For a viewport-only image, omit full-page mode or set it false as supported by the API. The source guide’s sample request uses the synchronous endpoint; for long-running or very large jobs, Urlbox also documents asynchronous rendering with polling or a webhook.

Do not assume a particular concurrency quota or completion time. The reviewed Urlbox documentation does not establish a fixed account-specific limit. Start with a small batch, check API responses and target-site behavior, then adjust pacing carefully.

3. Process a CSV with the Urlbox CLI

The Urlbox CLI guide describes looping through a CSV containing a URL and output name, invoking urlbox screenshot for each row, and writing PNG files to an output directory. This is convenient when the desired result is a folder of local files. Install and authenticate the CLI according to the Urlbox CLI documentation, then adapt this shell pattern to the CSV columns and command options supported by your installed CLI version:

mkdir -p screenshots

# Example input columns: url,filename
# Keep CSV parsing robust if fields can contain commas or quoted values.
# For simple, comma-free values only:
tail -n +2 urls.csv | while IFS=, read -r url filename; do
  [ -n "$url" ] || continue
  urlbox screenshot "$url" "screenshots/${filename}.png"
done

The loop is illustrative: shell’s basic comma splitting is not a general CSV parser. If URLs or names can contain commas, parse the file with a CSV library and invoke the CLI from that program. For a very long list, the CLI guide recommends an asynchronous rendering workflow rather than waiting on every render sequentially.

4. Choose full-page capture settings

Urlbox documents two full-page capture modes. The default stitch mode scrolls through the page to trigger lazy-loaded content, then joins screenshot sections; the docs describe it as more reliable and optimized for accuracy. The native mode uses browser-native capture and is faster, but may not work well on every site. Choose based on the pages you capture and whether you need the more accurate stitched result or want to try the faster native method. See the Urlbox full-page documentation for the current option syntax.

Need Approach Consideration
Capture the visible viewport Use a regular viewport screenshot Output does not include content below the fold.
Include a long page and trigger lazy content Full page with stitch mode Scrolling and stitching can take longer; inspect pages with sticky elements or unusual layouts.
Try a faster full-page capture Full page with native mode Urlbox notes it may not work well on every site; verify representative pages.
Capture unusually tall pages Consider PNG output Urlbox documents dimension limits for JPEG and WebP; PNG avoids those format-specific limits.

Urlbox documents maximum dimensions of 65,535 by 65,535 pixels for JPEG and 16,383 by 16,383 pixels for WebP. PNG is its suggested choice for full-page images because it does not have those format-specific limits. Very tall images can still be large files and may be awkward to view or process, so choose the output format with downstream storage and tooling in mind.

5. Handle large lists safely

For a handful of pages, a small script is usually easiest to customize. For hundreds or thousands, use bounded batches or an asynchronous workflow, and pay attention to both your API account’s limits and the load placed on the target sites. Urlbox’s guide calls out spacing requests for very large sites, describing 1000+ pages as a scale where pacing may be needed; this is operational guidance, not a throughput guarantee.

  • Bound concurrency. Avoid an unbounded Promise.all(urls.map(...)) for a large list. A fixed batch size makes request pressure easier to control.
  • Add pacing when needed. If you see rate limits or the target site struggles, reduce concurrency and add a delay between batches.
  • Use asynchronous jobs for long runs. Urlbox documents async POST rendering with polling or a webhook. This avoids holding a client request open while a large job completes.
  • Make retries selective. Record each source URL and error. Retry transient failures with a capped retry count and delay; do not retry authentication or invalid-parameter errors unchanged.
  • Persist outputs as they arrive. Write each successful image immediately. For large collections, organize them in cloud storage such as S3, as Urlbox’s guide recommends.
  • Keep an input-to-output manifest. Store the original URL, capture time, chosen options, output name, and result state alongside the files.

Urlbox does not publish a fixed concurrency number or processing-time guarantee in the sources reviewed for this guide. Treat concurrency as a setting to tune gradually, not a promised capacity.

6. Keep and organize the results

Urlbox’s quickstart says render URLs expire after 30 days. Download the image responses to your own storage or configure saving to a cloud bucket if you need longer retention. For repeatable jobs, keep a manifest and use stable filenames or versioned folders so a later capture does not silently overwrite an earlier one. See the Urlbox quickstart for render URL and storage details.

7. Troubleshoot common problems

Symptom Likely cause What to do
Authentication error The secret key is missing, incorrect, or sent in the wrong header. Check URLBOX_SECRET in the running process and use the documented Authorization: Bearer … header. Keep the key server-side.
One or more URLs fail while others work A target may be unavailable, slow, redirecting unexpectedly, or returning an error. Log the URL and response details, inspect that page independently, then retry only transient failures.
Rate limit responses Requests are arriving too quickly for the account or service. Reduce batch concurrency and space requests out. For large workloads, consider the documented async workflow.
Images miss content lower on the page The request captured only the viewport, or lazy content did not load as expected. Enable full-page capture. For lazy-loaded pages, try stitch mode and check representative output.
Full-page image is incomplete or awkward Native mode may not suit that page, or stitching may interact with sticky or dynamic content. Compare stitch and native on a small sample; retain the mode that captures the page correctly.
Output name collisions Different URLs produced the same path-derived filename. Include an index, hostname, or stable hash in filenames and keep a URL-to-file manifest.
A saved render link no longer works Urlbox says render URLs expire after 30 days. Download the image when the job completes or configure cloud storage for longer retention.
Very tall JPEG or WebP capture hits a dimension limit Urlbox documents maximum image dimensions for these formats. Use PNG for full-page captures where those format limits matter, and consider the storage implications of large files.

8. Understand performance, reliability, and cost

Total job time depends on the number of URLs, how long each page takes to render, the chosen full-page mode, and the concurrency you can use without causing rate limits or undue target-site load. The reviewed sources provide no independent speed benchmark or guaranteed throughput, so estimate from a small representative sample rather than extrapolating from a single page.

For reliability, treat each URL as its own result: save successful outputs immediately, log failures with enough detail to retry selectively, and make reruns safe. An async job with polling or a webhook can fit a long-running batch better than a single process waiting synchronously. Confirm current account quotas and pricing in Urlbox’s product documentation before estimating the cost of a large capture run; the sources used here do not establish a comparable per-job cost figure.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. A GET request can return PNG, JPEG, WebP, or PDF; its request parameters include those used by other screenshot APIs to make switching easier. See the ScreenshotNeo API documentation for options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie banners, popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
  • Bot checks, blank pages, timeouts, and failed loads are never billed, and response headers report the page verdict and billing status.
  • An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf.
  • 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000 screenshots.

Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.

FAQ

Can Urlbox screenshot every page on a site in one request?

No. The documented approach is to obtain a URL list and render each page separately, or use CaptureDeck for the no-code list workflow described in Urlbox’s guide.

How do I capture only pages from a sitemap?

Extract the page URLs from the sitemap into a list, review and deduplicate them, then feed that list into the API or CLI workflow above. Check robots.txt or the site’s footer for sitemap references if you do not know its location.

Should I use the API or CLI?

Use the API when you need custom result handling, batching, or integration into a larger job. Use the CLI when a CSV-to-local-files workflow is sufficient. For a no-code list process, Urlbox’s guide points to CaptureDeck.

Download each image when it is ready or configure cloud storage. Urlbox’s quickstart states that render URLs expire after 30 days.