ScreenshotNeo

BlogHow-to

How to Screenshot a List of URLs with Browserless

Capture a list of URLs with Browserless by sending concurrent POST requests, saving each image safely, and handling page options and failures.

By the ScreenshotNeo team4 October 20269 min read

To screenshot multiple URLs with Browserless, send one authenticated POST request to its regional /screenshot endpoint per URL, run a bounded number of requests concurrently, and save each image under a distinct filename. The endpoint accepts the URL and screenshot options as JSON and returns image bytes. A token is required as the token query parameter. Browserless screenshot API documentation describes the request; use the region and endpoint configured for your account.

JavaScript: capture a URL list

This Node.js example uses the built-in fetch API and node:fs/promises. It limits concurrent requests, checks HTTP status before saving, and reports individual failures while allowing other captures to finish. Set BROWSERLESS_TOKEN in the environment and replace the regional endpoint if your account uses another region.

import { mkdir, writeFile } from 'node:fs/promises';

const urls = [
  'https://example.com/',
  'https://example.org/',
];
const token = process.env.BROWSERLESS_TOKEN;
const endpoint = 'https://production-sfo.browserless.io/screenshot';
const outputDir = './screenshots';
const concurrency = 4;

if (!token) throw new Error('Set BROWSERLESS_TOKEN first.');
await mkdir(outputDir, { recursive: true });

async function capture(url, index) {
  const response = await fetch(`${endpoint}?token=${encodeURIComponent(token)}`, {
    method: 'POST',
    headers: { 'Content-Type': 'application/json' },
    body: JSON.stringify({
      url,
      options: { fullPage: true, type: 'png' },
    }),
  });
  if (!response.ok) {
    const detail = await response.text();
    throw new Error(`HTTP ${response.status} for ${url}: ${detail}`);
  }
  const bytes = Buffer.from(await response.arrayBuffer());
  const filename = `${String(index + 1).padStart(4, '0')}.png`;
  await writeFile(`${outputDir}/${filename}`, bytes);
  return { url, filename };
}

const results = new Array(urls.length);
let next = 0;
async function worker() {
  while (true) {
    const index = next++;
    if (index >= urls.length) return;
    try {
      results[index] = { ok: true, ...(await capture(urls[index], index)) };
    } catch (error) {
      results[index] = { ok: false, url: urls[index], error: String(error) };
    }
  }
}
await Promise.all(Array.from({ length: Math.min(concurrency, urls.length) }, worker));
console.log(results);
console.log(`${results.filter((r) => r.ok).length}/${urls.length} captures saved.`);

The numeric filenames are collision-safe for this input order and preserve the index-to-URL mapping in results. In a recurring workflow, write that mapping to a JSON or CSV manifest so failed URLs can be retried without guessing which file belongs to which page. For filenames based on host or path, sanitize characters and handle duplicate URLs explicitly.

cURL: capture one URL

For a single URL, send the request directly. Repeat it for each URL in a shell script or use the bounded-concurrency code above for a larger batch.

curl --fail-with-body --silent --show-error \
  -X POST 'https://production-sfo.browserless.io/screenshot?token=YOUR_TOKEN' \
  -H 'Content-Type: application/json' \
  --data '{"url":"https://example.com/","options":{"fullPage":true,"type":"png"}}' \
  --output example.png

Keep the token out of source control and logs. For repeated shell requests, read it from an environment variable and URL-encode it if constructing the query string manually.

Python: capture a URL list with bounded concurrency

Python’s ThreadPoolExecutor is suitable for this I/O-bound batch. Install the HTTP dependency with python -m pip install requests, set BROWSERLESS_TOKEN, then run:

import os
from concurrent.futures import ThreadPoolExecutor, as_completed
from pathlib import Path
import requests

URLS = ['https://example.com/', 'https://example.org/']
TOKEN = os.environ.get('BROWSERLESS_TOKEN')
ENDPOINT = 'https://production-sfo.browserless.io/screenshot'
OUTPUT = Path('screenshots')
CONCURRENCY = 4

if not TOKEN:
    raise RuntimeError('Set BROWSERLESS_TOKEN first.')
OUTPUT.mkdir(parents=True, exist_ok=True)

def capture(item):
    index, url = item
    response = requests.post(
        ENDPOINT,
        params={'token': TOKEN},
        json={'url': url, 'options': {'fullPage': True, 'type': 'png'}},
        timeout=(15, 120),
    )
    response.raise_for_status()
    filename = f'{index + 1:04}.png'
    (OUTPUT / filename).write_bytes(response.content)
    return {'url': url, 'filename': filename}

results = []
with ThreadPoolExecutor(max_workers=CONCURRENCY) as pool:
    futures = {pool.submit(capture, item): item for item in enumerate(URLS)}
    for future in as_completed(futures):
        item = futures[future]
        try:
            results.append({'ok': True, **future.result()})
        except Exception as error:
            results.append({'ok': False, 'url': item[1], 'error': str(error)})

for result in sorted(results, key=lambda row: row['url']):
    print(result)
print(f"{sum(row['ok'] for row in results)}/{len(URLS)} captures saved.")

The connect/read timeout tuple separates time allowed to establish the request from time allowed for the screenshot response. Choose values that fit your pages and job deadline; large full-page captures can take longer than ordinary viewport captures.

Choose capture options

Browserless accepts Puppeteer-style screenshot options under the JSON options object. The endpoint documentation covers available fields; confirm option names against the current REST docs when adapting a production integration.

Need Setting or approach What to consider
Entire document fullPage: true Captures beyond the visible viewport. Very long pages produce larger images and take more time.
Visible viewport Leave full-page capture off Useful for consistent previews and smaller files. Specify a viewport where supported when pages must be comparable.
Format type: PNG, JPEG, or WebP as supported PNG preserves crisp text and transparency where applicable; lossy formats can reduce file size. JPEG quality is relevant when using JPEG.
Specific element Selector-related screenshot option Use when the target is a component rather than the page. Verify the selector exists before capturing.
Specific rectangle options.clip Define a clip rectangle for a known region; ensure its dimensions fit the rendered page.
Page geometry Viewport and device scale options Use a fixed viewport and scale for consistent dimensions across URLs.
Delayed or lazy content Wait options; scrollPage: true for lazy-loading Wait for the page to render, and scroll before capture if content loads on scroll. Browserless documents scrollPage for this case.

Do not apply one setting blindly to every page. Choose viewport versus full-page capture, format, selector or clip, and readiness behavior based on the output you need. For lazy-loaded sections, scrolling can trigger loading, but pages that require user interaction may need a different automation workflow.

Scale a batch safely

  1. Validate inputs. Reject malformed URLs and unsupported schemes before submitting work. Keep the original URL with each result.
  2. Start with modest concurrency. Each REST request opens an independent browser session. Browserless demonstrates concurrent sessions, but its example does not establish a universally safe limit. Begin with a small worker count and adjust based on account limits, failures, and the pages being captured. Check the connection URL documentation for account and endpoint details.
  3. Set a deadline. Bound request duration in the client and the overall batch duration in your job runner. A page that hangs should not stall the whole list.
  4. Retry selectively. Retry transient network errors and appropriate server errors with a capped exponential backoff and jitter. Do not retry every failure indefinitely; authentication errors and invalid inputs need correction first.
  5. Persist progress. Record URL, filename, status, attempt count, and error. Write each successful image atomically where practical (temporary file then rename) so a process interruption does not leave a partial file looking complete.
  6. Respect target sites. A CAPTCHA, access-denied page, or blank result can reflect bot blocking or site policy. Browserless describes /unblock as a separate route for certain bot-detection cases; follow the site’s access rules and do not use it to bypass restrictions.

Parallel requests can shorten total wall-clock time compared with serial capture, but the actual duration depends on page load, image size, service capacity, and concurrency. Browserless’s concurrent-session example characterizes total time as roughly the slowest request; treat that as an illustrative pattern, not a guaranteed timing or service level.

REST screenshot endpoint or BrowserQL?

For a simple list where one URL should produce one image file, REST maps cleanly to the task: submit one request per URL and save its binary response. Browserless also documents screenshots through BrowserQL, which uses a GraphQL mutation and returns image data in a base64 field. Choose BrowserQL when the project already uses that interface or needs its query workflow; account for decoding the returned base64. See Browserless’s BrowserQL screenshot documentation.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. One GET request returns a PNG, JPEG, WebP, or PDF, and the same request pattern can be repeated for a URL list. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, newsletter popups, and chat widgets are removed before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Sign up free for 1,000 screenshots a month, no card required.

Troubleshooting

Symptom Likely cause Fix
401 or 403 response Missing, invalid, or misplaced token; wrong regional endpoint; account access issue. Pass the current token as the token query parameter and use the endpoint configured for your account. Keep credentials out of logs.
HTML or JSON saved with an image extension The request failed, but the client saved the error response body. Check the HTTP status before writing bytes. Log a short error body separately and only save successful responses as images.
Blank screenshot The target did not finish rendering, is blocked, or failed to load required content. Inspect the target’s accessibility and response behavior, adjust supported wait settings, and verify the URL manually. A bot check or access-denied page may be the actual rendered page.
Missing below-the-fold content Viewport-only capture or lazy content was not triggered. Enable full-page capture and use documented scrollPage: true where lazy loading requires scrolling.
Timeouts under load Pages are slow, captures are large, or too many sessions run at once. Reduce concurrency, set suitable client timeouts, and retry only transient failures with backoff.
Files overwrite each other Filename derived from a non-unique host or concurrent writes share a path. Use a stable index or unique identifier in the output name and retain a manifest mapping names to URLs.
Some files are missing after one error A batch-level rejection stopped processing or the process exited early. Capture errors per URL, persist completed work as it finishes, and rerun only failed entries.
Output dimensions differ between URLs Pages use different layout widths or device scale settings. Set consistent viewport and device scale options, and use viewport capture when fixed output dimensions matter.

Performance, reliability, and cost

Concurrency is the main control for throughput, but higher concurrency also means more simultaneous browser sessions and can increase timeouts or account-limit errors. Measure the completion and failure rate for your own URL set, then tune the worker count. Full-page captures, high device scale, and PNG output can increase transfer size and storage needs. If downstream consumers accept them, a lossy format can reduce image size.

For reliability, separate capture from downstream processing, persist a manifest, make retries idempotent by writing to deterministic paths, and retain failed URLs for inspection. Do not assume every successful HTTP response represents the intended page: inspect status and, where the workflow requires it, validate image content or dimensions before marking a capture complete.

Browserless usage and pricing depend on the account and plan; this guide does not assume a particular rate or quota. Check the current Browserless pricing documentation and account limits before estimating batch cost. Control costs by avoiding unnecessary recaptures, using cache or stored results in your own workflow where appropriate, and selecting viewport captures when full-page images are unnecessary.

FAQ

Does Browserless accept a list of URLs in one screenshot request?

The documented REST screenshot workflow is one URL per request. Submit one request for each URL and coordinate them in your client.

Can I preserve the input order when requests finish out of order?

Yes. Associate each task with its original index or URL, then sort the result records by that index before producing a manifest or report.

Should I use full-page capture for every URL?

No. Use it when the whole document is needed. Viewport captures are more appropriate for fixed-size previews and can reduce output size.

What if a page requires login?

The basic example only navigates to a URL. Pages requiring an authenticated browser session need an appropriate supported session or automation flow; do not place credentials in public logs or source code.

Can I capture pages blocked by their owners?

A CAPTCHA or access-denied response can be the target site’s access control. Respect site rules and applicable terms rather than treating every block as a technical error.