ScreenshotNeo

BlogHow-to

How to Bulk Screenshot URLs with Browserless and Playwright

Batch website screenshots with Browserless REST or remote Playwright. Save each result, limit concurrency, handle page readiness, and troubleshoot failed captures.

By the ScreenshotNeo team4 October 20269 min read

For a simple batch, send one POST request to Browserless’s /screenshot endpoint for each URL, then save each binary response to its own file. Use a bounded worker pool to control parallel requests. Choose remote Playwright when each page needs browser interactions, custom readiness checks, or page-level state before capture. The examples below show both approaches, plus cURL, Python, and Node.js. Keep your Browserless token in an environment variable or secret store.

1. Choose REST screenshots or remote Playwright

Need Use
Navigate to a URL and save an image with supported capture options Browserless REST /screenshot, one request per URL
Interact with the page, control navigation, wait for application state, or run page code Playwright connected to a Browserless remote browser

REST is simpler when the job is just URL to image. Remote Playwright gives you browser control but requires a compatible Browserless Playwright endpoint and uses a remote browser session. Browserless documents REST image responses and remote Playwright connections; check the current endpoint and account documentation for the correct connection details and capacity.

2. Bulk capture with the Browserless REST API

Send JSON with the target URL and screenshot options to the Browserless screenshot endpoint. The account token goes in the endpoint query string. The response is image bytes, so write it as binary data rather than trying to parse it as JSON. Browserless documents PNG, JPEG, and WebP output; PNG is used below.

cURL: capture one URL

export BROWSERLESS_TOKEN="YOUR_BROWSERLESS_TOKEN"
curl --fail --silent --show-error \
  -X POST "https://production-sfo.browserless.io/screenshot?token=${BROWSERLESS_TOKEN}" \
  -H "Content-Type: application/json" \
  --data '{"url":"https://example.com","options":{"fullPage":true,"type":"png"}}' \
  --output example.png

Use the endpoint hostname and token format provided for your Browserless account. This example illustrates the request shape; confirm your account’s current endpoint before running it.

Python: bounded concurrent batch

import os
import re
import time
from concurrent.futures import ThreadPoolExecutor, as_completed
from pathlib import Path

import requests

TOKEN = os.environ["BROWSERLESS_TOKEN"]
ENDPOINT = "https://production-sfo.browserless.io/screenshot"
URLS = [
    "https://example.com",
    "https://www.python.org",
]
OUTPUT = Path("screenshots")
OUTPUT.mkdir(exist_ok=True)
MAX_WORKERS = 3  # Set this to a value within your plan and target-site limits.


def filename_for(url: str) -> str:
    host = re.sub(r"[^a-zA-Z0-9.-]", "_", url.split("//", 1)[-1].split("/", 1)[0])
    return f"{host or 'page'}.png"


def capture(url: str) -> tuple[str, str, str | None]:
    path = OUTPUT / filename_for(url)
    payload = {
        "url": url,
        "options": {"fullPage": True, "type": "png"},
    }
    for attempt in range(3):
        try:
            response = requests.post(
                ENDPOINT,
                params={"token": TOKEN},
                json=payload,
                timeout=(10, 90),
            )
            response.raise_for_status()
            content_type = response.headers.get("content-type", "")
            if not content_type.startswith("image/"):
                raise ValueError(f"Expected image bytes, got Content-Type {content_type!r}")
            path.write_bytes(response.content)
            return url, "saved", str(path)
        except (requests.RequestException, ValueError) as exc:
            if attempt == 2:
                return url, "failed", str(exc)
            time.sleep(2 ** attempt)
    return url, "failed", "unreachable"


with ThreadPoolExecutor(max_workers=MAX_WORKERS) as pool:
    futures = [pool.submit(capture, url) for url in URLS]
    for future in as_completed(futures):
        url, status, detail = future.result()
        print(status, url, detail)

Set MAX_WORKERS conservatively. Browserless plan concurrency limits simultaneous sessions, and the destination site also has its own capacity and automation policies. For very large lists, feed URLs in chunks and persist each result as it finishes instead of holding all response data in memory.

Node.js: bounded concurrent batch

import { mkdir, writeFile } from "node:fs/promises";

const token = process.env.BROWSERLESS_TOKEN;
if (!token) throw new Error("Set BROWSERLESS_TOKEN");
const endpoint = "https://production-sfo.browserless.io/screenshot";
const urls = ["https://example.com", "https://www.python.org"];
const outputDir = "screenshots";
const concurrency = 3;
await mkdir(outputDir, { recursive: true });

async function capture(url) {
  const host = new URL(url).hostname.replace(/[^a-zA-Z0-9.-]/g, "_");
  const target = new URL(endpoint);
  target.searchParams.set("token", token);
  const response = await fetch(target, {
    method: "POST",
    headers: { "Content-Type": "application/json" },
    body: JSON.stringify({
      url,
      options: { fullPage: true, type: "png" },
    }),
    signal: AbortSignal.timeout(90000),
  });
  if (!response.ok) throw new Error(`HTTP ${response.status}: ${await response.text()}`);
  const type = response.headers.get("content-type") ?? "";
  if (!type.startsWith("image/")) throw new Error(`Expected image, received ${type}`);
  await writeFile(`${outputDir}/${host || "page"}.png`, Buffer.from(await response.arrayBuffer()));
  return { url, status: "saved" };
}

async function runPool(items, limit) {
  const results = new Array(items.length);
  let next = 0;
  async function worker() {
    while (true) {
      const index = next++;
      if (index >= items.length) return;
      try {
        results[index] = await capture(items[index]);
      } catch (error) {
        results[index] = { url: items[index], status: "failed", error: String(error) };
      }
    }
  }
  await Promise.all(Array.from({ length: Math.min(limit, items.length) }, worker));
  return results;
}

for (const result of await runPool(urls, concurrency)) console.log(result);

The pool reports an outcome for every input URL and keeps the number of active requests bounded. For production batches, also persist results to a database or a manifest file, and ensure output names are unique if different URLs can share a hostname.

3. Capture URLs with remote Playwright

Use this approach when a page must be interacted with before capture. Browserless documents a Playwright WebSocket endpoint and an example using chromium.connect. Its default endpoint speaks CDP; use the Playwright endpoint and connection method appropriate to your account rather than assuming every endpoint supports Playwright protocol connections.

import { chromium } from "playwright";
import { mkdir, writeFile } from "node:fs/promises";

const token = process.env.BROWSERLESS_TOKEN;
if (!token) throw new Error("Set BROWSERLESS_TOKEN");
const wsEndpoint = `wss://production-sfo.browserless.io/chromium/playwright?token=${token}`;
const urls = ["https://example.com", "https://www.python.org"];
await mkdir("screenshots", { recursive: true });

const browser = await chromium.connect(wsEndpoint);
try {
  for (const [index, url] of urls.entries()) {
    const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
    try {
      await page.goto(url, { waitUntil: "domcontentloaded", timeout: 60000 });
      // Prefer a page-specific selector when you know what marks the content ready.
      await page.locator("body").waitFor({ state: "visible", timeout: 15000 });
      await page.screenshot({ path: `screenshots/page-${index + 1}.png`, fullPage: true });
    } finally {
      await page.close();
    }
  }
} finally {
  await browser.close();
}

Install Playwright in your project using its official package instructions, and configure the Browserless endpoint from its current connection documentation. This loop processes sequentially within one connection. To parallelize Playwright jobs, use only the session concurrency your plan supports, and close pages and browser connections in finally blocks so sessions are released after errors.

Page readiness and lazy-loaded content

  • Prefer a selector that represents the actual content you need over an arbitrary delay.
  • Choose a navigation readiness event appropriate to the page; dynamic sites may continue changing after initial navigation.
  • A fixed wait can help when no reliable selector exists, but it adds time and may still miss late content.
  • Lazy-loaded images or sections may not exist until scrolled into view. Browserless advises scrolling the page together with full-page mode when the complete long page is needed.
  • For selector capture, confirm that the selector exists and is visible before taking the image.

4. Relevant capture options

Browserless’s screenshot API documents options for capture scope and output, viewport and device scale, clipping, selectors, navigation, and waiting. Exact accepted fields and defaults can change, so check the current API reference before depending on an option.

Decision Use it when Trade-off
Viewport or full page Viewport for the visible screen; full page for content below the fold Full-page files can be much larger and long pages may need scrolling to load lazy content
PNG, JPEG, or WebP Choose a format supported by your downstream use PNG is lossless; lossy formats can reduce file size, with quality settings where applicable
Viewport and device scale Match a browser size and pixel density for consistent captures Higher pixel density increases image dimensions and storage
Clip region Capture a defined rectangle Content outside the clip is omitted
Selector capture Capture a single element Missing or hidden elements need a readiness check and a clear failure outcome
Wait configuration Wait for navigation state or content readiness Long waits consume session time; overly short waits produce incomplete output
Scroll before full-page shot Long pages with lazy-loaded content Scrolling changes page state and can trigger more network requests

5. Reliability, concurrency, and cost

  • Bound parallel work. The number of simultaneous Browserless sessions is governed by your plan. The destination site’s capacity matters too; a client-side cap avoids sending a sudden burst.
  • Make batches resumable. Record each URL, output path, status, attempt count, and error. Skip completed URLs on restart.
  • Retry selectively. Retry transient connection errors and selected server errors with exponential backoff and a limit. Do not endlessly retry invalid URLs, authentication failures, missing selectors, or clear target-site blocks.
  • Check the payload. A successful HTTP status alone does not prove the expected page was captured. Check content type and inspect representative images for blank pages, CAPTCHA challenges, access-denied screens, or missing content.
  • Control memory and storage. Write each binary response as it completes. Use unique filenames or store a URL-to-file manifest to prevent collisions.
  • Estimate cost from work, not just URL count. Each REST request or remote browser session consumes service capacity. Large full-page captures also increase transfer and storage. Browserless plan limits and commercial details can change; verify current account documentation rather than assuming a fixed concurrency allowance.

There is no universal throughput figure: capture time depends on the target page, readiness condition, image size, network behavior, and available service concurrency. Run a representative small batch, then choose a concurrency cap that respects both your service plan and the sites you capture.

6. Troubleshooting

Symptom Likely cause Fix
401 or 403 response Missing, invalid, or unauthorized token; wrong endpoint Check the secret value and account-specific endpoint. Avoid printing tokens in logs.
429 or session-limit errors Too many simultaneous requests for available capacity Lower worker count, queue work, and check the current plan’s concurrency allowance.
Timeout or navigation error Slow page, long-running requests, or wait condition that never occurs Set a bounded timeout, use a page-specific ready selector, and report the URL as failed after limited retries.
Blank or white screenshot Page did not render as expected, navigation was premature, or automation was blocked Inspect the response image, adjust the readiness condition, and treat CAPTCHA or access-denied output as a target block rather than a successful capture.
CAPTCHA or access denied The target site is blocking or challenging automated access Do not assume a screenshot service can bypass it. Follow the site’s access rules and record the result as blocked.
Image file contains an error response Client saved a non-image response without checking headers or status Check HTTP status and Content-Type before writing the body as an image.
Missing lower-page images or sections Lazy loading did not run before capture Scroll the page before taking a full-page screenshot; verify relevant sections are present.
Playwright connection fails Using a CDP endpoint with a Playwright protocol connection, or incorrect WebSocket URL Use Browserless’s documented Playwright endpoint and matching connection method.
Files overwrite one another Output filenames are derived from a non-unique URL component Include a stable index or URL hash in the filename and keep a manifest mapping files to URLs.

7. Or skip the browser setup

For a direct URL-to-image request, ScreenshotNeo returns a screenshot or PDF from one GET request. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie banners are accepted and removed before the shot, along with known newsletter popups and chat widgets; each cleanup step can be turned off.
  • Bot checks, blank pages, failed loads, timeouts, and cache hits are never billed; response headers identify the page verdict and billing status.
  • An MCP server lets AI agents using Claude, Cursor, or another MCP client take screenshots, inspect page information, and capture PDFs.
  • The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots.

Sign up free for 1,000 screenshots a month with no card.

8. FAQ

Does the REST endpoint return JSON?

No. The documented screenshot response is image data. Save the response body as bytes and check the response status and content type first.

Can one request capture an entire URL list?

The workflow described here sends one screenshot request per URL. Your client can schedule those requests in bounded parallel batches.

Should I use full-page capture for every URL?

Only when below-the-fold content is needed. Viewport captures are smaller; full-page captures can require extra work to load lazy content.

Can Browserless guarantee a screenshot when a site shows a CAPTCHA?

No such guarantee is established by the reviewed documentation. A CAPTCHA or access-denied page is evidence that the target may be blocking automation.

Sources