Browserless vs Self-Hosted Puppeteer for Bulk Website Screenshots
Compare managed Browserless and self-hosted Puppeteer for screenshot batches, with runnable code, cost factors, queueing guidance, and a practical decision framework.
Short answer: use Browserless when you want a managed browser service and its REST screenshot endpoint or remote Puppeteer sessions fit your workload. Self-host Puppeteer when you need control over deployment and capacity and can take responsibility for browser workers, queues, Docker configuration, and host resources. Neither option is inherently faster or cheaper for every batch: measure your own URLs, capture settings, concurrency, failure rate, session duration, and operating effort.
For developers who want a one-request screenshot API without operating browser workers, ScreenshotNeo is the alternative to try first: it removes cookie banners, popups, and chat widgets before capture, and bills only clean shots.
How the two approaches differ
| Question | Browserless managed | Self-hosted Puppeteer |
|---|---|---|
| Who runs the browser fleet? | Browserless operates the managed browser service. | Your team deploys and operates the browser workers and their host environment. |
| How do you capture? | Use a REST screenshot endpoint for straightforward captures, or connect Puppeteer remotely over WebSocket for page-level automation. | Run Puppeteer against locally managed browser processes or workers and build the capture service and queue around them. |
| How is capacity controlled? | Managed session concurrency is associated with the plan. Parallel requests must be considered against that capacity. | You choose deployment capacity and configure concurrency and queue behavior; host sizing and operations are yours. |
| What does the supplied research establish about cost? | Browserless documents browser-time billing in 30-second increments, rounding partial increments upward. Current plan prices were not established here. | No workload-specific cost comparison is established. Include compute, memory, storage, networking, engineering, and on-call effort in your estimate. |
| What is the tradeoff? | Less browser infrastructure to operate, with managed usage and concurrency constraints to account for. | More deployment and capacity control, with more runtime responsibility. |
Choose based on the batch, not a generic speed claim
Browserless is worth evaluating when
- You want to avoid running and maintaining a browser fleet.
- A REST screenshot request covers the job without custom browser interactions.
- You need Puppeteer page operations but want the browser to run remotely.
- Your expected parallel sessions and session duration fit the plan after checking current capacity and pricing.
Self-hosting is worth evaluating when
- You need control over where browsers run, how capacity is provisioned, or how requests are queued.
- Your data or network requirements call for your own deployment.
- Your team can handle browser image updates, resource sizing, observability, restarts, and queue backpressure.
- You have measured enough workload data to compare infrastructure and operating effort with managed usage.
Browserless publishes an open-source Docker image with concurrency and queue settings. Its deployment guidance recommends 2 GB of shared memory for its example because Docker’s default 64 MB can cause Chrome crashes under load. Treat this as vendor deployment guidance, not a universal sizing formula; actual memory needs depend on pages, options, and parallel work.
Build a representative comparison
- Select representative URLs. Include fast and slow pages, pages with large images, redirects, and pages that sometimes fail. Use pages you are permitted to capture.
- Keep the capture identical. Use the same viewport, output format, full-page setting, navigation readiness rule, and timeout for both approaches.
- Test concurrency levels. Start with a small number of simultaneous jobs, then increase gradually. For Browserless, stay aware of plan concurrency and queueing. For self-hosting, watch host saturation and queue wait time.
- Record more than average duration. Track end-to-end latency, queue delay, successful outputs, retries, timeouts, browser crashes, CPU and memory use, and session duration.
- Include operational work. Record setup, deployment, upgrades, alerting, incident response, and the effort required to change capacity.
- Estimate cost from current inputs. Verify Browserless plan prices and concurrency directly. For self-hosting, estimate compute and storage from measured resource use, and include engineering and operations costs.
Browserless documentation describes parallel screenshot requests and gives an illustrative example where elapsed time is bounded roughly by the slowest request. That example is not a throughput benchmark for arbitrary websites, networks, or plan sizes. Do not use it to predict your own batch completion time.
Browserless: REST endpoint or remote Puppeteer?
Browserless offers both a REST screenshot API and remote Puppeteer sessions. Prefer the REST endpoint when each job is simply “open this URL and save the image.” Use remote Puppeteer when the workflow needs page-level interactions or existing Puppeteer logic. Its Puppeteer guide recommends puppeteer-core and replacing local launch() with a remote connect(). A local file path is not a path on the remote browser machine.
Minimal remote Puppeteer batch example
Install puppeteer-core in your Node project and set BROWSERLESS_WS_ENDPOINT to the WebSocket endpoint and token for your Browserless deployment. Keep credentials in an environment variable or secret store. The precise endpoint and authentication form depend on the deployment; consult the current Browserless connection guide.
import puppeteer from 'puppeteer-core';
const endpoint = process.env.BROWSERLESS_WS_ENDPOINT;
if (!endpoint) throw new Error('Set BROWSERLESS_WS_ENDPOINT');
const urls = [
'https://example.com',
'https://www.iana.org/',
];
const browser = await puppeteer.connect({ browserWSEndpoint: endpoint });
try {
// Keep this at or below the concurrency available to your account.
const results = await Promise.all(urls.map(async (url, index) => {
const page = await browser.newPage();
try {
await page.setViewport({ width: 1440, height: 900 });
await page.goto(url, { waitUntil: 'networkidle2', timeout: 60000 });
const path = `shot-${index}.png`;
await page.screenshot({ path, fullPage: true, type: 'png' });
return { url, path, ok: true };
} catch (error) {
return { url, ok: false, error: String(error) };
} finally {
await page.close();
}
}));
console.log(results);
} finally {
await browser.disconnect();
}
Promise.all above is appropriate only for a small batch known to fit the available concurrency. For a large list, use a bounded worker pool so the script does not open an unbounded number of pages or sessions. Ensure every page is closed and the browser connection is disconnected: Browserless says open sessions persist until timeout and continue to consume usage.
REST screenshot request with cURL
Browserless REST routes and endpoint host vary by deployment. The following shows the request shape; use the exact endpoint and authentication method in the current official API documentation for your account.
curl --fail --silent --show-error \
-X POST "$BROWSERLESS_SCREENSHOT_ENDPOINT" \
-H 'Content-Type: application/json' \
-d '{"url":"https://example.com","options":{"type":"png","fullPage":true}}' \
--output example.png
Python calling a REST screenshot endpoint
This uses the endpoint and authentication configured for your deployment. The response body is written as the image; check the HTTP status before treating it as a screenshot.
import os
import requests
endpoint = os.environ['BROWSERLESS_SCREENSHOT_ENDPOINT']
response = requests.post(
endpoint,
json={"url": "https://example.com", "options": {"type": "png", "fullPage": True}},
timeout=90,
)
response.raise_for_status()
with open('example.png', 'wb') as image_file:
image_file.write(response.content)
Node.js calling a REST screenshot endpoint
const endpoint = process.env.BROWSERLESS_SCREENSHOT_ENDPOINT;
if (!endpoint) throw new Error('Set BROWSERLESS_SCREENSHOT_ENDPOINT');
const response = await fetch(endpoint, {
method: 'POST',
headers: { 'content-type': 'application/json' },
body: JSON.stringify({
url: 'https://example.com',
options: { type: 'png', fullPage: true },
}),
signal: AbortSignal.timeout(90_000),
});
if (!response.ok) throw new Error(`Screenshot request failed: HTTP ${response.status}`);
await Bun.write('example.png', new Uint8Array(await response.arrayBuffer()));
The REST snippets deliberately leave endpoint and authentication details deployment-specific: do not copy a cloud URL into a self-hosted deployment or assume that cloud authentication applies to a private endpoint. Use the official docs for the exact current route and request schema.
Self-hosting: the pieces you must operate
A production bulk screenshot service needs more than a call to page.screenshot(). At minimum, plan for:
- Queue and backpressure: accept jobs separately from browser workers, cap active pages, and bound queued work so spikes do not exhaust memory.
- Worker lifecycle: create and close pages reliably, recycle unhealthy browser processes, and handle deployment updates.
- Resource sizing: monitor memory, CPU, shared memory, file descriptors, and temporary storage against real pages and concurrency.
- Timeouts and retries: set navigation and total-job deadlines. Retry transient navigation failures selectively; avoid retry storms and repeated expensive pages.
- Isolation: decide whether jobs may share a browser process or require stronger process/container separation. Untrusted pages can consume resources and should be treated as workload input.
- Output storage: write to durable object storage or a controlled output location, and define retention and cleanup.
- Observability: log job IDs, URL host, queue time, capture duration, browser exit reason, and output status without logging secrets or sensitive page contents.
Browserless’s Docker guidance provides concurrency and queue controls and recommends 2 GB shared memory for its example deployment. Apply that as a starting point for evaluation, then size and tune from observed workload behavior.
Bulk reliability and performance practices
Use bounded concurrency
For batch size N, a concurrency limit C prevents the job from launching all N browsers or pages at once. Choose C from available managed concurrency or measured self-host capacity. Raising C can reduce queue time only while the service and target sites can sustain the extra parallel work.
Define when a page is ready
Navigation completion does not guarantee every application is visually settled. Network-idle waits can also stall on pages with persistent connections. Decide whether your task needs a DOM state, a selector, a fixed delay, or a network-idle condition, and give the job a total deadline. Keep the readiness rule consistent in comparisons.
Handle retries and partial success
- Capture each URL as an independent job and save a result record even when it fails.
- Retry only transient errors, with a small retry limit and backoff.
- Do not retry invalid URLs, permanent authorization errors, or deterministic page failures without changing the input.
- Make output names or object keys unique per job so a retry cannot silently overwrite another URL’s screenshot.
- On a batch interruption, resume only missing or failed jobs rather than repeating successful captures.
Control screenshot size
Full-page images can be very tall and consume more browser memory and output storage than viewport captures. Use the smallest viewport and image format that meet the downstream need. Account for output transfer and storage as well as browser time when estimating total cost.
Cost and operational tradeoffs
Browserless documents browser-time billing in 30-second increments, rounding a partial increment up; its example says a 45-second session consumes two units. That rule means session duration and connection cleanup matter when estimating managed usage. It does not by itself establish the total price: verify current plan pricing, concurrency, and usage rules before committing.
Self-hosting has no universally comparable per-screenshot cost in the available evidence. Calculate from measured CPU and memory utilization, worker idle time, storage and network use, and the infrastructure needed for peak concurrency. Add engineering time for deployment, browser updates, monitoring, incident response, and capacity changes. The dossier does not establish which approach is cheaper or faster for a particular workload.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. One GET request can return a PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: HTTP ${res.status}`);
await Bun.write('shot.webp', new Uint8Array(await res.arrayBuffer()));
- Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
- Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing. Responses identify the page verdict and billing status in headers.
- An MCP server gives AI agents, including Claude, Cursor, and other MCP clients, tools to take screenshots, get page information, and capture PDFs.
- 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000. Every feature is on every plan.
Sign up for ScreenshotNeo and get 1,000 free screenshots a month with no card.
Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| Browserless connection fails or returns unauthorized | Wrong token, endpoint, or deployment host; cloud and private deployments may use different endpoints. | Confirm the deployment and use its current connection details. Keep the token out of source control and logs. |
| Requests queue or hit a concurrency limit | The batch exceeds available managed session concurrency or self-hosted worker capacity. | Bound client concurrency, inspect queue delay, and confirm current plan capacity or worker configuration. |
| Managed usage is higher than expected | Long-lived sessions remain open; browser time rounds up in 30-second increments. | Close pages and disconnect when finished, set deadlines, and measure session duration. Verify current billing rules. |
| Chrome crashes in Docker under load | Insufficient shared memory or host resources can cause instability. | Review the Browserless deployment guidance, including its 2 GB shared-memory recommendation for the example, then monitor and size for the actual workload. |
| Local files are missing in remote Puppeteer | The browser runs on the remote machine, where your local filesystem path does not exist. | Transfer required data through an appropriate remote mechanism or store outputs from the client side; do not assume local paths are shared. |
| Screenshot is blank or incomplete | Capture ran before client rendering, images, or other visible content finished loading. | Wait for a relevant selector or application-ready state, or adjust the navigation readiness condition and timeout. |
| Navigation hangs on network idle | The page maintains a persistent network connection or ongoing requests. | Use a selector or a bounded delay tied to the page’s behavior, plus an overall deadline. |
| Output file contains an error response | The script saved a failed HTTP response as image bytes. | Check the response status before writing output; use cURL’s --fail or the client’s status check. |
| Self-hosted workers slow down as batches grow | Concurrency may exceed CPU, memory, shared-memory, or target-site capacity. | Reduce worker concurrency, inspect resource metrics and queue delay, and rerun the representative workload at incremental levels. |
FAQ
Is Browserless just an HTTP screenshot API?
No. It also supports remote Puppeteer sessions over WebSocket. The REST route suits straightforward screenshot requests; remote sessions suit workflows needing Puppeteer operations.
Can the documentation prove which choice is faster?
No. The available material does not provide an independent comparative benchmark. Measure the same URLs, options, and concurrency on the actual deployment choices.
Does self-hosting guarantee a lower cost per screenshot?
No. Compute is only one input; capacity utilization and the work of operating the browser service also matter. Compare using measured workload and staffing inputs.
Should I use Browserless for a large batch?
It can be evaluated for bulk capture, but match parallel work to available concurrency and account for session duration. Confirm current plan limits before sizing a production batch.
