ScreenshotNeo

BlogComparisons

Web Capture SDK Options Explained

Compare browser SDKs, hosted REST APIs and persistent sessions, then choose the right screenshot workflow with runnable code and practical settings.

By the ScreenshotNeo team29 September 202610 min read

Web Capture SDK Options Explained

Direct answer: choose a browser automation SDK when screenshots are one step in a larger workflow that needs navigation, clicks, authentication, network interception, or a browser session you control. Choose a hosted REST screenshot API for a single capture without operating browser infrastructure. Choose a persistent browser connection when the page must remain open across several commands. The image settings still matter in every model: viewport versus full page, element or clip selection, output format, quality, device scale, background handling, and lazy-loaded content.

This guide explains those choices, shows runnable Puppeteer and Playwright examples, and covers the hosted REST approach. It also shows where ScreenshotNeo fits when you want a clean image from one request.

1. The three capture architectures

Browser automation library

Puppeteer is a JavaScript library for automating Chrome and Firefox through the Chrome DevTools Protocol and WebDriver BiDi. Its documented capabilities include screenshots, PDFs, page interaction, network interception, and performance analysis. Playwright likewise supports viewport, element, and full-scrollable-page screenshots. These libraries run in your application or CI environment, so you choose the browser version, launch flags, credentials, and session lifecycle.

Use this model when the screenshot depends on actions such as signing in, dismissing a dialog, opening a menu, scrolling a virtual list, or waiting for an application state. It is also a natural fit when screenshots are part of end-to-end tests or visual regression jobs.

Hosted REST capture

A hosted service accepts a URL and capture settings, starts or reuses a browser on its infrastructure, and returns an image or PDF. Browserless documents screenshots as a REST task and supports Puppeteer-style options such as full-page capture, clipping, viewport size, device scale factor, and selector-based capture. This avoids installing and patching browsers in your own workers.

A REST request is usually the shortest path for a one-off capture, scheduled thumbnail, report image, or URL list. Check the provider’s current authentication, limits, formats, pricing, and privacy terms in its documentation; the research for this article does not establish those details for Browserless.

Persistent browser connection

A WebSocket or protocol connection keeps a browser available while your code sends multiple commands. Browserless distinguishes this workflow from a one-shot REST request. It is useful when a task has a long interaction sequence, needs state between captures, or already uses Chrome DevTools Protocol (CDP) or another supported protocol. The trade-off is lifecycle management: you must handle connection failures, browser cleanup, idle timeouts, and concurrency.

2. Pick by workflow, not by brand

Need Start with Questions to answer
One URL, one image Hosted REST API How are URLs authenticated? Which formats, limits, and settings are supported?
Capture is part of a test or script Puppeteer or Playwright Which browser engines and language fit your existing test setup? Do you need clicks, cookies, or network control?
Several commands in one browser session Persistent connection How will you reconnect, close pages, isolate users, and enforce timeouts?
Long or lazy-loaded pages Any option, tested on the real page Does full-page mode include content below the fold? Is explicit scrolling required?

There is no universal winner in the reviewed documentation. Compare target browsers, programming language, session requirements, capture fidelity, and the amount of infrastructure your team wants to operate.

3. DIY capture with Puppeteer

Install Puppeteer in a new Node.js project. The package downloads a compatible browser during installation unless your environment is configured to use an existing executable.

Consent banners and overlays can change the pixels you capture; cleanup should be an explicit part of the workflow.
Consent banners and overlays can change the pixels you capture; cleanup should be an explicit part of the workflow.
npm install puppeteer

Create screenshot.mjs:

import puppeteer from 'puppeteer';

const browser = await puppeteer.launch({ headless: true });
try {
  const page = await browser.newPage();
  await page.setViewport({ width: 1440, height: 900, deviceScaleFactor: 1 });
  await page.goto('https://example.com', {
    waitUntil: 'networkidle2',
    timeout: 60_000
  });

  await page.screenshot({
    path: 'example-full.png',
    fullPage: true,
    type: 'png',
    omitBackground: false
  });
} finally {
  await browser.close();
}

Run it with node screenshot.mjs. The important options are:

  • fullPage captures the full scrollable page instead of only the viewport.
  • clip captures a rectangle with x, y, width, and height.
  • type can be PNG, JPEG, or WebP where supported by the installed Puppeteer version.
  • quality applies to JPEG and WebP; it does not apply to PNG.
  • omitBackground makes the page background transparent when the format supports transparency.
  • captureBeyondViewport controls whether a clipped region outside the current viewport may be captured.

Capture one element

const card = await page.locator('.pricing-card').boundingBox();
if (!card) throw new Error('pricing card is not visible');
await page.screenshot({ path: 'card.webp', type: 'webp', quality: 82, clip: card });

Element geometry can change after fonts or animations finish. Wait for the selector, make the element visible, and disable or pause animations with injected CSS when pixel stability matters.

Lazy-loaded content and long pages

Full-page mode does not guarantee that an application has loaded every image or virtualized row. Scroll in steps, wait for images, then capture:

await page.evaluate(async () => {
  await new Promise(resolve => {
    let y = 0;
    const step = 700;
    const timer = setInterval(() => {
      window.scrollBy(0, step);
      y += step;
      if (y >= document.body.scrollHeight) {
        clearInterval(timer);
        window.scrollTo(0, 0);
        resolve();
      }
    }, 100);
  });
});
await page.screenshot({ path: 'loaded.png', fullPage: true });

Scrolling behavior is page-specific. A service such as Browserless documents a scrollPage option for lazy-loaded content; do not assume that option exists in every wrapper.

4. DIY capture with Playwright

Playwright is another candidate when its supported browser engines, language bindings, and existing test tooling match your project.

npm install -D playwright
npx playwright install chromium
import { chromium } from 'playwright';

const browser = await chromium.launch();
try {
  const page = await browser.newPage({
    viewport: { width: 1440, height: 900 },
    deviceScaleFactor: 1
  });
  await page.goto('https://example.com', { waitUntil: 'networkidle', timeout: 60_000 });
  await page.screenshot({ path: 'page.webp', fullPage: true, type: 'webp', quality: 82 });
  await page.locator('.pricing-card').screenshot({ path: 'card.png' });
} finally {
  await browser.close();
}

Use a selector screenshot for one component and fullPage: true for the complete scrollable page. Validate behavior on your actual site because network-idle conditions, animations, iframes, and virtual lists can differ.

5. Hosted REST capture: what to send

A hosted endpoint generally needs an access key, a target URL, and capture parameters. Typical controls include viewport width and height, device scale, full-page mode, an element selector or clip rectangle, image type, JPEG/WebP quality, and background behavior. Some APIs also expose scrolling for lazy content. API wrappers may rename fields even when they forward Puppeteer-style options, so follow the provider’s schema.

For a production integration, record the response status, content type, request identifier, and provider-specific error body. Set a client timeout longer than the page’s expected load time, retry only transient failures, and avoid retrying authentication or invalid-URL errors.

6. Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns PNG, JPEG, WebP, or a PDF, while the service handles the browser work. See the ScreenshotNeo documentation for the full parameter list.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await Bun.write('shot.webp', bytes);

ScreenshotNeo supports full-page capture with lazy images loaded, CSS selector element capture, dark mode, twelve device presets or any viewport, retina scale, PDF paper sizes and margins, HTML/CSS to image, custom CSS and JavaScript, pre-capture clicks, hidden selectors, waits for a selector, delay or network idle, request and resource blocking, custom headers, cookies, user agent and Authorization, timezone, geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification.

Cookie banners, newsletter popups, and chat widgets are accepted or removed before capture through more than 60 known consent platforms and related overlays; each cleanup step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed. The response identifies the result with X-Page-Verdict and X-Billed headers.

For AI workflows, its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month without a card. Paid plans start at $5 for 3,000 shots; higher plans are Growth ($15/15,000), Pro ($39/60,000), Scale ($99/250,000), and Business ($249/1,000,000). Yearly billing gives two months free, and every feature is available on every plan. Create a free ScreenshotNeo account to start.

7. Settings that change the result

Setting Use it when Common mistake
Viewport You need a desktop, tablet, or mobile layout Capturing at a width that triggers an unexpected breakpoint
Device scale factor You need sharper retina-like pixels Comparing images with different pixel dimensions
Full page You need all scrollable content Assuming virtualized or lazy content is already loaded
Selector or clip You need one component or region Measuring before fonts and layout settle
Format and quality You need transparency, small files, or lossless output Setting JPEG quality for PNG
Background You need a transparent asset Using a format or CSS background that cannot preserve alpha
The same URL can produce different results depending on viewport, waits, scrolling, and capture scope.
The same URL can produce different results depending on viewport, waits, scrolling, and capture scope.

8. Reliability, performance, and cost

Reliability checklist

  • Set navigation and overall request deadlines.
  • Wait for a meaningful selector or application state instead of an arbitrary short delay.
  • Freeze animations and timestamps for visual regression comparisons.
  • Use isolated browser contexts for different users or cookie jars.
  • Capture response headers and error bodies so failures are diagnosable.
  • Retry network and provider-transient errors with bounded exponential backoff; do not blindly retry invalid URLs or 401 responses.

Performance considerations

Full-page images, high device scale factors, large PDFs, and long JavaScript waits consume more browser time and memory than a viewport screenshot. Blocking unnecessary ads, trackers, fonts, or media can reduce page work, but verify that blocked resources do not change the layout you need to document. Reuse a browser in a controlled worker or use a hosted service when browser startup dominates short jobs. Cache only when the page’s freshness requirements allow it.

Cost considerations

With self-hosted Puppeteer or Playwright, budget for browser CPU and memory, container images, patching, concurrency limits, queueing, and operational time. A hosted API moves those costs into a per-capture or subscription price; compare current vendor pricing and limits directly. ScreenshotNeo’s billing behavior is explicit: only clean shots are billed, while bot checks, blank pages, timeouts, failed loads, and cache hits are free, and headers report whether a response was billed.

9. Troubleshooting common failures

Browser fails to launch

Cause: missing browser binaries, sandbox restrictions, or incompatible system libraries. Fix: install the browser required by your package, use the documented container dependencies, and inspect the launch error before adding flags. Avoid copying production launch arguments from an unrelated environment.

Screenshot is blank or incomplete

Cause: the page is still loading, a client-side app has not mounted, or a bot challenge is blocking content. Fix: wait for a stable selector or network condition, inspect the page HTML and console output, and classify the URL as a challenge rather than retrying forever.

Full-page image misses lower sections

Cause: lazy images or virtualized content load only after scrolling. Fix: scroll in steps, wait for image completion, then capture; use a documented provider option such as Browserless’s scrollPage where available.

Element selector cannot be found

Cause: the element is inside an iframe, appears after an interaction, or uses a different responsive layout. Fix: wait for the frame and selector, interact with the page first, and confirm the viewport breakpoint.

Images differ between runs

Cause: animations, rotating content, fonts, ads, timestamps, or device scale differences. Fix: disable motion, block unstable resources, wait for fonts, fix the viewport and scale, and compare normalized output dimensions.

REST request returns an error

Cause: invalid credentials, malformed URL encoding, unsupported options, or a provider timeout. Fix: URL-encode the target, check authentication and content type, remove options one at a time, and retain the provider’s status and error payload for support.

10. A practical decision checklist

  1. Write down whether the job is one capture or a multi-step session.
  2. List required interactions: login, clicks, scrolling, downloads, frames, and network rules.
  3. Choose the browser engines and language your team already supports.
  4. Define output requirements: viewport, full page, selector, format, quality, transparency, and scale.
  5. Test two or three representative pages, including a long lazy-loaded page and a page with consent UI.
  6. Measure queue time, browser time, output size, failure rate, and operational effort using your own workload.
  7. Set timeouts, retry rules, cleanup, logging, and a cost ceiling before production.

FAQ

Should I use Puppeteer or Playwright for screenshots?

Use the library that matches your target browsers, language bindings, existing tests, and required session behavior. The reviewed documentation does not provide a controlled performance comparison.

Is a REST API the same as a persistent browser connection?

No. REST is a one-shot task request. A persistent connection keeps a browser and page available for continuing commands and state.

Why does full-page capture need extra testing?

Long pages may lazy-load images or virtualize rows only after scrolling. Confirm that all required content is present before saving the image.

Can I capture an element instead of the whole page?

Yes. Puppeteer and Playwright can capture an element or clipped rectangle, and hosted APIs may expose a selector or clip option.

When should I use ScreenshotNeo?

Use it when you want a single request, cleanup of consent banners and overlays, explicit verdict and billing headers, optional advanced capture controls, or MCP tools for AI agents. Start with the free 1,000-shot monthly plan without a card.