ScreenshotNeo

BlogHow-to

How to Build a JavaScript Screenshot API for URLs and Base64 Images

Build a Node.js screenshot API with Playwright, returning PNGs or base64 from URLs and image payloads. Includes runnable code, security limits, and troubleshooting.

By the ScreenshotNeo team29 September 202615 min read

How to Build a JavaScript Screenshot API for URLs and Base64 Images

A JavaScript screenshot API accepts a URL or image payload, opens or constructs a page in a controlled browser, captures it, and returns either image bytes or base64. A practical Node.js implementation can use Playwright: validate requests, isolate each capture in a fresh browser context, set explicit limits and timeouts, and return a structured response. For URL screenshots, Playwright can capture a full page or a selected element; for image input, validate and decode the data URI, then render it in a controlled document before capturing.

This guide builds a runnable POST /screenshot service with Express and Playwright. It covers the request contract, URL and base64 handling, output choices, production safeguards, errors, and operating costs. Playwright’s documentation notes that its screenshot API accepts parameters for image format, clip area, quality, and other capture behavior. Playwright Screenshots

1. Choose an API contract

Use one endpoint with a discriminated request body: callers provide either url or image, never both. Returning binary image bytes is efficient for clients that immediately save or display the result. Returning JSON with base64 is useful when the consumer requires JSON transport, but the encoded string is larger than the underlying bytes and must be decoded by the client.

Field Purpose Suggested validation
url Page to navigate to and capture. HTTPS by default; reject credentials, unsupported schemes, private destinations, and overlong values.
image Base64 or data:image/…;base64,… payload. Allow only approved image types; check encoded and decoded size; verify that decoded content is an image.
type png, jpeg, or webp. Allow only formats supported by the chosen browser and capture path.
fullPage Capture the full scrollable page instead of the viewport. Boolean; apply a maximum output/page dimension.
clip Capture a rectangle with x, y, width, and height. Finite nonnegative coordinates and positive dimensions within configured bounds.
quality Lossy output quality for JPEG or WebP where supported. Integer from 0 to 100; ignore or reject for PNG.
omitBackground Request a transparent background where supported. Boolean; most useful with PNG.
viewport Page width and height in CSS pixels. Bound each dimension and their product; optionally configure device scale.
waitUntil / selector Define when the page is ready for capture. Use a documented readiness policy and a bounded timeout.

Keep request options explicit. Never accept arbitrary browser flags or executable paths from callers. If a feature is not implemented, reject it instead of silently implying that it was honored.

2. Create the Node.js project

Use a current supported Node.js release. Create a directory, initialize a package, and install the server and browser library:

A screenshot endpoint validates a request, captures in an isolated browser context, and returns bytes or base64.
A screenshot endpoint validates a request, captures in an isolated browser context, and returns bytes or base64.
npm init -y
npm install express playwright
npx playwright install chromium

The browser install is a deployment dependency: install it in the build or container image, not on every request. This example uses Chromium. Playwright’s page screenshot method returns a buffer in Node.js, which can be sent directly or converted to base64. See the Playwright Page API and screenshot guide.

Save the following as server.mjs. It implements JSON/base64 responses, binary image responses, URL capture, and image-payload capture. The sample includes basic limits and validation; the production SSRF and concurrency controls discussed below still need to be completed for your deployment environment.

import express from 'express';
import { chromium } from 'playwright';

const app = express();
const PORT = Number(process.env.PORT || 3000);
const MAX_BODY = '8mb';
const MAX_IMAGE_BYTES = 5 * 1024 * 1024;
const NAV_TIMEOUT_MS = 20_000;
const CAPTURE_TIMEOUT_MS = 10_000;
const MAX_DIMENSION = 3000;

app.use(express.json({ limit: MAX_BODY }));

let browser;
const allowedTypes = new Set(['png', 'jpeg', 'webp']);
const mimeByType = { png: 'image/png', jpeg: 'image/jpeg', webp: 'image/webp' };

function bad(message) {
  const error = new Error(message);
  error.status = 400;
  return error;
}

function dimensions(value = {}) {
  const width = Number(value.width ?? 1280);
  const height = Number(value.height ?? 800);
  if (!Number.isInteger(width) || !Number.isInteger(height) ||
      width < 1 || height < 1 || width > MAX_DIMENSION || height > MAX_DIMENSION ||
      width * height > MAX_DIMENSION * MAX_DIMENSION) {
    throw bad('viewport dimensions are outside the allowed range');
  }
  return { width, height };
}

function parseImage(input) {
  if (typeof input !== 'string') throw bad('image must be a base64 string or data URI');
  let mediaType = 'image/png';
  let payload = input;
  const match = input.match(/^data:(image\/(?:png|jpeg|webp));base64,([A-Za-z0-9+/]*={0,2})$/);
  if (input.startsWith('data:')) {
    if (!match) throw bad('image data URI must be PNG, JPEG, or WebP base64');
    [, mediaType, payload] = match;
  } else if (!/^[A-Za-z0-9+/]*={0,2}$/.test(input)) {
    throw bad('image must contain standard base64 characters');
  }
  const bytes = Buffer.from(payload, 'base64');
  if (!bytes.length || bytes.length > MAX_IMAGE_BYTES) throw bad('decoded image is empty or too large');
  // Buffer.from is permissive: round-trip checking rejects malformed encodings.
  if (bytes.toString('base64').replace(/=+$/, '') !== payload.replace(/=+$/, '')) {
    throw bad('malformed base64 payload');
  }
  return { mediaType, bytes };
}

function validateUrl(value) {
  if (typeof value !== 'string' || value.length > 2048) throw bad('url is required and must be at most 2048 characters');
  let parsed;
  try { parsed = new URL(value); } catch { throw bad('url is invalid'); }
  if (parsed.protocol !== 'https:' && parsed.protocol !== 'http:') throw bad('only HTTP and HTTPS URLs are supported');
  if (parsed.username || parsed.password) throw bad('URLs containing credentials are not accepted');
  // Production must also resolve and block private, loopback, link-local, and metadata IP ranges,
  // and re-check every redirect and DNS resolution to mitigate SSRF and rebinding.
  return parsed.toString();
}

function outputOptions(body) {
  const type = body.type ?? 'png';
  if (!allowedTypes.has(type)) throw bad('type must be png, jpeg, or webp');
  if (body.quality !== undefined && (!Number.isInteger(body.quality) || body.quality < 0 || body.quality > 100)) {
    throw bad('quality must be an integer from 0 to 100');
  }
  return { type, quality: body.quality };
}

async function runCapture(body) {
  const { type, quality } = outputOptions(body);
  const viewport = dimensions(body.viewport);
  const context = await browser.newContext({ viewport });
  try {
    const page = await context.newPage();
    page.setDefaultNavigationTimeout(NAV_TIMEOUT_MS);
    page.setDefaultTimeout(CAPTURE_TIMEOUT_MS);
    if (body.url && body.image) throw bad('provide url or image, not both');

    if (body.url) {
      const url = validateUrl(body.url);
      const waitUntil = body.waitUntil ?? 'domcontentloaded';
      if (!['load', 'domcontentloaded', 'networkidle', 'commit'].includes(waitUntil)) throw bad('unsupported waitUntil value');
      await page.goto(url, { waitUntil, timeout: NAV_TIMEOUT_MS });
      if (body.selector) await page.locator(body.selector).waitFor({ state: 'visible', timeout: CAPTURE_TIMEOUT_MS });
    } else if (body.image) {
      const { mediaType, bytes } = parseImage(body.image);
      await page.setContent('<!doctype html><html><body style="margin:0;display:grid;place-items:center;min-height:100vh;background:transparent"><img id="source" style="max-width:100%;max-height:100%;object-fit:contain"></body></html>');
      await page.locator('#source').evaluate((img, data) => {
        img.src = `data:${data.mediaType};base64,${data.payload}`;
      }, { mediaType, payload: bytes.toString('base64') });
      await page.locator('#source').evaluate(img => img.decode());
    } else {
      throw bad('provide url or image');
    }

    const shot = { type, fullPage: Boolean(body.fullPage), timeout: CAPTURE_TIMEOUT_MS };
    if ((type === 'jpeg' || type === 'webp') && quality !== undefined) shot.quality = quality;
    if (body.omitBackground !== undefined) shot.omitBackground = Boolean(body.omitBackground);
    if (body.clip !== undefined) {
      const c = body.clip;
      if (!c || ![c.x, c.y, c.width, c.height].every(Number.isFinite) || c.x < 0 || c.y < 0 || c.width <= 0 || c.height <= 0 || c.width > MAX_DIMENSION || c.height > MAX_DIMENSION) {
        throw bad('clip must have nonnegative x/y and positive bounded width/height');
      }
      shot.clip = c;
    }
    return await page.screenshot(shot);
  } finally {
    await context.close();
  }
}

app.post('/screenshot', async (req, res) => {
  try {
    const bytes = await runCapture(req.body ?? {});
    const type = req.body.type ?? 'png';
    if (req.query.response === 'json') {
      res.json({ type, encoding: 'base64', data: bytes.toString('base64') });
    } else {
      res.type(mimeByType[type]).send(bytes);
    }
  } catch (error) {
    const status = error.status ?? (error.name === 'TimeoutError' ? 504 : 502);
    res.status(status).json({ error: error.message || 'capture failed' });
  }
});

const server = app.listen(PORT, async () => {
  browser = await chromium.launch({ headless: true });
  console.log(`Screenshot API listening on ${PORT}`);
});

async function shutdown() {
  server.close();
  if (browser) await browser.close();
}
process.on('SIGINT', shutdown);
process.on('SIGTERM', shutdown);

Start it with node server.mjs. The endpoint defaults to an image/png response. Add ?response=json to receive {"type":"png","encoding":"base64","data":"…"}. The data is raw base64, not a complete data URI. The implementation accepts a data URI or raw base64 for input; raw payloads default to PNG, so send a data URI when the source type is JPEG or WebP.

3. Call the endpoint

URL capture as a binary response

curl -sS -X POST http://localhost:3000/screenshot \
  -H 'Content-Type: application/json' \
  -d '{"url":"https://example.com","type":"png","fullPage":true}' \
  -o page.png

URL capture as base64 JSON

curl -sS -X POST 'http://localhost:3000/screenshot?response=json' \
  -H 'Content-Type: application/json' \
  -d '{"url":"https://example.com","type":"jpeg","quality":82}'

Base64 image input

For a data URI, send the media type and payload together. The service decodes it, renders it on a neutral document, waits for image decoding, and captures the render. For larger payloads, read from a file client-side and send the encoded content rather than manually copying a huge literal:

IMAGE_DATA="$(base64 -w 0 input.png)"
curl -sS -X POST 'http://localhost:3000/screenshot?response=json' \
  -H 'Content-Type: application/json' \
  --data-binary "{\"image\":\"data:image/png;base64,${IMAGE_DATA}\",\"type\":\"png\"}"

On systems where base64 -w 0 is unavailable, use that platform’s no-wrap option or strip line breaks. Ensure the client request body limit accommodates the encoded size; base64 adds roughly one third to the binary payload.

4. Return binary or base64

Binary is the better default for browser downloads, object storage uploads, and image-processing pipelines: it avoids JSON parsing and base64 expansion. Include a correct Content-Type, and consider Content-Disposition: attachment; filename="capture.png" for download-oriented clients. Base64 JSON is useful when an API consumer cannot accept a binary response, when the result must travel inside a JSON envelope, or when an agent tool expects text. Define whether the field is raw base64 or a complete data URI; this server uses raw base64 and identifies the format separately.

Do not log response bodies or base64 image inputs. Set response size limits and consider writing large captures to object storage and returning a short-lived authorized reference instead of a very large JSON response. That storage lifecycle and access policy are application-specific.

5. Capture options and readiness

Option When to use it Important behavior
fullPage Long articles, landing pages, and pages that scroll. May produce very tall images; cap document height and pixels, and expect lazy-loaded content to need scrolling or application-specific readiness logic.
clip A known region of the page. Use page-coordinate bounds and ensure the target region exists; clip and full-page combinations should be defined and validated by your API.
type and quality Choose PNG for sharp text/transparency, JPEG for photographic content, WebP where clients support it. Quality is relevant to lossy formats; output size depends on the page and encoding.
omitBackground Transparent output for diagrams or overlay assets. Use PNG and verify downstream alpha support.
viewport Reproduce a target layout or responsive breakpoint. Keep dimensions and total pixel area bounded; device scale multiplies pixel count and memory.
selector Wait until a known component appears. A visible selector wait is more reliable than guessing with a fixed sleep, but it can still time out if the page never renders it.
waitUntil Choose navigation readiness. networkidle may never occur on analytics, polling, or streaming pages; domcontentloaded is faster but does not guarantee images or application data are ready.

Wait policy is part of the API contract. A useful extension is an explicit delayMs with a strict maximum, or a caller-provided readiness selector. Do not permit unbounded delay or network-idle waiting. If a site lazy-loads images as the page scrolls, implement a bounded scroll-and-wait routine and enforce a maximum page height; simply requesting a full-page capture does not guarantee every site’s lazy content has loaded.

Element screenshots are often more stable than clipping guessed coordinates. Playwright supports taking an element screenshot through a locator. An API can add selector capture by locating a visible element and calling locator.screenshot(); validate the selector length and enforce the same timeout. Do not interpolate selectors or caller strings into executable JavaScript.

6. Accept and handle base64 images safely

A conventional image data URI looks like data:image/png;base64,<payload>. Split media type from payload, allow only formats your service supports, reject malformed or oversized encodings, and validate the decoded bytes as an actual image. Base64 syntax alone does not establish that a payload is safe or even an image. The sample round-trip check catches malformed encodings but is not a substitute for a robust image decoder and decompression-bomb limits.

The example displays the validated payload in a document created by page.setContent; it does not navigate to an arbitrary URL from the payload. A stricter implementation can use a route for a fixed local origin and serve only the decoded bytes with the verified content type. If you accept SVG, treat it as active document content: scripts, external references, and complex resource loads change the threat model. This example intentionally excludes SVG and only allows PNG, JPEG, and WebP.

For image input, distinguish “capture this image rendered at a chosen viewport” from “convert/re-encode this file.” A browser screenshot captures a rendering and may scale or letterbox according to CSS. If exact byte-preserving conversion is required, use a dedicated image-processing path instead of a browser screenshot.

7. Secure a URL screenshot service

A caller-controlled URL makes the browser an SSRF-capable network client. Scheme checks alone are insufficient. Before navigation, resolve hostnames and block loopback, private, link-local, multicast, and cloud metadata destinations for both IPv4 and IPv6. Re-check after redirects and protect against DNS rebinding; enforce egress network rules as a second boundary. The sample labels these checks but does not implement them because a correct policy depends on the deployment network and resolver.

  • Accept only HTTP and HTTPS, reject URL credentials, cap URL length, and normalize/parse with the platform URL parser.
  • Block private and special-use address ranges at connection time, including redirected hosts; account for IPv4-mapped IPv6 and alternate numeric IP forms.
  • Use a fresh browser context per request to prevent cookies, storage, permissions, and cache state from crossing tenants.
  • Set body, image bytes, viewport pixels, page height, navigation, readiness, screenshot, and total request limits.
  • Decide whether third-party scripts, fonts, images, and downloads are permitted. They affect fidelity, latency, and what data the browser discloses.
  • Run the browser as an unprivileged user in a restricted container, with no sensitive host mounts or credentials.
  • Limit concurrency and queue work so one very long page cannot exhaust memory or browser processes.
  • Log a request ID, sanitized hostname, duration, output type, and failure class. Never log auth headers, cookies, base64 payloads, or full sensitive URLs.

Do not expose custom headers, cookies, authorization, or JavaScript injection until you have defined whose credentials are used, what destinations are allowed, and how secrets are redacted. Each such feature changes the security boundary.

8. Performance, reliability, and cost

Browser startup is expensive relative to reusing an already launched browser, so the example launches Chromium once and creates an isolated context for each job. In production, use a bounded worker pool, monitor worker health, and recycle a browser after crashes or repeated resource growth. Do not share a context between unrelated requests. There are no universal latency or memory numbers: page complexity, scripts, fonts, image sizes, viewport, full-page dimensions, and network conditions all matter.

Set separate navigation, readiness, and screenshot limits plus an overall request deadline. Close contexts in a finally block even when navigation fails. Return stable error categories and a request ID so callers can retry appropriately. Retry transient browser failures only with a small bounded policy; avoid retrying invalid inputs, blocked targets, or deterministic timeouts without changing the request. If an operation is too long for a synchronous HTTP deadline, move it to a job queue and provide polling or a signed completion webhook.

Self-hosting costs include compute and memory for browser workers, container or VM time, network egress, storage for saved outputs, and operational work such as browser updates and abuse controls. Track captures by duration, output bytes, timeout class, and concurrency. Full-page and high-scale captures consume more resources than small viewport captures. Bound queues and reject overload clearly rather than accepting work that cannot complete.

9. Troubleshooting

Symptom Likely cause Fix
Browser executable missing Playwright package installed without its browser binary in the runtime image. Run npx playwright install chromium during image build and deploy the matching dependencies.
Navigation timeout Slow origin, never-ending requests, blocked egress, or an overly strict readiness condition. Use a bounded timeout and a readiness selector or less strict navigation state; inspect DNS and network access.
Screenshot looks incomplete Capture occurred before application rendering, fonts, images, or lazy content finished. Wait for a meaningful selector, image/font readiness where needed, and use bounded scrolling for lazy content.
Blank or transparent output Page has not rendered, screenshot is clipped outside content, or background omission is enabled. Check viewport/clip coordinates, wait for visible content, and disable transparency for an opaque capture.
Base64 decode error or corrupt image Line breaks, wrong padding, URL-safe alphabet, truncated body, or mismatched data-URI media type. Send standard base64, preserve padding, avoid truncation, and use a supported image/* data URI.
413 Payload Too Large Encoded input exceeds Express or proxy body limits. Reduce image size, raise limits carefully at every proxy layer, or upload through a separate controlled storage flow.
502 or browser crash under load Too many simultaneous pages or exceptionally large captures. Apply a semaphore/queue, lower pixel/page limits, observe memory, and recycle unhealthy workers.
Private host reachable Only scheme validation was implemented and SSRF address checks are missing. Block special-use ranges in the network path, validate each redirect and resolved address, and add outbound firewall rules.
Quality option has no effect Lossy quality supplied for PNG or unsupported type/engine behavior. Use JPEG/WebP with an allowed integer quality, or omit quality for PNG.

10. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. One GET request returns a PNG, JPEG, WebP, or PDF. Its capture flow accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each cleanup step can be turned off. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and whether it was billed. AI agents can use its MCP server tools: take_screenshot, get_page_info, and capture_pdf.

Consent overlays and promotional widgets can obscure the content a screenshot service is meant to capture.
Consent overlays and promotional widgets can obscure the content a screenshot service is meant to capture.

For a quick capture, use the documented API call; see the ScreenshotNeo API documentation for request options and setup.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const image = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', image));

Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed; an MCP server lets AI agents take screenshots; and 1,000 screenshots a month are free with no card, with paid plans starting at $5 for 3,000. Sign up free for ScreenshotNeo.

11. Frequently asked questions

Should an API return a data URI or raw base64?

Return raw base64 plus a separate format field when JSON is required. Return binary when the client can consume it. Use a complete data URI only when the immediate consumer expects one.

Can I screenshot an uploaded image without opening its URL?

Yes. Decode and validate the payload, then render it from a controlled document or local route. Do not navigate to arbitrary data supplied by a caller.

Is a full-page screenshot guaranteed to include every lazy-loaded image?

No. Lazy loading is site-specific. Scroll in bounded increments and wait for content when required, then enforce page-height and time limits.

Can one browser page serve every request?

Reuse a launched browser process for efficiency, but create an isolated context for each untrusted request and close it afterward to separate cookies and storage.

When should captures become asynchronous jobs?

Use jobs when page readiness or full-page work can exceed the HTTP request budget. Return a job identifier and define bounded retention and access controls for results.