ScreenshotNeo

BlogGuides

REST APIs for Screenshots, PDFs, and Scraping

Learn how REST APIs render webpages, create PDFs, scrape JavaScript content, and extract document data—with runnable examples and tool-selection guidance.

By the ScreenshotNeo team1 October 20268 min read

Short answer: choose a browser-rendering REST API when the page needs JavaScript, CSS, cookies, or a real viewport. Send a URL and rendering options, then receive a PNG, JPEG, WebP, PDF, rendered HTML, or selector-level data. Use a document extraction API when your input is an existing PDF and you need structured text, tables, images, or reading order.

This guide shows the request patterns, code, options, failure modes, and trade-offs for screenshot, PDF, scraping, and PDF-extraction APIs. It also shows when a static HTTP client is sufficient and when you need a full browser.

1. Pick the API type that matches the job

Job Input Best API class Typical output
Capture what a visitor sees URL Browser-rendering screenshot API PNG, JPEG, WebP
Print a rendered page URL or HTML Browser-rendering PDF API PDF
Read JavaScript-rendered content URL Browser content or scraping API HTML, selector JSON, extracted fields
Extract data from a PDF you already have Uploaded PDF PDF extraction API Structured JSON containing text, tables, images, and document structure
Run a multi-step workflow URL plus session actions Browser session API Results after navigation, clicks, login, or downloads

Cloudflare’s Browser Rendering REST API documents endpoints for screenshots, PDFs, HTML content, snapshots, and scraping. Browserless documents endpoints for screenshots, PDFs, rendered HTML, CSS-selector scraping, smart scraping, downloads, Lighthouse, and website unblocking. Adobe’s PDF Extract API is designed for structured extraction from native and scanned PDFs. See the Cloudflare Browser Rendering documentation, Browserless REST API documentation, and Adobe PDF Services documentation.

2. How a browser-rendering request works

  1. Your client authenticates with an API key, bearer token, or provider-specific credential.
  2. The service opens the URL in a managed browser.
  3. The browser applies viewport, device, locale, timezone, cookies, and headers.
  4. It waits for a load condition, selector, delay, or network idle state.
  5. Optional CSS, JavaScript, clicks, blocking rules, and authentication are applied.
  6. The service returns binary output or JSON, or queues an asynchronous job.

A plain HTTP fetch cannot execute client-side JavaScript or reproduce layout. If a page fills its content after loading, use a browser-rendering endpoint or a provider’s rendered-content endpoint.

3. Screenshot REST API request patterns

Minimal cURL request

curl -G "https://api.example.com/screenshot" \
  -H "Authorization: Bearer $API_TOKEN" \
  --data-urlencode "url=https://example.com" \
  -o page.png

Python with requests

import os
import requests

params = {"url": "https://example.com", "format": "png"}
r = requests.get(
    "https://api.example.com/screenshot",
    params=params,
    headers={"Authorization": f"Bearer {os.environ['API_TOKEN']}"},
    timeout=90,
)
r.raise_for_status()
with open("page.png", "wb") as f:
    f.write(r.content)

Node.js (built-in fetch)

const params = new URLSearchParams({
  url: 'https://example.com',
  format: 'png'
});

const res = await fetch(`https://api.example.com/screenshot?${params}`, {
  headers: { Authorization: `Bearer ${process.env.API_TOKEN}` }
});
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
const data = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('page.png', data));

Common screenshot options

Option Purpose Implementation notes
URL or HTML Select remote page or supplied markup HTML input avoids a public deployment for generated pages.
Viewport width and height Control responsive breakpoints Use the same dimensions as the device you want to represent.
Full-page Capture beyond the viewport Lazy-loaded images may require scrolling or provider support.
Device and scale Emulate a device and retina density Keep scale consistent when comparing screenshots.
Format and quality Choose PNG, JPEG, or WebP PNG preserves sharp text; JPEG and WebP usually reduce bytes.
Selector Capture one element Wait for the selector before capture when it is rendered asynchronously.
Delay, selector wait, or network idle Control readiness Prefer a deterministic selector over a large fixed delay.
CSS and JavaScript Modify the page before capture Useful for hiding consent banners or highlighting a region.
Headers, cookies, user agent Render authenticated or localized pages Do not log secrets in query strings or application logs.
Timezone, locale, geolocation Reproduce regional output Set all related values when testing a location-specific page.
Block rules Stop ads, trackers, requests, or resource types Blocking can change layout; verify that required assets still load.

4. Creating PDFs with a REST API

A browser PDF endpoint prints the rendered page, so CSS, fonts, images, and JavaScript readiness matter. Typical controls include paper size, margins, landscape orientation, headers and footers, and page ranges.

curl -X POST "https://api.example.com/pdf" \
  -H "Authorization: Bearer $API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/invoice/42",
    "paper": "A4",
    "landscape": false,
    "margin": {"top": "12mm", "right": "12mm", "bottom": "12mm", "left": "12mm"},
    "wait_for": "#invoice-ready"
  }' \
  -o invoice.pdf

For repeatable PDFs, include print CSS, wait for a page-ready marker, embed or reliably load fonts, and specify margins explicitly. A screenshot API is not automatically a PDF extraction API: rendering creates a document, while extraction reads an existing document’s structure.

5. Scraping JavaScript-rendered pages

Use a rendered HTML or scraping endpoint when data appears only after JavaScript executes. Selector scraping is useful for a small, known set of fields; smart scraping or browser scripting is better when the page structure varies.

curl -X POST "https://api.example.com/scrape" \
  -H "Authorization: Bearer $API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/catalog",
    "selectors": {
      "title": "h1",
      "price": ".price",
      "items": ".product-card"
    },
    "wait_for": ".product-card"
  }'

Scraping checklist

  • Confirm that automated access is permitted by the site’s terms and applicable rules.
  • Wait for a stable selector rather than guessing from a network delay.
  • Capture pagination state and deduplicate records.
  • Record the source URL and capture time with each result.
  • Handle missing selectors as data-quality errors, not empty success.
  • Protect credentials and personal data passed in cookies or headers.

6. Extracting text and tables from PDFs

When the input is an existing PDF, use a document extraction API. Adobe describes extraction of text, tables, images, and document structure into structured JSON, including native and scanned PDFs.

curl -X POST "https://pdf-services.adobe.io/operation/extractpdf" \
  -H "Authorization: Bearer $PDF_TOKEN" \
  -H "x-api-key: $PDF_CLIENT_ID" \
  -H "Content-Type: application/json" \
  -d '{
    "elementToExtract": ["text", "tables"]
  }'

The exact upload and job-polling sequence depends on the provider. Plan for an asynchronous response when documents are large or scanned. Scanned pages may require OCR, and table extraction should be checked against the original because merged cells, reading order, and handwritten marks can be ambiguous.

7. ScreenshotNeo: a one-call screenshot, PDF, or rendered capture

ScreenshotNeo is a website screenshot API and MCP server. Its GET endpoint accepts a URL and returns PNG, JPEG, WebP, or PDF. It can capture full pages, lazy-loaded images, or one CSS-selected element, and supports dark mode, device presets, custom viewports, retina scale, custom CSS and JavaScript, clicks, waits, blocking rules, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, caching, signed links, asynchronous jobs, webhooks, bulk capture, usage data, and an OpenAPI specification. Parameter names used by other screenshot APIs also work.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

Read the ScreenshotNeo API documentation for the complete parameter list. Before capture, ScreenshotNeo accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed. Response headers identify the page verdict and whether the request was billed.

8. Or skip the browser setup

Use the ScreenshotNeo call above when you want a managed browser without maintaining Playwright or Chromium infrastructure. Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed; and the MCP server lets Claude, Cursor, or another MCP client use take_screenshot, get_page_info, and capture_pdf. The Free plan includes 1,000 screenshots each month with no card, and paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

9. Reliability, performance, and cost

Reliability

  • Set a client timeout longer than the provider’s browser timeout and classify timeout responses separately from HTTP errors.
  • Retry transient 5xx and network failures with exponential backoff and a maximum attempt count.
  • Do not blindly retry authentication failures, invalid URLs, or selector-not-found errors.
  • For asynchronous jobs, persist the job ID and make webhook handling idempotent.
  • Store response metadata such as status, verdict, billed state, URL, and capture options.

Performance

  • Use a smaller viewport or element capture when a full page is unnecessary.
  • Block analytics, ads, and unused resource types when they do not affect the result.
  • Reuse cache entries for unchanged pages and choose a TTL that matches freshness needs.
  • Batch independent URLs when the provider supports bulk capture.
  • Prefer network-idle or selector readiness over excessive fixed delays.

Cost

Cost depends on the provider’s billing unit, browser time, output size, concurrency, and document limits. Compare those rules directly; the research sources do not provide an independent cross-vendor benchmark for latency, accuracy, anti-bot success, or total cost. ScreenshotNeo bills only clean shots and exposes billing status in response headers. Its plans are Free (1,000/month), Starter ($5 for 3,000), Growth ($15 for 15,000), Pro ($39 for 60,000), Scale ($99 for 250,000), and Business ($249 for 1,000,000); yearly billing gives two months free and every feature is included on every plan.

10. Troubleshooting

Symptom Likely cause Fix
Blank screenshot JavaScript has not finished or navigation failed Wait for a stable selector, inspect the page verdict, and retry transient failures.
Consent banner covers content Cookie platform was not handled Use a provider’s consent handling or click/hide the banner before capture.
Missing lazy images Images load only after scrolling Use full-page capture with lazy-image loading or scroll via browser scripting.
Wrong mobile layout Viewport or device emulation differs Set viewport width, height, device scale, user agent, and orientation together.
Selector not found Selector is incorrect or content is conditional Inspect the rendered DOM, wait longer, and handle the absent field explicitly.
PDF has clipped content Margins, print CSS, or page breaks are unsuitable Set paper and margins, add print styles, and test long tables.
401 or 403 Missing credentials or protected target Check API authentication, then provide permitted target headers or cookies.
Rate-limit response Concurrency exceeds the provider limit Throttle requests, honor retry headers, and queue asynchronous jobs.
Unexpected billing Cache, verdict, or retry behavior was not recorded Log billing headers and deduplicate retries with request IDs where supported.

11. Security and data handling

  • Keep API keys in environment variables or a secret manager.
  • Redact authorization headers, cookies, signed URLs, and personal data from logs.
  • Review retention, regional processing, encryption, upload limits, and deletion behavior before sending sensitive pages or PDFs.
  • Use allowlists for URLs supplied by users to reduce server-side request forgery risk.
  • Respect robots directives, terms, authentication boundaries, and applicable law.

12. FAQ

Can a screenshot API scrape a page?

Some providers expose rendered HTML or selector-scraping endpoints in addition to image output. A screenshot alone is pixels, not structured data.

Should I use a PDF API or a screenshot API?

Use a PDF endpoint when you need selectable, printable pages. Use an image endpoint for visual previews, thumbnails, or pixel comparison.

How do I handle a scanned PDF?

Choose a PDF extraction service that supports scanned documents and OCR, then validate text and tables against the source pages.

Is a browser required for every URL?

No. Static HTML can be fetched with a normal HTTP client, but JavaScript-rendered layouts require a browser or rendered-content service.

When should a job be asynchronous?

Use asynchronous jobs for large PDFs, many URLs, long pages, or workflows where webhook delivery and retries are easier than holding an HTTP connection open.