REST APIs for Screenshots, PDFs, and Scraping
Learn how REST APIs render webpages, create PDFs, scrape JavaScript content, and extract document data—with runnable examples and tool-selection guidance.
Short answer: choose a browser-rendering REST API when the page needs JavaScript, CSS, cookies, or a real viewport. Send a URL and rendering options, then receive a PNG, JPEG, WebP, PDF, rendered HTML, or selector-level data. Use a document extraction API when your input is an existing PDF and you need structured text, tables, images, or reading order.
This guide shows the request patterns, code, options, failure modes, and trade-offs for screenshot, PDF, scraping, and PDF-extraction APIs. It also shows when a static HTTP client is sufficient and when you need a full browser.
1. Pick the API type that matches the job
| Job | Input | Best API class | Typical output |
|---|---|---|---|
| Capture what a visitor sees | URL | Browser-rendering screenshot API | PNG, JPEG, WebP |
| Print a rendered page | URL or HTML | Browser-rendering PDF API | |
| Read JavaScript-rendered content | URL | Browser content or scraping API | HTML, selector JSON, extracted fields |
| Extract data from a PDF you already have | Uploaded PDF | PDF extraction API | Structured JSON containing text, tables, images, and document structure |
| Run a multi-step workflow | URL plus session actions | Browser session API | Results after navigation, clicks, login, or downloads |
Cloudflare’s Browser Rendering REST API documents endpoints for screenshots, PDFs, HTML content, snapshots, and scraping. Browserless documents endpoints for screenshots, PDFs, rendered HTML, CSS-selector scraping, smart scraping, downloads, Lighthouse, and website unblocking. Adobe’s PDF Extract API is designed for structured extraction from native and scanned PDFs. See the Cloudflare Browser Rendering documentation, Browserless REST API documentation, and Adobe PDF Services documentation.
2. How a browser-rendering request works
- Your client authenticates with an API key, bearer token, or provider-specific credential.
- The service opens the URL in a managed browser.
- The browser applies viewport, device, locale, timezone, cookies, and headers.
- It waits for a load condition, selector, delay, or network idle state.
- Optional CSS, JavaScript, clicks, blocking rules, and authentication are applied.
- The service returns binary output or JSON, or queues an asynchronous job.
A plain HTTP fetch cannot execute client-side JavaScript or reproduce layout. If a page fills its content after loading, use a browser-rendering endpoint or a provider’s rendered-content endpoint.
3. Screenshot REST API request patterns
Minimal cURL request
curl -G "https://api.example.com/screenshot" \
-H "Authorization: Bearer $API_TOKEN" \
--data-urlencode "url=https://example.com" \
-o page.png
Python with requests
import os
import requests
params = {"url": "https://example.com", "format": "png"}
r = requests.get(
"https://api.example.com/screenshot",
params=params,
headers={"Authorization": f"Bearer {os.environ['API_TOKEN']}"},
timeout=90,
)
r.raise_for_status()
with open("page.png", "wb") as f:
f.write(r.content)
Node.js (built-in fetch)
const params = new URLSearchParams({
url: 'https://example.com',
format: 'png'
});
const res = await fetch(`https://api.example.com/screenshot?${params}`, {
headers: { Authorization: `Bearer ${process.env.API_TOKEN}` }
});
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
const data = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('page.png', data));
Common screenshot options
| Option | Purpose | Implementation notes |
|---|---|---|
| URL or HTML | Select remote page or supplied markup | HTML input avoids a public deployment for generated pages. |
| Viewport width and height | Control responsive breakpoints | Use the same dimensions as the device you want to represent. |
| Full-page | Capture beyond the viewport | Lazy-loaded images may require scrolling or provider support. |
| Device and scale | Emulate a device and retina density | Keep scale consistent when comparing screenshots. |
| Format and quality | Choose PNG, JPEG, or WebP | PNG preserves sharp text; JPEG and WebP usually reduce bytes. |
| Selector | Capture one element | Wait for the selector before capture when it is rendered asynchronously. |
| Delay, selector wait, or network idle | Control readiness | Prefer a deterministic selector over a large fixed delay. |
| CSS and JavaScript | Modify the page before capture | Useful for hiding consent banners or highlighting a region. |
| Headers, cookies, user agent | Render authenticated or localized pages | Do not log secrets in query strings or application logs. |
| Timezone, locale, geolocation | Reproduce regional output | Set all related values when testing a location-specific page. |
| Block rules | Stop ads, trackers, requests, or resource types | Blocking can change layout; verify that required assets still load. |
4. Creating PDFs with a REST API
A browser PDF endpoint prints the rendered page, so CSS, fonts, images, and JavaScript readiness matter. Typical controls include paper size, margins, landscape orientation, headers and footers, and page ranges.
curl -X POST "https://api.example.com/pdf" \
-H "Authorization: Bearer $API_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com/invoice/42",
"paper": "A4",
"landscape": false,
"margin": {"top": "12mm", "right": "12mm", "bottom": "12mm", "left": "12mm"},
"wait_for": "#invoice-ready"
}' \
-o invoice.pdf
For repeatable PDFs, include print CSS, wait for a page-ready marker, embed or reliably load fonts, and specify margins explicitly. A screenshot API is not automatically a PDF extraction API: rendering creates a document, while extraction reads an existing document’s structure.
5. Scraping JavaScript-rendered pages
Use a rendered HTML or scraping endpoint when data appears only after JavaScript executes. Selector scraping is useful for a small, known set of fields; smart scraping or browser scripting is better when the page structure varies.
curl -X POST "https://api.example.com/scrape" \
-H "Authorization: Bearer $API_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"url": "https://example.com/catalog",
"selectors": {
"title": "h1",
"price": ".price",
"items": ".product-card"
},
"wait_for": ".product-card"
}'
Scraping checklist
- Confirm that automated access is permitted by the site’s terms and applicable rules.
- Wait for a stable selector rather than guessing from a network delay.
- Capture pagination state and deduplicate records.
- Record the source URL and capture time with each result.
- Handle missing selectors as data-quality errors, not empty success.
- Protect credentials and personal data passed in cookies or headers.
6. Extracting text and tables from PDFs
When the input is an existing PDF, use a document extraction API. Adobe describes extraction of text, tables, images, and document structure into structured JSON, including native and scanned PDFs.
curl -X POST "https://pdf-services.adobe.io/operation/extractpdf" \
-H "Authorization: Bearer $PDF_TOKEN" \
-H "x-api-key: $PDF_CLIENT_ID" \
-H "Content-Type: application/json" \
-d '{
"elementToExtract": ["text", "tables"]
}'
The exact upload and job-polling sequence depends on the provider. Plan for an asynchronous response when documents are large or scanned. Scanned pages may require OCR, and table extraction should be checked against the original because merged cells, reading order, and handwritten marks can be ambiguous.
7. ScreenshotNeo: a one-call screenshot, PDF, or rendered capture
ScreenshotNeo is a website screenshot API and MCP server. Its GET endpoint accepts a URL and returns PNG, JPEG, WebP, or PDF. It can capture full pages, lazy-loaded images, or one CSS-selected element, and supports dark mode, device presets, custom viewports, retina scale, custom CSS and JavaScript, clicks, waits, blocking rules, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, caching, signed links, asynchronous jobs, webhooks, bulk capture, usage data, and an OpenAPI specification. Parameter names used by other screenshot APIs also work.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
Read the ScreenshotNeo API documentation for the complete parameter list. Before capture, ScreenshotNeo accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed. Response headers identify the page verdict and whether the request was billed.
8. Or skip the browser setup
Use the ScreenshotNeo call above when you want a managed browser without maintaining Playwright or Chromium infrastructure. Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed; and the MCP server lets Claude, Cursor, or another MCP client use take_screenshot, get_page_info, and capture_pdf. The Free plan includes 1,000 screenshots each month with no card, and paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
9. Reliability, performance, and cost
Reliability
- Set a client timeout longer than the provider’s browser timeout and classify timeout responses separately from HTTP errors.
- Retry transient 5xx and network failures with exponential backoff and a maximum attempt count.
- Do not blindly retry authentication failures, invalid URLs, or selector-not-found errors.
- For asynchronous jobs, persist the job ID and make webhook handling idempotent.
- Store response metadata such as status, verdict, billed state, URL, and capture options.
Performance
- Use a smaller viewport or element capture when a full page is unnecessary.
- Block analytics, ads, and unused resource types when they do not affect the result.
- Reuse cache entries for unchanged pages and choose a TTL that matches freshness needs.
- Batch independent URLs when the provider supports bulk capture.
- Prefer network-idle or selector readiness over excessive fixed delays.
Cost
Cost depends on the provider’s billing unit, browser time, output size, concurrency, and document limits. Compare those rules directly; the research sources do not provide an independent cross-vendor benchmark for latency, accuracy, anti-bot success, or total cost. ScreenshotNeo bills only clean shots and exposes billing status in response headers. Its plans are Free (1,000/month), Starter ($5 for 3,000), Growth ($15 for 15,000), Pro ($39 for 60,000), Scale ($99 for 250,000), and Business ($249 for 1,000,000); yearly billing gives two months free and every feature is included on every plan.
10. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Blank screenshot | JavaScript has not finished or navigation failed | Wait for a stable selector, inspect the page verdict, and retry transient failures. |
| Consent banner covers content | Cookie platform was not handled | Use a provider’s consent handling or click/hide the banner before capture. |
| Missing lazy images | Images load only after scrolling | Use full-page capture with lazy-image loading or scroll via browser scripting. |
| Wrong mobile layout | Viewport or device emulation differs | Set viewport width, height, device scale, user agent, and orientation together. |
| Selector not found | Selector is incorrect or content is conditional | Inspect the rendered DOM, wait longer, and handle the absent field explicitly. |
| PDF has clipped content | Margins, print CSS, or page breaks are unsuitable | Set paper and margins, add print styles, and test long tables. |
| 401 or 403 | Missing credentials or protected target | Check API authentication, then provide permitted target headers or cookies. |
| Rate-limit response | Concurrency exceeds the provider limit | Throttle requests, honor retry headers, and queue asynchronous jobs. |
| Unexpected billing | Cache, verdict, or retry behavior was not recorded | Log billing headers and deduplicate retries with request IDs where supported. |
11. Security and data handling
- Keep API keys in environment variables or a secret manager.
- Redact authorization headers, cookies, signed URLs, and personal data from logs.
- Review retention, regional processing, encryption, upload limits, and deletion behavior before sending sensitive pages or PDFs.
- Use allowlists for URLs supplied by users to reduce server-side request forgery risk.
- Respect robots directives, terms, authentication boundaries, and applicable law.
12. FAQ
Can a screenshot API scrape a page?
Some providers expose rendered HTML or selector-scraping endpoints in addition to image output. A screenshot alone is pixels, not structured data.
Should I use a PDF API or a screenshot API?
Use a PDF endpoint when you need selectable, printable pages. Use an image endpoint for visual previews, thumbnails, or pixel comparison.
How do I handle a scanned PDF?
Choose a PDF extraction service that supports scanned documents and OCR, then validate text and tables against the source pages.
Is a browser required for every URL?
No. Static HTML can be fetched with a normal HTTP client, but JavaScript-rendered layouts require a browser or rendered-content service.
When should a job be asynchronous?
Use asynchronous jobs for large PDFs, many URLs, long pages, or workflows where webhook delivery and retries are easier than holding an HTTP connection open.


