How to Choose Page Content Formats in a Screenshot API
Choose screenshots, HTML, Markdown, or accessibility trees by what your next system needs, with API examples, trade-offs, and troubleshooting.
The right page format depends on what the next system must do with the result. Choose a screenshot when rendered appearance is the evidence, HTML when markup and document structure matter, Markdown when downstream processing is text oriented, and an accessibility tree when an agent needs semantic roles, labels, and hierarchy.
Before choosing, clarify what “format” means in your API. It may refer to the input (a URL, HTML, or Markdown), the page representation returned (content, Markdown, or an accessibility tree), or the encoding of a rendered image or document (PNG, JPEG, WebP, or PDF). These are separate decisions.
Choose by the consumer
| Representation | Use it when | Main trade-off |
|---|---|---|
| Screenshot | You need visual review, evidence of rendered appearance, visual regression input, or a pixel-level artifact. | An image does not itself provide semantic text structure. |
| HTML | You need markup, attributes, links, or document structure for a parser or transformation. | Your consumer must process HTML and its markup. |
| Markdown | You need a text-oriented representation, including downstream language-model processing. | It is content representation rather than a pixel-accurate visual record. |
| Accessibility tree | An agent needs semantic roles, labels, and hierarchy to interpret or navigate an interface. | It represents interface structure rather than complete visual appearance. |
There is no universal winner. The reviewed documentation describes capabilities and intended uses, not a head-to-head quality, latency, or cost result. Select the smallest representation that lets the next component do its job, and request more than one when you need independent visual and structural evidence.
Separate input format from output format
A screenshot service may accept a URL, raw HTML, or Markdown as input. That does not mean the response will be HTML or Markdown. For example, ScreenshotOne documents URL, HTML, and Markdown as input choices and documents a separate output format option. Its documentation also recommends a POST JSON body for large HTML or Markdown payloads because query strings are smaller. See the ScreenshotOne options documentation for its parameter definitions.
Use this checklist when reading an API reference:
- Source: What are you sending: URL, HTML, or Markdown?
- Representation: Does the service return content, Markdown, an accessibility tree, a screenshot, or several of these?
- Encoding: Is the visual artifact PNG, JPEG, WebP, or PDF? Is it returned as bytes, a URL, or base64?
- Cardinality: Can you request one format, or must you request a combination?
- Metadata: Are title, language, or other fields included?
When a screenshot is the correct format
Request a screenshot when appearance is the thing you need to preserve or inspect: a visual approval, a regression artifact, a rendered invoice, a social-card preview, or evidence that a page looked a particular way at a particular viewport. Screenshots include layout, typography, images, colors, and browser-rendered state in one artifact.
Do not treat a screenshot as a substitute for semantic extraction. OCR can recover some text, but it does not provide the original DOM relationships, link targets, or reliable roles. If a later step must answer “which button is labelled Submit?” or traverse headings, request HTML or an accessibility tree as well.
When HTML is the correct format
Use HTML when your consumer needs markup-oriented structure: extracting links, inspecting attributes, transforming a document, or running a DOM-aware pipeline. HTML keeps elements and nesting that a screenshot cannot expose.
Plan for page-specific complexity. HTML can include scripts, hidden elements, duplicated responsive markup, and presentation details that are irrelevant to a text consumer. If your parser only needs readable content, Markdown may reduce the amount of cleanup.
When Markdown is the correct format
Markdown fits text-oriented workflows such as summarization, indexing, retrieval, and language-model prompts. Cloudflare’s June 11, 2026 changelog describes Markdown as “a token-efficient representation of page content that LLMs can process directly, without parsing HTML markup.” That is the vendor’s description, not an independent benchmark.
Markdown is not a visual record. It can omit layout relationships, exact styling, and some interaction state. Preserve the screenshot too when visual evidence matters. Cloudflare’s API reference says its markdown field may include YAML frontmatter when page metadata is present, so parsers should tolerate frontmatter before the document body.
When an accessibility tree is the correct format
An accessibility tree is useful when an agent must interpret interface structure: roles, accessible names or labels, and parent-child hierarchy. It can be a better input for semantic navigation than raw HTML because it focuses on the interface exposed to assistive technologies.
Use it for tasks such as locating a navigation landmark, identifying a labelled form control, or deciding which button is actionable. It is not a complete representation of visual appearance, and receiving a tree does not by itself prove that a page conforms to every accessibility requirement.
Cloudflare Browser Run snapshot formats
Cloudflare Browser Run’s /snapshot documentation describes a multi-format endpoint. Its formats list accepts content, screenshot, markdown, and accessibilityTree. The current documentation says the endpoint requires at least two formats and defaults to HTML content plus a screenshot. If you need only one representation, use the corresponding single-format endpoint.
The API reference documents response fields for content, markdown, and screenshot; the screenshot is base64 encoded. A minimal request body follows the documented shape. Insert your account’s authentication and endpoint details as required by Cloudflare:
{
"url": "https://example.com",
"formats": ["content", "screenshot"]
}
For a combined visual and text workflow, request two representations and keep them associated with the same capture:
{
"url": "https://example.com/article",
"formats": ["screenshot", "markdown"]
}
For a semantic agent workflow, request the accessibility tree together with the representation needed for verification:
{
"url": "https://example.com/checkout",
"formats": ["accessibilityTree", "screenshot"]
}
Use the Cloudflare API reference for the exact authentication, request, and response envelope. Do not assume that a multi-format endpoint accepts a one-item list when the documentation requires at least two.
Implement a format decision in your application
A small explicit decision function prevents teams from treating every page as an image. The following Python example chooses a representation from the downstream task, then records why that choice was made.
from enum import Enum
class Need(str, Enum):
VISUAL = "visual"
MARKUP = "markup"
TEXT = "text"
SEMANTIC = "semantic"
VISUAL_AND_TEXT = "visual_and_text"
def formats_for(need: Need) -> list[str]:
if need is Need.VISUAL:
return ["screenshot"]
if need is Need.MARKUP:
return ["content"]
if need is Need.TEXT:
return ["markdown"]
if need is Need.SEMANTIC:
return ["accessibilityTree"]
if need is Need.VISUAL_AND_TEXT:
return ["screenshot", "markdown"]
raise ValueError(f"Unsupported need: {need}")
for need in Need:
print(need.value, formats_for(need))
If you are calling Cloudflare’s /snapshot, adjust single-format results to the documented single-format endpoint, or request at least two formats on /snapshot. The decision function describes your application requirement; the provider’s endpoint rules still apply.
cURL: save a screenshot response
curl -X POST "https://api.example.invalid/snapshot" \
-H "Authorization: Bearer YOUR_API_TOKEN" \
-H "Content-Type: application/json" \
--data '{"url":"https://example.com","formats":["content","screenshot"]}' \
-o snapshot.json
Replace the placeholder host and authentication with the provider documented for your account. The response envelope and base64 decoding rules are provider-specific.
Python: decode a base64 screenshot field
import base64
import json
import requests
payload = {
"url": "https://example.com",
"formats": ["content", "screenshot"],
}
response = requests.post(
"https://api.example.invalid/snapshot",
headers={
"Authorization": "Bearer YOUR_API_TOKEN",
"Content-Type": "application/json",
},
json=payload,
timeout=90,
)
response.raise_for_status()
data = response.json()
# Cloudflare documents screenshot as base64 encoded.
with open("page.png", "wb") as image_file:
image_file.write(base64.b64decode(data["screenshot"]))
print(data.get("content", ""))
Node.js: request and decode the response
const response = await fetch('https://api.example.invalid/snapshot', {
method: 'POST',
headers: {
'Authorization': 'Bearer YOUR_API_TOKEN',
'Content-Type': 'application/json'
},
body: JSON.stringify({
url: 'https://example.com',
formats: ['content', 'screenshot']
})
});
if (!response.ok) {
throw new Error(`Snapshot failed: ${response.status} ${await response.text()}`);
}
const data = await response.json();
const image = Buffer.from(data.screenshot, 'base64');
await Bun.write('page.png', image); // In Node.js, use fs.promises.writeFile instead.
console.log(data.content ?? '');
Combining formats without creating conflicting evidence
- Capture once when possible. Request the screenshot and structural representation from the same rendered page state so a later comparison is meaningful.
- Store provenance. Keep the URL, capture time, requested formats, viewport or browser settings, and provider response identifiers with the result.
- Use the screenshot to verify appearance. Use HTML, Markdown, or the accessibility tree to explain what is present and how it is structured.
- Handle disagreement explicitly. A visually hidden element can appear in HTML but not in the visual result. A Markdown converter can omit layout or controls. Treat each representation as evidence for its own question.
Performance, reliability, and cost considerations
The dossier contains no independent benchmark comparing these representations. Do not promise that Markdown is always faster, smaller, or more accurate. Measure the complete workflow you operate: browser render time, response size, parsing time, retries, and downstream processing.
- Response size: Screenshots and base64 encoding can make responses large; stream or persist them rather than logging them.
- Parsing: HTML requires an HTML parser; Markdown requires a Markdown-aware consumer; accessibility trees require code that understands roles and hierarchy.
- Reliability: Treat each requested field as optional until validated. Check HTTP status, required fields, base64 decoding, and content type before handing data to the next step.
- Cost: A multi-format request may be the right operational choice when one render supplies several consumers, but confirm how your provider meters requests and formats. The reviewed sources do not establish a universal pricing rule.
- Caching: Cache representations only when the page state, authentication, and freshness requirements permit it. A cached screenshot can be visually stale even when its HTML is still useful.
Common errors and fixes
| Error | Likely cause | Fix |
|---|---|---|
Validation error for formats |
The endpoint requires at least two formats. | Request two supported formats on /snapshot, or call the provider’s single-format endpoint. |
Expected markdown, received no field |
Markdown was not included in the requested formats, or the page produced no returned Markdown field. | Request markdown explicitly and inspect the documented response envelope. |
| Unreadable screenshot bytes | The screenshot field is base64 encoded or was decoded twice. | Decode exactly once and write binary bytes; do not treat the field as UTF-8 text. |
| Parser breaks on Markdown | YAML frontmatter appears before the Markdown body. | Use a frontmatter-aware parser or remove the frontmatter delimiter and metadata safely. |
| Agent cannot find a control | The selected representation lacks semantic or interaction structure. | Request an accessibility tree, or combine it with a screenshot for visual confirmation. |
| HTML and screenshot appear inconsistent | They were captured at different times or page states. | Capture them together when supported and record the render settings. |
| Large request rejected | Raw HTML or Markdown was placed in a query string. | Use a POST JSON body where the provider recommends it. |
Or skip the browser setup
ScreenshotNeo returns a clean screenshot or PDF from one GET request. It accepts the cookie or consent banner like a visitor, then removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.
It also supports full-page and element captures, dark mode, device presets, custom viewports, retina scale, PDF settings, custom CSS and JavaScript, clicks, waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, caching, signed links, asynchronous jobs, bulk capture, a usage API, and an OpenAPI specification. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
See the ScreenshotNeo API documentation for request options. The same call works from cURL, Python, or Node.js:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Choose the representation that matches your consumer: visual screenshot for appearance, HTML for markup, Markdown for text processing, and an accessibility tree for semantic navigation. If you only need a clean visual artifact, sign up for 1,000 free screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots.
FAQ
Can Markdown replace a screenshot?
No. Markdown represents page content for text processing; it does not preserve pixel-level layout or visual state.
Should I always request every format?
No. Request the formats your next system needs. Additional representations add response handling and storage work.
Is an accessibility tree the same as HTML?
No. An accessibility tree focuses on semantic roles, labels, and hierarchy exposed for interaction; HTML contains markup and attributes that may not map one-to-one to that tree.
Why does a snapshot endpoint require two formats?
Cloudflare’s documented /snapshot endpoint is designed for combined output and currently requires at least two formats. Use a single-format endpoint when only one representation is needed.
Can I use a screenshot API for HTML input?
Some services accept HTML input, but input and output formats are separate settings. Check the provider’s option definitions and use POST for large document payloads when recommended.


