APIs for Extracting Markdown, HTML, Text, and Proxy Data
Choose the right extraction output, JavaScript rendering, proxy mode, and structured data API for reliable web pipelines.
Web-extraction APIs fetch a URL or supplied markup and return the representation your application needs: Markdown, source HTML, plain text, structured JSON, or a proxied response. Choose the output first, then decide whether browser rendering, proxy routing, selectors, or automatic page classification are required.
Quick decision guide
| Need | Best starting output or feature | Why |
|---|---|---|
| LLM, RAG, or search indexing | Markdown | Preserves headings and links while removing most presentation markup. |
| Custom parser or archival fidelity | Raw HTML | Retains the source markup your parser expects. |
| Lightweight text processing | Plain text | Removes tags and reduces downstream parsing work. |
| Known fields such as article title and body | Structured JSON | Reduces selector and parser maintenance when the provider supports the page type. |
| Client-rendered content | Browser or JavaScript rendering | Executes page JavaScript before extraction. |
| Access from a different network location | Proxy mode | Changes the access route; it does not itself determine the output format. |
How the extraction pipeline works
- Submit a URL or markup. Most services accept a URL; some also accept an HTTP response body, browser HTML, or caller-supplied HTML/plain text.
- Choose the access path. Use a direct HTTP fetch for static pages. Add browser rendering when content appears only after JavaScript runs. Use proxy routing when geography, network access, or rate limits require it.
- Select the representation. Request Markdown, source HTML, plain text, or a structured schema.
- Validate the result. Check status, content type, extracted length, title, and required fields before storing or sending the result to an LLM.
- Apply policy controls. Respect robots directives, site terms, authentication rules, rate limits, and applicable law.
Markdown extraction
Markdown is usually the most convenient format for an LLM or RAG pipeline. It keeps semantic structure such as headings, lists, and links while discarding much of the navigation and presentation markup. Firecrawl positions its Scrape product as turning any URL into clean Markdown or structured data for AI agents. ScrapingBee documents a return_page_markdown option that returns the main page content with HTML tags and unnecessary information stripped.
When Markdown is a good fit
- Chunking documentation or articles for retrieval.
- Sending page content to an LLM with fewer tokens than raw HTML.
- Preserving headings and links for citations or navigation.
Markdown edge cases
- Tables, embedded media, code blocks, and unusual widgets may be represented differently by each provider.
- Repeated navigation can still appear on sites with weak article detection.
- JavaScript-rendered text is absent unless rendering is enabled.
Raw HTML extraction
Request raw or source HTML when your own parser, sanitizer, or archival process needs the original markup. ScrapingBee documents return_page_source; Zyte distinguishes an HTTP response body from browser HTML and caller-supplied HTML. Raw HTML is larger and more variable than Markdown, so store the response content type and encoding with the payload.
Plain-text extraction
Plain text is useful for simple search, language detection, or low-overhead preprocessing. It removes tags but also removes structural signals. Keep the source URL and extraction timestamp beside the text so downstream systems can trace it back to the page.
Structured JSON extraction
Structured extraction is useful when you need fields rather than a document. Diffbot says its Extract product uses computer vision and natural language processing to read a page and return clean, structured JSON; its Article extractor targets news articles, blog posts, and other text-heavy pages, including clean body text. Automatic page classification can reduce parser maintenance, but you should still validate required fields because page layouts and classifications vary.
Use structured extraction when
- Your records have a stable schema such as
title,author,published_at, andbody. - You want a provider to identify page types instead of maintaining selectors for every site.
- You can reject or quarantine responses missing required fields.
JavaScript rendering and browser HTML
A plain HTTP request sees the server response. A browser-rendered request executes JavaScript and can expose content inserted after load. ScrapingBee documents JavaScript rendering; Zyte distinguishes browserHtml from httpResponseBody and notes that browser HTML typically improves quality when rendering is needed. Firecrawl explicitly markets coverage for JavaScript-heavy sites.
Enable rendering only for targets that need it. It generally adds browser startup and page-load work, and it can trigger consent dialogs, bot checks, or login flows that a static request never encounters.
Diagnosing a JavaScript-dependent page
- Fetch the URL with a normal HTTP client.
- Search the response for the text or data you expect.
- If it is missing but visible in a browser, retry with browser rendering.
- Set a bounded wait condition and validate that the expected selector or field exists.
Proxy mode: what it does and does not do
Proxy support changes how a request reaches the target. ScrapingBee and Zyte document proxy modes separately from their content extraction options. Evaluate proxy geography, authentication, rotation, rate limits, and site permissions independently from whether the response is Markdown, HTML, text, or JSON.
Zyte documents extraction at https://api.zyte.com/v1/extract and proxy use through https://api.zyte.com:8011. The extraction endpoint and proxy endpoint represent different access paths; select the one that matches your operational need.
Provider capabilities at a glance
| Provider | Documented strengths | Check before production |
|---|---|---|
| Firecrawl | Markdown or structured data for AI agents; coverage positioned for JavaScript-heavy, gated, and region-specific sites. | Current schema, rendering controls, limits, and pricing. |
| ScrapingBee | Markdown, text, source HTML, JavaScript rendering, premium proxies, CSS/XPath rules, AI extraction, and proxy front end. | Which rendering, proxy, and extraction options apply to your plan. |
| Zyte API | Extraction from HTTP response body, browser HTML, or user HTML; separate proxy endpoint. | Source selection, browser requirements, and current request schema. |
| Diffbot Extract | Automatic page classification and structured JSON; accepts supplied HTML or plain text when the caller can access markup. | Page-type coverage and field validation for your URLs. |
There is no common benchmark in the reviewed official documentation. Do not publish an accuracy, latency, or cost ranking without testing representative URLs with the same requirements.
DIY baseline: fetch and convert a page locally
A local baseline helps you determine whether a hosted API is necessary. The example below fetches server-rendered HTML and converts it to Markdown. It does not execute JavaScript, rotate proxies, or bypass access controls.
Python
pip install requests beautifulsoup4 markdownify
import requests
from bs4 import BeautifulSoup
from markdownify import markdownify
url = "https://example.com"
r = requests.get(url, timeout=30, headers={"User-Agent": "extractor/1.0"})
r.raise_for_status()
soup = BeautifulSoup(r.text, "html.parser")
for node in soup(["script", "style", "noscript"]):
node.decompose()
markdown = markdownify(str(soup), heading_style="ATX")
print(markdown[:5000])
cURL
curl --fail --location --max-time 30 \
-H 'User-Agent: extractor/1.0' \
'https://example.com' \
-o page.html
Node.js
const res = await fetch('https://example.com', {
headers: { 'User-Agent': 'extractor/1.0' }
});
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
const html = await res.text();
console.log(html.slice(0, 5000));
For production extraction, add retries with backoff, response-size limits, content-type checks, deduplication, and a dead-letter path for pages that fail validation.
Reliability, performance, and cost checklist
- Timeouts: Use separate connect and total timeouts where your client supports them.
- Retries: Retry transient network and 5xx failures with exponential backoff; avoid retrying deterministic 4xx errors.
- Idempotency: Key stored results by canonical URL plus extraction options and content version.
- Rendering budget: Route only JavaScript-dependent pages through a browser.
- Payload size: Compress or store raw HTML separately from normalized Markdown or JSON.
- Rate limits: Use a queue and per-host concurrency limits.
- Validation: Require minimum content length and required JSON fields.
- Cost: Compare request pricing, browser-rendering charges, proxy charges, storage, and retries using your real URL mix.
- Compliance: Confirm collection is allowed by the site, your agreements, and applicable law.
Common errors and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Empty article or shell HTML | Content is inserted by JavaScript. | Enable browser rendering and wait for a meaningful selector or field. |
| Markdown contains menus and cookie text | The provider could not identify the main content. | Use CSS/XPath rules, an article extractor, or post-process boilerplate. |
| 403 or repeated bot challenge | The site blocks the access path. | Check permission and terms, slow the request rate, and evaluate an allowed proxy or browser path. |
| Structured fields are missing | Page type does not match the selected extractor. | Validate the schema, choose a suitable extractor, or fall back to Markdown/HTML. |
| Requests time out | Slow origin, heavy assets, or browser startup. | Set bounded timeouts, reduce concurrency, and avoid rendering when it is unnecessary. |
| Wrong language or region | Origin varies by IP, headers, or cookies. | Set the required locale controls where supported and record them with the result. |
Or skip the browser setup
When your deliverable is a visual capture rather than extracted content, ScreenshotNeo provides a single screenshot API request for PNG, JPEG, WebP, or PDF output. Its clean-shot flow accepts consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.
See the ScreenshotNeo API documentation for the complete option list.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Features include full-page capture with lazy images loaded, CSS-element capture, dark mode, device presets and custom viewports, retina scale, PDF paper and page controls, custom CSS and JavaScript, clicks, waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, selectable cache TTLs, signed links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification. Parameter names used by other screenshot APIs also work, which helps when switching.
The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free.
Create a free ScreenshotNeo account and get 1,000 screenshots each month with no card.
FAQ
Should I store Markdown or HTML?
Store Markdown for retrieval and language-model workflows. Keep raw HTML when you need to re-parse, audit, or preserve markup.
Does a proxy make a page JavaScript-rendered?
No. Proxy routing changes the network path. Browser rendering is a separate capability.
When should I use structured JSON?
Use it when the provider’s page classifier and schema match your records and you can validate required fields.
How can I compare providers fairly?
Build a representative URL set, request the same output and rendering mode, record failures and field completeness, and calculate cost from current plans.
Can ScreenshotNeo extract Markdown?
ScreenshotNeo is designed for screenshots and PDFs. Use a content-extraction API when you need Markdown, HTML, text, or structured JSON; use ScreenshotNeo when the required result is a clean visual capture.


