How to Scrape Cloudflare-Protected Websites with an API
Learn what Cloudflare’s APIs can render or crawl, why they cannot bypass bot checks, and how to build an authorized extraction workflow.

Short answer: an API can render, crawl, or extract content from a Cloudflare-protected website only when the site permits that access. Cloudflare’s Browser Rendering API provides separate workflows for crawling pages, fetching JavaScript-rendered HTML, and extracting selected elements. It identifies itself as a bot and does not bypass Cloudflare bot detection or CAPTCHAs. For a reliable integration, confirm permission, check robots.txt and Content Signals, choose the smallest workflow that returns the data you need, and pace requests responsibly.
This guide shows how to select and call Cloudflare’s /crawl, /content, and /scrape workflows, handle asynchronous jobs, extract structured fields, and troubleshoot failures. It also explains when a screenshot API is the better fit.
What “Cloudflare-protected” means for an API
Cloudflare can protect a site with a web application firewall, rate limits, bot detection, and CAPTCHA challenges. A rendering API does not automatically grant access through those controls. Cloudflare’s own changelog states: “the /crawl endpoint cannot bypass Cloudflare bot detection or captchas, and self-identifies as a bot.” Cloudflare’s March 10, 2026 Browser Rendering announcement documents that boundary.
Therefore, treat scraping as an authorized data-access project:
- Use a documented API supplied by the site when one exists.
- Obtain written permission for crawling or extraction when you operate outside your own properties.
- Review
robots.txt, Content Signals, terms, and any published crawl policy. - Honor rate limits and crawl delays. Do not try to evade controls by rotating user-agent strings.
- Store only the content and personal data your use case requires.
Choose the right Cloudflare Browser Rendering workflow
| Need | Workflow | What it does |
|---|---|---|
| Discover and process many pages | /crawl |
Starts an asynchronous crawl, follows permitted links, and returns HTML, Markdown, or JSON results. |
| One JavaScript-heavy page | /content |
Executes JavaScript and returns the page’s rendered HTML. |
| Specific fields or repeated elements | /scrape |
Uses CSS selectors to extract headings, links, prices, metadata, or other selected elements. |
| Static HTML only | /crawl with rendering disabled |
Cloudflare documents render: false for static content, avoiding browser time. |
Make the choice using five questions: Do you have permission? Is the content static or rendered by JavaScript? Do you need one page or link discovery? Which output format and schema do you need? What page, browser-time, and pacing limits apply?

Prerequisites and access
- Create a Cloudflare API token with the Browser Rendering permissions required by the workflow. Cloudflare documents that
/contentuses a REST API token with Browser Rendering Edit permission, or Workers Bindings. - Set your account identifier and token as environment variables. Keep tokens on the server; never place them in browser-side JavaScript.
- Read the target site’s robots policy and Content Signals. A crawl job can be rejected when its declared purpose or content-use level is disallowed.
- Define a narrow page limit, depth, selector set, and request schedule before sending traffic.
export CF_ACCOUNT_ID="your-account-id"
export CF_API_TOKEN="your-api-token"
# Set this to the API base shown in Cloudflare's Browser Rendering documentation.
export CF_API_BASE="https://api.cloudflare.com/client/v4"
Run an authorized crawl with /crawl
/crawl is asynchronous: submit a crawl request, receive a job ID, then retrieve the results. It supports HTML, Markdown, and JSON output. Cloudflare applies per-domain rate limiting, respects a site’s crawl-delay, and otherwise documents a default 0.5-second delay between requests to the same domain when no delay is specified. These are behaviors of this endpoint, not a universal rule for other crawlers. See the crawl endpoint documentation for the current request schema and limits.
Submit the job with cURL
curl -X POST "$CF_API_BASE/accounts/$CF_ACCOUNT_ID/browser-rendering/crawl" \
-H "Authorization: Bearer $CF_API_TOKEN" \
-H "Content-Type: application/json" \
--data '{
"url": "https://example.com",
"render": true,
"limit": 25,
"depth": 2,
"formats": ["markdown", "json"]
}'
Use the exact field names and output options supported by your account’s current documentation. The response contains a job identifier; save it with the target URL, policy decision, and your own correlation ID.
Submit and poll in Python
import os
import time
import requests
base = os.environ["CF_API_BASE"]
account = os.environ["CF_ACCOUNT_ID"]
headers = {
"Authorization": f"Bearer {os.environ['CF_API_TOKEN']}",
"Content-Type": "application/json",
}
payload = {
"url": "https://example.com",
"render": True,
"limit": 25,
"depth": 2,
"formats": ["markdown", "json"],
}
start = requests.post(
f"{base}/accounts/{account}/browser-rendering/crawl",
headers=headers,
json=payload,
timeout=30,
)
start.raise_for_status()
job = start.json()
job_id = job["result"]["id"]
while True:
status = requests.get(
f"{base}/accounts/{account}/browser-rendering/crawl/{job_id}",
headers=headers,
timeout=30,
)
status.raise_for_status()
result = status.json()["result"]
state = result.get("status")
if state in {"complete", "completed", "failed"}:
print(result)
break
time.sleep(2)
Submit and poll in Node.js
const base = process.env.CF_API_BASE;
const account = process.env.CF_ACCOUNT_ID;
const headers = {
Authorization: `Bearer ${process.env.CF_API_TOKEN}`,
'Content-Type': 'application/json'
};
const start = await fetch(
`${base}/accounts/${account}/browser-rendering/crawl`,
{
method: 'POST',
headers,
body: JSON.stringify({
url: 'https://example.com',
render: true,
limit: 25,
depth: 2,
formats: ['markdown', 'json']
})
}
);
if (!start.ok) throw new Error(`submit failed: ${start.status}`);
const job = await start.json();
const id = job.result.id;
for (;;) {
const response = await fetch(
`${base}/accounts/${account}/browser-rendering/crawl/${id}`,
{ headers }
);
if (!response.ok) throw new Error(`poll failed: ${response.status}`);
const result = (await response.json()).result;
if (['complete', 'completed', 'failed'].includes(result.status)) {
console.log(JSON.stringify(result, null, 2));
break;
}
await new Promise(resolve => setTimeout(resolve, 2000));
}
Fetch one rendered page with /content
Use /content when you need the final HTML of one page after its JavaScript has run. This is useful for client-rendered product details, documentation, or dashboards where the initial response contains little usable markup. It is not a bypass mechanism: a bot check or CAPTCHA can still prevent the page from being returned. Read Cloudflare’s content endpoint documentation for the current request and response fields.
curl -X POST "$CF_API_BASE/accounts/$CF_ACCOUNT_ID/browser-rendering/content" \
-H "Authorization: Bearer $CF_API_TOKEN" \
-H "Content-Type: application/json" \
--data '{"url":"https://example.com/app"}'
Save the returned HTML, record the timestamp and URL, and parse it with an HTML parser. Avoid regex for nested markup. If the page is static, use a non-rendering crawl where appropriate to reduce browser use.
Extract selected fields with /scrape
/scrape is appropriate when you need a stable set of elements instead of an entire document. Define selectors for each field, such as a title, price, metadata value, or repeated link. Keep selectors specific enough to avoid navigation and boilerplate. Changing the user-agent parameter does not bypass protection.
curl -X POST "$CF_API_BASE/accounts/$CF_ACCOUNT_ID/browser-rendering/scrape" \
-H "Authorization: Bearer $CF_API_TOKEN" \
-H "Content-Type: application/json" \
--data '{
"url": "https://example.com/catalog/item-1",
"selectors": {
"title": "h1",
"price": "[data-price]",
"links": "main a"
}
}'
Validate the response schema before writing it to a database. A selector can match zero, one, or many nodes as the site changes. Treat an empty value as a data-quality event and alert on it instead of silently storing a false value.
Respect robots.txt, Content Signals, and rate limits
Cloudflare’s crawl service respects robots.txt and crawl-delay. Its documentation also describes Content Signals that can reject a job when the declared crawl purpose or content-use level is disallowed. Check these policies before scheduling a job, and retain the policy decision with your run record.
Site owners can use Cloudflare WAF rate-limiting rules to constrain scraping-style patterns, such as repeated price lookups or requests keyed by a product ID. The rate-limiting guidance is defensive configuration advice. Do not turn it into an evasion plan.
Reliability and performance practices
- Bound work: set page limits and depth for every crawl. A link graph can grow far beyond the pages you intended.
- Use asynchronous processing: queue crawl jobs and poll with backoff rather than holding a request open.
- Make jobs restartable: persist the job ID, request payload, and last successful result. Retry only transient failures and avoid duplicate writes with an idempotency key in your own system.
- Separate browser and parser time: rendering consumes browser time; parse and normalize results outside the browser where possible.
- Cache responsibly: store results for a defined freshness window when the use case permits, and honor deletion or update requests.
- Measure quality: track empty selector matches, unexpected status transitions, document size, and policy rejections.
- Stay within budgets: Cloudflare documents browser-time limits, including a Workers Free allowance of 10 minutes of browser use per day. Check the current limits before choosing crawl size.
Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| 401 or 403 | Missing, expired, or under-permissioned token | Create or update a token with the required Browser Rendering permission and keep it server-side. |
| Crawl job rejected immediately | robots.txt or Content Signals disallow the declared purpose | Stop the job, review the site policy, and obtain permission or change the use case. |
| CAPTCHA or bot-check HTML | The site challenged the request | Do not attempt to evade it. Use the site’s documented API or request authorized access. |
| Empty scrape fields | Selector mismatch, changed markup, or content loaded later | Inspect rendered HTML, update selectors, and add a data-quality alert. |
| Static pages consume too much browser time | Rendering was enabled unnecessarily | Use the documented non-rendering option for static content. |
| Slow or incomplete crawl | Large depth or limit, crawl-delay, per-domain rate limiting, or browser-time budget | Reduce scope, split jobs, honor delays, and monitor completion status. |
| Duplicate records after retry | Client retried after a timeout without deduplication | Persist job IDs and deduplicate by canonical URL plus crawl version. |

Or skip the browser setup
If your deliverable is a visual capture rather than HTML or structured fields, ScreenshotNeo provides a single website screenshot API call. It removes cookie and consent banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Responses identify the page verdict and billing state with X-Page-Verdict and X-Billed headers. ScreenshotNeo also offers an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
See the ScreenshotNeo API documentation for all options. This request returns a WebP image:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo supports full-page and element captures, dark mode, device presets, custom viewports, retina scale, PDFs, custom CSS and JavaScript, clicks, waits, blocked resources, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, caching, signed links, asynchronous webhooks, bulk capture, and a usage API. Every feature is on every plan. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots.
Start with 1,000 free ScreenshotNeo screenshots per month—no card required.
FAQ
Can an API scrape a Cloudflare-protected site without permission?
No. Rendering an accessible page is different from bypassing an access control. Use a documented API or obtain permission.
Does changing the user agent defeat Cloudflare bot detection?
No. Cloudflare’s documentation does not present user-agent changes as a bypass, and the crawl endpoint identifies itself as a bot.
When should I use /content instead of /scrape?
Use /content when you need the complete JavaScript-rendered HTML. Use /scrape when a small, defined set of CSS-selected fields is enough.
Is /crawl synchronous?
No. Submit the crawl, retain its job ID, and retrieve results after processing completes.
Can ScreenshotNeo return HTML instead of an image?
ScreenshotNeo is a screenshot and PDF API. Use Cloudflare’s rendering workflows when your output is HTML or structured page data.
Operational checklist
- Permission and site policy reviewed.
- robots.txt, crawl-delay, and Content Signals checked.
- Workflow selected for static, rendered, extracted, or multi-page content.
- Page limit, depth, selectors, and pacing configured.
- Asynchronous job IDs and retries are persisted.
- CAPTCHAs and bot checks treated as access boundaries.
- Empty fields, policy rejections, and browser-time usage monitored.
- Output retention and deletion rules documented.
With those controls in place, an API can make authorized Cloudflare-site extraction predictable without confusing rendering capability with permission to defeat a site’s defenses.