How to Handle Bot Detection in Headless Pyppeteer
Diagnose bot checks in headless Pyppeteer, collect useful evidence, and choose authorized fixes without relying on brittle stealth tricks.
Direct answer: treat a bot-check page as a compatibility, policy, or access problem first. Confirm that you are allowed to automate the site, record exactly what the browser receives, reproduce with a current Chrome build, and compare Pyppeteer runtime modes only to isolate rendering differences. A changed user agent, headful Chrome, a stealth plugin, or a proxy is not a guaranteed fix.
Pyppeteer is an unofficial Python port of Puppeteer, and its repository says it is unmaintained and has been outside minor changes for a long time. Browser-version drift is therefore part of the diagnosis, not an unusual edge case. Read the Pyppeteer repository notice.
1. What “bot detection” actually means
Modern defenses combine several signal classes. Cloudflare documents heuristic checks, known malicious fingerprints, invisible JavaScript detections, and machine-learning scoring based on headers, session characteristics, and browser signals. Its documentation summarizes the design as: “Cloudflare uses multiple detection engines because different bot types require different detection strategies.” Cloudflare bot detection engines
| Signal class | What you may observe | Useful diagnostic |
|---|---|---|
| Heuristics and fingerprints | 403, block page, or an interstitial before your application loads | Save status, headers, redirect chain, and response body |
| JavaScript detection | The HTML loads, then a cookie or challenge result appears asynchronously | Capture console messages, cookies, and network requests after load |
| Session behavior | One request works but bursts, repeated logins, or inconsistent cookies fail | Reuse a coherent context and slow requests to the published rate |
| Browser/runtime signals | Headless mode is challenged while another runtime renders differently | Compare supported headless modes with identical navigation steps |
| Policy or access controls | Every runtime receives a challenge or denial | Ask the site owner for an API, allowlist, or verified-bot route |
Cloudflare’s JavaScript Detections run invisibly on HTML responses and are separate from interactive Challenge Pages and Turnstile. A successful page load does not prove that your automation is permitted or that a later request will be accepted. JavaScript Detections documentation
2. Start with authorization
- Read the target’s terms and
robots.txt. - Obtain written permission for authenticated, high-volume, or sensitive workflows.
- Prefer the site’s API, export, test environment, allowlist, or verified-bot program.
- Do not rotate identities, solve challenges, or disguise traffic to defeat a block.
Cloudflare’s sample bot terms describe automated access as something a site may permit or restrict. Cloudflare sample terms
3. Build a reproducible Pyppeteer diagnostic
Record the Python and Pyppeteer versions, browser revision or executable path, operating system, launch flags, target URL, timestamp, and whether the failure occurs only in headless mode. The following script logs the first response, redirects, cookies, console errors, title, and a screenshot of the page you actually received.
import asyncio
import json
import os
import platform
from pathlib import Path
from pyppeteer import launch
URL = os.environ.get("TARGET_URL", "https://example.com")
MODE = os.environ.get("PYPPETEER_MODE", "new")
async def main():
launch_options = {
"headless": False if MODE == "headful" else ("shell" if MODE == "shell" else True),
"args": ["--no-sandbox", "--disable-setuid-sandbox"],
"dumpio": True,
}
browser = await launch(launch_options)
page = await browser.newPage()
page.setDefaultNavigationTimeout(60_000)
events = {"responses": [], "requests": [], "console": [], "page_errors": []}
page.on("response", lambda response: events["responses"].append({
"url": response.url,
"status": response.status,
"headers": response.headers,
}))
page.on("request", lambda request: events["requests"].append({
"method": request.method,
"url": request.url,
"resource_type": request.resourceType,
}))
page.on("console", lambda message: events["console"].append({
"type": message.type,
"text": message.text,
}))
page.on("pageerror", lambda error: events["page_errors"].append(str(error)))
try:
response = await page.goto(URL, {"waitUntil": "domcontentloaded", "timeout": 60_000})
await asyncio.sleep(5)
result = {
"python": platform.python_version(),
"platform": platform.platform(),
"mode": MODE,
"final_url": page.url,
"title": await page.title(),
"first_status": response.status if response else None,
"cookies": await page.cookies(),
"html_prefix": (await page.content())[:1000],
"events": events,
}
Path("bot-diagnostic.json").write_text(json.dumps(result, indent=2, default=str))
await page.screenshot({"path": "bot-page.png", "fullPage": True})
print(json.dumps({k: result[k] for k in ("mode", "final_url", "title", "first_status")}, indent=2))
finally:
await browser.close()
asyncio.get_event_loop().run_until_complete(main())
Run the same script three times:
PYPPETEER_MODE=new python diagnose.py
PYPPETEER_MODE=shell python diagnose.py
PYPPETEER_MODE=headful python diagnose.py
Use headless=True, headless='shell' where your installed compatibility layer supports it, and headless=False as separate controlled experiments. Puppeteer documents these as different runtime choices; changing modes can expose a rendering or compatibility issue but does not guarantee acceptance by a protected site. Puppeteer headless modes
4. Read the evidence before changing code
Identify the response
- 403 or 401 on the first document: access policy, reputation, credentials, or a perimeter rule may be responsible.
- 200 with a challenge title: the browser rendered a challenge document instead of the application.
- Redirect loop: cookies, clock, locale, or session state may not persist.
- Application HTML followed by a challenge: inspect JavaScript detection, cookies, and asynchronous requests.
Compare the right variables
Hold URL, account, network, viewport, locale, and navigation code constant. Change one variable at a time: browser executable, Pyppeteer revision, headless mode, or wait condition. Save response headers and the complete redirect chain for every run.
Check session coherence
Use one browser context for a logical workflow, persist only the cookies you are authorized to persist, avoid needless parallel bursts, and honor rate limits. A user-agent string alone changes one request attribute while defenses can evaluate many others.
5. Common fixes that do not solve the root cause
| Attempt | Why it often fails | Safer next step |
|---|---|---|
Change User-Agent |
Detection can also use headers, session behavior, JavaScript, and browser signals. | Compare complete diagnostics and request an allowlist or API. |
| Switch to headful Chrome | It changes the runtime, but a protected site may still identify or restrict automation. | Use it to isolate compatibility and rendering differences only. |
Patch navigator.webdriver |
It addresses one observable and can create an inconsistent fingerprint. | Keep the runtime supportable and use an authorized automation route. |
| Install a stealth wrapper | Detection vendors change; no package guarantees acceptance. | Prefer a supported browser build, test endpoint, or managed browser. |
| Rotate proxies or identities | It can violate terms, break sessions, and make behavior less reproducible. | Obtain permission and use the site’s intended access path. |
6. Troubleshooting checklist
“Pyppeteer downloads an old Chromium”
Pyppeteer can download its pinned Chromium revision. Record that revision, install a current stable Chrome or Chromium where your environment permits, and pass its executable path explicitly. Re-run the same diagnostic before changing page code. The unmaintained repository status means browser drift should be expected.
“It works locally but fails in CI”
Compare OS image, Chrome executable, sandbox permissions, fonts, timezone, locale, outbound IP, and environment variables. Save the first response and screenshot from both systems. Do not infer that CI failure is caused by headless mode alone.
“Navigation times out”
Distinguish a slow application from a challenge that never finishes. Use a bounded navigation timeout, wait for a specific application selector when possible, and save the HTML on timeout. Avoid infinite retries; use exponential backoff and a maximum attempt count.
“The page is blank”
Check console errors, failed requests, JavaScript exceptions, viewport size, and whether the page requires a later network call. A blank document, a blocked document, and an application that has not hydrated are different outcomes.
“A CAPTCHA or Turnstile appears”
Stop and verify authorization. Ask the owner for a test key, allowlist, API, or verified-bot treatment. Do not automate solving or bypassing the challenge.
“The challenge appears only in headless mode”
Run the three controlled modes above and compare status, title, cookies, and redirects. If only one mode differs, report the exact browser version and launch options to the site owner or your browser provider. A headful result is not proof that production automation is permitted.
7. Reliability, performance, and cost
- Reuse carefully: reusing a browser and context reduces startup time, but recycle unhealthy sessions and isolate unrelated accounts.
- Wait for intent: prefer a known selector or a bounded network-idle wait over arbitrary multi-minute sleeps.
- Limit concurrency: small, measured batches are easier to debug and less likely to trigger rate controls than bursts.
- Capture evidence: retain status, headers, title, cookies, console errors, and a redacted screenshot for each failed run.
- Budget browser work: browser startup, rendering, retries, storage, and egress all affect cost. An API or managed browser can trade local maintenance for usage-based pricing and vendor dependency.
- Plan for change: defenses, browser releases, and Pyppeteer compatibility evolve. Pin versions, schedule dependency reviews, and keep an alternate authorized access path.
8. When a managed browser is appropriate
If maintaining Chrome, fonts, sandboxing, regional routing, and session cleanup is not your product’s core job, evaluate an authorized managed browser. Cloudflare Browser Run documents Puppeteer support through CDP and describes hosted headless browsers for automation, screenshots, PDFs, and testing. It still identifies automation and does not promise that a protected third-party site will allow every workflow. Browser Run overview · Puppeteer over CDP
9. Or skip the browser setup
For a screenshot workflow, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns PNG, JPEG, WebP, or PDF. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo API documentation for parameters and response details. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing; response headers identify the page verdict and whether the request was billed. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. Free usage includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots.
Create a free ScreenshotNeo account.
10. FAQ
Should I abandon Pyppeteer?
For new work, evaluate a maintained browser library or a managed browser. For an existing script, first pin and record its browser revision, then decide whether migration effort is justified by reliability and support needs.
Can a bot score tell me exactly why I was blocked?
No. A score or challenge result is an outcome, not a complete explanation. Your most useful evidence remains the response, redirects, cookies, console output, and authorized site-owner feedback.
Is headful Chrome safer?
It can help distinguish rendering differences, but it is not an acceptance guarantee. Treat it as a diagnostic mode.
What is the fastest legitimate fix?
Use the target’s API, test endpoint, allowlist, or verified-bot process. If you only need screenshots, use a screenshot API that reports failed and non-billable outcomes clearly.


