How to Save Bot-Protected Web Pages as PDFs
Save a bot-protected page as a PDF after authorized access, with browser, Puppeteer, Cloudflare and ScreenshotNeo workflows.

Short answer: If you are allowed to access the page, open it in a normal browser, complete the bot challenge or sign-in, wait until the real content has rendered, then use Print → Save as PDF. For an authorized repeatable workflow, use a browser-rendering API or Puppeteer with an explicit wait condition and print settings. A changed user agent or headless browser is not a legitimate way to defeat bot protection.
A challenge is an access-control decision made by the site owner. It is not a PDF-format problem. Render the page only after you have legitimate access, and use an official export or ask the site owner when access is blocked.
1. Save the page manually in a browser
This is the simplest method for a one-off PDF. It also gives you the best chance of completing an interactive challenge, login, multi-factor prompt, or consent step correctly.
- Open the URL in a standard browser. Use Chrome, Edge, Firefox, or Safari with the account and network you are authorized to use.
- Complete the challenge or login. Do not start printing while a Cloudflare challenge, interstitial, or sign-in form is still visible.
- Verify the actual page. Confirm that the article, dashboard, or application data you need is visible. A page shell with a spinner is not a completed load.
- Dismiss overlays. Close cookie notices, newsletter forms, chat widgets, and modal dialogs. Expand accordions or “read more” sections that must appear in the PDF.
- Wait for dynamic content. Let images, charts, tables, fonts, and client-rendered text finish loading. Scroll through a long page if that triggers lazy loading.
- Print. Press
Ctrl+Pon Windows/Linux orCmd+Pon macOS, then select Save as PDF. - Choose print settings. Set paper size, portrait or landscape orientation, margins, scale, page range, and whether background graphics should be included.
- Reopen the PDF. Check the page count, text search, links, images, headings, and the edges of tables or code blocks.
Settings that commonly change the result
| Setting | Use it when | Common issue |
|---|---|---|
| Background graphics | The page uses colored panels, charts, or a dark theme | Disabled backgrounds make sections unreadable |
| Margins | You need more horizontal space for tables | Large margins cause wrapping and extra pages |
| Scale | A table or code block is clipped | Too much shrinking makes text hard to read |
| Landscape | The document has wide tables or dashboards | Portrait mode clips columns |
| Page range | You need only selected sections | Headers or context may be omitted |
If the site offers a reader view, print view, export button, or official report download, prefer that output. It usually has fewer navigation elements and more stable pagination.
2. Why bot-protected pages fail to print
Bot protection can show a challenge before the page, replace the response with a block page, or allow the document shell while withholding API data. Printing captures whatever is currently rendered, including an interstitial or an empty application.

Cloudflare documents separate policies for Search, Agent, and Training bot categories. A challenge or block therefore reflects a site policy, not a missing PDF capability. If you own the protected zone, you can create an appropriate WAF skip rule for authorized Browser Run traffic; custom rules using Bot Management fields require Enterprise access according to Cloudflare’s guidance.
Do not assume that changing a user agent solves this. Cloudflare states: The userAgent parameter does not bypass bot protection. Requests from Browser Run will always be identified as a bot.
See the Cloudflare Browser Rendering documentation and the bot category policy documentation.
3. Authorized automation with Puppeteer
Use automation only for pages you own or have permission to process. A browser session must still satisfy the site’s access controls. Puppeteer is useful when you need the same PDF layout repeatedly and can provide an authenticated session or an allowlisted rendering path.
Install and run
npm install puppeteer
const puppeteer = require('puppeteer');
(async () => {
const browser = await puppeteer.launch({
headless: true,
args: ['--no-sandbox']
});
const page = await browser.newPage();
await page.setViewport({ width: 1440, height: 1000, deviceScaleFactor: 1 });
// Supply credentials or an approved session before navigating when required.
await page.goto('https://example.com/article', {
waitUntil: 'networkidle2',
timeout: 90000
});
// Replace this selector with a stable element that proves the real page loaded.
await page.waitForSelector('article', { timeout: 30000 });
// Allow fonts, charts, or late client rendering to settle.
await page.evaluate(() => document.fonts.ready);
await new Promise(resolve => setTimeout(resolve, 1000));
await page.pdf({
path: 'page.pdf',
format: 'A4',
printBackground: true,
margin: { top: '16mm', right: '14mm', bottom: '16mm', left: '14mm' },
displayHeaderFooter: false,
preferCSSPageSize: true
});
await browser.close();
})();
waitUntil: 'networkidle2' is a useful baseline for JavaScript-heavy pages, but it is not proof that the application is ready. Analytics, ads, or long polling can keep the network busy, while a page can become visually complete before the network is idle. Combine a network condition with a meaningful selector, a known application state, or a short settling delay.
Print CSS versus screen CSS
Puppeteer generates PDFs using the print CSS media type. That means a site’s @media print rules can hide navigation, change colors, or reflow columns. When the screen layout is the required output, set the media type explicitly:
await page.emulateMediaType('screen');
await page.pdf({
path: 'screen-layout.pdf',
format: 'Letter',
printBackground: true
});
Use print media for a document-like result and screen media for dashboards or branded layouts. The Puppeteer API documents createPDFStream, page formats, margins, headers and footers, backgrounds, and CSS media behavior at pptr.dev.
Authenticated pages and cookies
For a permitted private page, establish the session before navigation. Avoid putting passwords in URLs or source files. Load secrets from your deployment’s secret store, and clear the browser profile after the job.
await page.setCookie({
name: 'session',
value: process.env.SESSION_COOKIE,
domain: 'example.com',
path: '/',
httpOnly: true,
secure: true
});
await page.goto('https://example.com/account/report', {
waitUntil: 'networkidle2',
timeout: 90000
});
If the site requires an interactive challenge, complete it in an approved normal browser session or use an owner-approved Browser Rendering workflow. Do not try to automate challenge solving or present a renderer as a bypass.
4. Cloudflare Browser Rendering PDF API
Cloudflare’s Browser Rendering /pdf Quick Action accepts a URL or custom HTML and returns a rendered PDF. It supports custom CSS, print backgrounds, headers and footers, page dimensions, resource allow/deny lists, and gotoOptions such as networkidle0 and networkidle2. The API also exposes controls for authorization, Browser Rendering write permission, timeout, format, scale, and resource types.
Use it when you control the Cloudflare account or have explicit permission from the site owner. Configure the zone’s WAF policy to allow the authorized traffic where appropriate. Start with a selector or application-ready condition, then add PDF settings such as paper size, margins, landscape mode, and print backgrounds. The primary references are the Cloudflare PDF Quick Action and Browser Rendering API reference.
5. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. Its PDF endpoint can render an authorized URL with one GET request. Use the ScreenshotNeo documentation for the complete option list.

curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-d format=pdf \
-o page.pdf
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={
"access_key": "YOUR_API_KEY",
"url": "https://stripe.com",
"format": "pdf"
},
timeout=90
)
r.raise_for_status()
open("page.pdf", "wb").write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com',
format: 'pdf'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const fs = require('node:fs');
fs.writeFileSync('page.pdf', Buffer.from(await res.arrayBuffer()));
ScreenshotNeo accepts options for paper size, margins, landscape mode, page ranges, custom CSS and JavaScript, waits for a selector, delay, or network idle, custom headers and cookies, user agent, Authorization, timezone, geolocation, blocked resource types, and caching with a TTL you choose. It can also capture a single CSS-selected element, load lazy images for full-page captures, click an element before capture, hide selectors, and create signed links or asynchronous jobs with signed webhooks.
For clean output, cookie and consent banners are accepted and more than 60 known consent platforms, newsletter popups, and chat widgets are removed before capture; each step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. Response headers report the page verdict and billing state through X-Page-Verdict and X-Billed. ScreenshotNeo also provides an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
Plans include 1,000 screenshots per month free with no card, then Starter at $5 for 3,000, Growth at $15 for 15,000, Pro at $39 for 60,000, Scale at $99 for 250,000, and Business at $249 for 1,000,000. Yearly billing gives two months free, and every feature is included on every plan. Create a free ScreenshotNeo account to start with 1,000 screenshots a month and no card.
6. Troubleshooting incomplete or blocked PDFs
| Symptom | Likely cause | Fix |
|---|---|---|
| The PDF contains a challenge page | The session never reached the protected content | Complete the challenge in an authorized browser, then print; ask the owner for an export or allowlist |
| Text is missing | Client rendering was still running | Wait for a stable selector, fonts, and application data before generating the PDF |
| Images are blank | Lazy loading, blocked resources, or a premature capture | Scroll to trigger lazy loading, allow image resources, and wait for image completion |
| Tables are clipped | Portrait layout, large margins, or fixed-width CSS | Use landscape, reduce margins, set scale carefully, or add print CSS |
| Colors disappear | Background graphics are disabled | Enable print backgrounds or set the renderer’s print-background option |
| Puppeteer times out | Long polling, ads, or a slow dependency | Use a meaningful selector, a bounded timeout, and resource blocking for nonessential types |
| Only a blank shell is saved | Authentication or API calls were not available | Load an approved session, verify response status, and inspect console/network errors |
| A user-agent change has no effect | Bot protection identifies automation separately | Do not treat user-agent changes as a bypass; use owner-approved access |
7. Reliability, performance, and cost considerations
- Make readiness explicit. A selector such as
articleor[data-rendered="true"]is more reliable than an arbitrary sleep alone. - Bound every job. Set navigation and total job timeouts so a stalled third-party resource cannot consume a worker indefinitely.
- Reduce unnecessary work. Block ads, trackers, videos, and fonts only when doing so does not change the document you need.
- Reuse carefully. Browser reuse improves throughput, but isolate cookies and storage between users and tenants.
- Cache stable pages. A chosen TTL can reduce repeated rendering. Do not cache pages containing private or rapidly changing data without a clear retention policy.
- Validate the artifact. Check HTTP status, content type, file size, page count, and a small text extraction sample before marking a job successful.
- Prefer asynchronous jobs for batches. For many URLs, queue work, retry transient failures with backoff, and record the final verdict and billing state.
For ScreenshotNeo, cache hits and failed or unusable captures are not billed, and the response headers identify the result. Bulk capture supports up to 100 URLs per call, while signed webhooks let an asynchronous workflow receive completion notifications.
8. What you may and may not automate
Automation is appropriate for your own site, a customer-approved integration, a public page whose terms permit capture, or an official API/export. It is not appropriate to defeat a challenge on an unrelated site, evade rate limits, or collect content that the owner intentionally restricts. If access is denied, request permission, use the site’s export, or ask for an accessible copy.
FAQ
Can I print a page that still shows a bot check?
You can print what is visible, but the result will be the challenge page. Complete authorized access first or obtain an official export.
Does Puppeteer bypass Cloudflare?
No. Puppeteer renders a browser session; it does not grant permission or defeat a challenge. Use it after authorized access or with an owner-approved rendering setup.
Why does the browser show the page but my script gets a blank PDF?
The interactive browser may have cookies, tokens, completed challenges, or a different wait state. Reproduce the approved session, wait for application data, and inspect network and console errors.
Should I use print CSS or screen CSS?
Use print CSS for a document-like report. Use screen CSS when preserving a dashboard or branded layout matters more than compact pagination.
Can I save a JavaScript-heavy page reliably?
Yes, when the page is authorized and you wait for a meaningful readiness condition, loaded assets, and any required user interaction before generating the PDF.


