How to Download a PDF with Puppeteer by Clicking Its Download Button
Use Puppeteer to click a PDF download button, wait for the right event, verify the file, and handle Chrome’s PDF viewer reliably.
To download a PDF after clicking a button in Puppeteer, configure a writable download directory, allow downloads with DownloadBehavior, wait for the correct navigation or request event before clicking, then verify the completed file on disk. A click can either start a normal download or navigate to a PDF document, and those cases require different handling.
Puppeteer is a JavaScript library for controlling Chrome or Firefox through DevTools Protocol or WebDriver BiDi. See the official Puppeteer documentation for the current API.
Complete working example: click a button and save the PDF
The following Node.js script handles a normal browser download. It creates a directory, enables downloads, waits for a stable locator, clicks the control, waits for the download request to finish, and checks that a non-empty PDF exists.
import puppeteer from 'puppeteer';
import fs from 'node:fs/promises';
import path from 'node:path';
const downloadPath = path.resolve('./downloads');
const targetUrl = 'https://example.com/reports';
await fs.mkdir(downloadPath, { recursive: true });
const browser = await puppeteer.launch({
headless: true
});
try {
const page = await browser.newPage();
// DownloadBehavior is required for an unattended download.
const client = await page.createCDPSession();
await client.send('Browser.setDownloadBehavior', {
behavior: 'allow',
downloadPath
});
await page.goto(targetUrl, { waitUntil: 'domcontentloaded' });
const before = new Set(await fs.readdir(downloadPath));
const downloadRequest = page.waitForRequest(request => {
const resourceType = request.resourceType();
return resourceType === 'document' || /pdf/i.test(request.url());
}, { timeout: 30000 }).catch(() => null);
await page.locator('button[data-download="pdf"]').click();
await downloadRequest;
// Chrome may write a temporary .crdownload file first.
const deadline = Date.now() + 30000;
let pdfPath;
while (Date.now() < deadline) {
const files = await fs.readdir(downloadPath);
const candidate = files.find(file =>
file.toLowerCase().endsWith('.pdf') && !before.has(file)
);
const temporary = files.some(file => file.endsWith('.crdownload'));
if (candidate && !temporary) {
const fullPath = path.join(downloadPath, candidate);
const stat = await fs.stat(fullPath);
if (stat.size > 0) {
pdfPath = fullPath;
break;
}
}
await new Promise(resolve => setTimeout(resolve, 250));
}
if (!pdfPath) {
throw new Error('The click completed, but no non-empty PDF appeared');
}
console.log(`Saved PDF to ${pdfPath}`);
} finally {
await browser.close();
}
The selector is deliberately specific. Replace button[data-download="pdf"] with the actual control on your page. Puppeteer locators wait for an element to exist and be usable before interacting with it; the page interactions guide shows this pattern.
Install and run it
mkdir puppeteer-pdf-download
cd puppeteer-pdf-download
npm init -y
npm install puppeteer
node download-pdf.mjs
If your project uses CommonJS, replace the import with const puppeteer = require('puppeteer'); and use a .cjs file.
Choose the correct event before clicking
The most common cause of flaky scripts is clicking first and waiting afterward. Register the wait before the click so a fast navigation or request cannot be missed.
When the click navigates to a PDF URL
Use Promise.all with page.waitForNavigation(). Puppeteer documents this pattern because a separate wait after click() can race.
const [response] = await Promise.all([
page.waitForNavigation({ waitUntil: 'networkidle2' }),
page.locator('button[data-download="pdf"]').click()
]);
if (!response) {
throw new Error('The button did not produce a navigation response');
}
console.log('Navigated to:', response.url());
This tells you that the browser navigated; it does not by itself guarantee that a file was saved. Inspect the response, stream the bytes, or use a browser download flow if the site marks the response as an attachment.
When the click starts a normal download
A normal download may not create a new page navigation. Observe requests or request completion and then inspect the configured directory. Relevant page events include request, response, and requestfinished; see the PageEvent API.
const finished = page.waitForRequest(request =>
/\.pdf(?:$|\?)/i.test(request.url()),
{ timeout: 30000 }
);
await page.locator('button[data-download="pdf"]').click();
const request = await finished;
console.log('PDF request:', request.url());
Some applications use a generated URL, a POST request, or a URL without a .pdf suffix. In those cases, match on a stable path, query parameter, response header, or the route known from your application instead of relying only on the extension.
Configure downloads explicitly
Headless browsers need an explicit download policy and path for unattended jobs. Puppeteer’s DownloadBehavior API requires downloadPath when the policy is allow or allowAndName. Configure it before navigating or clicking.
const client = await page.createCDPSession();
await client.send('Browser.setDownloadBehavior', {
behavior: 'allow',
downloadPath: '/absolute/path/to/downloads'
});
| Setting | Use |
|---|---|
allow |
Allow downloads and let Chrome choose the resulting filename. |
allowAndName |
Allow downloads while using browser-assigned names; provide a path and account for the resulting naming behavior. |
deny |
Block downloads. This is useful for tests that must prove a click does not download. |
Use an absolute, writable directory. In containers, verify that the directory belongs to the browser user and is not mounted read-only.
Verify the file before using it
Do not assume that a click means a usable PDF exists. Chrome can leave a temporary .crdownload file while bytes are arriving, and a server can return an HTML error page with a PDF-looking URL.
- Wait until a new file appears.
- Wait until no matching
.crdownloadfile remains. - Check that the final name ends in
.pdfwhen the site uses normal extensions. - Check that the file size is greater than zero.
- Optionally inspect the first five bytes for the PDF signature
%PDF-.
const bytes = await fs.readFile(pdfPath);
const header = bytes.subarray(0, 5).toString('ascii');
if (header !== '%PDF-') {
throw new Error('The downloaded file is not a PDF');
}
Filename conventions are application-specific. If several downloads can happen in parallel, create one directory per job or record the directory contents before the click and compare afterward.
When the button opens Chrome’s PDF viewer
A PDF viewer navigation is different from a normal download. The browser may display the document in a tab without emitting the filesystem download you expected. Treat this as navigation or response handling: capture the response URL and headers, or use the site’s underlying PDF endpoint directly if your application is allowed to do so.
Puppeteer documents that headless shell mode does not support navigation to a PDF document. If the site relies on viewer navigation, use a browser mode that supports the behavior or handle the PDF response at the network layer. Do not wait forever for a download file that the browser was never asked to create.
const pdfResponsePromise = page.waitForResponse(response => {
const contentType = response.headers()['content-type'] || '';
return contentType.toLowerCase().includes('application/pdf');
}, { timeout: 30000 });
await page.locator('button[data-download="pdf"]').click();
const pdfResponse = await pdfResponsePromise;
if (!pdfResponse.ok()) {
throw new Error(`PDF request failed with ${pdfResponse.status()}`);
}
const body = await pdfResponse.body();
await fs.writeFile('./downloads/report.pdf', body);
Whether response bodies are available depends on how the page and browser expose the request. If the response is a cross-origin viewer navigation or is not retained, use the endpoint and authentication flow documented by the application.
Do not confuse downloading with generating
Page.pdf() prints the current rendered page to a new PDF. It does not click a button or retrieve the server’s existing PDF. Puppeteer’s PDF guide summarizes this distinction: “For printing PDFs use Page.pdf().”
await page.goto('https://example.com/invoice', {
waitUntil: 'networkidle2'
});
await page.pdf({
path: './downloads/invoice-rendered.pdf',
format: 'A4',
printBackground: true
});
| Requirement | Correct approach |
|---|---|
| Retrieve the PDF supplied by a download button | Configure downloads and handle the download request or response. |
| Handle a button that opens a PDF document | Wait for navigation or an application/pdf response and save the bytes. |
| Create a PDF from rendered HTML | Use page.pdf(). |
Selectors and interaction patterns
Prefer a stable accessible name, data attribute, or application-specific selector. Avoid selecting the third button on a page or relying on generated class names.
// Accessible role and name
await page.getByRole('button', { name: /download pdf/i }).click();
// Stable attribute
await page.locator('[data-testid="download-pdf"]').click();
// Link that points to a PDF
await page.locator('a[href*="pdf"]').click();
If the control is disabled until a report finishes generating, wait for the enabled state or for a completion marker. If it is inside an iframe, obtain the frame first:
const frame = page.frames().find(frame => frame.url().includes('/reports/embedded'));
if (!frame) throw new Error('Report frame was not found');
await frame.locator('button[data-download="pdf"]').click();
Common failures and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| No file appears | Downloads are denied, the path is missing, or the directory is not writable. | Set Browser.setDownloadBehavior with allow and an absolute path; create and permission the directory first. |
| Timeout after a successful click | The wait started after the click or the click caused navigation rather than a file download. | Register the wait first; use Promise.all for navigation and response/request observation for downloads. |
| Wrong button is clicked | A broad selector matches several controls. | Use an accessible name, stable data attribute, or a selector tied to the report. |
Only .crdownload exists |
The transfer is still in progress or failed. | Wait for the temporary file to disappear, then verify size and PDF signature; investigate server or network errors if it remains. |
| Downloaded file is HTML | Authentication expired, a redirect led to a login page, or the server returned an error document. | Check response status, final URL, content type, cookies, and authorization before accepting the file. |
| Button opens a viewer | The site navigates to a PDF instead of issuing a browser download. | Handle PDF navigation/response behavior and account for headless-shell PDF limitations. |
| Click has no effect | The element is covered, disabled, inside an iframe, or requires a prior user action. | Wait for visibility and enabled state, inspect the frame, and reproduce the required interaction sequence. |
| Several jobs overwrite files | All jobs share one download directory and names collide. | Use an isolated directory per job or rename only after identifying the newly completed file. |
Reliability and performance checklist
- Use
domcontentloadedfor the initial page when full network idle is unnecessary; wait for the report control or completion marker explicitly. - Use
networkidle2for navigation-triggering clicks only when the page genuinely needs background requests to settle. - Set bounded timeouts and include the target URL, selector, and directory in error messages.
- Capture browser console, page errors, request failures, and response status codes for diagnosing intermittent jobs.
- Keep one browser process per worker and reuse pages when safe; isolate download directories between concurrent tasks.
- Retry only idempotent navigation or download steps, and avoid clicking twice unless the application guarantees duplicate downloads are safe.
- Clean temporary files after success and preserve failed-job artifacts long enough to diagnose them.
Puppeteer’s official sources do not publish a universal download speed, success rate, or timing benchmark. Measure your own target pages, network, browser version, and concurrency instead of assuming a fixed number.
Or skip the browser setup
If the goal is simply a clean screenshot or PDF of a URL, ScreenshotNeo provides a single API request. Its PDF capture supports paper size, margins, landscape mode, and page ranges, so you do not need to maintain a browser download directory for that use case. Read the ScreenshotNeo API documentation for the request options.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo returned ${res.status}`);
const bytes = Buffer.from(await res.arrayBuffer());
require('node:fs').writeFileSync('shot.webp', bytes);
ScreenshotNeo accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. It also has an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is included on every plan. Create a free ScreenshotNeo account.
FAQ
Should I wait for navigation or a download event?
Wait for navigation when the click changes the page URL. Wait for a request, response, or completed file when the click starts a background download.
Can I use page.pdf() to save the clicked PDF?
No. page.pdf() renders the current page into a new PDF. Use download or response handling to retrieve an existing PDF.
Why does the script work headed but fail headless?
Headless mode can differ in download policy and PDF viewer behavior. Set DownloadBehavior explicitly and handle PDF navigation as a response rather than assuming a viewer creates a file.
How do I know the download is complete?
Wait for the temporary download file to disappear, then verify that the final file is non-empty and, where appropriate, begins with %PDF-.
What if the website requires login?
Authenticate before the click, preserve the page’s cookies and authorization state, and check the final response status and content type so a login page is not saved as a PDF.


