How to Control File Downloads in Puppeteer
Set Puppeteer’s download policy and destination, then use a bounded, content-aware check to confirm that each download has finished.
To control file downloads in Puppeteer, set a download policy and an absolute, writable destination on the browser context before opening the page. Then use a separate, bounded completion check: the policy sets permission and location, but it does not tell your script that a transfer has finished.
The current Puppeteer API supports allow, deny, default, and allowAndName. The API requires downloadPath when the policy is allow or allowAndName. See the DownloadBehavior API reference.
Set a download policy and destination
Create a context with its own download directory for each job. Use an absolute path that the process can write to, and create the directory before starting the browser workflow. Isolating jobs avoids confusing a new transfer with an old file of the same name.
import puppeteer from 'puppeteer';
import { mkdir } from 'node:fs/promises';
import path from 'node:path';
import os from 'node:os';
const downloadPath = path.join(os.tmpdir(), `puppeteer-download-${Date.now()}`);
await mkdir(downloadPath, { recursive: true });
const browser = await puppeteer.launch({ headless: true });
try {
const context = await browser.createBrowserContext({
downloadBehavior: {
policy: 'allow',
downloadPath,
},
});
const page = await context.newPage();
await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
await page.locator('a[download]').click();
// Add a completion check appropriate to the file and protocol; see below.
console.log(`Download directory: ${downloadPath}`);
await context.close();
} finally {
await browser.close();
}
Replace the example URL and selector with the page and control that initiate your download. If you need to inspect or process the downloaded file, do so before closing the context or removing its directory.
Choose the policy
| Policy | Use | Filename behavior |
|---|---|---|
allow |
Permit downloads to the configured destination. | Uses the browser’s regular filename behavior. |
allowAndName |
Permit downloads with protocol-defined GUID-based names. | Do not expect the suggested filename; map files using the relevant notification or other reliable metadata. |
deny |
Prevent downloads in the context. | No downloaded file should be expected. |
default |
Use the browser’s default behavior where available. | Behavior depends on the browser and protocol. |
Both allow and allowAndName require a downloadPath. Check the documentation for the exact Puppeteer version and browser/protocol you deploy, particularly if you use WebDriver BiDi: the practical guide notes that allowAndName is unsupported with BiDi. The CDP protocol schema snapshot lists the four policies and marks the older Page-domain command as deprecated; do not treat that snapshot as a version-independent guarantee (protocol schema snapshot).
Wait for a download to finish
Setting the policy is not a completion signal. A pre-existing file with the expected name, a nonzero file size, or the appearance of a temporary .crdownload file alone does not establish that this run completed successfully.
For a known filename and expected content, use a fresh directory and poll for the final file with a deadline. Then validate the content, such as by checking the expected byte count or parsing the file. This example is a bounded filesystem check, not a universal Puppeteer download event:
import { readdir, readFile, stat } from 'node:fs/promises';
import path from 'node:path';
async function waitForExpectedFile(directory, filename, {
timeoutMs = 30_000,
expectedBytes,
} = {}) {
const target = path.join(directory, filename);
const deadline = Date.now() + timeoutMs;
while (Date.now() < deadline) {
const names = await readdir(directory);
if (names.includes(filename)) {
try {
const info = await stat(target);
if (info.isFile() && info.size > 0 &&
(expectedBytes === undefined || info.size === expectedBytes)) {
// Read the file to ensure it is accessible, then validate its format
// or contents for your application before treating it as complete.
await readFile(target);
return target;
}
} catch {
// The browser may still be writing or moving the file; check again.
}
}
await new Promise(resolve => setTimeout(resolve, 200));
}
throw new Error(`Download did not produce a validated ${filename} within ${timeoutMs} ms`);
}
Call the helper only after triggering the download, and give it the server’s expected filename and, when known, expected size. If the server chooses arbitrary filenames or content length is unknown, prefer a notification mechanism documented for the exact browser/protocol setup, then validate the resulting file and enforce a deadline. Handle duplicate names and concurrent downloads explicitly.
Do not assume page.on('download') is a universal Puppeteer event. The available notification workflow depends on the browser and protocol; consult the matching protocol documentation and the Puppeteer Page API for your version.
Configure downloads through Puppeteer settings
Puppeteer configuration can set browser download options, but the setup path matters. The official configuration guide says configuration files and environment variables are ignored by puppeteer-core. It also says to rerun browser installation after changing browser download options in configuration so the settings take effect. If you use puppeteer-core, configure the browser/context through the API and verify behavior against the browser you provide.
Common problems and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| The browser rejects the download setup. | downloadPath is missing for allow or allowAndName, or the path is not usable. |
Provide an absolute, writable path and ensure its directory exists. |
| No file appears. | The policy denies downloads, the page did not trigger one, or the browser/protocol uses different behavior. | Confirm the triggering action and policy, then check the documentation matching your Puppeteer and browser versions. |
| The script proceeds while the file is incomplete. | The script treats the policy, a temporary file, or any nonzero size as completion. | Use a deadline, wait for the final expected file or documented notification, and validate bytes or content. |
| The expected filename is missing. | allowAndName uses GUID-based names, or the server suggested a different name. |
Use regular allow when the filename matters, or correlate the GUID with protocol metadata and validate the file. |
| Configuration changes have no effect. | The process uses puppeteer-core, which ignores Puppeteer config files and environment variables, or the browser installation was not rerun after a configuration change. |
Use the API with puppeteer-core; with Puppeteer’s configuration flow, rerun browser installation as its guide specifies. |
| Two jobs inspect the same file or overwrite each other’s output. | They share a directory or filename. | Give each job a fresh directory and define an explicit collision policy. |
| Behavior differs under BiDi. | CDP and BiDi do not have identical download support; allowAndName is noted as unsupported with BiDi in the practical guide. |
Check the deployed protocol’s documentation and select a supported policy. |
Reliability, performance, and cost
- Reliability: isolate each job, set a finite deadline, validate file content, and clean up temporary files after successful processing or failure. Preserve enough error detail to identify the URL, browser/protocol, and stage that failed.
- Performance: polling adds filesystem checks; use a modest interval and a deadline appropriate to the expected transfer. For high concurrency, avoid a single shared directory and avoid tight polling loops.
- Storage: downloads consume local disk space. Set limits and cleanup policies for large or repeated jobs, and account for partial files left by timeouts.
- Cost: Puppeteer itself does not provide a per-download billing model in the cited guidance. Your costs depend on compute, browser runtime, storage, and any infrastructure you operate.
Or skip the browser setup
If your task is to capture a webpage as an image or PDF rather than download a file from an arbitrary site, ScreenshotNeo is a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. See the API documentation for its parameters and response details.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
- Cookie banners are accepted and removed before capture; newsletter popups and chat widgets are also removed. Each step can be turned off.
- Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing. Response headers report the page verdict and billing status.
- An MCP server lets Claude, Cursor, and other MCP clients use
take_screenshot,get_page_info, andcapture_pdf. - The free plan includes 1,000 screenshots a month with no card. Paid plans start at $5 for 3,000 screenshots; every feature is on every plan.
Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.
FAQ
Does Puppeteer automatically download files to the directory I choose?
Only when the context’s policy and destination are configured appropriately, and the path is writable. Browser and protocol support can affect behavior.
Is allowAndName compatible with every Puppeteer protocol?
No. Check the documentation for the exact setup; the cited practical guide says it is unsupported with WebDriver BiDi.
Can I use Puppeteer downloads for webpage screenshots?
Downloads and screenshots are different tasks. Puppeteer can automate a browser, while ScreenshotNeo provides a direct screenshot and PDF API for webpage captures.


