How to Download and Manage Files with Puppeteer
Configure Puppeteer’s browser download policy and save path, then manage completion, validation, naming, and cleanup in your Node.js application.
Puppeteer can configure whether the browser allows downloads and where it saves them. Your Node.js application still has to detect when a file is ready, validate it, give it a safe name, and clean it up. Puppeteer’s Files guide states: “Currently, Puppeteer does not offer a way to handle file downloads in a programmatic way.” So there is no built-in Puppeteer download-completion or file-lifecycle API to rely on.
This guide configures Chromium’s download behavior for a browser context, then uses Node.js filesystem checks as an application-level pattern. It also distinguishes downloading from uploading: Puppeteer’s documented uploadFile() example is for sending a local file through an HTML file input, not retrieving a download.
1. Install Puppeteer and choose a browser
For a local browser managed by the package, install puppeteer:
npm install puppeteer
The package downloads Chrome for Testing and chrome-headless-shell by default. The 25.12.0 installation guide gives approximate download sizes of 170 MB on macOS, 282 MB on Linux, and 280 MB on Windows; these are platform-specific estimates, not fixed requirements. puppeteer-core does not download Chrome. Choose it when you manage the executable yourself or connect to a remote browser. See the official installation guide.
If your package manager blocks install scripts, Puppeteer’s browser download may be skipped and launch can fail because Chrome is missing. Install it manually with:
npx puppeteer browsers install
Alternatively, configure your package manager to allow the Puppeteer postinstall script. Check the current installation guide for the package manager you use.
2. Configure a download destination and policy
Use a dedicated directory per job or browser context where practical. Ensure the process can write to it, avoid sharing a directory between untrusted jobs, and decide whether the downstream process can accept browser-generated GUID filenames.
The Puppeteer DownloadBehavior reference documents these policies:
| Policy | Effect | Use when |
|---|---|---|
allow |
Permits downloads. | You want files saved under the configured path. |
deny |
Blocks downloads. | A workflow must not write downloaded files. |
default |
Uses browser default behavior. | You intentionally want the browser’s default policy. |
allowAndName |
Permits downloads and names files by download GUID. | Your application can map GUID-named files to jobs. |
The API reference says downloadPath is required with allow and allowAndName. With allowAndName, do not expect the server-provided filename to be preserved.
Puppeteer 25.12.0 exposes downloadBehavior in ConnectOptions to configure behavior for the context. This example connects to an already-running browser using its WebSocket endpoint. Configure that endpoint through your deployment; do not expose it publicly.
import puppeteer from 'puppeteer';
import { mkdir } from 'node:fs/promises';
import path from 'node:path';
const downloadPath = path.resolve('downloads', crypto.randomUUID());
await mkdir(downloadPath, { recursive: true });
const browser = await puppeteer.connect({
browserWSEndpoint: process.env.PUPPETEER_WS_ENDPOINT,
downloadBehavior: {
policy: 'allow',
downloadPath,
},
});
try {
const page = await browser.newPage();
await page.goto('https://example.com/report', { waitUntil: 'domcontentloaded' });
await page.locator('a[href="/files/report.csv"]').click();
// The click starts the browser download. Completion detection and validation
// are application responsibilities; see the next section.
} finally {
await browser.disconnect();
}
Save this as an ES module, set PUPPETEER_WS_ENDPOINT to the browser’s WebSocket endpoint, and replace the example URL and selector with the target page. This configures browser behavior; it does not provide a Puppeteer download event or guarantee the click triggered a download. Match the option shape to your installed Puppeteer version and browser protocol.
Locally launched browsers and context-level configuration
The documentation references downloadBehavior on ConnectOptions. If you launch a local browser rather than connect to one, verify whether your installed Puppeteer version exposes the option on the launch configuration you use. The BrowserContextOptions reference also lists it, but that page is explicitly marked Next and should not be treated as a stable-release guarantee.
Older examples often call Page.setDownloadBehavior. Do not copy that into a current setup without checking compatibility: API location and browser-protocol support can vary. Prefer the API documented for your installed release.
3. Detect and manage the downloaded file in Node.js
Because Puppeteer does not provide programmatic download handling, completion checks and file management belong to your application. A common filesystem pattern is to wait for a new file, ignore temporary partial-download files, require the file size to remain stable for a short interval, and enforce a deadline. This is a practical pattern, not a Puppeteer-prescribed algorithm; validate it with the browser, server, and file types in your workflow.
import { readdir, stat, rename, rm } from 'node:fs/promises';
import path from 'node:path';
async function waitForDownload(directory, {
timeoutMs = 60_000,
stableMs = 1_000,
pollMs = 250,
maxBytes = 100 * 1024 * 1024,
} = {}) {
const deadline = Date.now() + timeoutMs;
let previous = new Map();
while (Date.now() < deadline) {
const names = await readdir(directory);
const candidates = names.filter(name =>
!name.endsWith('.crdownload') &&
!name.endsWith('.part') &&
!name.endsWith('.tmp')
);
const current = new Map();
for (const name of candidates) {
const fullPath = path.join(directory, name);
try {
const info = await stat(fullPath);
if (!info.isFile()) continue;
if (info.size > maxBytes) throw new Error(`Download exceeds size limit: ${name}`);
current.set(name, { size: info.size, mtimeMs: info.mtimeMs, fullPath });
} catch (error) {
if (error.code === 'ENOENT') continue; // File changed during directory scan.
throw error;
}
}
for (const [name, file] of current) {
const old = previous.get(name);
if (old && old.size === file.size && old.mtimeMs === file.mtimeMs) {
return file.fullPath;
}
}
previous = current;
await new Promise(resolve => setTimeout(resolve, pollMs));
}
throw new Error(`Timed out waiting for a completed download in ${directory}`);
}
const filePath = await waitForDownload(downloadPath, { timeoutMs: 90_000 });
const info = await stat(filePath);
if (info.size === 0) throw new Error('Downloaded file is empty');
// Use a name chosen by your application, not an unchecked server filename.
const managedPath = path.join(downloadPath, 'report.csv');
await rename(filePath, managedPath);
// After downstream processing, remove the job directory when safe.
// await rm(downloadPath, { recursive: true, force: true });
The snippet is runnable as a Node.js module, but it is intentionally a filesystem heuristic rather than a definitive browser signal. A stable file might still be incomplete if the server pauses mid-transfer. For stronger guarantees, use a known expected size, checksum, file-format parser, or server-provided job status where available. Also account for multiple downloads: associate each with a unique job directory, or maintain an explicit mapping instead of accepting whichever file appears first.
Validate, name, and clean up safely
- Validate: check size bounds and expected content type or file signature. Do not trust only the extension or the page’s displayed filename.
- Name: generate the destination name yourself. If using GUID filenames, persist the mapping from browser job to GUID. Sanitize any server-provided name to prevent path traversal.
- Handle collisions: use unique per-job directories or exclusive file creation so parallel runs cannot overwrite one another.
- Clean up: delete temporary files after successful processing and on failure. Use a retention policy for abandoned job directories after crashes.
- Bound resources: set timeouts, file-size limits, and a concurrency limit. Large or stalled downloads otherwise consume disk space and worker capacity.
- Protect data: downloaded files may contain credentials or personal data. Restrict directory permissions and avoid logging file contents or sensitive URLs.
4. Uploading is a different operation
For upload, Puppeteer interacts with a file input on the page. The official Files guide’s example waits for input[type=file] and calls uploadFile(). That sends a local file to the page; it does not configure or manage a download.
const input = await page.waitForSelector('input[type="file"]');
await input.uploadFile('./path-to-local-file');
5. cURL, Python, and Node.js for a screenshot of a download page
If your goal is to capture what a download page looks like, a screenshot API can avoid installing and managing a browser in your application. It captures the rendered page; it does not download the linked file or replace Puppeteer’s file workflow. ScreenshotNeo is a screenshot API and MCP server for developers. Its one-call API returns an image or PDF, and its request options include custom headers and cookies for pages that need them. See the ScreenshotNeo API documentation.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.com/downloads \
-o download-page.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com/downloads"},
timeout=90,
)
r.raise_for_status()
open("download-page.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://example.com/downloads',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) =>
writeFile('download-page.webp', Buffer.from(await res.arrayBuffer()))
);
Or skip the browser setup
For screenshots of a download page, ScreenshotNeo accepts one GET request with the page URL and returns an image or PDF. Cookie and consent banners, newsletter popups, and chat widgets are removed before capture; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.com/downloads \
-o download-page.webp
Get a free API key at ScreenshotNeo sign-up.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Chrome executable missing at launch | Install scripts were blocked or the package is puppeteer-core without a managed browser. |
Run npx puppeteer browsers install, permit the postinstall script, or configure the executable/remote browser for your setup. |
| Download is blocked | Policy is deny or default behavior does not permit it. |
Configure allow with a writable downloadPath; verify the installed API version and context. |
| Download path is ignored or configuration errors | The policy/path combination is invalid, the path is missing, or the option is not supported at that API location. | Supply downloadPath for allow/allowAndName and check your version’s API reference. |
| File never appears | The click did not start a download, the page opened a new tab, navigation replaced the page, or the server returned an error page. | Check the selector and network response; allow for a new target where relevant; inspect the page and use a deadline so workers do not wait forever. |
| Wait returns too early or times out | Polling can mistake an old file for a new one, or a slow transfer can pause between writes. | Use an empty per-job directory, require stable size, increase timeout appropriately, and validate expected size or content before processing. |
| Filename is a GUID | The configured allowAndName policy uses the download GUID. |
Keep a job-to-GUID mapping if available in your integration, or choose a policy and naming workflow that suits your application; do not assume the server name survives. |
| Parallel jobs overwrite or consume excess disk | Jobs share paths or have no quotas and cleanup policy. | Use isolated directories, unique names, concurrency and size limits, and cleanup after processing plus stale-directory retention. |
| Upload example does not retrieve a file | uploadFile() sends a local file through an input element. |
Configure browser download policy for downloads, then manage the resulting filesystem entry in Node.js. |
Performance, reliability, and cost
- Startup: Puppeteer’s managed browser download adds initial setup size and time. Reuse browser processes where your isolation model permits, instead of starting one for every small job.
- Concurrency: downloads use disk and browser resources. Limit simultaneous jobs and give each job its own destination to avoid races.
- Reliability: distinguish a navigation finishing from a file download finishing. Use bounded waits and validate the actual file; filesystem stability alone is a heuristic.
- Retries: retry only when the action is safe and the failure is transient. Use a fresh job directory on retry so a partial or previous file is not mistaken for success.
- Cost: Puppeteer itself is a library, but operating it requires compute, browser storage, and disk capacity in your environment. ScreenshotNeo has a free allowance of 1,000 shots per month; paid plans are $5 for 3,000, $15 for 15,000, $39 for 60,000, $99 for 250,000, and $249 for 1,000,000. Yearly billing gives two months free; all listed features are on every plan. A screenshot API is for page captures, not retrieval of arbitrary downloaded files.
FAQ
Can Puppeteer return the downloaded bytes directly?
The documented Files guide says programmatic download handling is not offered. Treat the browser as writing to disk and let your application validate and process the file.
Does setting a download path preserve the website’s filename?
A path selects the save directory. With allowAndName, the documented naming uses download GUIDs, so plan for application-side naming.
Can I use Puppeteer to download a PDF?
If the site serves it as a browser download, configure download behavior and manage the saved file. If you need a PDF rendering of a webpage, Puppeteer’s page-to-PDF workflow is a different operation.
Does a screenshot capture download the file?
No. ScreenshotNeo captures the rendered page as an image or PDF. Use Puppeteer or an HTTP client to retrieve the actual file.


