How to Manage Browser Downloads and Executables with Puppeteer
Learn to save files from Puppeteer pages, install and pin Puppeteer’s browser, and launch an external Chrome executable.
“Browser downloads” in Puppeteer can mean two different things: saving a file from a page that Puppeteer controls, or installing the browser executable that Puppeteer launches. Configure those workflows separately. For page downloads, set the browser’s download behavior and destination. For the browser itself, choose Puppeteer-managed installation or provide a compatible external executable.
1. Save files downloaded by a page
Puppeteer’s DownloadBehavior contract defines a policy and a download path. The path is required when the policy is allow or allowAndName. With allowAndName, downloaded files use download GUIDs as their names. Check the API reference for the Puppeteer version installed in your project: DownloadBehavior.
The following example uses the Chrome DevTools Protocol to allow downloads to a chosen directory, triggers a page download, and waits for the expected file. It assumes Node.js with puppeteer installed and a page with a link whose selector is #download. Change the URL and selector to match your page.
import puppeteer from 'puppeteer';
import { access, mkdir } from 'node:fs/promises';
import path from 'node:path';
const downloadPath = path.resolve('downloads');
await mkdir(downloadPath, { recursive: true });
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
const client = await page.createCDPSession();
await client.send('Page.setDownloadBehavior', {
behavior: 'allow',
downloadPath,
});
await page.goto('https://example.com/downloads', {
waitUntil: 'domcontentloaded',
});
await page.locator('#download').click();
// Replace this with the actual expected filename and a deadline suitable
// for the file and network. A directory watcher is preferable if names vary.
const expectedFile = path.join(downloadPath, 'report.csv');
const deadline = Date.now() + 30_000;
while (Date.now() < deadline) {
try {
await access(expectedFile);
break;
} catch {
await new Promise((resolve) => setTimeout(resolve, 250));
}
}
await access(expectedFile);
console.log(`Downloaded ${expectedFile}`);
} finally {
await browser.close();
}
The wait loop is application-specific: downloads may have different names or take longer. In production, watch the destination directory and check that the final file is complete before consuming it. Chromium may create temporary partial files while a download is in progress.
Choose a download policy and filename strategy
| Setting | Use | Details |
|---|---|---|
deny |
Prevent page downloads | Useful when the automation should never write downloaded files. |
allow |
Save downloads to a known directory | Set downloadPath. The page or server determines the filename. |
allowAndName |
Allow downloads with GUID-based names | Set downloadPath; map GUIDs to application-level names if needed. |
The documented interface describes the behavior and destination, but it does not provide a universal high-level helper for waiting on each download or resolving a GUID to the original filename. Use the version-specific API and protocol behavior for your installed release, and make your own completion and file-identification logic explicit.
Operational details
- Create the destination directory before allowing downloads, and use an absolute path to avoid ambiguity about the process working directory.
- In containers or CI, ensure the browser process can write to that directory and that the directory is available to the consuming process.
- Use unique per-job directories for parallel downloads so jobs cannot overwrite files or mistake another job’s output for their own.
- Validate file size, type, and content before trusting a download. A successful click does not establish that the expected file was saved.
- Set an explicit deadline and handle navigation, authentication, and server errors separately from download completion.
2. Install the browser executable Puppeteer launches
Installing puppeteer downloads a compatible Chrome for Testing browser and chrome-headless-shell by default. puppeteer-core does not download Chrome; it is intended for remote or externally managed browsers. The default browser cache is $HOME/.cache/puppeteer starting with Puppeteer 19.0.0. See the official installation guide for configuration and environment variables.
Default installation
npm install puppeteer
Then launch the bundled browser:
import puppeteer from 'puppeteer';
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
console.log(await page.title());
} finally {
await browser.close();
}
The install downloads browser binaries, so package installation requires the install step to be allowed and the environment to have access to the configured browser source. If a package manager blocks install scripts, follow the manual browser installation instructions in Puppeteer’s installation guide and confirm the browser is present in the configured cache.
Use puppeteer-core with an external browser
Choose puppeteer-core when your environment manages the browser separately, or when connecting to a remote browser. Provide an executable path or an appropriate installed-browser channel. Puppeteer guarantees compatibility only with its bundled browser; for a custom executable, the documentation recommends also specifying the browser type. Arbitrary browser versions are not guaranteed to work.
npm install puppeteer-core
import puppeteer from 'puppeteer-core';
const browser = await puppeteer.launch({
executablePath: '/usr/bin/google-chrome',
browser: 'chrome',
headless: true,
});
try {
const page = await browser.newPage();
await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
console.log(await page.title());
} finally {
await browser.close();
}
Replace the example path with the executable path on the target machine. The launch options reference documents executablePath and the browser option. A channel can be used when launching an installed Chrome channel supported by Puppeteer; consult the installed version’s launch options for valid values.
3. Pin browser versions and control installation
For reproducible builds, pin Puppeteer in your package lockfile and use the browser revision managed for that Puppeteer version. The installation guide also documents setting a browser build ID and cache directory when managing browser installation directly. Avoid relying on an unpinned system browser when repeatable automation matters.
The @puppeteer/browsers install options include the browser, build ID, platform, and cache directory. An optional expected SHA-256 hash makes installation fail if the downloaded archive does not match. If you omit it, installation proceeds without integrity verification. Custom providers are not officially supported; compatibility, testing, and maintenance are your responsibility. See the browser management API.
Use the official package’s browser installation workflow when possible. If you choose a custom provider, pin the build, record its source and hash in your build configuration, and test the exact executable with the Puppeteer version you deploy.
4. Choose the right headless mode
headless: true selects current headless Chrome. headless: 'shell' selects chrome-headless-shell. The shell does not completely match regular Chrome, though it may be chosen for performance when its reduced feature match is acceptable. Validate the specific download and page workflows in the mode you will run; do not assume both modes behave identically. See the official headless modes guide.
// Current headless Chrome
const browser = await puppeteer.launch({ headless: true });
// chrome-headless-shell
const shell = await puppeteer.launch({ headless: 'shell' });
5. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Page download does not start | Download behavior is not set to allow, the click did not target the download action, or the page requires authentication. | Set the policy and path before triggering the download. Confirm the page reached the expected state and the action works with the same session. |
| Download path error or no file appears | The path is missing for an allow policy, is relative to an unexpected working directory, or is not writable by the browser process. | Set an absolute downloadPath, create it first, and check permissions in the runtime/container. |
| File has an unexpected name | The server chose a filename, or allowAndName uses GUID names. |
Use the documented policy intentionally and add directory monitoring or application logic to map completed files. |
Could not find Chrome after installing puppeteer-core |
puppeteer-core does not download a browser. |
Install and manage a browser separately, then pass its executable path or supported channel. |
| Executable path does not exist | The path is for another OS, image, or user, or the browser was never installed. | Inspect the deployed environment and configure the actual executable location. Do not assume a developer machine’s path exists in CI. |
| Browser launches locally but fails in deployment | Different browser build, missing runtime dependencies, permissions, or incompatible Puppeteer/browser versions. | Pin package and browser versions, install required runtime dependencies for the deployment image, and test the deployed image. The bundled browser is the compatibility baseline. |
| Install succeeds but browser binary is absent | Install scripts may have been disabled by the package manager or the cache was redirected. | Check Puppeteer’s configured cache and environment. Use the official manual installation path if install scripts are blocked. |
| Unexpected difference in headless behavior | Current headless Chrome and chrome-headless-shell are not behavior-identical. | Run the workflow in the selected mode and switch modes only after validating the features you need. |
6. Performance, reliability, and cost
- Installation cost: Browser downloads add network and storage work to setup and CI cache misses. Cache the configured Puppeteer browser directory between builds when your CI system supports it, and keep the cache key tied to the Puppeteer/browser version.
- Runtime cost: Reuse a browser process for related work when isolation requirements permit, while using separate pages or contexts as appropriate. Close browsers in a
finallyblock so failures do not leave processes running. - Download reliability: Treat the download as an asynchronous operation. Use a deadline, identify the expected output, verify completion, and isolate concurrent jobs in separate directories.
- Reproducibility: Lock dependencies and browser builds. A system browser may simplify centrally managed installations, but it adds version coordination and carries the custom-executable compatibility caveat.
- Integrity: Where you manage browser archives directly, use the optional expected hash if you have a trusted expected SHA-256 value. Without it, the documented installer does not perform that integrity verification.
7. Or skip the browser setup
If your goal is to capture a page as an image or PDF rather than automate its download flow, ScreenshotNeo provides a website screenshot API and MCP server. A single GET request returns PNG, JPEG, WebP, or PDF. The call below saves a WebP screenshot; see the ScreenshotNeo API documentation for parameters and options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed, and response headers identify the page verdict and billing status. An MCP server lets AI agents use screenshot, page information, and PDF capture tools. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, with no card required.
FAQ
Does Puppeteer download Chrome when I install it?
The puppeteer package downloads its compatible browser by default. puppeteer-core does not.
Can Puppeteer use the Chrome already installed on my machine?
Yes. Launch with its executable path or a supported installed-browser channel. Compatibility is only guaranteed with Puppeteer’s bundled browser.
Should I use allow or allowAndName?
Use allow when the server-provided filename is useful. Choose allowAndName when GUID filenames suit your workflow and your code can track the associated download.
Is headless shell the same as headless Chrome?
No. Puppeteer documents that the modes do not completely match, so validate the one you deploy.


