ScreenshotNeo

BlogHow-to

How to Upload and Download Files with Puppeteer

Upload files through a page’s file input or chooser. For downloads, configure Chrome’s save policy and handle file completion and validation in your Node.js script.

By the ScreenshotNeo team4 October 202610 min read

Puppeteer can upload files directly through an <input type="file">, or respond to a page’s file chooser with local file paths. Downloads are different: Puppeteer can configure Chrome’s download policy and destination, but its Files guide says it does not handle downloaded files programmatically. Your Node.js script must separately detect completion, validate the result, and consume or clean up the file.

The examples below target Node.js with Puppeteer 25.12.0 and Chrome. Check the API reference for your installed version when using chooser or download behavior APIs, since the current references show slightly different version numbers for these methods.

1. Install Puppeteer and prepare a local file

Install Puppeteer in a Node.js project:

npm install puppeteer

The puppeteer package downloads a compatible Chrome build by default. puppeteer-core does not download Chrome and is intended for a browser you manage or connect to remotely. See the official installation guide.

Paths are paths on the machine where Chrome runs. If your script connects to remote Chrome, a path on your laptop is not automatically available to that browser; stage the file where the browser can access it and use an absolute path.

2. Upload through a standard file input

When the page exposes a file input, wait for it and call uploadFile. This example is a complete script against a replaceable form URL and selector:

import puppeteer from 'puppeteer';
import path from 'node:path';

const formUrl = 'https://example.com/upload';
const filePath = path.resolve('fixtures/report.pdf');

const browser = await puppeteer.launch({ headless: true });
try {
  const page = await browser.newPage();
  await page.goto(formUrl, { waitUntil: 'domcontentloaded' });

  const input = await page.waitForSelector('input[type="file"]', {
    timeout: 15_000,
  });
  if (!input) throw new Error('File input did not appear');

  await input.uploadFile(filePath);

  // Select the form's submit control and wait for the app's own success state.
  await page.locator('button[type="submit"]').click();
  await page.waitForSelector('[data-upload-status="success"]', {
    timeout: 30_000,
  });
} finally {
  await browser.close();
}

The selector for success is application-specific: replace it with a confirmation element, URL change, or other signal that means the server accepted the upload. uploadFile places the file in the input; it does not submit the form or prove that the server stored the file. Puppeteer documents this direct route in its Files guide.

Multiple files

Pass multiple paths as an array when the input allows multiple files:

await input.uploadFile(
  path.resolve('fixtures/first.pdf'),
  path.resolve('fixtures/second.pdf'),
);

If the input is not configured with multiple, the page may accept only one selection. Check the form’s actual behavior and assert the resulting filenames or upload status.

3. Upload through a page-triggered file chooser

Some pages hide the file input behind an “Upload” button. Register waitForFileChooser() before clicking; waiting after the click can miss the chooser event:

import puppeteer from 'puppeteer';
import path from 'node:path';

const browser = await puppeteer.launch({ headless: true });
try {
  const page = await browser.newPage();
  await page.goto('https://example.com/upload', {
    waitUntil: 'domcontentloaded',
  });

  const [chooser] = await Promise.all([
    page.waitForFileChooser({ timeout: 15_000 }),
    page.locator('#upload-file-button').click(),
  ]);

  const paths = [path.resolve('fixtures/report.pdf')];
  if (chooser.isMultiple()) {
    paths.push(path.resolve('fixtures/appendix.pdf'));
  }
  await chooser.accept(paths);

  await page.waitForSelector('[data-upload-status="success"]', {
    timeout: 30_000,
  });
} finally {
  await browser.close();
}

The chooser API’s accept(paths) does not validate that paths exist. Relative paths resolve from the Node process’s current working directory; use absolute paths, especially with remote Chrome. Only one chooser can be open at a time, and each must be accepted or canceled before another chooser can appear. The official FileChooser reference documents these limits.

waitForFileChooser() must be called before the chooser opens. It does not intercept DOM picker APIs such as window.showOpenFilePicker(). In headful mode, Puppeteer’s chooser wait suppresses the native picker UI rather than displaying it to a person. See the Page.waitForFileChooser reference.

4. Configure Chrome to allow downloads

Download configuration controls browser policy and destination. It does not give your script a downloaded file object or prove the download completed. The API defines deny, allow, allowAndName, and default. A downloadPath is required with allow and allowAndName; allowAndName uses download GUIDs as filenames. See the DownloadBehavior reference.

For Chrome through Node.js, Puppeteer’s current ConnectOptions includes downloadBehavior. This runnable setup connects to an existing Chrome DevTools endpoint supplied in BROWSER_WS_ENDPOINT:

import puppeteer from 'puppeteer';
import { mkdir } from 'node:fs/promises';
import path from 'node:path';

const endpoint = process.env.BROWSER_WS_ENDPOINT;
if (!endpoint) throw new Error('Set BROWSER_WS_ENDPOINT');

const downloadDir = path.resolve('downloads');
await mkdir(downloadDir, { recursive: true });

const browser = await puppeteer.connect({
  browserWSEndpoint: endpoint,
  downloadBehavior: {
    policy: 'allow',
    downloadPath: downloadDir,
  },
});
try {
  const page = await browser.newPage();
  await page.goto('https://example.com/export', {
    waitUntil: 'domcontentloaded',
  });
  await page.locator('#download-report').click();
  // Detect and validate the resulting file separately (next section).
} finally {
  await browser.disconnect();
}

Start or obtain the Chrome endpoint according to your environment. The cited option is Chrome-specific; the API reference marks its use with puppeteer.connect() experimental. Puppeteer’s official ConnectOptions reference describes the setting. A browser download can also have a server-controlled filename, so do not assume the final name from the link text.

Policy Effect Path requirement
default Use the browser’s default behavior. No explicit path requirement documented for this policy.
deny Deny download requests. No explicit path requirement documented for this policy.
allow Allow downloads. downloadPath required.
allowAndName Allow downloads and name files using download GUIDs. downloadPath required.

5. Wait for and validate a downloaded file

The download policy is only browser setup. Your application needs its own completion and validation strategy. A basic filesystem poll can work for a controlled flow if you know the expected filename and Chrome’s naming behavior. The helper below waits for a nonempty file whose size remains unchanged for two polls; it also times out and surfaces a useful error. Treat this as an application-level heuristic and verify it against your Chrome version and download flow.

import { access, readdir, stat } from 'node:fs/promises';
import path from 'node:path';

const sleep = ms => new Promise(resolve => setTimeout(resolve, ms));

async function waitForStableFile(directory, {
  timeoutMs = 60_000,
  stablePollsRequired = 2,
} = {}) {
  const deadline = Date.now() + timeoutMs;
  let previous = new Map();
  let stablePolls = 0;

  while (Date.now() < deadline) {
    const names = await readdir(directory);
    const candidates = names.filter(name =>
      !name.endsWith('.crdownload') && !name.endsWith('.tmp')
    );
    const current = new Map();
    for (const name of candidates) {
      const file = path.join(directory, name);
      try {
        const info = await stat(file);
        if (info.isFile() && info.size > 0) current.set(name, info.size);
      } catch {
        // The browser may still be creating or renaming this file.
      }
    }

    if (current.size > 0 && [...current].every(([name, size]) =>
      previous.get(name) === size
    )) {
      stablePolls += 1;
      if (stablePolls >= stablePollsRequired) {
        const [name] = current.keys();
        return path.join(directory, name);
      }
    } else {
      stablePolls = 0;
    }
    previous = current;
    await sleep(500);
  }
  throw new Error(`No stable completed download appeared in ${directory}`);
}

const downloadedPath = await waitForStableFile(downloadDir);
const result = await stat(downloadedPath);
if (result.size === 0) throw new Error('Downloaded file is empty');
console.log(`Download saved: ${downloadedPath} (${result.size} bytes)`);

This polling helper is not a Puppeteer download API. It can mistake an unrelated file in a shared directory for the target, and a stable size alone does not prove the content is valid. For reliable automation, use a fresh dedicated directory per job, establish the expected file type/name where possible, validate file contents with the appropriate parser or checksum, and clean up in a finally block. If exact download completion events or naming are important, confirm the supported mechanism for your precise Puppeteer and Chrome versions before relying on it.

6. cURL, Python, and Node.js alternatives for downloading

Puppeteer is useful when the download requires browser state or interaction. If the endpoint can be fetched directly, use an HTTP client instead. These examples assume you already know the authorized download URL; they do not bypass login, access controls, or a site’s terms.

cURL

curl --fail --location --output report.pdf 'https://example.com/files/report.pdf'

Python

import requests

url = 'https://example.com/files/report.pdf'
with requests.get(url, stream=True, timeout=(10, 90)) as response:
    response.raise_for_status()
    with open('report.pdf', 'wb') as output:
        for chunk in response.iter_content(chunk_size=1024 * 64):
            if chunk:
                output.write(chunk)

Node.js

import { createWriteStream } from 'node:fs';
import { Readable } from 'node:stream';
import { pipeline } from 'node:stream/promises';

const response = await fetch('https://example.com/files/report.pdf');
if (!response.ok || !response.body) {
  throw new Error(`Download failed: HTTP ${response.status}`);
}
await pipeline(Readable.fromWeb(response.body), createWriteStream('report.pdf'));

Direct HTTP downloads may need the same authentication cookies or headers as the browser session. Do not copy secrets into logs, and make sure the endpoint and credentials are permitted for your use.

7. Troubleshooting uploads and downloads

Symptom Likely cause Fix
waitForSelector times out for the file input The input is rendered later, inside a frame, or the page uses a chooser button instead. Wait for the page’s actual state; inspect whether the input is in an iframe; use chooser handling only when a click launches a supported file chooser.
Chooser wait times out The click did not trigger a chooser, or the wait was registered after the click. Register waitForFileChooser() before the click using Promise.all; confirm the selector triggers a native chooser.
File is missing even though accept() resolved accept() does not check that the path exists; a relative path may resolve from an unexpected working directory. Resolve the path with path.resolve, check it with Node’s filesystem APIs, and ensure the browser host can access it.
Remote browser cannot upload a local file The path belongs to the Node machine, not the remote Chrome host. Stage the file on the browser host or use an upload workflow supported by that remote environment; pass its absolute path.
window.showOpenFilePicker() is not intercepted Puppeteer’s chooser interception does not support that DOM API. Use a page flow exposing a file input if available, or implement a supported application-level route.
Second chooser never appears An earlier chooser remains unresolved. Accept or cancel each chooser before opening another.
Download is blocked or missing Chrome policy is default/deny, the download path is absent, or the browser cannot write there. Set an allowed policy and a writable absolute path; ensure the directory exists and is on the browser host.
Script returns before the file is ready Click completion is not download completion. Wait separately for the file, enforce a timeout, then validate size and contents.
Installed Puppeteer cannot find Chrome Package install scripts may have been blocked, skipping browser download. Follow the official installation guide to install the compatible browser or explicitly configure the managed browser path.

8. Performance, reliability, and cost

  • Keep the browser alive across related tasks. Launching Chrome is a significant setup step; reuse a browser for a batch when isolation requirements allow, and always close or disconnect it in cleanup code.
  • Bound waits. Set timeouts for selector, chooser, navigation, and download waits. A missing page state should fail clearly rather than hang indefinitely.
  • Isolate download directories. A per-job directory makes it easier to identify the right file, avoid collisions, and clean up safely.
  • Limit concurrency deliberately. Parallel pages and large files consume browser, memory, disk, and network resources. Begin with a bounded worker count and observe your own workload.
  • Plan for retries carefully. Retrying an upload or download may duplicate a server-side action. Retry only when the operation is idempotent or the application can detect an already completed result.
  • Cost depends on your runtime. Puppeteer itself is an open-source library, but the machines, browser hosting, storage, bandwidth, and engineering time used to run it can have costs. The official installation guide notes that the default install downloads Chrome binaries.

Or skip the browser setup

If your task is to capture a web page as an image or PDF rather than upload or download a file, ScreenshotNeo provides a one-request screenshot API and an MCP server for AI agents. It removes cookie banners, popups, and chat widgets before the shot. Bot checks, blank pages, and failed loads are never billed. AI agents can take screenshots through its MCP server. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

See the ScreenshotNeo API documentation for options and setup, then sign up free for 1,000 screenshots a month with no card.

FAQ

Can Puppeteer upload a file without opening a visible picker?

Yes. For a standard file input, call uploadFile with a path. For a supported page-triggered chooser, intercept it with waitForFileChooser() and accept paths.

Can Puppeteer return the downloaded file’s bytes?

The current Files guide says Puppeteer does not offer programmatic download handling. Configure Chrome to save the file, then use Node.js filesystem code or an HTTP client to read and validate it.

Does waitForFileChooser() work with every file picker?

No. The API reference specifically excludes dialogs triggered through DOM APIs such as window.showOpenFilePicker().

Which machine needs access to the file path?

The machine running Chrome. For remote Chrome, make the upload file available to that host and use an absolute path.