How to Download Files with Browser Automation
Learn how to trigger, await, verify, and save browser downloads with Playwright, Selenium, Puppeteer, cURL, Python, and Node.js.
How do I download a file with browser automation? Treat the click and the byte transfer as two separate jobs. Your script must trigger the browser control, wait for the framework’s download completion signal, and copy the finished file to an application-owned path. In Playwright, register the download wait before clicking, await the resulting Download object, then call saveAs().
import { chromium } from 'playwright';
import path from 'node:path';
const browser = await chromium.launch();
const context = await browser.newContext({ acceptDownloads: true });
const page = await context.newPage();
await page.goto('https://example.com/reports');
const destination = path.resolve('downloads/report.pdf');
const downloadPromise = page.waitForEvent('download');
await page.getByRole('link', { name: 'Download report' }).click();
const download = await downloadPromise;
await download.saveAs(destination);
console.log(`Saved ${download.suggestedFilename()} to ${destination}`);
await browser.close();
This sequence follows the Playwright Download API. The listener is installed before the action, so fast downloads are not missed. saveAs() waits for the transfer to finish and persists it before the browser context is closed.
1. Decide whether you need a browser
If you already know the file URL and can supply the required cookies, bearer token, or other authorization, an HTTP client is usually simpler and easier to verify. Selenium’s official guidance recommends using Selenium to locate the link and obtain the relevant cookies, then using an HTTP library such as cURL to retrieve the file. A browser is necessary when JavaScript creates the URL, a click starts a POST or generated export, a consent step must be accepted, or the authenticated session cannot be reproduced with a direct request.
| Situation | Good approach |
|---|---|
| Stable public URL | cURL, Python requests, or Node fetch |
| File appears after a UI action | Playwright download event, then saveAs() |
| Login/session is browser-only | Use the browser for login, then transfer with the framework or an authenticated HTTP client |
| Need to assert transfer completion | Use a framework completion API or control the HTTP response directly |
2. Playwright: the complete download workflow
Install and run
npm install playwright
npx playwright install chromium
Save the first example as download.mjs and run node download.mjs. Create the destination directory before saving, or use fs.mkdir() in the script.
import { chromium } from 'playwright';
import { mkdir } from 'node:fs/promises';
import path from 'node:path';
const browser = await chromium.launch();
const context = await browser.newContext({ acceptDownloads: true });
const page = await context.newPage();
try {
await mkdir('downloads', { recursive: true });
await page.goto('https://example.com/reports', { waitUntil: 'domcontentloaded' });
const downloadPromise = page.waitForEvent('download');
await page.getByRole('button', { name: /export/i }).click();
const download = await downloadPromise;
const filename = download.suggestedFilename();
const safeName = filename.replace(/[^a-zA-Z0-9._-]/g, '_');
const destination = path.resolve('downloads', safeName);
await download.saveAs(destination);
const failure = await download.failure();
if (failure) throw new Error(`Download failed: ${failure}`);
console.log(destination);
} finally {
await context.close();
await browser.close();
}
Why the order matters
page.waitForEvent('download')creates a pending promise.- The click triggers the transfer.
- The promise resolves with the
Downloadobject. saveAs()copies the completed file to a controlled path.
Do not start waiting after the click. A small or cached file can emit the event before your listener exists.
Useful Playwright download properties
download.url(): the URL used by the transfer.download.suggestedFilename(): a name derived fromContent-Dispositionor the HTMLdownloadattribute. It is a suggestion and can differ between browsers.download.path(): waits for completion and returns a temporary path. The API documentation says this throws when the browser is connected remotely.download.createReadStream(): exposes the payload stream when you need to upload or hash it without choosing a final path.download.saveAs(path): copies the file and waits for completion.download.failure(): reports a failed or canceled transfer.
Context downloads are temporary by default and are deleted when the context closes. Persist required files with saveAs() before cleanup. You can also configure a browser launch downloadsPath, but an explicit, per-job destination is easier to reason about in workers.
Waiting for a generated export
const downloadPromise = page.waitForEvent('download');
await page.getByRole('button', { name: 'Create CSV' }).click();
const download = await downloadPromise;
await download.saveAs('/var/tmp/jobs/job-123/export.csv');
If the button first starts a server-side job, wait for the job status or download link to appear, then install the download listener immediately before the final click. Prefer a selector wait, response wait, or framework completion method over an arbitrary sleep.
3. Python Playwright
pip install playwright
playwright install chromium
from pathlib import Path
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch()
context = browser.new_context(accept_downloads=True)
page = context.new_page()
page.goto("https://example.com/reports")
with page.expect_download() as download_info:
page.get_by_role("link", name="Download report").click()
download = download_info.value
destination_dir = Path("downloads")
destination_dir.mkdir(parents=True, exist_ok=True)
safe_name = "".join(c if c.isalnum() or c in ".-_" else "_" for c in download.suggested_filename)
destination = destination_dir / safe_name
download.save_as(str(destination))
if download.failure():
raise RuntimeError(download.failure())
print(destination)
context.close()
browser.close()
Use the asynchronous API when your application already runs an asyncio event loop. The lifecycle is the same: expect the download before the action, await it, and save it.
4. Selenium: use the browser for context, HTTP for bytes
Selenium’s file-download guidance says its API does not expose download progress, making it less suitable for tests that must verify downloaded files. The documented pattern is to find the link and required cookies with Selenium, then use an HTTP library such as cURL.
from pathlib import Path
import requests
from selenium import webdriver
from selenium.webdriver.common.by import By
browser = webdriver.Chrome()
try:
browser.get("https://example.com/reports")
href = browser.find_element(By.CSS_SELECTOR, "a[data-download]").get_attribute("href")
cookies = {cookie["name"]: cookie["value"] for cookie in browser.get_cookies()}
response = requests.get(href, cookies=cookies, stream=True, timeout=90)
response.raise_for_status()
destination = Path("downloads/report.pdf")
destination.parent.mkdir(parents=True, exist_ok=True)
with destination.open("wb") as file:
for chunk in response.iter_content(chunk_size=1024 * 1024):
if chunk:
file.write(chunk)
finally:
browser.quit()
Keep cookies and authorization scoped to the intended origin. Validate the final response status, content type, size, and file signature before handing the file to another system.
5. Puppeteer: know the current limitation
The current official Puppeteer Files guide states: “Currently, Puppeteer does not offer a way to handle file downloads in a programmatic way.” Do not copy Playwright’s waitForEvent('download') API into Puppeteer code. Puppeteer and Chromium versions can expose browser or protocol-specific mechanisms, but those are version-dependent; check the exact official documentation for the version you deploy before selecting one.
If you control a stable file URL, use Node’s HTTP client instead of trying to treat a Puppeteer click as a completed download:
import { createWriteStream } from 'node:fs';
import { mkdir } from 'node:fs/promises';
import { pipeline } from 'node:stream/promises';
await mkdir('downloads', { recursive: true });
const response = await fetch('https://example.com/report.pdf');
if (!response.ok || !response.body) throw new Error(`HTTP ${response.status}`);
await pipeline(response.body, createWriteStream('downloads/report.pdf'));
6. Direct HTTP downloads with cURL, Python, and Node.js
cURL
curl -fL --retry 3 --connect-timeout 10 --max-time 90 \
-o downloads/report.pdf \
https://example.com/report.pdf
For a browser session, export the relevant cookies and pass them with --cookie, or send an authorization header. Avoid writing session secrets into shell history or logs.
Python requests
import requests
with requests.get(
"https://example.com/report.pdf",
stream=True,
timeout=(10, 90),
) as response:
response.raise_for_status()
with open("downloads/report.pdf", "wb") as file:
for chunk in response.iter_content(1024 * 1024):
if chunk:
file.write(chunk)
Node.js fetch
import { createWriteStream } from 'node:fs';
import { mkdir } from 'node:fs/promises';
import { pipeline } from 'node:stream/promises';
await mkdir('downloads', { recursive: true });
const response = await fetch('https://example.com/report.pdf');
if (!response.ok || !response.body) throw new Error(`HTTP ${response.status}`);
await pipeline(response.body, createWriteStream('downloads/report.pdf'));
7. Safe filenames, integrity, and cleanup
- Generate a unique directory or filename per job to avoid collisions between workers.
- Treat a suggested filename as untrusted input. Remove path separators and control characters.
- Check status,
Content-Type, expected size limits, and a magic-byte signature where practical. A server error page can otherwise be saved as.pdf. - Use finite navigation, download, and HTTP deadlines. Cancel and delete partial files after timeout.
- Hash the completed file when downstream systems require integrity or deduplication.
- Delete temporary browser contexts and abandoned job directories.
- Never forward cookies or authorization headers to a different origin unless that transfer is explicitly required.
8. Common errors and fixes
| Error | Likely cause | Fix |
|---|---|---|
| Download event timeout | The click did not trigger a download, the listener was installed too late, or a popup blocked the action. | Register the wait before clicking; verify the selector and inspect whether the action opens a new page or displays an error. |
| File disappears after the script exits | It remained in Playwright’s temporary context directory. | Call saveAs() before closing the context. |
| Saved file is HTML or JSON | Authentication expired, a redirect reached a login page, or the server returned an error. | Check status, final URL, content type, and a short body preview before accepting the file. |
| Suggested filename creates an unsafe path | The server supplied separators or unexpected characters. | Sanitize the name and resolve it beneath a fixed destination directory. |
download.path() fails remotely |
The browser runs on another machine. | Use saveAs() to a path accessible to the client, or stream the file through your remote execution system. |
| Selenium test cannot tell when transfer finished | Selenium’s API does not expose download progress. | Use the documented HTTP-client approach or monitor the destination directory with an application-level deadline. |
| Puppeteer example has no download event | Puppeteer’s official Files guide does not provide a programmatic download API. | Use a direct HTTP request when possible and consult the version-specific Puppeteer/browser protocol documentation for browser-controlled cases. |
| Zero-byte or partial file | The process ended before the transfer completed or the connection failed. | Await saveAs() or the HTTP stream pipeline, check failure state, and remove incomplete output on errors. |
9. Performance, reliability, and cost
- Prefer HTTP for known URLs. It avoids browser startup, rendering, and download-directory management.
- Reuse browser contexts carefully. Reusing a browser reduces startup work, while separate contexts isolate cookies and downloads between jobs.
- Stream large files. Do not buffer multi-gigabyte responses in memory; write chunks to disk or object storage.
- Bound concurrency. Too many simultaneous browsers or transfers can exhaust CPU, memory, sockets, or disk.
- Retry selectively. Retry transient connection failures and 5xx responses, but do not blindly retry authentication failures or invalid URLs.
- Make jobs idempotent. Use a job ID and deterministic destination so a retry cannot overwrite an unrelated file.
- Measure the right completion point. A click means the browser started an action; a resolved download or closed HTTP stream means bytes were persisted.
10. Or skip the browser setup
If your goal is a clean visual capture of a page rather than downloading an attachment, ScreenshotNeo provides a single request for PNG, JPEG, WebP, or PDF output. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
import { writeFile } from 'node:fs/promises';
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
await writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and each response reports its page verdict and billing status in X-Page-Verdict and X-Billed headers. Its MCP server lets Claude, Cursor, and other MCP clients use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots.
Create your free ScreenshotNeo account.
11. FAQ
How do I wait for a file download to finish?
Use the framework’s completion primitive. In Playwright, await the Download event and then await download.saveAs() or download.path(). Do not rely on a fixed sleep.
Can I choose the browser’s download folder?
Playwright supports a browser launch downloadsPath, but downloads are still temporary in the context lifecycle. Copy important files with saveAs() to an application-owned location.
Why is my downloaded filename different between browsers?
The suggested name can come from the response’s Content-Disposition header or the HTML download attribute, and browsers may interpret those hints differently. Sanitize and assign your own final name.
Should I use Selenium or Playwright?
Choose Playwright when you need a first-class download event and an awaited save operation. Selenium’s published guidance favors using Selenium for browser context and an HTTP client for the actual file transfer.
Can a remote browser save directly to my local machine?
Not necessarily. A remote browser’s filesystem is different from the client’s. Transfer the bytes through the remote execution layer or use a save operation supported by that deployment.


