ScreenshotNeo

BlogHow-to

How to Download Files with Selenium and Python

Learn how to download files with Selenium and Python, configure local browsers, wait for completion, transfer authenticated downloads, and retrieve files from Selenium Grid.

By the ScreenshotNeo team4 October 202613 min read

Short answer: If you need to test the file’s contents, use Selenium to reach the page and discover the download link, then fetch the file with an HTTP client such as Python’s requests. If the browser download interaction itself is what you are testing, configure that browser’s download directory and check for completion separately: Selenium does not expose browser download progress. With Selenium Grid, the file starts on the remote machine; enable managed downloads and retrieve it through the Remote WebDriver API.

This guide covers local Chrome, Firefox, and Edge setup; authenticated HTTP retrieval; completion checks; and remote Grid downloads. Browser preferences vary, so use the browser-specific configuration that matches the browser and driver versions in your environment. See Selenium’s official guidance on file downloads.

1. Choose the right download method

Method Use it when Where the file lands Key limitation
HTTP client after Selenium navigation You need to validate bytes, file format, or contents. The output path chosen by your Python test. Authentication, cookies, redirects, and streaming behavior are application-specific.
Browser download to a configured directory The browser’s download interaction is part of the scenario. The machine running the browser. A click starts a download; it does not confirm completion.
Grid managed download The browser runs remotely and the test client needs the file. Retrieved to a client-side directory through Selenium. Enable support on the Grid node and in the session; file listings are snapshots.

Selenium recommends the HTTP-client approach when the goal is to test a downloaded file. WebDriver does not provide an API for download progress, so a successful click alone is not evidence that the complete file exists.

2. Install Selenium and prepare a local output folder

Install Selenium and, for the HTTP method, Requests:

python -m pip install selenium requests

Create a known folder and make sure it exists before starting the browser:

from pathlib import Path

DOWNLOAD_DIR = Path("downloads").resolve()
DOWNLOAD_DIR.mkdir(parents=True, exist_ok=True)
print(DOWNLOAD_DIR)

The directory is on the machine that runs the browser. For a local browser, that is normally the machine running the test. For Remote WebDriver, it is on the remote machine. Selenium’s Remote WebDriver documentation explains this distinction.

3. Configure a browser download directory

Selenium does not define one universal preference dictionary for all browsers. Chrome, Edge, and Firefox have browser-specific options and behavior. Confirm the preferences against the Selenium binding and browser versions your project uses.

Chrome

This example sets Chrome’s download directory and disables the first-download prompt for an unattended local run:

from pathlib import Path
from selenium import webdriver

folder = Path("downloads").resolve()
folder.mkdir(parents=True, exist_ok=True)

options = webdriver.ChromeOptions()
options.add_experimental_option("prefs", {
    "download.default_directory": str(folder),
    "download.prompt_for_download": False,
    "download.directory_upgrade": True,
})

driver = webdriver.Chrome(options=options)
try:
    driver.get("https://example.com")
    # Locate and click the download control for your application.
finally:
    driver.quit()

These are Chrome preferences, not Selenium-wide settings. Selenium’s Python ChromeOptions API also exposes an enable_downloads property for supported use cases. Chrome and ChromeDriver major versions must match; see Selenium’s Chrome guidance.

Firefox

Firefox uses its own preference names and MIME handling. The example below configures a known download directory and instructs Firefox to save common file types without opening an external application. Adjust the MIME types for the file your application serves.

from pathlib import Path
from selenium import webdriver

folder = Path("downloads").resolve()
folder.mkdir(parents=True, exist_ok=True)

options = webdriver.FirefoxOptions()
options.set_preference("browser.download.folderList", 2)
options.set_preference("browser.download.dir", str(folder))
options.set_preference("browser.download.useDownloadDir", True)
options.set_preference(
    "browser.helperApps.neverAsk.saveToDisk",
    "application/pdf,text/csv,application/octet-stream",
)
options.set_preference("pdfjs.disabled", True)

driver = webdriver.Firefox(options=options)
try:
    driver.get("https://example.com")
    # Locate and click the download control for your application.
finally:
    driver.quit()

Firefox’s preferences and download behavior are distinct from Chrome’s. Consult the current Selenium Python Firefox Options API and Firefox guidance. Selenium’s documented Firefox minimum for Selenium 4 is Firefox 78; the guidance recommends the latest geckodriver.

Edge

Edge is Chromium-based, and commonly uses Chromium download preferences. Use Edge’s options class, and verify that the preferences are honored by the specific Edge version and environment you run:

from pathlib import Path
from selenium import webdriver

folder = Path("downloads").resolve()
folder.mkdir(parents=True, exist_ok=True)

options = webdriver.EdgeOptions()
options.add_experimental_option("prefs", {
    "download.default_directory": str(folder),
    "download.prompt_for_download": False,
    "download.directory_upgrade": True,
})

driver = webdriver.Edge(options=options)
try:
    driver.get("https://example.com")
    # Locate and click the download control for your application.
finally:
    driver.quit()

Do not assume every Chrome preference or behavior is guaranteed to work identically in Edge. Validate the configuration in the browser build used by your CI or workstation.

4. Wait for a browser download to finish

A browser-triggered download needs an explicit completion check. Avoid a fixed sleep: download time depends on file size, network speed, and server behavior. A practical local check is to poll for the expected file and reject common temporary partial-download suffixes.

import time
from pathlib import Path


def wait_for_download(folder: Path, filename: str, timeout: float = 60) -> Path:
    """Wait for a named download to exist and stop changing size."""
    target = folder / filename
    deadline = time.monotonic() + timeout
    previous_size = None
    stable_checks = 0

    while time.monotonic() < deadline:
        # Chrome commonly uses .crdownload; Firefox commonly uses .part.
        partials = list(folder.glob("*.crdownload")) + list(folder.glob("*.part"))
        if target.is_file() and not partials:
            size = target.stat().st_size
            if size == previous_size and size > 0:
                stable_checks += 1
                if stable_checks >= 2:
                    return target
            else:
                previous_size = size
                stable_checks = 0
        time.sleep(0.25)

    raise TimeoutError(f"Download {filename!r} did not complete in {timeout}s")

# Call after clicking the download control:
# downloaded = wait_for_download(Path("downloads").resolve(), "report.csv")

This is a filesystem heuristic, not a browser-provided completion signal. A temporary suffix can vary by browser, and a stable non-empty file does not prove the content is valid. Where possible, have the application expose a reliable completion signal, then validate the file’s expected size, type, checksum, or contents.

When file retrieval or contents are under test, Selenium can navigate through the UI while Requests handles the bytes. Transfer only the cookies and headers the application requires. The following is a runnable pattern; replace the example URL and link locator with the application’s actual values.

from pathlib import Path
from urllib.parse import urljoin

import requests
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait

PAGE_URL = "https://example.com/reports"
DOWNLOAD_DIR = Path("downloads").resolve()
DOWNLOAD_DIR.mkdir(parents=True, exist_ok=True)

# Configure login state here if the site requires it. The browser must be
# able to reach the page and expose the download link.
driver = webdriver.Chrome()
try:
    driver.get(PAGE_URL)

    link = WebDriverWait(driver, 20).until(
        EC.presence_of_element_located((By.CSS_SELECTOR, "a.download-link"))
    )
    download_url = urljoin(driver.current_url, link.get_attribute("href"))

    # Copy browser cookies into a Requests session. Some sites also require
    # selected headers, a CSRF token, or a particular referer; add only what
    # the application actually requires.
    session = requests.Session()
    for cookie in driver.get_cookies():
        session.cookies.set(
            cookie["name"], cookie["value"],
            domain=cookie.get("domain"), path=cookie.get("path", "/"),
        )

    response = session.get(download_url, stream=True, timeout=(10, 90))
    response.raise_for_status()

    output = DOWNLOAD_DIR / "report.csv"
    with output.open("wb") as file:
        for chunk in response.iter_content(chunk_size=1024 * 64):
            if chunk:
                file.write(chunk)

    if output.stat().st_size == 0:
        raise ValueError("The downloaded file is empty")
    print(f"Saved {output} ({output.stat().st_size} bytes)")
finally:
    driver.quit()

Use urljoin because the link may be relative. stream=True avoids loading the entire response into memory. The connect and read timeouts are separate values. Choose values appropriate for the application and expected file size.

Authentication and redirects

  • Copy cookies only after Selenium has reached the authenticated state. Cookies may be scoped to a domain and path; preserve those fields when transferring them.
  • Some applications require an authorization header, CSRF token, referer, or a short-lived signed URL in addition to cookies. Reproduce only the requirements observed for that application.
  • Requests follows redirects by default. If the final response is an HTML login page, inspect response.url, the status, and the response content type before treating it as a file.
  • For a one-time signed URL, request it promptly. Do not log credentials, cookie values, or signed query strings.
  • For large files, stream chunks to disk as in the example. Consider writing to a temporary filename and renaming it after validation, so a failed transfer is not mistaken for a completed file.

The cookie-transfer pattern is not universal: single sign-on, client certificates, browser-bound tokens, and application-specific download endpoints can need a different flow. Selenium’s official guidance recommends using an HTTP library but leaves these application details to the site and client.

6. Download files from Selenium Grid

With Remote WebDriver, the browser runs on a Grid node. Its normal download directory is remote, so a path on the test client will not directly contain the browser’s file. Selenium Grid’s managed-download feature can transfer session files back to the client.

  1. Start the Grid node or standalone server with managed downloads enabled, for example --enable-managed-downloads true.
  2. Request managed downloads for the session using the se:downloadsEnabled capability. In current Selenium Python options, check the browser options’ enable_downloads property and confirm the capability serialization for your Grid version.
  3. Trigger the download and wait until the application or browser-side completion condition indicates it has finished.
  4. List the files available to the session, then retrieve the selected file to a client-side directory.
from pathlib import Path
from selenium import webdriver
from selenium.webdriver.common.by import By

folder = Path("grid-downloads").resolve()
folder.mkdir(parents=True, exist_ok=True)

options = webdriver.ChromeOptions()
options.enable_downloads = True

driver = webdriver.Remote(
    command_executor="http://localhost:4444",
    options=options,
)
try:
    driver.get("https://example.com/reports")
    driver.find_element(By.CSS_SELECTOR, "a.download-link").click()

    # The list is an immediate snapshot. Poll or wait for an application
    # completion signal before assuming the desired file is available.
    files = driver.get_downloadable_files()
    if "report.csv" not in files:
        raise TimeoutError(f"report.csv is not ready; available files: {files}")

    driver.download_file("report.csv", str(folder))
    print(folder / "report.csv")
finally:
    driver.quit()

The code assumes managed downloads are enabled on the server and supported by the active Selenium, browser, and Grid versions. The current Python Remote WebDriver API documents get_downloadable_files(), download_file(file_name, target_directory), and delete_downloadable_files(). The Grid documentation lists Chrome, Firefox, and Edge support and explains that its file list is an immediate snapshot. Download storage is session-scoped and is cleaned up when the session ends or times out. See Grid CLI options and the Python Remote WebDriver API.

7. Validate the downloaded file

A successful HTTP response or a file appearing in a folder does not by itself prove the expected artifact was downloaded. Validate at least the checks that matter to your test:

  • Status: call raise_for_status() for HTTP retrieval and check the final URL when redirects are possible.
  • Non-empty output: reject an empty file if the expected artifact must contain data.
  • Format: inspect the expected signature or parse the file with its intended reader. Do not rely only on the filename extension or response Content-Type.
  • Contents: assert a meaningful value, header, row, or document property rather than only asserting that a path exists.
  • Integrity: when the application provides a checksum or expected size, compare it.

8. Performance, reliability, and cost

  • Use the lightest tool for the assertion. If the test concerns file bytes, an HTTP request after Selenium discovery avoids using the browser to transfer the file. Keep the browser when the UI interaction itself is under test.
  • Stream large responses. Write chunks to disk so memory use does not grow with the whole file size. Use bounded timeouts and handle interrupted transfers.
  • Make completion observable. Prefer an application completion signal over a fixed sleep. For local downloads, poll for the expected file and account for temporary partial files. For Grid, remember that the file listing is only a snapshot.
  • Keep sessions isolated. Use a unique directory per test or clean the directory before the run, so a leftover file cannot make a new download appear successful.
  • Protect credentials. Treat cookies, authorization headers, and signed URLs as secrets. Avoid writing them to logs or shared artifacts.
  • Control the environment. Pin compatible browser and driver versions in CI. Selenium’s Chrome guidance says Chrome and ChromeDriver major versions must match. Its Firefox guidance recommends the latest geckodriver and documents Firefox 78 as the Selenium 4 minimum.
  • Account for infrastructure cost. Browser sessions consume more setup and runtime resources than a direct HTTP transfer. Grid managed downloads also depend on Grid node configuration and session storage. No universal timing or cost figure applies; measure in your own environment.

The Selenium downloads page listed Python binding 4.49.0, released September 9, 2026, at the time of the research for this guide. Verify the current binding and browser compatibility for your environment in Selenium’s downloads and releases page.

9. Troubleshooting

Symptom Likely cause What to check or change
The click succeeds but no file is present yet. The download is still in progress, or the click only initiated it. Wait on a completion condition, check the configured folder, and account for browser-specific partial-file suffixes.
The file is saved somewhere unexpected. The configured path was relative, invalid, ignored, or applied to a different browser machine. Use an absolute path, create the directory first, inspect the browser options, and remember Remote WebDriver writes on the remote host.
A PDF opens in the browser instead of saving. The browser’s PDF handling or MIME preferences open it inline. Configure the selected browser’s PDF/download behavior and verify the response’s content disposition and MIME type.
Requests saves an HTML page instead of the file. The transferred session is unauthenticated, a redirect returned a login page, or a signed link expired. Inspect status, final URL, content type, and a safe preview of the response. Transfer the required session state or obtain a fresh URL.
Requests returns 401 or 403. The endpoint needs authentication, a CSRF token, a referer, or another application-specific condition. Compare the browser’s actual download request requirements. Add only the needed cookies or headers; do not assume cookies alone are sufficient.
Requests returns 404 for a link Selenium found. The link is relative, temporary, session-bound, or has expired. Resolve relative links with urljoin, use the current authenticated session, and reacquire short-lived links just before requesting.
The saved file is empty or truncated. The transfer was interrupted, an error page was saved, or the write was treated as complete too early. Check HTTP status before writing, stream with a read timeout, write to a temporary path, and validate size or content before renaming.
get_downloadable_files returns an empty list. The download is not finished, or managed downloads were not enabled on the node or session. Check the Grid startup flag and session capability, then wait for a completion signal and poll the snapshot.
Grid lists the file but retrieval fails. The session expired, the file was removed, or the browser/Grid versions do not support the workflow. Retrieve before quitting the session; verify managed-download support and inspect Grid logs and version compatibility.
Browser startup fails after adding download preferences. A preference is invalid for that browser/version, or browser and driver versions are incompatible. Reduce to browser-specific options, validate them against the browser’s current Selenium API reference, and align the browser/driver versions.

10. Or skip the browser setup

If you need a screenshot of the page around the download flow rather than the downloaded file itself, ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request returns a PNG, JPEG, WebP, or PDF. It does not download arbitrary files from a page; it captures the rendered page.

For example, use the API to capture a page as WebP. See the ScreenshotNeo documentation for options.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
  • Cookie banners are accepted and removed before capture; more than 60 known consent platforms, newsletter popups, and chat widgets can be removed, with each step configurable.
  • Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Responses include X-Page-Verdict and X-Billed headers.
  • An MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs.
  • 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000 screenshots.

Sign up free for 1,000 screenshots a month, with no card required.

11. FAQ

Yes. A download can be triggered by a button, script, form submission, or navigation. The test still needs to configure the browser if it expects browser-managed saving, or discover the resulting request and fetch it with an HTTP client.

Can I save a download to an arbitrary local path in a remote session?

A browser preference points to a path on the remote machine. Use Grid managed downloads to retrieve a file to the client, or configure a shared location that both the test and browser host can access.

Should I use Selenium for every file download test?

No. Use Selenium when browser interaction is part of the behavior under test. For assertions about the retrieved artifact, Selenium’s guidance favors discovering the link and using an HTTP client.

Does this approach work for every authentication scheme?

No single cookie-copy snippet covers every site. Browser-bound credentials, client certificates, expiring tokens, and application-specific headers can require a site-specific request flow.