Automating Browsers with Python
Learn how to automate browsers with Python using Playwright or Selenium, with setup, runnable scripts, debugging, CI guidance, and screenshot options.
Direct answer: use Playwright when you want a modern Python API for Chromium, Firefox, and WebKit with either synchronous or asynchronous code. Use Selenium when your project already uses WebDriver or needs its established browser and grid workflow. Both are valid choices; the right one depends on browser and operating-system coverage, async needs, test tooling, and existing infrastructure.
What browser automation means in Python
Browser automation drives a real browser (or a browser engine) through code. A script can open a URL, find elements, click, type, upload files, wait for navigation, read page content, take screenshots, and generate test assertions. Typical uses include end-to-end tests, smoke checks, data-entry workflows, regression testing, and collecting rendered output.
Automation is different from an HTTP client. Requests made with requests do not execute JavaScript or reproduce browser rendering. Choose browser automation when the page behavior, DOM, cookies, or visual result matters.
Playwright or Selenium?
| Question | Playwright | Selenium |
|---|---|---|
| Python API | Sync and async interfaces | WebDriver bindings |
| Browser engines | Chromium, Firefox, WebKit; branded Chrome and Edge channels are documented | Chrome, Edge, Firefox, Safari, WebKitGTK, and WPEWebKit in the current Python API |
| Best starting point | New end-to-end tests or browser workflows | Existing WebDriver code, grid, or team conventions |
| Browser installation | Install the Python package and then browser binaries with the Playwright CLI | Modern Selenium uses Selenium Manager for driver setup on most supported platforms |
| Async integration | First-class async API for asyncio |
Commonly synchronous WebDriver calls; confirm the current API for your integration |
Playwright’s documentation says it “was created specifically to accommodate the needs of end-to-end testing.” The same documentation describes it as a general-purpose tool for web applications with both sync and async Python APIs. Selenium is a strong fit when a project already depends on WebDriver protocols, drivers, or remote browser infrastructure. The official sources do not establish a universal speed or reliability winner.
Install Playwright
- Create and activate a virtual environment.
- Install the Python package.
- Install the browser binaries with the Playwright CLI.
python -m venv .venv
# macOS/Linux
source .venv/bin/activate
# Windows PowerShell: .venv\\Scripts\\Activate.ps1
python -m pip install --upgrade pip
pip install playwright
playwright install
Playwright browser versions track library releases. After upgrading Playwright, run playwright install again when the required browser revision changes. See the library installation guide and browser management documentation for deployment details.
Your first Playwright script (synchronous)
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page(viewport={"width": 1280, "height": 800})
page.goto("https://example.com", wait_until="domcontentloaded")
print(page.title())
print(page.locator("h1").inner_text())
page.screenshot(path="example.png", full_page=True)
browser.close()
Use locators such as get_by_role, get_by_label, and get_by_text when possible. CSS and XPath selectors are useful for stable technical hooks, but avoid selectors tied to generated class names.
Async Playwright with asyncio
import asyncio
from playwright.async_api import async_playwright
async def main():
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True)
page = await browser.new_page()
await page.goto("https://example.com", wait_until="networkidle")
print(await page.title())
await page.screenshot(path="async-example.webp", full_page=True)
await browser.close()
if __name__ == "__main__":
asyncio.run(main())
Prefer the async API when your application already uses asyncio or runs many independent browser tasks under an async scheduler. Do not call synchronous Playwright APIs from an active event loop.
Playwright end-to-end tests with pytest
For pytest-based end-to-end tests, Playwright recommends its pytest plugin.
pip install pytest-playwright
playwright install
# test_home.py
def test_homepage(page):
page.goto("https://example.com")
assert page.get_by_role("heading").first.is_visible()
Run it with pytest. Keep test data isolated, use deterministic URLs, and capture a trace or screenshot when a failure needs investigation. Consult the current pytest documentation for fixture and CI options.
Install and use Selenium
python -m venv .venv
source .venv/bin/activate
pip install selenium
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
options = webdriver.ChromeOptions()
options.add_argument("--headless=new")
options.add_argument("--window-size=1280,800")
driver = webdriver.Chrome(options=options)
try:
driver.get("https://example.com")
heading = WebDriverWait(driver, 15).until(
EC.visibility_of_element_located((By.TAG_NAME, "h1"))
)
print(driver.title)
print(heading.text)
driver.save_screenshot("selenium.png")
finally:
driver.quit()
Current Selenium documentation lists Python 3.10+ and browser support including Chrome, Edge, Firefox, Safari, WebKitGTK, and WPEWebKit. Selenium Manager handles driver installation in modern Selenium versions for most supported browser and platform combinations. A driver still connects WebDriver to the browser you select; enterprise policies, remote grids, and unusual browser builds may need additional configuration.
Common browser actions
Navigation and waiting
page.goto("https://app.example/login", wait_until="domcontentloaded")
page.get_by_label("Email").fill("person@example.com")
page.get_by_label("Password").fill("secret")
page.get_by_role("button", name="Sign in").click()
page.wait_for_url("**/dashboard")
page.get_by_role("heading", name="Dashboard").wait_for()
Wait for a meaningful condition such as a URL, selector, or response. Fixed sleeps are a last resort for animations or third-party widgets.
Files, cookies, and JavaScript
page.get_by_label("Upload").set_input_files("report.pdf")
context = browser.new_context(storage_state="auth.json")
page = context.new_page()
page.evaluate("document.body.dataset.automated = 'true'")
Save authenticated state only in protected CI storage. Never commit cookies, tokens, or passwords.
Capturing one element or the full page
page.locator("#invoice").screenshot(path="invoice.png")
page.screenshot(path="full.png", full_page=True)
Browser configuration checklist
- Headless mode: use headless in CI; run headed locally when diagnosing layout or focus issues.
- Viewport and device: set a known viewport or use a documented device preset.
- Timeouts: set a sensible action and navigation timeout; avoid unbounded waits.
- Locale, timezone, and permissions: configure them when formatting or geolocation affects behavior.
- Proxy and certificates: configure explicitly for corporate networks and test environments.
- Browser channel: use Playwright’s documented Chrome or Edge channels only when the branded browser is required; enterprise policies can change behavior.
- Network control: block analytics or stub unstable third-party calls when the test does not need them.
- Artifacts: retain screenshots, traces, console logs, and network logs only for failed or sampled runs.
Reliability and performance
- Reuse a browser process and create isolated contexts or sessions instead of launching a new browser for every page.
- Keep selectors semantic and stable. Add application-level test IDs for controls that have no reliable accessible name.
- Wait on observable state, not arbitrary delays. Account for lazy loading, service workers, WebSockets, and animations.
- Limit concurrency to what the host can support. More workers increase CPU, memory, file-descriptor, and network pressure.
- Pin Python and library versions in CI, then update browser binaries deliberately.
- Use retries only for known transient infrastructure failures; retries can hide real product defects.
- For long pages, expect full-page screenshots to load lazy images and consume more memory than viewport captures.
There is no benchmark in the reviewed documentation that proves one library is universally faster. Measure your own pages, browser matrix, and CI hardware if throughput is a requirement.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
Executable doesn't exist in Playwright |
Browser binaries were not installed or do not match the package. | Run playwright install in the same environment and check the package version. |
| Browser starts locally but fails in CI | Missing system dependencies, sandbox restrictions, display server, or different architecture. | Follow the Playwright or Selenium deployment guide for the CI image; use headless mode and install required dependencies. |
| Timeout waiting for a locator | Wrong selector, delayed application state, iframe, popup, or a blocked request. | Inspect the DOM, target the correct frame, wait for a meaningful state, and capture a trace or screenshot. |
| Element is covered or not clickable | Consent banner, modal, animation, sticky header, or overlay. | Handle the overlay, wait for it to disappear, scroll the element into view, or click the intended control explicitly. |
| Selenium driver error | Browser, driver, Selenium, or policy mismatch. | Upgrade Selenium, let Selenium Manager resolve the driver, and verify the installed browser and enterprise policy. |
| Intermittent blank or incomplete screenshots | Capture occurred before fonts, images, or client rendering finished. | Wait for a selector or network condition, disable animations where appropriate, and ensure lazy content is triggered. |
| Authentication disappears between tests | Each test uses a fresh context or profile. | Persist storage state securely or perform a deterministic login fixture. |
When a hosted screenshot API is simpler
If your goal is a rendered image or PDF rather than interaction, a hosted API can remove browser installation and maintenance from your application. ScreenshotNeo is the first service to try: it produces clean shots, bills only clean shots, and its paid plans start at $5.
Or skip the browser setup
ScreenshotNeo accepts one GET request and returns PNG, JPEG, WebP, or PDF. Before capture it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing result.
See the ScreenshotNeo API documentation for all options.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" \\
-d access_key=YOUR_API_KEY \\
--data-urlencode url=https://stripe.com \\
-o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const data = Buffer.from(await res.arrayBuffer());
require('fs').writeFileSync('shot.webp', data);
ScreenshotNeo also supports full-page capture with lazy images loaded, CSS-element capture, dark mode, device presets or custom viewports, retina scale, PDF paper and margin settings, custom CSS and JavaScript, clicks, selector waits, delays, network-idle waits, request blocking, headers, cookies, user agents, Authorization, timezone, geolocation, transparent backgrounds, resizing, selectable-TTL caching, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification. Its parameter names are compatible with those used by other screenshot APIs.
An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Create a free ScreenshotNeo account.
Cost and operational notes
- Self-hosted Playwright or Selenium costs the compute, storage, browser updates, and CI maintenance required to run them.
- A hosted API shifts browser operations to a per-capture plan. Cache repeated URLs where freshness allows it.
- For ScreenshotNeo, only clean shots are billed; failed loads, bot checks, blank pages, timeouts, and cache hits are free. Inspect
X-Page-VerdictandX-Billedin responses when reconciling usage. - Use asynchronous jobs and signed webhooks for long captures, and bulk capture for batches rather than opening hundreds of independent connections.
FAQ
Can Python browser automation run without a visible desktop?
Yes. Playwright and Selenium both support headless operation. CI images still need the browser and system dependencies required by the selected engine.
Should I learn Playwright before Selenium?
Choose Playwright for a new project that benefits from its sync or async API and multi-engine workflow. Start with Selenium when your organization already standardizes on WebDriver or a Selenium Grid.
Does browser automation replace API tests?
No. API tests are usually simpler for service contracts; browser tests cover rendered behavior, integration, and user-visible flows. Use each at the layer it verifies.
Can ScreenshotNeo interact with a page like Playwright?
ScreenshotNeo is a capture API with options for waits, clicks, custom JavaScript, headers, cookies, and related capture controls. Use Playwright or Selenium when you need a long interactive session with arbitrary application logic.


