ScreenshotNeo

BlogHow-to

Browser Automation with Python

Learn browser automation in Python with Playwright and Selenium, including waits, drivers, CI, troubleshooting, and screenshot workflows.

By the ScreenshotNeo team1 October 20268 min read

Python browser automation means driving a real browser from code to navigate pages, locate elements, click and type, wait for asynchronous content, collect data, run tests, or capture screenshots. The two main choices are Playwright and Selenium WebDriver.

Choose Playwright when you want a bundled, version-matched browser, built-in auto-waiting, and Chromium, Firefox, and WebKit support. Choose Selenium when you need WebDriver-standard tooling, broad browser-driver integrations, or an existing Selenium test suite. Both work well in CI when browser versions, dependencies, timeouts, and artifacts are pinned and monitored.

What you need before automating

  • Python 3.10 or newer for the current Selenium Python bindings.
  • A virtual environment for project dependencies.
  • A decision about headed versus headless execution.
  • Stable locators, explicit waits, and cleanup code.
  • CI system dependencies and browser binaries installed during the build.
python -m venv .venv
# macOS/Linux
source .venv/bin/activate
# Windows PowerShell
# .venv\\Scripts\\Activate.ps1

Playwright: the quickest complete setup

Playwright provides synchronous and asynchronous Python APIs and can automate Chromium, Firefox, and WebKit. Install the package and then download the browser binaries required by your Playwright version. See the Playwright Python introduction and browser installation guide.

pip install playwright
playwright install
# Linux CI may also need:
# playwright install --with-deps

Minimal synchronous script

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    page.goto('https://example.com', wait_until='domcontentloaded')
    print(page.title())
    browser.close()

Async Playwright

import asyncio
from playwright.async_api import async_playwright

async def main():
    async with async_playwright() as p:
        browser = await p.chromium.launch(headless=True)
        page = await browser.new_page()
        await page.goto('https://example.com', wait_until='domcontentloaded')
        print(await page.title())
        await browser.close()

asyncio.run(main())

Locators, actions, and assertions

from playwright.sync_api import sync_playwright, expect

with sync_playwright() as p:
    browser = p.chromium.launch()
    page = browser.new_page()
    page.goto('https://example.com')
    heading = page.get_by_role('heading')
    expect(heading).to_be_visible()
    print(heading.inner_text())
    browser.close()

Prefer role, label, text, and test-id locators over brittle CSS paths. Playwright actions wait for an element to be actionable; add an explicit wait only for a known application state.

Waiting for dynamic pages

page.goto('https://example.com/dashboard')
page.get_by_role('button', name='Load data').click()
page.locator('[data-testid="results"]').wait_for(state='visible')
page.wait_for_load_state('networkidle')
page.wait_for_timeout(500)  # only when a documented animation needs it

Use domcontentloaded for a fast first render, load when subresources matter, and networkidle only when the site actually becomes quiet. A permanently open analytics connection can make network-idle waits hang.

Contexts, devices, and downloads

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch()
    context = browser.new_context(
        viewport={'width': 1280, 'height': 800},
        device_scale_factor=2,
        color_scheme='dark',
        locale='en-US',
        timezone_id='America/New_York'
    )
    page = context.new_page()
    page.goto('https://example.com')
    with page.expect_download() as event:
        page.get_by_text('Download').click()
    download = event.value
    download.save_as('report.pdf')
    context.close()
    browser.close()

A browser context isolates cookies, local storage, permissions, and pages. Create one context per test or user session rather than launching a new browser process for every URL.

Selenium WebDriver: standards-based automation

Selenium supplies Python bindings over WebDriver. Its documentation lists Chrome, Edge, Firefox, Safari, WebKitGTK, and WPEWebKit support. Modern Selenium commonly uses Selenium Manager to obtain compatible drivers when a WebDriver is instantiated. WebDriver is a W3C Recommendation; Selenium also documents WebDriver BiDi for bidirectional events such as network requests, console messages, and JavaScript errors. See the Selenium documentation.

pip install selenium

Minimal Selenium script

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

options = webdriver.ChromeOptions()
options.add_argument('--headless')
options.add_argument('--window-size=1280,800')

driver = webdriver.Chrome(options=options)
try:
    driver.get('https://selenium.dev')
    print(driver.title)
    heading = WebDriverWait(driver, 15).until(
        EC.visibility_of_element_located((By.TAG_NAME, 'h1'))
    )
    print(heading.text)
finally:
    driver.quit()

Explicit waits and reliable locators

wait = WebDriverWait(driver, 20, poll_frequency=0.2)
button = wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, '[data-testid="submit"]')))
button.click()
wait.until(EC.url_contains('/success'))

Do not use fixed sleeps as the main synchronization strategy. Wait for visibility, clickability, a URL change, a DOM property, or an application-specific condition. Keep a single explicit wait policy per operation so timeout behavior is predictable.

Playwright or Selenium?

Decision point Playwright Selenium
Browser engines Installs version-matched Chromium, Firefox, and WebKit. Uses browser-specific WebDriver implementations across major browsers.
Python API Sync and async APIs with high-level locators and auto-waiting. WebDriver sessions with explicit waits and expected conditions.
Setup playwright install downloads supported binaries. Selenium Manager commonly resolves drivers automatically; explicit drivers remain possible.
Protocol Playwright’s own browser automation API. W3C WebDriver plus WebDriver BiDi event capabilities.
Best fit New end-to-end tests, scraping workflows, and multi-engine projects. Existing WebDriver infrastructure, standards-based grids, and broad vendor tooling.
CI concern Pin Playwright and install its matching browsers and OS dependencies. Pin Selenium, browser versions, and driver or Selenium Manager behavior.

Neither tool removes the need to design stable selectors, isolate test data, collect failure artifacts, and maintain browser compatibility. Choose the tool your team can update consistently.

Common automation patterns

Login with a saved session in Playwright

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch()
    context = browser.new_context(storage_state='auth.json')
    page = context.new_page()
    page.goto('https://example.com/account')
    print(page.locator('h1').inner_text())
    context.close()
    browser.close()

Create auth.json in a separate authenticated setup job, protect it as a secret, and never commit it.

Capture a screenshot for debugging

page.screenshot(path='artifacts/failure.png', full_page=True)

Intercept or block resources

def route_handler(route):
    if route.request.resource_type in {'image', 'font'}:
        route.abort()
    else:
        route.continue_()

page.route('**/*', route_handler)

Blocking resources can speed up data-only tasks, but it can also change layout and break applications that depend on fonts or images. Apply it only when the reduced page is valid for your purpose.

CI, reliability, and performance

  • Pin Python, Playwright or Selenium, browser versions, and OS images.
  • Launch one browser and reuse contexts where isolation permits.
  • Set navigation, action, and assertion timeouts explicitly.
  • Retry only transient failures; retries must preserve logs and screenshots.
  • Save HTML, console output, network errors, and screenshots on failure.
  • Run tests in parallel only after removing shared accounts, files, and mutable server state.
  • Use a headed run locally when diagnosing focus, viewport, permissions, or rendering issues.

Browser startup is expensive. Reusing a process and creating isolated contexts usually gives better throughput than launching a browser per page. Network latency, third-party scripts, animations, and server-side rate limits often dominate total runtime. Cost depends on your compute provider, CI minutes, parallel workers, and any external browser grid; neither Playwright nor Selenium has a universal per-screenshot price.

Troubleshooting

Symptom Likely cause Fix
Browser executable not found (Playwright) Browser binaries were not installed for this package version. Run playwright install; on Linux CI use playwright install --with-deps when required.
Driver or session creation failure (Selenium) Browser and driver versions or permissions do not match. Update Selenium, let Selenium Manager resolve the driver, or pin a compatible explicit driver.
Timeout waiting for an element Wrong locator, delayed app state, iframe, overlay, or navigation not finished. Inspect the DOM, target the frame, wait for the real state, and capture a screenshot and console log.
Element is covered or not clickable Cookie banner, modal, animation, or sticky header intercepts the click. Close the overlay, wait for it to disappear, scroll into view, and avoid force-click unless the condition is understood.
Works locally but fails in CI Missing OS libraries, different viewport, slower CPU, sandbox restrictions, or timezone. Install dependencies, set viewport and timezone, increase evidence-rich timeouts, and use the same container image.
Blank or incomplete screenshot Lazy content has not loaded, page is still rendering, or the viewport is wrong. Wait for a selector or image state, scroll when needed, and verify the target URL and viewport.
Intermittent failures Race conditions, shared state, animations, or third-party requests. Replace sleeps with state waits, disable animations where safe, isolate data, and record traces or logs.

Or skip the browser setup

If your goal is a clean website image or PDF rather than interactive browser control, ScreenshotNeo provides a single GET request. It accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

See the ScreenshotNeo API documentation for all options.

curl -G 'https://api.screenshotneo.com/v1/shot' \\
  -d access_key=YOUR_API_KEY \\
  --data-urlencode url=https://stripe.com \\
  -o shot.webp
import requests

r = requests.get(
    'https://api.screenshotneo.com/v1/shot',
    params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'},
    timeout=90,
)
r.raise_for_status()
open('shot.webp', 'wb').write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const data = Buffer.from(await res.arrayBuffer());
await fs.promises.writeFile('shot.webp', data);

You can also set full-page capture, lazy-image loading, CSS element capture, dark mode, device presets, custom viewports, retina scale, PDF paper and page ranges, custom CSS or JavaScript, clicks, selector or network waits, request blocking, headers, cookies, user agent, authorization, timezone, geolocation, transparent backgrounds, resizing, caching TTL, signed links, async webhooks, bulk capture of up to 100 URLs, and usage reporting. The parameter names used by other screenshot APIs also work when switching.

Every feature is included on every plan: 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

FAQ

Can Python automate a browser without Selenium?

Yes. Playwright is the other central Python choice and includes its own browser automation API and browser binaries.

Should I use a real browser in production?

Use a real browser when JavaScript execution, layout, authentication, or user interaction affects the result. For a static HTTP response, a browser may be unnecessary overhead.

How do I automate Chrome headless?

Playwright uses p.chromium.launch(headless=True). Selenium uses ChromeOptions with a headless argument, as shown above.

What is WebDriver BiDi?

It is Selenium’s bidirectional browser protocol capability for streaming events such as network activity, console messages, and JavaScript errors.

How should I choose waits?

Wait for the application state you need: a visible locator, URL, DOM property, download, or network condition. Avoid arbitrary sleeps except for a known animation or delayed external system.