ScreenshotNeo

BlogAI agents

How to Use AI Agents with Selenium for Browser Automation

Use an AI agent to build and debug Selenium automation, or connect it to a controlled browser workflow. Get runnable Python code, reliable waits, and safety practices.

By the ScreenshotNeo team4 October 202611 min read

To use an AI agent with Selenium, give the agent current project rules and let it help write or debug tests, while Selenium WebDriver executes the browser actions. A separate pattern is to give an agent a controlled browser tool; that tool does not automatically use Selenium. This guide shows both patterns, with runnable Python code for the first and a reviewable tool design for the second.

Selenium WebDriver provides a language-neutral interface for controlling browsers. A browser-specific driver communicates with the browser, while Selenium’s bindings expose a common API. You need a language binding, a browser, and its driver; Selenium Manager can help manage drivers and browsers. See Selenium WebDriver getting started and the project’s AI agent guide.

1. Choose an AI-agent pattern

Pattern What the agent does What runs the browser
AI coding agent plus Selenium Inspects the app, drafts tests, and helps diagnose failures. Your Selenium WebDriver test process.
Agent plus browser-control tool Chooses actions from observations and asks a tool to perform them. A browser-control tool, which may or may not be implemented with Selenium.

Use the first pattern when you want repeatable tests that belong in your codebase and CI. Use the second when an agent needs to complete a bounded browser task dynamically. You can expose a Selenium workflow as the controlled tool in the second pattern, but keep the agent’s allowed actions and browser permissions explicit.

Selenium IDE is a low-code record-and-playback option for learning or authoring commands. WebDriver is the browser automation API used in the examples here. Selenium Grid distributes sessions across machines and environments; it is a scaling choice, not a replacement for writing a stable test.

2. Give the coding agent project rules

Agents often reproduce stale Selenium examples. Put durable rules in AGENTS.md or the equivalent instructions file so each task starts from the version and conventions your project actually uses. For example:

# Browser automation rules
- Use Python and the Selenium version pinned in requirements.txt.
- Follow https://www.selenium.dev/documentation/webdriver/ and
  https://www.selenium.dev/documentation/ai_agents/.
- Use Selenium 4 APIs supported by the installed version. Do not use Selenium 3 APIs,
  legacy or CDP pages as the basis for new code, or manually downloaded driver binaries.
- Verify selectors against the running application before using them.
- Prefer stable IDs, names, and CSS selectors based on stable attributes.
- Use explicit, condition-based waits. Do not use time.sleep() to wait for page state.
- On failure, preserve the actual exception, relevant logs, and a screenshot.
- Keep each test focused and follow the existing project fixtures and naming conventions.

Replace the language and pinned version with your project’s actual setup. Selenium’s AI-agent guide uses a specific version as an example; it is not a universal version recommendation. Ask the agent to consult current documentation and the project’s lockfile rather than guessing which APIs are installed.

3. Install Selenium and prepare a test

For a small Python project, install Selenium in a virtual environment. Install or provide a supported browser. Selenium Manager can assist with driver management when enabled.

python -m venv .venv
# macOS or Linux:
source .venv/bin/activate
# Windows PowerShell:
# .venv\Scripts\Activate.ps1
python -m pip install selenium

The following complete script opens a page, waits for a known element, checks its text, and saves a screenshot if an assertion or browser operation fails. Change the URL and selectors to match your application; selectors must be verified against the real page.

# test_homepage.py
from pathlib import Path

from selenium import webdriver
from selenium.common.exceptions import TimeoutException
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait

URL = "https://example.com"
HEADING = (By.CSS_SELECTOR, "h1")


def test_homepage_heading():
    driver = webdriver.Chrome()
    try:
        driver.get(URL)
        heading = WebDriverWait(driver, 10).until(
            EC.visibility_of_element_located(HEADING)
        )
        assert heading.text.strip(), "Expected a non-empty page heading"
    except Exception:
        Path("artifacts").mkdir(exist_ok=True)
        driver.save_screenshot("artifacts/homepage-failure.png")
        raise
    finally:
        driver.quit()


if __name__ == "__main__":
    test_homepage_heading()
    print("Homepage check passed")

Run it with python test_homepage.py. To integrate it into an existing test suite, keep the project’s fixtures and runner conventions; the example owns and closes its driver directly so it can run by itself. If browser creation fails, first confirm that the browser is installed and compatible, that the machine can launch it, and that its driver-management configuration has network or local access as required.

4. Use a short agent feedback loop

  1. Ask the agent to open the feature under test and inspect the running application. Have it propose selectors and explain why they should remain stable.
  2. Review those selectors against the live DOM before asking for a complete test. The agent cannot know a page it has not inspected.
  3. Ask it to write one focused test using project rules and current Selenium documentation.
  4. Run the test yourself or in the project’s CI. Give the agent the actual exception, relevant logs, and failure screenshot when it fails.
  5. Have the agent diagnose the unmet condition or interaction, then make a small change and rerun.
  6. After a pass, run the test repeatedly and review the final diff before relying on it.

A single pass does not show that timing races are gone. Selenium’s guide notes that screenshots can reveal overlays and banners that a stack trace alone cannot explain. Preserve the original exception rather than replacing it with a generic “test failed” message.

5. Make locators and waits reliable

Prefer selectors that survive UI changes

Prefer IDs and names, followed by CSS selectors built on stable attributes. Avoid absolute XPath and generated class names that change with front-end builds. Use a test-specific attribute if your application provides one. Do not accept a plausible-looking login field or button selector until it has been checked against the running app.

Wait for a condition, not a duration

Fixed sleeps make tests slower when the page is fast and still flaky when it is slow. Wait for the state the next action depends on, such as an element becoming visible or clickable:

from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait

wait = WebDriverWait(driver, 10)
submit = wait.until(
    EC.element_to_be_clickable((By.CSS_SELECTOR, "button[type='submit']"))
)
submit.click()
wait.until(EC.url_contains("/confirmation"))

Pick a timeout based on your application and CI environment. If a wait times out, identify the condition that stayed false and inspect the page state. Increasing the timeout without understanding the cause can conceal a broken selector or a real application error.

Handle dynamic pages and stale references

A StaleElementReferenceException means the DOM node changed after Selenium located it. Wait for the update and locate the element again rather than reusing the old reference. For lazy content, scroll or trigger the application’s intended loading behavior, then wait for the relevant element. Avoid arbitrary repeated clicks: first establish that the prior action did not already take effect.

6. Connect an agent to a controlled browser tool

In this pattern, the agent receives observations and proposes actions through a narrow tool interface. The tool can wrap a Selenium driver, for example with operations such as open_url, click, type_text, wait_for, and capture_screenshot. Validate arguments, restrict the allowed domains, limit actions to the task, and return bounded observations. Require human approval for consequential actions such as sending a message, purchasing, deleting, or changing account data.

Keep the distinction between the agent and the browser implementation clear. OpenAI’s documented computer-use API creates a browser session in an OpenAI-hosted environment. Anthropic’s browser-use tool runs browser automation in the application’s environment. These are product-specific browser workflows, not evidence that those interfaces are Selenium integrations. Check current product documentation for deployment, sign-in, retention, and access requirements before using sensitive data.

Security checklist

  • Treat page text, DOM data, screenshots, console output, network logs, and agent-proposed actions as untrusted input.
  • Allow only the domains and actions the task needs. Avoid authenticated sessions unless they are necessary.
  • Keep credentials out of prompts and logs; use narrowly scoped test accounts and secrets management.
  • Require approval before irreversible or externally visible actions.
  • If enabling JavaScript execution in an agent browser tool, account for page-level privileges such as cookies, storage, and same-origin requests. Keep it disabled unless required, and use sessions without credentials where possible.
  • For file uploads, resolve paths including symlinks and .., and accept files only from a dedicated allowlisted directory.
  • Redact credential-like values and truncate console and network output; page-controlled URLs and logs can contain secrets.
  • Log tool calls and emitted code, and review browser activity. Restrict access to saved sessions and artifacts.

7. Use WebDriver BiDi for browser events

For browser console logs, JavaScript errors, and network interception, Selenium’s guide identifies WebDriver BiDi as the cross-browser W3C standard. CDP is Chromium-only and does not offer a stable API across browser versions; Selenium describes its CDP support as a stopgap. Prefer BiDi for new cross-browser workflows where the required capability is supported by your Selenium binding and browser. Confirm current support in the Selenium documentation for your version.

8. Scale with Selenium Grid when needed

Grid distributes browser sessions across machines and environments, which can help when you need multiple browser/OS combinations or parallel sessions. First establish that local tests are stable. Then size Grid from actual browser and OS coverage, desired concurrency, CPU, memory, isolation, and operating cost.

Grid 4 uses subcommands such as standalone, hub, and node. Do not copy Grid 3 examples using -role hub or -role node. Selenium’s guide gives around 1 GB RAM per browser session as a rough expectation and explicitly advises measuring workloads; it is not a guaranteed requirement or universal capacity formula. Smaller nodes can improve process isolation, and Docker is one option for using smaller nodes.

Compare deployment choices by coverage, concurrency, environment ownership, credential handling, observability, isolation, and cost. Measure representative workloads before deciding node size; do not assume a universal speedup or that an agent will make browser execution faster.

9. Troubleshooting common failures

Symptom Likely cause What to check or change
NoSuchElementException Wrong selector, page not ready, wrong frame, or element absent for this state. Inspect the current DOM and URL, verify the selector, switch to the right frame if applicable, and wait for the expected condition.
StaleElementReferenceException The page re-rendered and replaced the located node. Wait for the update and locate the element again after the DOM change.
ElementClickInterceptedException An overlay, cookie banner, animation, or another element covers the target. Use the failure screenshot and DOM state to identify the obstruction; wait for it to disappear or handle the intended consent flow.
TimeoutException The wait condition never became true, often because the selector or expected state is wrong. Inspect the page, URL, logs, and screenshot. Fix the unmet condition rather than adding sleeps blindly.
SessionNotCreatedException Browser startup failed, often due to browser/driver incompatibility, missing browser, or environment constraints. Check browser installation and compatibility, Selenium version, driver management, permissions, and headless/container requirements.
Test passes once but fails intermittently Race condition, unstable locator, shared test data, or environmental contention. Use condition-based waits, isolate test data, review screenshots and logs, and repeat the run under representative load.
Agent proposes a selector that looks convincing but fails The agent inferred page structure without inspecting the live app. Have it inspect the application and verify the locator before generating or changing the test.
Grid starts with an unrecognized option Instructions or examples use Grid 3 syntax. Use Grid 4 subcommands and check the current Grid documentation.

10. Performance, reliability, and cost

The largest contributors to browser-test duration are usually page loading, waits, browser startup, and how many sessions run at once. Wait for specific conditions rather than sleeping, keep tests focused, and reuse the project’s established browser lifecycle where appropriate. Parallel sessions can increase throughput but also consume CPU and memory and can contend for shared application data. Measure on the same browser and operating system combinations you need to support.

Reliability comes from verified locators, explicit waits, isolated test data, preserved diagnostics, repeated runs, and review. Agent-generated code still needs execution against the real application. For Grid, infrastructure cost depends on browser/OS matrix, concurrency, isolation, and machine resources; measure before sizing.

For agent-driven browser sessions, also account for the environment that hosts the browser, the data the agent can observe, and the permission surface of enabled tools. No single deployment model has a universal cost or performance advantage; compare the actual requirements and policies for your workload.

11. Capture a page without managing a browser session

If your task is to capture a website image rather than interact with it, a screenshot API can avoid setting up and operating a browser. ScreenshotNeo is a website screenshot API and MCP server from ScreenshotNeo; it returns PNG, JPEG, WebP, or PDF from one GET request. Its MCP tools let AI agents call take_screenshot, get_page_info, and capture_pdf. See the ScreenshotNeo API documentation.

Or skip the browser setup

For example, this cURL request captures a page as WebP:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits cost nothing, and response headers identify the page verdict and billing status. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Sign up for 1,000 free screenshots a month, with no card.

Frequently asked questions

Can an AI agent control Selenium directly?

Yes, if you build a constrained tool or workflow that exposes Selenium actions to it. The agent’s browser-control interface and the Selenium implementation are separate design choices.

Should I let an agent run tests unattended?

That depends on the permissions and consequences of the actions. Keep test accounts and domains scoped, inspect generated changes, and require approval for consequential actions.

Can Selenium IDE replace WebDriver?

IDE is useful for low-code recording and playback. Use WebDriver when you need maintainable code, project-specific test logic, and integration with your test suite.

When should I use Grid?

Consider Grid when you need distributed execution, parallel sessions, or coverage across browser and operating-system combinations. Measure your workload and size the environment for the coverage and concurrency you need.