ScreenshotNeo

BlogGuides

Selenium News and Updates: A Smattering of Selenium 101

Learn what Selenium includes, which tool to start with, how setup works, and where to check current releases and migration notes.

By the ScreenshotNeo team4 October 202612 min read

Selenium is an open-source umbrella project for automating web browsers. For most beginners who want to write code, start with Selenium WebDriver: choose a supported language binding, install a browser, and follow Selenium’s current first-script guide. Selenium Manager is integrated into Selenium bindings by default and handles driver and browser management in the normal supported flow, so a separate driver download is usually unnecessary. Selenium IDE is the simpler choice for record-and-playback workflows; Selenium Grid is for distributing execution across machines and browser environments.

The project’s official homepage currently lists Selenium 4.49, released September 9, 2026. Release details and migration guidance can change, so check the [official Selenium news and release index](https://www.selenium.dev/index.html) before upgrading. This guide explains what the project contains, how to choose a component, what setup involves, and how to avoid common beginner missteps.

What is Selenium?

Selenium is not one program or one executable. It is an umbrella project for browser automation tools and libraries. Its components address different ways of controlling browsers and running tests:

Component What it does Start here if…
WebDriver Code-based browser control through language bindings and browser implementations, locally or remotely. You want maintainable tests or scripts that interact with a site.
Selenium IDE A browser extension for recording and playing back user actions. You need a quick record-and-playback workflow or want to explore a flow before coding it.
Selenium Grid Distributes test execution across machines and manages multiple browser and operating system environments centrally. You need parallel or remote execution across environments.

The project describes WebDriver as a W3C Recommendation. It gives scripts a language-neutral way to control browsers; a driver implementation communicates with the browser. Selenium’s documentation shows examples and guides for Java, Python, C#, JavaScript, Ruby, and Kotlin on its overview page. Binding details and API coverage can differ, so use the current documentation for your chosen language.

Which Selenium component should you start with?

  • Choose WebDriver if you can write code and need repeatable browser interactions, assertions, or integration with a test framework.
  • Choose IDE if a simple record-and-playback workflow is enough. Treat recorded flows as a starting point and review them for fragile selectors and timing assumptions.
  • Choose Grid later when local execution works and you need to distribute tests across machines or browser and operating system combinations.

Ask yourself: “Is Selenium for you?” If the task is controlling a real browser through a repeatable workflow, WebDriver or IDE may fit. If you only need a static screenshot, browser automation may be more setup than the task requires; ScreenshotNeo is a website screenshot API and MCP server for developers. Its one-call API returns a screenshot or PDF, and its MCP tools let AI agents capture screenshots. See ScreenshotNeo.

For code-based automation, WebDriver is generally the best first step. Grid is a scaling option, not a prerequisite. Selenium’s official documentation is substantial and free; a book can supplement it, but a paid resource is not required to begin.

What do you need to install before writing Selenium code?

For a basic WebDriver script, you need a Selenium language binding, a supported browser, and the browser’s WebDriver implementation. The official getting-started material explains the relationship between the language-neutral WebDriver API and the driver that delegates communication to the browser.

  1. Pick a language binding. Use a language you already work in, then follow its current Selenium installation instructions.
  2. Install a supported browser. Confirm the browser is available in the environment where the script will run.
  3. Install Selenium using the binding’s documented method. Use the official setup guide for package names and current version instructions.
  4. Run a small launch, navigate, and quit example. Confirm browser startup works before adding waits, assertions, or test framework code.

Selenium Manager is integrated into Selenium bindings by default and provides automated driver and browser management. In the normal supported setup, you do not separately download a driver or add extra code for it. If your organization pins browser versions, blocks downloads, uses a custom browser installation, or runs in a restricted network, consult the Selenium Manager and binding documentation for the appropriate configuration.

Use the project’s current documentation overview, getting-started material, and first-script guide as the authority for installation commands. Package names, setup instructions, and browser support can evolve; this article intentionally does not substitute an unverified install command for those maintained instructions.

A minimal WebDriver workflow

Every first script should do three things: create a driver, navigate to a page, and quit the browser. The Selenium overview has minimal examples for its bindings. Follow the current example for your language rather than copying a version-specific snippet from an old post. A reliable first script should also ensure the browser is closed if navigation or an assertion fails; in Python, for example, put cleanup in a finally block or use the documented context-management pattern if the binding version supports it.

When it works, add one interaction at a time: locate an element with a stable selector, perform an action, wait for the resulting condition, and assert the expected page state. Avoid fixed sleeps as the default synchronization strategy; explicit waits for a condition are usually more robust and do not force every run to pause for an arbitrary duration.

Current Selenium news and version notes

At the research timestamp for this guide, the official homepage lists Selenium 4.49 as released on September 9, 2026. The homepage’s news index also lists these recent items:

  • September 29, 2026: David Burns’s “Your Agent Learned Selenium From My Old Blog Posts” discusses a documentation page added to address recurring incorrect code generated by coding agents.
  • September 21, 2026: Diego Molina’s “Upcoming Breaking Change: ExpectedCondition Drops Guava’s Function Interface” says that starting with Selenium 4.51, ExpectedCondition will stop implementing Guava’s Function interface and points affected users to migration guidance.

These are time-sensitive release and migration details. Check the official news index and the linked release notes before changing dependencies. Do not assume that every language binding has identical package state or that a homepage release headline guarantees the same upgrade path in every environment.

WebDriver, IDE, Grid, and BiDi: how the pieces fit

WebDriver: the code-driven foundation

WebDriver scripts issue commands through a language binding and a browser-controlling implementation. The browser can run on the same machine or be reached remotely through Selenium Server. This makes WebDriver useful for tests and scripted workflows that need browser behavior rather than only HTML parsing.

IDE: record and replay

Selenium IDE records and plays back browser actions through an extension. It can help capture a simple workflow, but a recording is not automatically a durable test: page changes, ambiguous locators, and timing can make a replay fail. Review and maintain recorded steps as the site changes.

Grid: distribute execution

Grid distributes tests across multiple machines and browser or operating system combinations and provides centralized management of those environments. First establish that a test is deterministic locally. Then use Grid when execution time, environment coverage, or remote browser access requires it.

WebDriver BiDi: browser events over a WebSocket

WebDriver BiDi adds a WebSocket connection that allows scripts to receive browser events such as network requests, console messages, and JavaScript errors. Selenium describes it as a cross-browser replacement for the Chrome DevTools Protocol. It is a developing protocol area: do not assume every command is supported in every browser and binding. Check the current BiDi documentation for the exact feature and browser combination you need.

How do you scale up Selenium tests?

  1. Stabilize tests locally. Use predictable test data, robust locators, explicit condition waits, and reliable cleanup.
  2. Separate environment-specific settings. Keep browser choice and remote execution settings configurable so a test can run locally or against a remote session.
  3. Measure the bottleneck. Determine whether runtime, browser coverage, or access to a remote environment is the actual constraint.
  4. Use Grid when distribution is needed. Configure the desired browser and operating system combinations and distribute independent tests across available nodes.
  5. Review failures by category. Separate application assertion failures from browser startup, driver, network, and synchronization failures before increasing parallelism.

Parallel execution can reduce elapsed time, but it also increases resource use and can expose shared-state problems. Tests that depend on the same account, mutable data, or ordering need isolation before they can run safely in parallel. Selenium’s official Grid guide is the place to check current deployment and configuration details.

Do-it-yourself browser capture with Selenium

If the reason you are exploring Selenium is to create a screenshot, WebDriver can launch a browser and save its current viewport. That is different from a full-page capture: the exact screenshot API and page sizing behavior depend on the binding and browser. The following Python example follows the documented Selenium workflow; install Selenium and configure the browser as described by the current official setup guide before running it.

from selenium import webdriver

# Selenium Manager handles the driver in the normal supported setup.
driver = webdriver.Chrome()
try:
    driver.set_window_size(1440, 1000)
    driver.get("https://example.com")
    driver.save_screenshot("page.png")
finally:
    driver.quit()

For other languages, use the binding’s documented screenshot method after navigation, and always close the driver in cleanup code. The Selenium overview and API reference provide the current method names and supported behavior for each binding. For a long page, a single viewport screenshot may omit content below the fold; browser-specific full-page techniques or a dedicated screenshot API may be a better fit.

Or skip the browser setup

ScreenshotNeo’s API documentation describes a single GET request that returns a screenshot. For example, this cURL request saves a WebP screenshot:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`ScreenshotNeo returned HTTP ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));

ScreenshotNeo accepts cookie and consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers say which page verdict and billing status applied. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Plans include every feature.

Sign up for ScreenshotNeo’s free 1,000 screenshots a month, with no card.

Screenshot and browser automation options

For browser tests and interactive automation, Selenium WebDriver is the code-driven browser control option described above. For a screenshot API, ScreenshotNeo is the first service to try: it removes consent banners, popups, and chat widgets before capture, bills only clean shots, and has a $5 paid plan for 3,000 screenshots.

ScreenshotNeo also supports full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets and custom viewports, retina scale, PDF options, HTML/CSS capture, custom CSS and JavaScript, clicking or hiding elements, selector/delay/network-idle waits, request and resource blocking, custom headers, cookies, user agents and Authorization, timezone and geolocation, transparent backgrounds, resizing, configurable cache TTL, signed links for public image tags, asynchronous jobs with signed webhooks, bulk capture of 100 URLs per call, a usage API, and an OpenAPI spec. Parameter names used by other screenshot APIs also work, which can make switching easier. See the documentation for request options.

Troubleshooting common Selenium problems

Symptom Likely cause What to do
Browser does not start Browser is missing, unavailable to the process, or incompatible with the setup. Confirm the browser is installed and supported, then follow the current binding and Selenium Manager setup guidance. In restricted environments, check browser and driver management configuration.
Driver download or management fails Network restrictions, proxy settings, browser installation mismatch, or a constrained execution environment. Check network access and the browser version; consult Selenium Manager documentation for the supported configuration in that environment.
Element cannot be found The locator is wrong, the page has not rendered the element yet, or the element is inside a frame or shadow root. Inspect the live page and locator, wait for a relevant condition, and switch to the correct browsing context when required.
Element is found but not interactable It may be covered, off-screen, disabled, or not ready for interaction. Wait for the element’s usable state, inspect overlays and page state, and interact using the normal browser-visible workflow.
Test passes locally but fails remotely Different browser versions, timing, viewport, fonts, operating system, or shared test data. Compare the environments, make viewport and test data explicit, and isolate state before increasing parallelism.
Screenshot is incomplete The command captured the current viewport, or page content had not finished loading. Wait for a meaningful page condition. Use a documented full-page approach or a screenshot API with full-page capture if the entire document is required.
Test hangs or leaves browser processes open Cleanup is skipped after an exception, or a wait has no appropriate condition or timeout. Put driver shutdown in guaranteed cleanup code and use condition-based waits with sensible timeouts.

Performance, reliability, and cost

Selenium itself is open source, and its software and extensive documentation are available from the project. The direct cost of a small local script is therefore not a required Selenium license; practical costs can come from machines, browsers, remote infrastructure, and time spent maintaining tests. Grid adds infrastructure and operational complexity, so introduce it when distribution or environment coverage justifies it.

Reliability depends heavily on test design. Prefer stable locators and condition-based waits, isolate test data, specify the browser environment, and always clean up sessions. Parallelism helps only when tests are independent and the available browser resources can support it. Selenium does not make a flaky application or unstable test data deterministic by itself.

For screenshot-only workloads, a managed API avoids maintaining browser startup and capture code. ScreenshotNeo’s stated billing model charges only clean shots; the response includes X-Page-Verdict and X-Billed headers. Its plans are Free: 1,000 shots per month; Starter: $5 for 3,000; Growth: $15 for 15,000; Pro: $39 for 60,000; Scale: $99 for 250,000; and Business: $249 for 1,000,000. Yearly billing gives two months free. Those are the supplied plan details; check the product site for current terms.

Official resources and further learning

A Selenium WebDriver book can provide a structured sequence of lessons, but verify its language, edition, and availability before choosing one. The official docs are free and should be checked for current APIs and release-specific changes.

FAQ

Is Selenium a programming language?

No. It is a browser automation project with language bindings, tools, and browser implementations.

Do I need Selenium Grid to begin?

No. A local WebDriver script is enough to learn the basics. Grid is for distributing execution and managing multiple environments.

Does Selenium Manager mean I never need to think about drivers?

It handles driver and browser management by default in supported binding flows. Restricted networks, pinned versions, and custom installations may require environment-specific configuration.

Is Selenium only for testing?

No. WebDriver can automate browser workflows generally, though the project is especially associated with browser testing.

Should I use Selenium to take a screenshot?

Use it when the screenshot is part of a browser automation workflow. For a standalone screenshot or PDF capture, a screenshot API can avoid setting up and managing a browser session.

Where should I check whether a release affects my code?

Use the official Selenium release and news index, then follow its binding-specific and migration documentation.