What Is the Actions Class in Selenium and How to Use It
Learn how Selenium Actions composes keyboard, pointer, and wheel input, with runnable examples, browser constraints, input cleanup, and troubleshooting.
Selenium’s Actions API lets you send low-level keyboard, pointer, and scroll-wheel input to a browser. Use it when a test needs a gesture such as hovering, dragging, holding a modifier key while typing, or scrolling by a controlled amount. In Java, create an Actions object, chain operations, then call perform(). Python uses ActionChains; JavaScript uses driver.actions(). For a regular click or text entry, prefer Selenium’s simpler element methods when they express the interaction clearly.
The Actions API is a “low-level interface for providing virtualized device input actions to the web browser,” as described in Selenium’s official guide. It supports keyboard, pointer (mouse, pen, or touch), and wheel input sources. Wheel input was introduced in Selenium 4.2. The convenience methods below compose these lower-level inputs into sequences.
1. How do I use Actions in Selenium?
- Find the element or coordinates where the interaction should occur.
- Create a sequence with the binding’s Actions API.
- Call
perform()to send the sequence to the browser. - Wait for the application’s resulting state and assert on that state.
- Release any keys or pointer buttons left held down.
Here is a complete Java example using Selenium Manager’s default driver setup. It opens Selenium’s interaction demo, hovers over a target, and clicks it. Use a Java project with the Selenium Java dependency installed; current dependency instructions are in the Selenium installation guide.
import org.openqa.selenium.By;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.WebElement;
import org.openqa.selenium.chrome.ChromeDriver;
import org.openqa.selenium.interactions.Actions;
public class ActionsExample {
public static void main(String[] args) {
WebDriver driver = new ChromeDriver();
try {
driver.get("https://www.selenium.dev/selenium/web/mouse_interaction.html");
WebElement target = driver.findElement(By.id("clickable"));
new Actions(driver)
.moveToElement(target)
.click()
.perform();
} finally {
driver.quit();
}
}
}
Methods can be chained because they build a sequence. The browser receives the sequence when you call perform(); constructing a chain alone does not execute it. Read the official Actions API documentation and your language binding’s API reference alongside this guide. Method names and signatures vary by binding.
2. Choose the right interaction API
| Need | Usually use | Why |
|---|---|---|
| Click a button or link | element.click() |
Direct element interaction is simpler. |
| Enter text in a field | element.send_keys(...) |
Targets the field directly. |
| Hover, drag, click-and-hold, or use offsets | Actions | These require pointer movement and button state. |
| Hold a modifier while typing | Actions | Key-down and key-up can bracket other input. |
| Scroll by a delta or scroll an element into view | Wheel Actions | Expresses scroll input as part of the sequence. |
Actions is not a universal replacement for element interactions. Use it when the test needs device-level gestures, sequencing, or timing. Selenium’s mouse guide, keyboard guide, and wheel guide document the convenience operations and constraints.
3. Common pointer actions
Hover over an element
Moving to an element places the pointer at its in-view center by default. The target must be in the viewport. A hover may trigger an application event, but it does not guarantee that asynchronous content has finished appearing; wait for the resulting element or state.
WebElement menu = driver.findElement(By.id("account-menu"));
new Actions(driver)
.moveToElement(menu)
.perform();
Click and hold, then release
clickAndHold() presses the pointer button and leaves it depressed. Add release() when the gesture should end, or use click() for a press-and-release click.
WebElement handle = driver.findElement(By.id("slider-handle"));
new Actions(driver)
.moveToElement(handle)
.clickAndHold()
.pause(java.time.Duration.ofMillis(300))
.release()
.perform();
Drag and drop
For a simple drag, use Selenium’s convenience method. If the page needs intermediate pointer movement or a custom pause, build the sequence explicitly. Drag behavior depends on the page’s event handlers and browser; confirm the resulting application state rather than assuming the gesture succeeded.
WebElement source = driver.findElement(By.id("drag-source"));
WebElement destination = driver.findElement(By.id("drop-target"));
new Actions(driver)
.dragAndDrop(source, destination)
.perform();
Double-click, context-click, and offsets
WebElement target = driver.findElement(By.id("target"));
new Actions(driver)
.moveToElement(target)
.doubleClick()
.perform();
new Actions(driver)
.contextClick(target)
.perform();
// Move relative to the element's in-view center, then click.
new Actions(driver)
.moveToElement(target, 10, 5)
.click()
.perform();
Pointer offsets still need to land in the viewport and on the intended hit target. If the target is partly clipped, scroll it into view first and verify its visible bounds.
4. Keyboard input and modifier chords
Use keyDown and keyUp to hold and release a modifier around text entry. The following selects all text in a field and replaces it. Keys constants represent special keys; ordinary characters can be sent as strings.
import org.openqa.selenium.Keys;
import org.openqa.selenium.WebElement;
import org.openqa.selenium.interactions.Actions;
WebElement field = driver.findElement(By.name("query"));
new Actions(driver)
.click(field)
.keyDown(Keys.CONTROL)
.sendKeys("a")
.keyUp(Keys.CONTROL)
.sendKeys("new search")
.perform();
On macOS, use the platform’s Command modifier where the application expects it. Keep the modifier held only for the intended chord and release it in the same sequence. If an action is interrupted, clear or reset input state before the next interaction.
5. Scrolling with wheel input
Wheel actions let a sequence scroll by vertical and horizontal deltas or scroll an element into view. Selenium’s wheel documentation identifies these actions as Chromium-only; check the current documentation and target browser before relying on them. Wheel actions do not automatically bring every target into view before pointer interaction, so scroll explicitly when necessary.
WebElement footer = driver.findElement(By.id("footer"));
new Actions(driver)
.scrollToElement(footer)
.perform();
// Or scroll by horizontal and vertical deltas from the viewport.
new Actions(driver)
.scrollByAmount(0, 600)
.perform();
For a nested scroll container, provide the container as the origin where supported by the binding and Selenium version, then verify that the intended container moved. A page-level scroll may not affect an inner panel.
6. Python and JavaScript examples
Python: ActionChains
Install the Selenium Python package in your environment, then run this script. Selenium Manager can handle the driver setup for supported browsers.
from selenium import webdriver
from selenium.webdriver import ActionChains
from selenium.webdriver.common.by import By
driver = webdriver.Chrome()
try:
driver.get("https://www.selenium.dev/selenium/web/mouse_interaction.html")
target = driver.find_element(By.ID, "clickable")
ActionChains(driver) \
.move_to_element(target) \
.click() \
.perform()
finally:
driver.quit()
Python uses ActionChains, with snake_case method names such as move_to_element and click_and_hold. A keyboard chord can use key_down and key_up. To clear action state, the official guide demonstrates ActionBuilder(driver).clear_actions().
JavaScript: WebDriver actions
Install the selenium-webdriver package in a Node.js project. This CommonJS script demonstrates the binding’s asynchronous calls:
const { Builder, By } = require('selenium-webdriver');
(async function run() {
const driver = await new Builder().forBrowser('chrome').build();
try {
await driver.get('https://www.selenium.dev/selenium/web/mouse_interaction.html');
const target = await driver.findElement(By.id('clickable'));
await driver.actions()
.move({ origin: target })
.click()
.perform();
} finally {
await driver.quit();
}
})();
JavaScript uses driver.actions() and awaits the sequence’s perform(). The JavaScript API documents synchronized ticks for input sources; when composing asynchronous multi-device sequences, add pauses where needed to coordinate timing. Do not assume every binding exposes identical methods or scheduling behavior.
7. Timing, viewport, and input state
Use pauses only for deliberate timing
pause() inserts a duration into the sequence. It can represent a gesture that intentionally holds for a while or allow a short interval between input actions. It is not a substitute for waiting until an application condition is true. For dynamic pages, use an explicit wait for the expected state; fixed delays can make a test slow or flaky.
new Actions(driver)
.moveToElement(target)
.pause(java.time.Duration.ofMillis(250))
.click()
.perform();
Keep the target and pointer in view
Element-based pointer movement can fail when an element is outside the viewport. Scroll the page or element into view first, then locate it again if the page rerendered. Offsets also have viewport constraints. Sticky headers, overlays, and nested scrolling regions can change which point receives the input.
Release held input
The driver remembers input state across sequences. Creating a new Actions object does not by itself guarantee that a previously depressed key or pointer button is released. End a gesture with keyUp or release when appropriate. If a sequence fails midway, use the binding’s clear/reset mechanism. Selenium documents Java’s RemoteWebDriver.resetInputState() and Python’s ActionBuilder(driver).clear_actions(); consult the binding documentation for its equivalent.
8. Troubleshooting common Actions errors
| Symptom | Likely cause | Fix |
|---|---|---|
| Move target out of bounds or element not interactable | The target or requested pointer location is outside the viewport, clipped, or obscured. | Scroll it into view, inspect overlays and sticky headers, and use an in-view point. Check the element again after scrolling. |
| Click appears to do nothing | The pointer hit an overlay or a different part of the element; the app may also need time to update. | Confirm the hit target and wait for the expected application state before asserting. |
| Drag starts but does not complete | The page may require a particular drag path or duration, or the pointer button remained held. | Build the press, movement, and release steps explicitly; add a pause only if the gesture requires it; verify the drop result. |
| Later typing is unexpectedly shifted or a click behaves like a hold | A modifier key or pointer button remained depressed from an earlier sequence. | Send the matching key-up or release action, or clear/reset the driver’s input state. |
| Wheel method is unsupported | The chosen browser or binding version does not support the documented wheel action. | Check Selenium’s current wheel guide and browser support; use a supported browser or a suitable alternative interaction. |
| Actions method name is missing | Examples from another language binding or Selenium version were copied. | Use the binding-specific reference: Java camelCase, Python snake_case, JavaScript API methods and async calls. |
| Test is flaky after a fixed pause | A fixed delay does not match variable application load time. | Wait for a specific visible, enabled, or changed condition instead of increasing the pause blindly. |
9. Performance, reliability, and cost
Actions batches input operations into a sequence, but each WebDriver interaction still involves browser automation communication. Keep a chain focused on one logical gesture and avoid inserting long pauses unless the interaction needs them. Prefer condition-based waits for application readiness, since arbitrary delays add runtime while still failing to guarantee readiness.
For reliability, use stable locators, ensure pointer targets are visible, release held inputs, and assert the page’s resulting state. Browser support can differ, especially for wheel input, so run gesture-dependent checks in the browsers your application supports. Selenium itself is the browser automation layer; infrastructure and browser execution costs depend on where and how you run it.
10. Capture the result without managing a browser
If your goal is to save a website screenshot rather than test a user gesture, ScreenshotNeo provides a website screenshot API and MCP server. Its API accepts a URL in one GET request and can return PNG, JPEG, WebP, or PDF. See ScreenshotNeo and the API documentation.
Or skip the browser setup
After the Selenium method above, use a direct screenshot request when you need the page image without setting up a WebDriver session. This cURL example saves a WebP image:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Equivalent Python request:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
Equivalent Node.js request:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the capture; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up free for 1,000 screenshots a month, with no card required.
11. FAQ
Is the Actions class available in Python?
Python exposes the related API as ActionChains, rather than a class named Actions.
Does Actions automatically scroll an element into view?
Do not rely on that behavior. Selenium’s wheel guide says Actions does not automatically bring target elements into view; scroll explicitly when needed.
Can I combine keyboard and pointer actions?
Yes. Actions sequences can coordinate multiple input sources. Account for tick timing and release any keys or buttons the sequence leaves held.
When was wheel input added?
The Selenium Actions guide identifies wheel input as introduced in Selenium 4.2.


