ScreenshotNeo

BlogHow-to

How to Perform Mouse Actions in Selenium WebDriver

Learn to click, hover, right-click, double-click, and drag elements with Selenium WebDriver Actions, with runnable Java, Python, and JavaScript examples.

By the ScreenshotNeo team4 October 20269 min read

Selenium WebDriver performs mouse gestures through its Actions API. Locate the element, add the gesture to an action chain, and call perform() to send it to the browser. Use convenience methods for common gestures such as click, hover, context-click, double-click, and drag-and-drop; use lower-level pointer actions when you need finer control.

The examples below use Selenium’s Java, Python, and JavaScript bindings. They assume the driver is already created and the page is loaded. The Actions API is a low-level interface for virtualized device input; it supports key, pointer, and wheel input sources. Selenium Actions API documentation and the binding references describe the available methods and signatures.

1. Set up a target element

Find a stable target before performing a gesture. Prefer a unique ID or a locator tied to the element’s role or purpose over a brittle position in the DOM. For the examples, suppose the page has an element with id="target" and a draggable element with id="source".

// Java
WebElement target = driver.findElement(By.id("target"));
WebElement source = driver.findElement(By.id("source"));
WebElement destination = driver.findElement(By.id("destination"));
# Python
from selenium.webdriver.common.by import By

target = driver.find_element(By.ID, "target")
source = driver.find_element(By.ID, "source")
destination = driver.find_element(By.ID, "destination")
// JavaScript (selenium-webdriver)
const { By } = require('selenium-webdriver');

const target = await driver.findElement(By.id('target'));
const source = await driver.findElement(By.id('source'));
const destination = await driver.findElement(By.id('destination'));

Make sure the element exists and is ready for interaction. If the page renders asynchronously, wait for the relevant state before building the action chain. A found element can still be covered by an overlay, outside the viewport, or replaced by a later render.

2. Perform common mouse gestures

In Java, create an Actions object, add a gesture, then call perform(). Python uses ActionChains. JavaScript uses the driver’s action builder and ends with perform(). Method names vary slightly by binding.

Click

An element click moves to the element and clicks it. Use the Actions API when you want the gesture to be part of a longer pointer sequence; ordinary WebDriver element clicking may be simpler for a standalone click.

// Java
new Actions(driver).click(target).perform();

// Python
from selenium.webdriver.common.action_chains import ActionChains
ActionChains(driver).click(target).perform()

// JavaScript
await driver.actions().click(target).perform();

Click and hold

clickAndHold presses the left button on the target without releasing it. Use it when the application responds to a held press or as part of a custom drag sequence. If a later step fails while the button is held, release it or reset the input state before continuing.

// Java
new Actions(driver).clickAndHold(target).perform();

// Python
ActionChains(driver).click_and_hold(target).perform()

// JavaScript
await driver.actions().move({ origin: target }).press().perform();

Right-click (context-click)

Selenium names the right-click convenience method contextClick in Java and context_click in Python. It moves to the target and presses and releases the right mouse button.

// Java
new Actions(driver).contextClick(target).perform();

// Python
ActionChains(driver).context_click(target).perform()

// JavaScript
await driver.actions().move({ origin: target }).press(2).release(2).perform();

For JavaScript’s low-level example, button 2 is the right mouse button in the WebDriver Actions API. A page may suppress the browser context menu or handle the event itself, so a successful gesture does not guarantee that a native menu will appear.

Double-click

Use the binding’s double-click helper to send two left-button clicks at the target.

// Java
new Actions(driver).doubleClick(target).perform();

// Python
ActionChains(driver).double_click(target).perform()

// JavaScript
await driver.actions().move({ origin: target }).doubleClick().perform();

Hover (move to an element)

Hovering moves the pointer to the in-view center of the element. The target must be in the viewport; otherwise Selenium can report an error. Hover often triggers menus or tooltips, but the application may require a short wait before the resulting content is ready.

// Java
new Actions(driver).moveToElement(target).perform();

// Python
ActionChains(driver).move_to_element(target).perform()

// JavaScript
await driver.actions().move({ origin: target }).perform();

Move by an offset

Use offsets when the interaction depends on a particular point rather than an element’s center. Java’s moveByOffset moves relative to the current pointer location; moveToElement accepts offsets relative to an element. Python offers corresponding offset methods. JavaScript’s low-level pointer API lets you specify an origin and offsets.

// Java: move to a point 30px right and 10px up from the current pointer
new Actions(driver).moveByOffset(30, -10).perform();

// Java: move to a point offset from the element's top-left
new Actions(driver).moveToElement(target, 30, 10).perform();

// Python: move relative to the element
ActionChains(driver).move_to_element_with_offset(target, 30, 10).perform()

// JavaScript: element origin plus offsets
await driver.actions().move({ origin: target, x: 30, y: 10 }).perform();

Positive X moves right and positive Y moves down. The pointer must remain in the viewport for coordinate moves. Check the exact offset origin and method signature in your binding’s reference, since conventions differ.

Drag and drop

The drag-and-drop gesture is a press and hold at the source, movement to the destination, then release. Selenium provides a convenience helper when both endpoints are elements.

// Java
new Actions(driver).dragAndDrop(source, destination).perform();

// Python
ActionChains(driver).drag_and_drop(source, destination).perform()

// JavaScript
await driver.actions().dragAndDrop(source, destination).perform();

To move by a distance instead of to another element, use the offset helper:

// Java: drag from source by 120px right and 40px down
new Actions(driver).dragAndDropBy(source, 120, 40).perform();

// Python
ActionChains(driver).drag_and_drop_by_offset(source, 120, 40).perform()

Some pages implement drag behavior in ways that a simple drag helper does not trigger reliably. For more control, compose the press, movement, and release explicitly, and add pauses only when the page needs time between stages.

3. Chain gestures and control pointer actions

Action chains are useful when a workflow involves multiple related steps. Build the sequence first and execute it with perform(). A pause can give the page time to react between pointer movements; avoid arbitrary sleeps when an explicit page condition can be awaited.

// Java: hover, pause, then click
new Actions(driver)
    .moveToElement(target)
    .pause(Duration.ofMillis(300))
    .click()
    .perform();

// Python
from selenium.webdriver.common.action_chains import ActionChains

ActionChains(driver).move_to_element(target).pause(0.3).click().perform()

// JavaScript
await driver.actions()
  .move({ origin: target })
  .pause(300)
  .click()
  .perform();

For interactions not covered by convenience methods, the low-level Actions API exposes pointer movement, button down, button up, and other device actions. When coordinating multiple input devices at this level, the caller is responsible for synchronizing their action sequences. If an action fails after a button press, release or reset the input state as appropriate before reusing the session.

4. Choose element targets or coordinates

Approach Use it when Watch for
Element-based action The page identifies the control as a DOM element and its center is a suitable target. The element may be outside the viewport, covered, stale after a rerender, or have an interaction area that is not centered.
Element-relative offset A specific point within a known element is required. Offset origins and method signatures vary by binding; verify the coordinate reference.
Pointer-relative offset The gesture needs to continue from the current pointer position. The pointer’s starting location matters, and coordinates must stay in the viewport.
Low-level pointer sequence A convenience helper cannot express the timing or movement sequence required. You must sequence and release input correctly; multiple device sources require synchronization.

5. Make interactions more reliable

  1. Wait for the actual interaction state. Wait for visibility or application readiness before acting. Presence in the DOM alone does not mean the target is usable.
  2. Keep targets in view. Hover and coordinate movement depend on viewport position. Scroll the element into view when necessary, then recalculate any coordinates that depend on layout.
  3. Use stable locators. Avoid selecting by a fragile position when a stable ID, accessible role, or other durable locator is available.
  4. Account for overlays and animation. A cookie banner, modal, loading layer, or animation can intercept the pointer. Wait for it to disappear or interact with it intentionally.
  5. Use pauses selectively. Add a pause between actions only if the page needs time to respond. Prefer a wait for a visible state over a fixed delay when possible.
  6. Recover from held input. If a sequence stops after pressing a button, release it or clear/reset action state using the mechanism available in the binding and driver.

6. Troubleshoot common failures

Symptom Likely cause Fix
Element not interactable or click intercepted The target is hidden, covered, animating, or not in a usable state. Wait for visibility and readiness, handle the overlay, and confirm the intended target is clickable.
Hover or offset move errors The element or requested pointer coordinate is outside the viewport. Scroll the target into view and keep the resulting pointer location within the viewport.
Drag-and-drop does not move the item The page’s drag behavior does not respond to the helper’s default press-and-move sequence, or an overlay intercepts it. Verify source and destination, ensure the source is visible, and compose press, movement, and release steps with suitable timing if needed.
Context menu does not appear The site handles or suppresses the context-menu event, or the wrong target received the gesture. Confirm the target and test the page’s expected context-click behavior; the browser menu itself may be suppressed by the application.
Double-click triggers only one action The target moved, rerendered, or did not accept two clicks at the expected location. Wait for stable layout, target the intended element, and check whether the application handles double-clicks at all.
Later actions behave as if a button is held A prior chain pressed a mouse button and failed before release. Send a release or clear/reset the input state using the binding and driver mechanism.
Method name or argument error Examples differ across language bindings or Selenium releases. Use the matching binding’s current API reference and its method spelling and argument order.

7. Performance, reliability, and cost

Mouse actions are browser commands, so each additional action and wait adds time to an automation flow. Keep chains focused, avoid fixed delays where a state-based wait is available, and do not add coordinate calculations when an element target expresses the intent more robustly. A longer chain can reduce round trips, but a failed intermediate interaction still needs diagnosis and recovery.

Reliability depends on page state, viewport geometry, overlays, and the site’s event handling. Element-based actions are generally easier to understand when the target is a real control; coordinates are appropriate when the interaction is inherently spatial, such as a canvas. Selenium itself does not charge per mouse action; costs depend on your browser infrastructure and any third-party services you use.

Or skip the browser setup

If your goal is to capture the page after automation, ScreenshotNeo provides a website screenshot API and MCP server for developers. Its API returns a PNG, JPEG, WebP, or PDF from one GET request. It accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents and MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Every feature is on every plan. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Try ScreenshotNeo and sign up for 1,000 free screenshots a month, with no card.

Frequently asked questions

Does Selenium move a physical mouse?

No. The Actions API sends virtualized input actions to the browser through WebDriver.

Should I use an Actions click or WebElement click?

Use an Actions click when it belongs in a pointer sequence or requires pointer positioning. For a straightforward standalone control activation, a regular element click may be enough.

Can Selenium perform touch or pen input with the same API?

The pointer input source covers mouse, pen, or touch input, while the convenience methods shown here describe common mouse gestures.

Why does a hover menu close before I can click it?

The pointer may leave the menu’s hover region between commands, or the page may need time to render it. Move through the intended hover path and wait for the menu’s ready state before clicking.