ScreenshotNeo

BlogHow-to

How to Run Visual Tests on Android Apps with Appium

Add screenshot assertions to Android Appium tests with UiAutomator2 and the Images plugin. Set up repeatable screens, choose a comparison mode, and maintain reliable baselines.

By the ScreenshotNeo team4 October 202610 min read

A visual test for an Android app uses Appium to reach a known screen, captures the rendered screen, and compares it with an approved reference image. Appium’s Images plugin provides three comparison modes: getSimilarity for equal-size images, matchTemplate to find a smaller image inside a screenshot, and matchFeatures for images that may be transformed. Choose the mode based on the question your test asks; no single score or threshold is right for every screen.

This guide uses Appium’s maintained UiAutomator2 driver and its Images plugin. The exact client-library syntax varies. The Python example below uses Selenium’s Python client with Appium’s documented WebDriver endpoint; treat it as an adaptable integration example and check your client’s current examples for session setup and screenshot calls.

1. Understand the visual assertion

A screenshot comparison is an assertion added to a normal UI test:

  1. Start a session on the intended Android target.
  2. Navigate to the screen and establish a predictable app state.
  3. Capture the current screenshot and load the approved baseline.
  4. Compare the images using the mode that fits the assertion.
  5. Fail according to a threshold or match rule chosen for that screen.
  6. Keep the screenshot and useful comparison output as failure artifacts.

Appium Inspector can help during development: it can issue commands, show the app hierarchy, and display screenshots. Baseline naming, storage, approval, and update policy are decisions for your test workflow; the plugin does not automatically create or approve baselines.

2. Install and start the Android session

Install Appium and the UiAutomator2 driver according to the official setup instructions, then start the Appium server. The Appium ecosystem lists UiAutomator2 as a maintained driver for Android native, hybrid, and web automation. The Images plugin must also be installed and enabled on the server before its comparison endpoint is available.

Appium-specific capabilities use the appium: namespace when written as W3C capabilities. The key capabilities for a basic Android session are:

Capability Purpose
appium:automationName Selects the driver, commonly UiAutomator2.
appium:app Points to an installable application path.
appium:udid Selects a particular device when more than one target is available or a specific device is required.

For example, the capability payload conceptually looks like this; configure it in the format your Appium client expects:

{
  "platformName": "Android",
  "appium:automationName": "UiAutomator2",
  "appium:app": "/absolute/path/to/app.apk",
  "appium:udid": "emulator-5554"
}

The udid can identify an emulator or a connected physical device. Omit it if your setup has a single unambiguous target and your client/server configuration selects it appropriately. Consult the [Appium Session Capabilities documentation](https://appium.io/docs/en/latest/guides/caps/) and your chosen client’s examples for complete setup syntax.

3. Make the screen repeatable

A comparison is useful only when the test reaches a comparable state. Keep the navigation path and test data stable, and wait for the visual state under test rather than relying on a fixed pause alone.

  • Use deterministic test accounts and fixtures where possible.
  • Control or exclude time-sensitive content, rotating promotions, and remote data when they are not the behavior being tested.
  • Wait for a meaningful screen element or state before taking the screenshot.
  • Keep device configuration consistent, including screen size, orientation, font/display settings, and app build.
  • Decide explicitly whether system bars, animations, cursor/focus state, and transient notifications belong in the assertion.

These are workflow recommendations: Appium does not automatically normalize app state, mask changing regions, or choose a baseline for you. If different device configurations legitimately render different layouts, keep separate baselines or assert an invariant such as the presence and location of a key element.

4. Capture a screenshot and compare it

The Images plugin documents POST /session/:sessionId/appium/compare_images. The request supplies a comparison mode, two base64-encoded image files, and mode-specific options. In a test, the current screenshot can come from the Appium client; the expected image can be read from versioned test data.

Below is a Python example of the integration shape. It assumes an Appium server at http://127.0.0.1:4723, an Images plugin enabled on that server, a working Android target, and an existing baseline at baselines/home.png. It uses Selenium’s Python WebDriver client for the session and standard Python libraries to call the documented plugin endpoint. Adapt session construction to the client and server version in your project. This is an example, not a claim that it was executed.

import base64
import json
import os
from pathlib import Path
from urllib.request import Request, urlopen

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait

APPIUM = "http://127.0.0.1:4723"
BASELINE = Path("baselines/home.png")
ARTIFACTS = Path("artifacts")
ARTIFACTS.mkdir(exist_ok=True)

options = webdriver.ChromeOptions()
options.platform_name = "Android"
options.set_capability("appium:automationName", "UiAutomator2")
options.set_capability("appium:app", os.environ["ANDROID_APP_PATH"])
# Set this when selecting a particular emulator/device:
# options.set_capability("appium:udid", "emulator-5554")

driver = webdriver.Remote(command_executor=APPIUM, options=options)
try:
    wait = WebDriverWait(driver, 30)
    # Replace with an element that proves the target screen is ready.
    wait.until(EC.visibility_of_element_located((By.ID, "com.example.app:id/home_title")))

    actual_png = driver.get_screenshot_as_png()
    (ARTIFACTS / "home-actual.png").write_bytes(actual_png)
    expected_png = BASELINE.read_bytes()

    payload = {
        "mode": "getSimilarity",
        " firstImage": base64.b64encode(actual_png).decode("ascii"),
        "secondImage": base64.b64encode(expected_png).decode("ascii"),
    }
    # The Images plugin endpoint is scoped to the active session.
    request = Request(
        f"{APPIUM}/session/{driver.session_id}/appium/compare_images",
        data=json.dumps(payload).encode("utf-8"),
        headers={"Content-Type": "application/json"},
        method="POST",
    )
    with urlopen(request, timeout=30) as response:
        result = json.load(response)

    print("Image comparison:", result)
    # The endpoint returns a similarity score. Set the accepted rule for
    # this screen based on reviewed baseline variation and test purpose.
    score = result.get("value", {}).get("score")
    if score is None:
        raise AssertionError(f"No similarity score in plugin response: {result}")
    minimum_score = float(os.environ.get("MIN_SIMILARITY", "0.98"))
    if score < minimum_score:
        raise AssertionError(f"Similarity {score} is below {minimum_score}")
finally:
    driver.quit()

In the JSON above, use the exact parameter names required by the Images plugin version installed on your server; its endpoint documentation is authoritative. The sample’s illustrative similarity cutoff is deliberately configurable and is not a universal recommendation. A real test should confirm the response shape and option names against the installed plugin documentation.

For a direct template-search call, the request concept is:

POST /session/SESSION_ID/appium/compare_images
Content-Type: application/json

{
  "mode": "matchTemplate",
  "firstImage": "BASE64_SCREENSHOT",
  "secondImage": "BASE64_SMALL_REFERENCE",
  "options": {
    "threshold": 0.8
  }
}

Use the documented request schema for your plugin version. Save the actual screenshot when the assertion fails; it makes layout drift and state mistakes much easier to diagnose.

5. Choose the comparison mode

Mode Best fit Result and cautions
getSimilarity Compare two equal-dimension images as a whole. Returns a similarity score. It does not answer where a local mismatch occurred. Keep image dimensions aligned.
matchTemplate Find a smaller reference, such as an icon or visual marker, inside a larger screenshot. Returns a match rectangle and score. Supports a threshold, multiple matches, and optional visualization. The documented score range is 0.0–1.0 and the documented default threshold is 0.5.
matchFeatures Match images that may be rotated, scaled, or otherwise modified. Uses OpenCV feature detector and descriptor matcher options. Choose it when geometric changes are part of the question, rather than for an ordinary pixel-stable whole-screen regression.

Examples of distinct assertions:

  • “Did the checkout screen remain visually close to its approved rendering?” Use whole-image similarity, with controlled dimensions and a reviewed acceptance rule.
  • “Is this warning icon visible somewhere in the current screen?” Use template matching and inspect the returned match location and score.
  • “Can this logo be recognized after a scale or rotation change?” Consider feature matching and configure its detector/matcher options for the transformation involved.

Do not use the template threshold of 0.5 as a blanket pass rule. It is a documented default parameter, not evidence that it is suitable for your screen or test intent. Similarly, establish a meaningful similarity cutoff by reviewing stable runs and known regressions; avoid tuning it only until the current build passes.

6. Keep baselines and failures maintainable

  • Name baselines with the screen, app version or fixture, device profile, and orientation when those affect rendering.
  • Store expected images under version control or another reviewable artifact process.
  • Review changes to baselines deliberately. Do not replace the expected image automatically after every failure.
  • On failure, retain the actual image, the expected image, comparison result, device/configuration details, and relevant app logs.
  • Use plugin visualizations when available for template matching to understand why a region matched or did not.
  • Separate visual behavior from incidental content. Prefer stable fixtures or a targeted template assertion if the whole screen contains irrelevant dynamic areas.

Where device rendering differs, either maintain a baseline per target configuration or write an assertion around the visual property that must remain invariant. A single baseline should not be expected to represent materially different screen layouts.

7. Troubleshooting

Symptom Likely cause What to do
Session creation fails or the driver is not found. UiAutomator2 is missing, the Appium server is not running, or the automation capability is malformed. Install/configure the driver, check server logs, and provide appium:automationName with the expected value.
The wrong emulator or phone receives commands. Several targets are available and no specific device was selected. Set appium:udid to the intended device identifier and confirm it is connected/available.
Comparison route returns 404 or is unavailable. The Images plugin is not installed, enabled, or exposed by the running server. Check the plugin installation and server startup configuration, then verify the endpoint for the installed plugin version.
getSimilarity rejects images or results are misleading. The images have different dimensions, or the test state/device profile changed. Capture at the same screen size and orientation, verify the correct baseline, and inspect both images directly.
Template match misses a visible target. The template differs in scale, crop, antialiasing, or appearance, or the threshold is too strict. Use a clean representative crop, inspect output and visualization, and tune the threshold against known positive and negative cases.
Template matching reports false positives. The target is visually generic or threshold is too permissive. Use a more distinctive template, constrain expected location if supported by your workflow, raise the acceptance bar based on reviewed cases, or assert another property too.
Similarity fluctuates between runs. Dynamic text/data, animation, asynchronous loading, system UI, or differing device settings affect pixels. Make inputs deterministic, wait for the ready state, stabilize animation/content, and align the target configuration.
The test fails before comparison because the screen is not ready. A fixed sleep is too short or synchronization checks the wrong element. Wait for a visible, screen-specific readiness condition and capture only after it is satisfied.

8. Performance, reliability, and cost

Image comparison adds work after screenshot capture. Whole-image similarity processes the full pair of images; larger captures generally carry more data through the client and plugin. Keep visual assertions scoped to the question: a targeted element check can be more stable and easier to inspect than a full-screen comparison when only one visual element matters. The cited Appium documentation does not establish performance benchmarks, so measure the added time in your own test environment.

Reliability depends on repeatable state, consistent target configuration, and diagnosable artifacts. Run comparisons against known-good and known-bad examples while choosing acceptance rules, and watch for false positives and false negatives as the app evolves. Treat baseline updates as reviewed test changes.

The Appium references do not define a price for the Images plugin or a hosted visual-testing service. Your costs depend on the devices or emulator infrastructure, CI runtime, and artifact storage you choose. An emulator is a valid target; a physical phone is optional when the test specifically needs hardware behavior.

9. Or skip the browser setup

Appium visual testing is for Android app screens. For website screenshots, [ScreenshotNeo](https://screenshotneo.com) offers a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF; use Appium for the native app test and ScreenshotNeo when the target is a webpage.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the [ScreenshotNeo API documentation](https://screenshotneo.com/docs/) for request options. ScreenshotNeo removes cookie/consent banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. [Create a free ScreenshotNeo account](https://screenshotneo.com/account/sign-up/).

10. FAQ

Does the Images plugin compare Android app screenshots directly?

Yes. Capture an image from the Appium session and provide it with the reference image to the plugin’s comparison endpoint. Your test code decides how the result becomes a pass or failure.

Is a physical Android phone required?

No. An emulator can be the session target. Use a physical device when the behavior under test depends on hardware or a particular device configuration.

Should I use visual comparison instead of accessibility or element assertions?

Use each for the property it can establish. Element and accessibility assertions test structure and semantics; screenshot comparisons test rendered appearance. Many suites benefit from both.

Can one threshold work for every screen?

Usually that is a poor assumption. Screens differ in dynamic content, visual density, and the consequence of a mismatch. Choose and review rules per assertion.

Primary references