ScreenshotNeo

BlogGuides

Getting Started with Visual UI Testing for Android Apps

Build reliable Android UI tests by combining deterministic test data, behavior assertions, and screenshot comparisons across the devices your users rely on.

By the ScreenshotNeo team4 October 20269 min read

Visual UI testing for Android combines two checks: a behavior test verifies that the app responds correctly to user actions, and a screenshot comparison checks whether the rendered screen still matches an approved image. Start with one important user journey, make its data deterministic, assert its behavior with the framework suited to your UI, then capture and compare the relevant screen. A screenshot by itself does not prove that the app works correctly or that it is accessible.

1. Choose one user journey and make it repeatable

Begin with a single flow such as signing in, saving an item, or completing a checkout step. A small, repeatable test is easier to diagnose than a broad tour through the app.

  1. Choose a screen and the user action that matters.
  2. Decide what observable result means the action worked, such as a confirmation label appearing or a saved item being shown.
  3. Arrange stable test data. Use a fake repository or replace external dependencies where possible so network responses, account state, and content do not change between runs.
  4. Choose the screen state to capture. Fix inputs such as text, date, and selected options so the baseline remains meaningful.

Android instrumented tests belong in the module’s src/androidTest/java source set. Android describes UI tests as launching an app or part of it, simulating interactions, and checking the response. Its UI testing guide also recommends an architecture that makes dependencies replaceable for tests. See Android’s UI testing guide.

2. Choose the Android UI testing framework

Need Good starting point Notes
Interact with classic View-based screens in the app Espresso Provides view actions and assertions. It waits for the main message queue, AsyncTask work, and configured idling resources to become idle.
Test Compose screens and components Compose UI testing APIs Use the Compose test APIs to find nodes, perform actions, and check semantics.
Operate outside the app process, across apps, or on system UI UI Automator Useful for flows involving permissions, notifications, launchers, or another app. The modern 2.4 API is documented as under development; check its current status before adopting it.
Run suitable UI tests on the JVM Robolectric Can shorten feedback loops for tests that work in its environment. Keep device-specific behavior checks on an emulator or device.
Compare appearance with an approved image Screenshot testing workflow Capture a known screen state and compare the result with a reviewed baseline. Select a library or workflow compatible with the app and build setup.

For a first in-app interaction test, Espresso is a sensible starting point for Views. Compose apps should use Compose testing APIs for Compose content. Espresso synchronization helps with common asynchronous UI work, but custom background work may require an idling resource or an explicit test design that waits for the relevant state. Official references: Espresso, Compose testing, and UI Automator.

3. Add a behavior assertion before the visual check

The following is a runnable Espresso test once the example app supplies a button with resource ID save_button and a confirmation view with ID saved_message. Replace MainActivity and those IDs with the app’s actual activity and resources. The assertion checks behavior; it is not a screenshot comparison.

package com.example.app

import androidx.test.ext.junit.runners.AndroidJUnit4
import androidx.test.rule.ActivityTestRule
import androidx.test.espresso.Espresso.onView
import androidx.test.espresso.action.ViewActions.click
import androidx.test.espresso.assertion.ViewAssertions.matches
import androidx.test.espresso.matcher.ViewMatchers.isDisplayed
import androidx.test.espresso.matcher.ViewMatchers.withId
import org.junit.Rule
import org.junit.Test
import org.junit.runner.RunWith

@RunWith(AndroidJUnit4::class)
class SaveItemTest {
    @get:Rule
    val activityRule = ActivityTestRule(MainActivity::class.java)

    @Test
    fun savingItemShowsConfirmation() {
        onView(withId(R.id.save_button)).perform(click())
        onView(withId(R.id.saved_message)).check(matches(isDisplayed()))
    }
}

The example uses AndroidX Test’s ActivityTestRule for brevity. If the project uses a newer test setup, use the activity launch rule already configured by that project. Keep the test dependency versions aligned with the app’s AndroidX Test configuration rather than copying arbitrary version numbers.

Compose alternative

For a Compose screen, use its test rule and semantics rather than searching for View resource IDs. This illustrative test assumes the screen exposes a button labeled “Save” and a confirmation text “Saved”.

@get:Rule
val composeRule = createComposeRule()

@Test
fun savingItemShowsConfirmation() {
    composeRule.setContent { SaveScreen() }
    composeRule.onNodeWithText("Save").performClick()
    composeRule.onNodeWithText("Saved").assertIsDisplayed()
}

Import the Compose testing APIs and JUnit rule that match the project’s Compose test dependencies. If the app uses a semantic label instead of visible text, query that semantics property so the test reflects the intended user-facing contract.

4. Capture and compare a screenshot

A screenshot test has two distinct steps: capture the screen in a controlled state, then compare the captured image with an approved reference. Saving an image for manual inspection is useful, but it is not an automated regression comparison until a comparison step checks it against a baseline.

  1. Run the behavior test and move the UI into the state to review.
  2. Capture the screen using the chosen screenshot testing library or Android tooling.
  3. Review the initial capture and approve it as a baseline only after confirming the intended layout and content.
  4. On later runs, compare the new capture with that baseline. Inspect the diff and the full images before accepting a changed baseline.
  5. Keep a behavior assertion for important interactions even when a screenshot comparison exists.

Android’s screenshot testing guidance defines this approach as capturing UI and comparing it with a previously approved image. UI Automator can also capture a full screen, window, or element and attach artifacts to Android Studio test results; a raw capture still needs a comparison or review process to serve as a visual regression check. See Android screenshot testing and UI Automator.

Keep baselines useful

  • Use stable fake content and predictable app state.
  • Capture at a defined device size, density, orientation, and locale.
  • Decide how dynamic content such as clocks, rotating promotions, and user-specific text is controlled or excluded.
  • Review rendering changes deliberately. Font rendering, OS versions, and device density can create legitimate pixel differences.
  • Do not treat a passing image comparison as proof of accessibility, correct navigation, or business logic.

5. Expand coverage across relevant configurations

Android devices vary by API level, locale, orientation, and form factor. Choose a focused set based on the app’s actual audience and risk; there is rarely value in running every possible combination on every change.

Configuration What it can expose Practical choice
API level / OS version Platform behavior and rendering changes Include supported versions that matter to the audience, plus a current version when useful.
Locale Longer translations, text direction, formatting, and missing strings Prioritize shipped locales and layouts sensitive to text length or direction.
Orientation Resizing, layout constraints, and state restoration Test both orientations if the app supports them or has orientation-sensitive screens.
Form factor Tablet, foldable, or other large-screen layouts Include the form factors the app claims to support.
Hardware Behavior tied to sensors, camera, biometrics, or vendor-specific behavior Use a physical device when the flow depends on hardware or emulator behavior is insufficient.

Android Studio emulators and local devices are useful during development. Robolectric can run suitable tests on the JVM. Firebase Test Lab can run scripted Espresso or UI Automator instrumentation tests on selected physical and virtual device configurations, and returns results with screenshots, videos, and logs. It also offers Robo exploration for an initial code-free pass. Physical-device runs may reveal issues not seen on emulators. Review the Firebase Test Lab documentation and current pricing and quotas before planning cloud usage; project use through the Firebase console requires the Blaze pay-as-you-go plan linked to Cloud Billing according to its Get Started guide.

6. Run locally, then use a device matrix

  1. Run the instrumented test on one emulator or device while developing.
  2. Inspect the assertion failure and captured artifacts before changing the test or baseline.
  3. Run a focused set of important configurations in continuous integration.
  4. Schedule a broader matrix when the change touches shared layouts, platform-specific behavior, or supported device classes.
  5. Use Firebase Test Lab or another workflow only after checking current setup, limits, and pricing.

Firebase documentation states maximum test durations of 45 minutes on physical devices and 60 minutes on virtual devices. These are service limits, not recommended test lengths; check the current Firebase documentation for updated limits.

7. Troubleshooting common failures

Symptom Likely cause Fix
Espresso cannot find a view The view is not on screen yet, an ID changed, or the test is on the wrong screen. Check navigation and the actual resource ID; wait for the app’s real loading state rather than adding an arbitrary long sleep.
Test fails intermittently around asynchronous content Work continues outside Espresso’s synchronization signals. Expose an idling resource for app work or assert an observable loaded state. Avoid fixed delays unless a timed behavior itself is under test.
Compose node query matches nothing The text or semantics differ from the query, or content has not been composed. Inspect the Compose semantics tree, query a stable semantic property, and verify the expected screen is displayed.
Screenshot differs on every run Dynamic data, animations, system bars, time, or asynchronous loading changes the capture. Use deterministic data, settle the UI, control animation where possible, and standardize device and system configuration.
Many pixels differ after an OS or emulator update Rendering, fonts, system UI, or density changed. Compare the old and new full images and diff, confirm the intended rendering, then update baselines intentionally for the affected configuration.
Test Lab run fails but local emulator passes Device, OS, locale, permissions, or hardware behavior differs. Use returned logs, video, and screenshots to identify the failing configuration; reproduce locally with the closest available configuration.
Screenshot artifact exists but no regression failure occurs The workflow captured an image but did not compare it with a baseline. Add a compatible baseline comparison step or treat the artifact as a manual review aid, not an automated visual test.

8. Performance, reliability, and cost

Run a small, deterministic suite frequently and reserve broader device matrices for changes that warrant them. UI tests cost more time than isolated unit tests because they launch and render the app; stable setup, focused journeys, and avoiding needless repeated navigation keep feedback manageable.

Reliability depends on controlling state and waiting for meaningful conditions. Avoid external live services for core UI assertions when a fake dependency can provide the same state. Keep behavior checks and image comparisons separate enough that a baseline update cannot hide a broken interaction.

Local emulator and device runs use local development resources. Cloud device execution has service-specific quotas and billing terms that can change; consult the provider’s current pricing documentation before scaling runs. Firebase Test Lab’s console workflow requires a Blaze plan linked to Cloud Billing, so do not assume cloud runs are universally free.

Or skip the browser setup

Android app screenshot tests still need an Android test runner and a device or emulator. For a related web page—such as a website or web-based flow your app displays—ScreenshotNeo can capture the page through one API request. It is a website screenshot API and MCP server, not a replacement for native Android UI testing. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
  • Cookie banners are accepted like a visitor and removed along with 60+ known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off.
  • Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Response headers identify the page verdict and billing status.
  • An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
  • The free plan includes 1,000 screenshots a month with no card. Paid plans start at $5 for 3,000 screenshots.

Sign up for ScreenshotNeo’s free 1,000 screenshots a month, with no card required.

FAQ

Does screenshot testing replace Espresso or Compose assertions?

No. A screenshot comparison checks rendered appearance against an approved image. Keep behavior assertions to verify interactions and outcomes.

Can I use UI Automator for a normal in-app screen?

It can interact with UI, but Espresso or Compose testing APIs are generally the more direct starting point for in-app content. UI Automator is especially useful when the test crosses process or system UI boundaries.

Do I need a physical Android phone?

No. Emulators are useful for local development and cloud workflows can provide device configurations. A physical device is useful when the feature depends on hardware or when you need to check behavior on target hardware.

How many device combinations should I test?

Start with configurations tied to your supported audience and the risks in the changed code. Add API levels, locales, orientations, and form factors where they can change the result.