Mobile App Testing: How to Get It Right
Learn how to test mobile apps with a risk-based workflow for user journeys, devices, automation, accessibility, and release checks.
To test a mobile app well, start with the user journeys whose failure would matter most, then exercise them on the device, operating system, account, language, network, and app states your users actually encounter. Combine repeatable automated checks with manual exploration, accessibility testing, platform release reports, and feedback from real users. No single test run proves an app is free of defects: its results describe only the conditions it exercised.
This guide gives you a practical workflow for deciding what to test, how to cover Android and iOS devices, where automation helps, how to test accessibility, and how to triage findings before release.
1. Decide what matters before choosing tests
Testing every possible combination of device, operating system, language, network, and app state is rarely practical. Begin with risk: identify the user journeys and conditions where a failure would have the greatest effect, then build coverage around them.
Map critical user journeys
For your app, list the flows that would most hurt users or the business if they broke. Typical candidates include:
- First launch, onboarding, and permission requests.
- Account creation, sign-in, sign-out, and session recovery.
- The app’s primary task or content flow.
- A purchase or other high-value action, if the app has one.
- Recovery from errors, offline states, and interrupted work.
- Updating from a previous app version while retaining expected data.
For each journey, record what the test depends on: account state, permissions, network access, backend services, test data, and any expected recovery behavior. A journey map turns “the app launches” into concrete checks that reflect how people use the product.
Choose conditions based on your users
Decide which platform versions, device models, screen sizes, locales, and network conditions matter from your audience and your technical risks. Include conditions that may change the behavior being tested: a fresh install versus an upgrade, a signed-in versus signed-out user, a permission granted versus denied, or a strong versus unreliable connection.
A test result is useful only when you know what it covered. Record the device and OS, app build, locale, account and app state, network conditions, and the journey performed. A successful run on one device does not establish that another device or state behaves the same way.
2. Build a layered test workflow
Use different testing methods for different risks. Automation is good at repeating stable checks after changes; people are better at exploring unclear behavior, awkward transitions, and interactions that a script does not model well.
- Run fast automated checks on meaningful changes. Keep stable, high-value unit and UI checks repeatable so regressions are easier to spot.
- Explore manually. Walk through critical journeys and try interruptions, unusual inputs, denied permissions, recovery paths, and other edge states.
- Use platform reports and targeted device tests. Treat automated crawling as a useful signal, then fill gaps that matter for your users.
- Test accessibility throughout. Combine assistive technology, analysis tools, automation, and feedback from people with disabilities.
- Triage, fix, and rerun. Retest the affected journey and related regression checks under the conditions relevant to the fix.
This combination gives repeatable coverage without mistaking a passing automated run for proof that every defect has been found.
3. Choose device and OS coverage deliberately
Start with the platform versions and device sizes that match your audience and technical risks. Add focused coverage for any important device, locale, or behavior that your general release checks do not represent.
For Android, Google Play’s pre-launch report selects devices for breadth. Its device mix can vary and takes factors such as popularity, crash frequency, resolution, manufacturer, and OS version into account. That makes the report a useful baseline sample, not a guarantee that every important device or configuration is covered. See Google’s guidance on using a pre-launch report and understanding its limits.
| Condition | Why include it | Example question |
|---|---|---|
| OS version | Platform behavior and compatibility can differ. | Does the critical flow work on the versions important to our users? |
| Device and screen size | Hardware, resolution, and layout can affect behavior. | Are controls visible and usable on our common device sizes? |
| Locale and language | Text, content, and layout may change. | Does the localized screen remain readable and complete? |
| Network | Requests may be slow, interrupted, or unavailable. | Can users understand and recover from a failed request? |
| Account and app state | New, returning, and partially configured users follow different paths. | Does the app handle an upgrade or expired session as intended? |
| Permissions and interruptions | Denied permissions, dialogs, or interruption can change a journey. | Can the user continue or recover after denying access or returning? |
You do not need a physical device for every useful check. Platform tools can extend coverage, but when a feature depends on actual hardware behavior, add tests on relevant physical devices.
4. Use Google Play’s pre-launch report as an Android release signal
After you upload an app bundle or save a production release, Play Console can install the app on a set of test devices and crawl it for several minutes. The crawler performs basic actions such as typing, tapping, and swiping. The report can flag stability, compatibility, performance, and accessibility issues.
Help the crawler reach useful screens
- If sign-in is required, provide a dedicated test account. Google recommends not supplying official credentials.
- Add a Robo script for a known user path when the general crawl does not reach it.
- For an OpenGL app or game, use a game loop where applicable.
- Add deep links for additional entry points; Google’s help page allows up to three.
- Set language preferences when language-specific content or layout matters.
- Review device-level screenshots, videos, stack traces, and performance details where available.
Know what the report cannot establish
Reports are generated automatically, subject to test-lab capacity. Some apps need changes to be crawlable: country checks or install validation, for example, can interfere. Test devices are located in the United States, so location-restricted content may not match what users elsewhere see. The crawler cannot make purchases, which means subscription or in-app product flows may be only partly exercised.
Google explicitly cautions that the pre-launch report cannot guarantee that tests will identify every issue. Treat a finding as something to investigate and a clean report as one useful signal, not a certification. See Google’s explanation of pre-launch report coverage and limitations.
5. Add targeted Android cloud tests when there is a coverage gap
Firebase Test Lab’s Robo test powers Google Play’s pre-launch report. Firebase documents using Robo tests to target particular devices, locales, and Android versions, and to run tests for longer durations. Use those controls when the broad pre-launch report leaves a meaningful gap—for example, when an important locale or OS version needs focused coverage. Read Firebase’s guide to testing beyond pre-launch reports.
Targeted cloud tests can broaden the conditions you exercise, but they do not replace physical-device checks when hardware behavior is central to the app. Choose the combination that matches the feature and risk you are validating.
6. Test accessibility with more than automated checks
Accessibility belongs throughout the quality workflow. Android’s guidance describes four complementary approaches:
- Manual interaction: use accessibility services such as TalkBack to experience the app with assistive technology.
- Analysis tools: use tools to find potential accessibility improvements and investigate what they report.
- Automated UI tests: test accessible properties and interactions. For Compose, testing APIs can locate elements, check attributes, and perform actions through the semantics tree that accessibility services read.
- User testing: learn from people who use the app, including people with disabilities. Their feedback can reveal usability problems that automated checks do not describe.
Automated accessibility results are signals to investigate. They do not replace using assistive technology or observing people using the app. Android’s guide, Test your app’s accessibility, explains these approaches and why testing from the user’s perspective can expose missed usability issues.
For iOS, Apple’s documentation describes accessibility testing using accessibility settings and assistive technologies, as well as accessibility audits in UI tests. Consult Apple’s accessibility testing documentation for current platform guidance.
7. Triage findings and verify fixes
Prioritize a finding by its effect on users, how broadly it reproduces, and whether it blocks a critical journey. A crash on sign-in deserves more attention than a cosmetic issue on an uncommon screen, though both may merit follow-up.
- Capture the failing build, device and OS, app and account state, locale, network conditions, and reproduction steps.
- Use available screenshots, videos, logs, stack traces, and performance details to narrow down the cause.
- Fix the issue and rerun the affected journey under the conditions where it occurred.
- Run related regression checks and record exactly what the new result covers.
Google Play groups pre-launch findings by severity and can provide device details and diagnostic material such as screenshots, videos, and stack traces. Those details help investigation; the practical goal is still to reproduce, fix, and retest the affected behavior.
8. Capture screenshots for visual checks and bug reports
Screenshots help teams compare a screen before and after a change, document a visual defect, or attach evidence to a bug report. A screenshot is a record of one rendered state, so pair it with the device, OS, locale, app build, and app state that produced it. It cannot show interactions or establish that a whole journey works.
You can capture screens on your own device or use your existing platform workflow. For repeatable web-page captures used in documentation, test fixtures, or issue reports, ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. It is useful for web content in an app’s workflow; it does not replace native app testing on devices.
Or skip the browser setup
For a web page used in a test fixture or report, one GET request returns an image or PDF. See the ScreenshotNeo API documentation for options and response details.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const image = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', image));
ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are never billed, and response headers identify the page verdict and billing status. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. Sign up for 1,000 free screenshots a month, with no card required.
9. Keep test runs useful, reliable, and affordable
Testing costs time to build, maintain, and investigate, as well as any service or device costs. The research sources here do not establish current prices or service-level promises for Google Play or Firebase Test Lab, so check their current terms when comparing costs or turnaround.
- Spend effort where failure matters. Automate stable checks for critical journeys and use exploratory testing for behavior that is uncertain or difficult to script.
- Keep test setup controlled. Use dedicated accounts and known data so failures are easier to reproduce. Avoid relying on personal or production credentials.
- Make failures diagnosable. Save build and environment details plus useful screenshots, video, logs, or stack traces when available.
- Account for variability. Device selection and test-lab capacity can vary; rerun and investigate findings rather than interpreting a single run as complete coverage.
- Choose device access to fit risk. Cloud tools can add targeted coverage. Use physical devices when the behavior depends on real hardware.
Common mobile app testing problems and fixes
| Problem | Likely cause | What to do |
|---|---|---|
| The crawler never reaches a signed-in screen. | The test flow requires authentication or a known path. | Provide a dedicated test account and add a Robo script or deep link for the journey. |
| Location-specific content is missing or different. | Pre-launch report devices are located in the United States, or the app restricts content by country. | Check whether location or install validation blocks crawling, then add focused tests for relevant conditions. |
| A purchase flow appears untested. | The pre-launch crawler cannot make purchases. | Test the purchase journey through an appropriate controlled test path; do not infer purchase correctness from the report. |
| A report has no result yet. | Automated report generation is subject to test-lab capacity. | Check report status and allow for processing; use another targeted check for a time-sensitive release decision. |
| A test passes, but users still report a failure. | The run did not exercise the affected device, OS, locale, network, account, or app state. | Collect reproduction conditions and add a focused test for that combination. |
| An accessibility scanner reports no issues, but a screen remains hard to use. | Automated checks cannot capture every assistive-technology or usability problem. | Try the flow with assistive technology and include feedback from people with disabilities. |
| A fix appears to work locally but fails in a release report. | The reported environment or state differs from the local setup. | Compare build, device, OS, locale, account state, and network; reproduce under the report’s conditions. |
Frequently asked questions
How do I test a mobile app before launch?
Map critical journeys, automate repeatable checks, explore edge states manually, test accessibility, and use platform reports or targeted device coverage. Review failures, fix them, and rerun the affected paths with the relevant conditions recorded.
What should I test in a mobile app?
Test the journeys that matter to your users: onboarding, account access, the primary task, high-value actions, error recovery, and updates where relevant. Include the devices, OS versions, locales, permissions, networks, and app states that can change those journeys.
Can automated mobile app testing catch every bug?
No. Automation can repeat defined checks and platform crawlers can reveal useful issues, but a result only covers the conditions exercised. Use manual exploration and user feedback to find problems outside those paths.
How do I test an app on different devices?
Select devices and OS versions based on your audience and technical risks. Use broad platform sampling as a baseline, add targeted tests for missing important conditions, and include physical devices when hardware behavior matters.
How do I test mobile app accessibility?
Use assistive technology manually, run analysis and automated UI checks, and include feedback from people with disabilities. Treat tool findings as prompts to investigate rather than a complete measure of usability.


