ScreenshotNeo

BlogHow-to

How to Test Mobile App Security Workflows on Real Devices

Build a repeatable Android and iOS security-testing workflow on real devices, mapped to OWASP MASVS and MASTG, with practical checks and reporting guidance.

By the ScreenshotNeo team4 October 202611 min read

To test mobile app security on real devices, define the app’s threat model and authorized scope first, map applicable OWASP MASVS controls to verification work in the OWASP MASTG, then exercise the app and its backend on representative Android and iOS devices. Combine repeatable automated checks with manual verification, and record enough device, build, account, and evidence detail for another person to reproduce each result.

MASVS provides security requirements; MASTG provides testing guidance and resources to verify them. They can be used together or separately depending on the assessment objective. Select tests based on the app’s architecture, data sensitivity, platform features, and threat model; a checklist is a planning aid, not a claim that every test applies to every app.

1. Set authorization, scope, and test conditions

Before installing a build, agree on written authorization and boundaries. A mobile assessment can touch production services, personal data, identity systems, or third-party integrations. Use a designated test environment and accounts wherever possible.

  • Build: record version, build identifier, signing/build variant, and source revision if available.
  • Devices: record make/model, OS version, stock or modified state, and relevant hardware features.
  • Accounts: prepare test users for each role, including limited-privilege and administrator roles where relevant. Use synthetic data.
  • Backend: document environment, test endpoints, data reset procedure, and actions that must not be performed.
  • Instrumentation: state whether debugging, traffic inspection, rooted/jailbroken devices, or other instrumentation is allowed.
  • Out-of-scope systems: list third-party services, production tenants, and user accounts that must not be touched.

Keep credentials and sensitive evidence out of issue trackers and screenshots. Redact tokens, personal data, and secrets before sharing a report.

2. Build a control-to-test plan

Start with the MASVS control groups that fit the app: storage, cryptography, authentication and authorization, network communication, platform interaction, code quality, resilience, and privacy. Use the MASTG guide and checklist to select concrete verification cases. The OWASP Mobile Application Security Cheat Sheet offers additional practical guidance, but it does not replace app-specific verification.

Area Questions to turn into checks
Storage What sensitive information persists, where does it persist, and what protects it?
Cryptography Are platform-appropriate key and cryptographic APIs used for the data and threat model?
Authentication and authorization Can expired, revoked, altered, or lower-privilege sessions still perform sensitive actions?
Network Are remote exchanges protected for confidentiality and integrity, with appropriate certificate validation?
Platform interaction Can links, intents, extensions, widgets, shortcuts, or app-to-app boundaries expose sensitive actions or data?
Code quality and resilience Are debug settings, dependencies, build configuration, and relevant tampering defenses appropriate to the threat model?
Privacy Does the app collect, retain, expose, or share more information than the workflow requires?

For each selected test, write down its expected result, preconditions, device/platform applicability, and evidence to collect. Mark tests as applicable, not applicable, not tested, or blocked; do not turn an untested item into a pass.

3. Choose representative real devices

There is no universal phone or fixed device count that proves an app secure. Select devices that reflect supported operating systems and the app’s actual user population. For Android, account for manufacturer and OS variation. Secure-hardware availability differs across devices, and some devices run older Android versions. A single recent flagship cannot establish behavior across that range.

Choose candidate devices against these practical factors:

  • Supported OS versions and planned upgrade coverage.
  • Manufacturer and OS variation for Android.
  • Hardware-backed key storage, biometrics, or other security features the app uses.
  • App-specific capabilities such as NFC, camera, eSIM, or connected accessories.
  • Stock versus rooted or jailbroken state, and whether instrumentation is needed.
  • Whether the same device state and build can be preserved to repeat a test.

Emulators can help with repeatable development checks, but they do not demonstrate behavior that depends on real hardware. Record the exact device and OS for every observation, including any non-default security state.

4. Exercise local data and privacy behavior

Use synthetic test data and inspect what remains on the device during normal use, after interruptions, and after logout. Focus on data the app handles, such as tokens, profile details, messages, payment settings, or documents.

  1. Start from a documented app and device state.
  2. Perform one workflow, such as signing in, viewing sensitive data, editing it, or uploading a file.
  3. Inspect storage and exposure points available to your authorized test setup: files, databases, preferences, logs, caches, keyboard suggestions, screenshots or background snapshots, backups, and data shared through platform mechanisms.
  4. Repeat after backgrounding, force-closing, restarting, logging out, switching accounts, and using any relevant backup or device-lock flows.
  5. Record where the data appeared, the action that created it, whether it was protected, and the relevant MASVS control and MASTG test.

Check whether sensitive values are minimized and protected with platform-appropriate storage and key APIs. Consider cloud backup, keyboard caches, lost-device access, inter-process communication, and data left after account changes. A value absent from one inspected location is not proof that it is absent everywhere.

5. Test identity, sessions, and backend authorization

Test security boundaries at the server as well as in the app UI. A client-side screen or button restriction can often be bypassed; the backend must enforce access to protected operations and data.

Exercise the flows that apply to the app:

  • Login, logout, session renewal, and session expiry.
  • App restart, device lock and unlock, and account switching.
  • Biometric unlock and the fallback path.
  • Role changes and account deactivation.
  • Reauthentication for sensitive operations, such as changing credentials or payment settings.
  • Requests with missing, expired, revoked, or altered session credentials against the authorized test backend.

Verify that revocable tokens and secure session handling behave as intended, and that a lower-privilege user cannot perform a privileged operation by calling the relevant backend path directly. Do not use real users’ accounts or production data for these checks.

6. Inspect network behavior

Capture app traffic only through an approved test setup. Verify encrypted transport, certificate validation, and the handling of sensitive request and response data. Exercise relevant network changes and failures, such as disconnecting during a request or retrying after a timeout, to see whether the app exposes data or leaves an operation in an unsafe state.

A proxy that cannot observe traffic does not by itself prove the connection is secure. Pinning, mutual TLS, or test-environment configuration can affect visibility. If the app uses certificate pinning, assess it against the threat model and operational requirements; pinning is not automatically required for every app. Treat any test instrumentation or control bypass as a scoped verification technique, not as the goal of the assessment.

7. Probe platform entry points and app boundaries

Inventory the integrations the app actually uses, then test their security assumptions. Depending on the platform and product, these may include permissions, deep links, universal links, Android intents, iOS app extensions, widgets, shortcuts, Siri integrations, and app-to-app data exchange.

  • Try relevant links and actions when logged out, after session expiry, and while the device is locked.
  • Check whether URL or intent parameters can select an unauthorized account, object, or action.
  • Verify that sensitive screens do not become available through widgets, shortcuts, notifications, or extensions without the intended checks.
  • Inspect whether another app or an unintended component can trigger functionality or read data through inter-process communication.
  • Review permissions in the context of actual app behavior and user-facing workflows.

Do not test platform features the app does not expose as though they were universal requirements. Hybrid and web-based apps can still have native components and platform boundaries worth checking.

8. Review code, integrity, and resilience

Review build configuration, dependencies, debug settings, binary integrity, and any applicable tampering defenses. Check that release builds do not expose development-only behavior or diagnostics that reveal sensitive information. Map each finding to the relevant control and the app’s stated threat model.

Root or jailbreak detection and anti-tamper controls are threat-specific defenses. Their presence does not prove an app is secure, and their absence does not alone establish a vulnerability. Evaluate what threat they address, how the app behaves when a check triggers, and whether the control creates operational or accessibility problems.

9. Combine automated checks with manual verification

Automate stable checks that can be repeated reliably, such as build configuration checks or deterministic workflow assertions. Use manual testing for context-dependent behavior, platform entry points, role boundaries, and unexpected states. OWASP describes using MASTG and its checklist as a baseline for manual assessment or as a template for automated tests.

For every result, retain:

  • Build identifier and test date.
  • Device model, OS version, and stock or modified state.
  • Account role and sanitized test-data description.
  • Preconditions and exact reproduction steps.
  • Observed result and appropriately redacted evidence.
  • Impact and the applicable MASVS control and MASTG test reference.
  • Status: passed, failed, not applicable, not tested, or blocked.

Keep evidence proportionate: a finding should be reproducible without retaining actual user data or unnecessary secrets. Attach sanitized logs, request details, or screenshots only where they clarify the observation. If you need a screenshot of a public web page as supporting documentation, ScreenshotNeo is a website screenshot API and MCP server; it captures web pages, not native mobile app screens.

10. Make the workflow repeatable

  1. Freeze the scope, build, accounts, backend, and device matrix for a test run.
  2. Reset app and backend state using documented procedures.
  3. Run the selected automated checks and record their versions and outputs.
  4. Follow the manual scenarios on each applicable real device.
  5. Reproduce any failure from a clean state and capture sanitized evidence.
  6. Map results to MASVS and MASTG references, then review gaps and blocked items.
  7. After fixes, retest the same scenario and affected neighboring flows on the relevant platform versions.

Preserving the build and device state matters: a changed OS, account role, backend configuration, or app build can change the result. If state cannot be preserved, document the difference and avoid presenting the rerun as an exact reproduction.

Performance, reliability, and cost considerations

Real-device testing takes longer than a single emulator run because devices, OS versions, accounts, and backend state all need coordination. Keep the first device matrix small and tied to supported platforms and risk; expand it when a test depends on manufacturer behavior, hardware-backed storage, biometrics, or app-specific peripherals. The sources do not prescribe a universal device count.

  • Reduce wasted runs: stabilize test data and reset steps, and separate environment failures from app findings.
  • Improve reliability: record exact device, OS, build, and account details; repeat important failures from a clean state.
  • Control evidence costs: collect only evidence needed to reproduce and explain the result, and sanitize it before sharing.
  • Plan coverage intentionally: prioritize controls based on sensitivity, architecture, and threat model instead of running every possible test on every device.
  • Account for backend behavior: coordinate test windows and data cleanup with the service owner so a client test does not cause unintended changes.

Troubleshooting common problems

Symptom Likely cause What to do
Traffic is invisible to the proxy Pinning, mutual TLS, or the test environment affects inspection. Confirm the approved setup and app configuration. Treat invisibility as an observation, not proof of secure transport; use an authorized verification approach suited to the threat model.
A test passes on one Android device but fails on another Manufacturer, OS, or hardware-backed storage differences. Record both device configurations, determine whether the behavior is supported, and expand representative coverage where the app promises support.
A protected action is blocked in the UI but succeeds through a request Authorization may exist only in the client. Reproduce only against the authorized test backend, document the request and account role, and route the finding to the backend authorization owner.
Local data appears after logout or app restart Data may remain in a cache, database, backup, log, or platform surface. Identify the exact location and lifecycle, test account switching and restart behavior, and verify intended deletion or protection.
Biometric testing behaves inconsistently Device enrollment, OS behavior, fallback rules, or test state differs. Record enrollment and lock state, repeat the same flow, and test fallback and reauthentication separately.
A result cannot be reproduced later Build, OS, backend, account, or device state changed. Use the recorded identifiers and preconditions; if exact state is unavailable, label the retest accordingly and capture the differences.
Automated checks report noisy failures Checks may depend on timing, network state, data, or platform version. Stabilize prerequisites, classify environmental failures separately, and manually verify security-relevant results before reporting them as findings.
Test evidence contains secrets or personal information Real data or unredacted credentials entered the test flow. Stop sharing the artifact, follow the organization’s data-handling process, replace the fixture with synthetic data, and capture sanitized evidence.

Or skip the browser setup

For screenshots of web pages used as supporting documentation, ScreenshotNeo provides a one-request website capture. It does not capture native mobile app screens.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo API documentation

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

ScreenshotNeo accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up for 1,000 free screenshots a month, with no card required.

Frequently asked questions

Can an emulator replace real-device testing?

Use emulators for repeatable checks, but include real devices when behavior depends on hardware, vendor or OS differences, or the app’s actual supported device surfaces.

Do I need to run every MASTG test?

No. Select applicable verification work based on the app’s architecture, data sensitivity, features, and threat model, and label exclusions and untested items clearly.

Does certificate pinning have to be enabled?

The cited guidance treats pinning as a consideration, not a universal requirement. Evaluate it against the app’s threat model and operational needs.

What makes a finding useful to the development team?

A reproducible step sequence, precise build and device details, sanitized evidence, impact, and a mapping to the selected MASVS control and MASTG test.

Primary OWASP resources