ScreenshotNeo

BlogHow-to

How to Test UI Localization Across Languages and Regions

A practical workflow for catching internationalization bugs early and validating localized interfaces across languages, regions, scripts, and devices.

By the ScreenshotNeo team4 October 202612 min read

Test UI localization in stages: first check that the code and source interface are ready for internationalization, then run a pseudolocalized build, choose a risk-based locale matrix, and validate real localized builds for functionality, visual fit, language quality, and market behavior. Automate repeatable checks, but include human visual review and native-language review for important content. Pseudolocalization finds engineering defects; it cannot prove that a real translation is accurate or culturally appropriate.

There is no universally correct number of locales to test. Choose language-region combinations that exercise the scripts, directions, formatting rules, devices, and product features your supported markets require.

1. Separate internationalization testing from localization testing

Internationalization (i18n) testing asks whether the product can support different languages, scripts, directions, and regional conventions without code changes or broken behavior. It should begin before translations are ready. Localization (l10n) testing checks whether each translated, market-specific version is complete, correct, usable, and appropriate for its intended audience.

Stage What it catches What it cannot establish
Code and source UI review Hard-coded text, fragile string assembly, fixed-size layouts, encoding and locale assumptions Whether future translations fit every screen
Pseudolocalized build Unextracted strings, clipping, expansion, wrapping, some font and direction problems Translation quality, terminology, cultural suitability
Real localized build Functional parity, language accuracy, visual fit, regional behavior Every possible locale and device combination unless the matrix covers them

Microsoft recommends code and source-interface review, pseudolocalization, and then starting localization with a small set of languages. Pseudolocalized text can expose truncation, concatenation, and strings missed by localization extraction. Microsoft’s internationalization testing guidance and its pseudolocalization overview describe these uses.

2. Review internationalization readiness before translation

Inspect the code and the source-language product before spending time reviewing translated screens. A defect in string handling or layout can affect every locale.

  • String resources: confirm all user-visible text comes from localizable resources, including validation errors, accessibility labels, notifications, emails, empty states, menus, and text embedded in images.
  • Sentence construction: avoid assembling sentences from fragments. Word order, agreement, punctuation, and grammatical gender can differ by language. Use complete translatable messages with named placeholders where the framework supports them.
  • Encoding and storage: handle Unicode end to end in source files, APIs, databases, search, sorting, and exports. Test supplementary characters and combining marks where relevant; a visible character may consist of more than one code point.
  • Layout: do not assume labels fit one line, text has a fixed length, or controls can grow only horizontally. Check wrapping, truncation, flexible containers, clipping, and text scaling.
  • Fonts: verify glyph coverage, fallback behavior, shaping, line height, and weight for the scripts you support. Check that downloaded or embedded fonts render the required characters.
  • Locale-sensitive logic: use platform or library formatters for dates, times, numbers, currencies, units, plural forms, and collation rather than hand-built patterns.
  • Input and data models: support the names, addresses, phone numbers, postal codes, keyboards, and input methods relevant to supported regions. Avoid restrictive validation that assumes one writing system or one field order.
  • Language and direction metadata: set the document or view language and base text direction correctly. W3C recommends UTF-8, language declarations, local formats for forms, and use of dir="rtl" for right-to-left HTML. See the W3C Internationalization Quick Tips.

3. Build a risk-based locale matrix

List the language and region combinations your product actually permits. Language and region may be selected independently, so test supported combinations rather than assuming that one language maps to one country. CLDR provides locale data and algorithms used by many platforms and applications for conventions such as dates, numbers, and units; see the Unicode CLDR project.

Risk axis What to include Example test focus
Direction At least one supported RTL locale if RTL is in scope, plus mixed-direction content Navigation order, alignment, icons, punctuation, embedded URLs and numbers
Script and fonts Representative non-Latin scripts and scripts with shaping or glyph-coverage demands Font fallback, combining marks, line height, search and input
Text length Languages likely to expand or contract relative to the source Buttons, narrow columns, dialogs, navigation, error messages
Formatting Regions with different date, number, currency, unit, calendar, or week conventions Display, parsing, sorting, boundaries, round trips
Market behavior Regions with different payment, address, legal, support, or availability flows Feature flags, eligible methods, copy, validation and fallback
Platform configuration Relevant OS, browser, app language, and region combinations System UI, locale fallback, language switching, persisted preference

Document why each locale is included and which risks it covers. A useful matrix maps locales to critical product journeys, not just to a list of translation files. Include high-risk flows such as account creation, checkout, recovery, error handling, and support if they exist in your product.

4. Run pseudolocalization early

Generate a pseudo language by accenting or expanding source strings, adding visible delimiters, and optionally inserting unique markers. For example, a source label like Save changes might become [Šááṽëë čħááñĝëš··]. The exact transformation is tool-specific; treat this as an illustration, not a required encoding.

  1. Enable pseudo resources or the platform’s pseudo locale in a development build.
  2. Make strings visibly different from source text so untranslated hard-coded content stands out.
  3. Add expansion sufficient to stress narrow controls and dialogs; include characters from scripts your font stack must handle.
  4. Exercise critical flows, including empty, error, loading, and authenticated states.
  5. Record clipping, missing delimiters, broken interpolation, layout overflow, and strings that remain untranslated.
  6. If RTL is supported, test an actual RTL locale or a pseudo mode that accurately models your product’s bidirectional behavior.

Apple documents accented pseudolanguages for checking whether views adapt to text with high and low ascenders in Preparing your interface for localization. Microsoft also describes pseudo locales that exercise mirrored and East Asian-like cases. Do not assume a mirrored pseudo mode accurately validates real RTL behavior; test the real direction and mixed-direction strings.

5. Test locale and region behavior

Run the same input and user journey under the locale and region settings the product supports. Check both presentation and parsing: an output can look plausible while the app interprets user input incorrectly.

  • Dates and times: display and parse local conventions, time zones, daylight-saving transitions where applicable, 12/24-hour preferences, and date boundaries. Check sorting and filters as well as labels.
  • Numbers and currency: verify decimal and grouping separators, negative values, rounding, currency placement, and currency selection. Ensure parsing does not silently convert or truncate values.
  • Units and calendars: confirm units, conversions, calendar assumptions, month names, and week starts where the product exposes them.
  • Pluralization: test zero, one, few, many, and other forms as relevant to the locale rules; do not assume English singular/plural logic generalizes.
  • Names and addresses: test alternate order, optional fields, long values, local scripts, and data that does not fit a first-name/last-name template.
  • Input methods: enter text using the relevant keyboard or IME, including composition, suggestions, paste, correction, and validation.
  • Language switching and fallback: change language at runtime if supported, persist the choice, verify the selected region, and check the fallback chain when a resource is missing.
  • Mixed-language content: combine localized labels with user names, addresses, numbers, URLs, or content in another script. Check bidirectional isolation and punctuation.

Use the platform’s locale-aware formatting APIs or a library backed by locale data. Avoid tests that assert one hard-coded punctuation pattern across all locales; assert semantic values and the expected formatted result for the selected locale.

6. Validate actual localized builds

Pseudolocalization is an engineering screen. For each supported locale, or a documented risk-based sample, test real localized resources and run representative user journeys from start to finish.

Functional parity

  • Every source-language feature and state is available where the market supports it.
  • Buttons, links, menus, forms, validation, search, sorting, and navigation behave correctly.
  • Localized strings use the intended resource and placeholders receive the right values.
  • Locale changes do not break saved preferences, cached content, URLs, or API data.
  • Regional differences in payments, availability, support, or policy are implemented correctly.

Visual completeness and fit

  • Check every critical screen for missing strings, clipping, overlaps, unexpected wrapping, truncation, and text hidden behind fixed controls.
  • Inspect typography, glyph coverage, line breaks, alignment, spacing, icons, and layout direction.
  • Check small and large viewports, text scaling, and relevant device orientations.
  • Review screenshots with a human; automated diffs can flag changes but may be noisy when fonts or platform rendering differ.

Linguistic and market review

  • Have a qualified reviewer check accuracy, terminology, tone, spelling, grammar, and consistency.
  • Review onboarding, payments, errors, legal content, and support text with extra care.
  • Check dates, examples, imagery, units, symbols, and market-specific assumptions for local appropriateness.
  • Give reviewers screen context and the source string or resource identifier so they can distinguish a translation issue from a layout or implementation issue.

Microsoft’s guidance distinguishes functional, visual, linguistic, and market-specific validation. A passing screenshot comparison cannot establish that wording is natural or correct.

7. Automate parity checks and capture evidence

Automate the checks that are deterministic and repeatable in your stack, then save evidence that helps reproduce failures.

  • Resource integrity: compare resource keys across locales, detect empty values, flag unexpected source-language leftovers, and validate placeholders.
  • Formatting tests: test date, number, currency, unit, plural, and parsing behavior with locale-aware expected results.
  • Journey tests: run the same critical workflow under selected language-region configurations and verify behavior, accessibility names, and key content presence.
  • Visual regression: capture the same route, state, viewport, and locale; compare against a baseline while allowing for known platform and font differences.
  • Bug records: save locale, region, OS/browser, app version, screen, viewport, steps, expected and actual behavior, and a screenshot. Label whether the root cause is shared code or locale-specific content.

For web screens, a screenshot API can capture rendered states after your test sets the locale, route, and viewport. ScreenshotNeo is a website screenshot API and MCP server; its options include viewport and device presets, full-page capture, element capture, custom CSS and JavaScript, waits, and custom headers. Use it as one evidence-capture option, not as a substitute for locale-aware test setup or human language review. Its documentation lists request parameters. For native app UI, use your platform’s UI test and screenshot facilities; Apple’s Xcode localization screenshot guide describes generating screenshots per localization and attaching string-to-frame context.

For screenshot comparisons, keep the viewport, browser or OS version, font availability, app data, and UI state consistent. A visual difference can come from rendering environment changes rather than a localization regression. Keep the original captures with test metadata so reviewers can reproduce the result.

8. Set release gates by risk

Define release criteria before the final localization pass. There is no source-backed universal locale count or numerical visual threshold, so set gates according to supported-market risk and defect severity.

  • Block release: broken critical flows, unreadable or missing critical text, incorrect amount/date interpretation, inaccessible controls, or high-severity market-specific errors.
  • Require disposition: clipping or truncation in noncritical areas, untranslated low-visibility strings, or minor terminology inconsistencies.
  • Record accepted risk: list the affected locale, screen, user impact, owner, and reason for deferral.
  • Re-run shared-code fixes broadly: a layout or formatting fix can affect many locales, so retest the locales that exercise the same risk axis.

9. Troubleshooting common failures

Symptom Likely cause What to do
Text is cut off or overlaps Fixed dimensions, no wrapping, or expansion not considered Allow flexible sizing and wrapping; test pseudo-expanded strings and small viewports.
Some screens remain in the source language Hard-coded strings, missing resource keys, or fallback masking gaps Use localization debugging or pseudo markers; audit resources and user-visible states.
A sentence sounds broken in translation Sentence assembled from separately translated fragments or placeholders in the wrong grammatical position Localize whole messages and give translators context for placeholders.
RTL screen looks partly reversed or punctuation moves Direction applied only with CSS, mixed-direction text is not isolated, or physical left/right assumptions remain Set semantic base direction, use direction-aware layout properties, and test real RTL plus mixed content.
Currency or date displays correctly but parses incorrectly Formatting and parsing use different locale assumptions or hard-coded patterns Use locale-aware formatters and parsers; test round trips and ambiguous input.
Characters become boxes or detach from marks Missing glyphs, unsuitable fallback font, or shaping/rendering issue Check font coverage and shaping on real target platforms and input paths.
Automated screenshot diff is noisy Different fonts, OS/browser versions, dynamic data, animation, or timing Stabilize the environment and state, wait for content, mask intentionally variable regions, and review diffs manually.
Locale test runs with unexpected language or region Application locale, OS locale, browser preferences, and region are configured independently Set and log each relevant setting explicitly; verify the app’s resolved locale and fallback.
Text input fails for a supported script Restrictive validation, normalization assumptions, or incomplete database/API support Test entry, composition, persistence, search, and round trips using representative script data.

10. Performance, reliability, and cost

Localization matrices multiply test work: if every route is captured in every locale, device, and viewport, the run can grow quickly. Prioritize high-risk combinations, parallelize independent runs where the test infrastructure allows it, and run a broader matrix on release candidates than on every small change.

  • Reuse deterministic test data and avoid changing locale settings midway through a test unless language switching itself is under test.
  • Wait for a meaningful stable condition, such as a key element or completed network activity, rather than relying on a short fixed delay alone.
  • Separate failures caused by shared code from translation defects so a broad rerun is targeted and useful.
  • Keep screenshots and metadata only as long as needed by your review and release process; visual evidence can contain private or user-like data, so use synthetic test data.
  • When using a capture service, account for the number of URLs, retries, and locales in your workflow and check its billing rules. ScreenshotNeo states that only clean shots are billed; bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses include X-Page-Verdict and X-Billed headers. Its plans include 1,000 shots per month free with no card, Starter at $5 for 3,000, Growth at $15 for 15,000, Pro at $39 for 60,000, Scale at $99 for 250,000, and Business at $249 for 1,000,000; yearly billing gives two months free and every feature is on every plan. See the ScreenshotNeo site for the product and plans.

11. Capture web localization evidence with code

The following examples show a minimal capture of a web page. Set the target application to the locale and region under test before capture, or capture an already configured test URL. A screenshot records the rendered state; it does not configure your application locale by itself. Replace the URL with your own test route.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com \
  -o localized-page.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
r.raise_for_status()
open("localized-page.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.com',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('localized-page.webp', bytes));

See the ScreenshotNeo API documentation for the full set of parameters. Its API supports options such as viewport and device presets, full-page or selector capture, waits, custom CSS and JavaScript, headers, cookies, user agent, and output formats. Use a stable test URL and configure the application itself to render the intended locale before capturing.

12. Or skip the browser setup

ScreenshotNeo makes a one-call capture from a URL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. 1,000 screenshots a month are free with no card, and paid plans start at $5 for 3,000. See the API documentation for options, then sign up for the free plan.

FAQ

Can I test localization if I cannot read the language?

Yes, for engineering and layout issues: use pseudo strings, compare resource completeness, run the same journeys, and inspect screenshots for clipping or broken controls. Use a qualified language reviewer for accuracy, tone, terminology, and cultural fit.

Is pseudolocalization enough before release?

No. It helps find readiness defects before translation, but real localized builds are needed to validate language, regional conventions, and market behavior.

How many locales should a test suite cover?

Choose a risk-based set from supported markets and document what each locale exercises. No universal locale count applies to every product.

Should language and region always match?

Not necessarily. If your product lets users choose them independently, test supported combinations independently, including formatting and fallback behavior.

Can screenshot comparison verify a translation?

No. It can help identify visual differences and missing content, but it cannot determine whether wording is correct or appropriate.