ScreenshotNeo

BlogGuides

Common Usability Testing Mistakes and How to Avoid Them

Avoid usability studies that waste participant time or lead to weak decisions. Plan focused questions, recruit relevant users, facilitate neutrally, and retest changes.

By the ScreenshotNeo team4 October 202611 min read

Usability testing is most useful when it answers a focused design question with evidence from people who reflect the intended users. To avoid common mistakes, plan the decision the study should inform, recruit for relevant needs and behavior, write realistic non-leading tasks, let participants work without coaching, account for accessibility, and interpret observations within the limits of the method and sample.

A study can be carefully moderated and still mislead if it recruits the wrong people, gives away the intended path, or treats a few participants’ results as population statistics. Use the lifecycle below as a planning and review checklist.

1. Start with a focused research question

“Test the app” is too broad to guide useful tasks or interpretation. Name the decision the study should inform, the users and context involved, and the uncertainty you need to resolve. For example: “Can first-time customers find the return deadline and understand what they need to do next?”

For each question, decide what evidence would be useful: task completion, a recurring error, uncertainty expressed during the attempt, or a mismatch between what someone believes happened and what actually happened. Avoid combining unrelated goals in one session. The Office for Health Improvement and Disparities (OHID) and Nielsen Norman Group (NN/g) both advise choosing clear goals; adding goals can dilute what the study reveals about each one.

  • Before recruiting: write the decision, the target users, the relevant context, and the questions the session can realistically answer.
  • Before each task: connect it to one research question and say what you will observe.
  • After the study: check that findings support a decision, rather than simply describing a collection of comments.

See the [OHID qualitative testing guidance](https://www.gov.uk/guidance/usability-testing) and NN/g’s [Checklist for Planning Usability Studies](https://www.nngroup.com/articles/plan-usability-study/).

2. Recruit people who reflect actual or likely users

Convenient participants—colleagues, friends, family, or product experts—may have different goals, experience, language, expectations, or access needs from ordinary users. Recruit based on behaviors and needs that matter to the research question, not just demographic labels or availability. A screening questionnaire can check relevant experience without revealing the exact task or desired answer.

Plan recruitment with access in mind. Consider recruitment channel, timing, location, compensation, communication preferences, and whether people can use their own devices and assistive technology. These choices affect who can take part. Provide suitable support and enough lead time when recruiting participants with specific access needs. Avoid repeatedly asking the same participants, since familiarity with the product or study can change their behavior.

For research involving people with disabilities, include people with relevant disabilities and assistive technology use in the target group. Do not infer the experience of an entire disability group from one person. Usability sessions can identify barriers, but they do not replace an accessibility conformance evaluation against the applicable standards. See [GOV.UK guidance on finding participants](https://www.gov.uk/service-manual/user-research/find-user-research-participants) and [Section 508’s usability testing tips](https://www.section508.gov/test/508-testing/).

3. Match the sample size and method to the purpose

There is no universal “five users” rule. Small qualitative studies can reveal usability problems and guide iterative design; they do not estimate population-wide performance precisely. A quantitative benchmark needs a larger sample, consistent tasks and conditions, and a plan for analyzing performance. Distinct user groups may need to be studied separately if their goals or contexts differ.

Study purpose Guidance in the sources How to apply it
Iterative qualitative testing OHID suggests 5 to 6 participants. NN/g recommends 5 for a traditional qualitative study. Use the number as practical guidance for a round, then recruit again as needed to examine new issues or distinct user groups.
Usability benchmarking GOV.UK’s benchmarking guidance targets 30 to 60 actual or likely users. Use a larger sample and consistent tasks and conditions to compare performance meaningfully.
Quantitative studies or eyetracking NN/g says at least 20 to 30 participants may be needed in each target user group. Plan sample size around the measure, variation, and groups you need to compare.

These recommendations describe different methods and purposes, so they are not interchangeable prescriptions. OHID describes an EPIC HIV project with 29 participants across four testing rounds as an example of iterative refinement and contextual recruitment; it is an illustration, not a universal sample target. Review the original context in [OHID’s qualitative guidance](https://www.gov.uk/guidance/usability-testing), [GOV.UK benchmarking guidance](https://www.gov.uk/service-manual/user-research/usability-benchmarking), and [NN/g’s planning checklist](https://www.nngroup.com/articles/plan-usability-study/).

4. Write believable tasks that do not give away the answer

A task should describe a goal a participant could plausibly have, without naming the button, menu, route, or sequence you want them to use. “Find out whether this jacket can be returned after 30 days” gives a goal. “Open the Returns menu and click the 30-day policy” gives away a path.

  • Present one task at a time, in neutral, consistent wording.
  • Use a believable situation and enough context to make the goal understandable.
  • Do not use internal product language the participant would not know.
  • Avoid implying the task is easy, or that a particular action is expected.
  • Pilot the wording with someone who was not involved in writing it. Check that they understand the goal without being told where to go.

For benchmarks, GOV.UK suggests no more than five tasks per participant and up to ten minutes per task as rules of thumb. A task can be challenging enough to reveal friction without being so long or artificial that fatigue becomes the main result. See [GOV.UK’s guidance on moderated usability testing](https://www.gov.uk/service-manual/user-research/using-moderated-usability-testing).

5. Facilitate without steering the participant

Participants can feel that they are being evaluated. At the start, explain that you are evaluating the service or prototype, not their ability. Tell them that confusion and failure are useful evidence, and that they may stop at any time. If recording, explain what will be recorded, why, who will see it, and how it will be handled; obtain informed consent before recording.

During a task, give the participant room to act. OHID’s advice is: “Give the participant a task and then let them complete it. Try to resist influencing how they engage with the prototype or giving them too many instructions.” (OHID, Usability testing: qualitative studies, 2020.)

When you need to follow up, ask about what you observed without suggesting an explanation or solution:

Leading prompt More neutral follow-up
“Did you see the blue button?” “What are you looking for now?”
“Was that easy?” “What was that part like for you?”
“Would a filter have helped?” “What, if anything, did you expect to happen?”
“You’re finished, right?” “What would you do next?”

Do not praise a particular route in a way that nudges later tasks. A note-taker can help capture behavior and context while the moderator listens and keeps the session moving. Use consistent instructions across participants when comparing results. Digital.gov offers a sample script and practical prompts in its [guide to conducting a usability test](https://digital.gov/guides/usability-testing/).

6. Choose a setting that reflects the real context

A lab, remote session, or unmoderated study can each be useful. Choose based on the question, the participant’s needs, and what must be observed.

Format Useful when Watch for
Moderated You need to clarify what happened or ask follow-up questions. Moderator prompts can change behavior; keep them neutral.
Unmoderated You need faster collection or participants who are harder to schedule. You cannot clarify an instruction or probe an unexpected action in the moment.
In person Physical context, subtle cues, or hands-on interaction matter. A lab can remove environmental factors that affect actual use.
Remote Participants need to use their own environment, devices, or assistive setup. Technical problems and reduced visibility may make behavior harder to interpret.

When a configured assistive technology setup is relevant, it may be difficult to reproduce in a lab. Let participants use their usual tools and devices where practical. If the environment materially affects the task, test in the real context or document how the study setup differs. These methods have different strengths; none is best for every question. See [OHID’s qualitative testing guidance](https://www.gov.uk/guidance/usability-testing) and [NN/g’s study planning checklist](https://www.nngroup.com/articles/plan-usability-study/).

7. Treat accessibility as part of the study design

Accessibility should shape recruitment, setup, task materials, and session timing from the start. Ask about access arrangements in a respectful way, allow participants to use the tools they rely on, and provide appropriate communication support. Budget time for recruitment and setup rather than treating accommodations as last-minute exceptions.

One participant’s experience can reveal a concrete barrier and suggest a design question; it cannot represent everyone who shares a diagnosis or access need. Combine usability research with evaluation against applicable accessibility requirements. The [GOV.UK moderated testing guidance](https://www.gov.uk/service-manual/user-research/using-moderated-usability-testing) and [Section 508 guidance](https://www.section508.gov/test/508-testing/) cover involving disabled participants and assistive technology.

8. Measure outcomes without overstating precision

Observe what participants do as well as what they say. A person may report success while making a mistake, abandon a task without saying why, or complete it only after a workaround. Record enough context to interpret the outcome.

  • Task outcome: completed, partially completed, failed, abandoned, or believed complete when it was not.
  • Behavior: errors, backtracking, hesitation, help requests, and unexpected routes.
  • Time: useful for comparison when tasks and conditions are consistent; interpret it alongside task success.
  • Participant account: expectations, uncertainty, and comments about the experience.
  • Context: device, assistive technology, environment, and any technical interruption that affected the attempt.

In qualitative discovery, counts can help organize recurring issues, but a small sample’s percentages should not imply population-wide certainty. In benchmarking, define success and timing rules before sessions and keep tasks and conditions consistent enough for comparison. GOV.UK’s [benchmarking guidance](https://www.gov.uk/service-manual/user-research/usability-benchmarking) discusses task measures and analysis.

9. Protect participant information and interpret evidence carefully

Tell participants what recording is for and obtain informed consent. Store recordings and notes with appropriate access controls, and avoid collecting personal information you do not need. Real user data can make a task more contextual, but use it only if the service can handle it securely; otherwise, prepare realistic dummy data. Explain limitations in the research readout, including recruitment, setup, sample, and any interrupted sessions.

Use recordings, direct observation, participant comments, and relevant analytics as complementary evidence. A recording does not make an interpretation self-evident, and analytics may show where behavior changed without explaining why. Separate observed events from your explanation of those events. GOV.UK’s [moderated testing guidance](https://www.gov.uk/service-manual/user-research/using-moderated-usability-testing) covers personal data and session planning.

10. Turn findings into design changes and retest

After sessions, group related problems by task and user goal. Give priority to recurring failures, high-impact errors, and barriers that prevent people from completing an important task. Keep evidence and interpretation distinct: note what happened, how often it appeared in this study, the relevant context, and what you think the design team should investigate or change.

Share findings in a form the team can act on: the task, observed issue, consequence, supporting evidence, limitation, and a possible design opportunity. Then test meaningful changes with relevant users. For benchmark rounds, keep tasks and conditions consistent enough to compare, and review them when the service or user behavior changes. Iteration is part of the research, not a substitute for documenting what each round can establish.

Usability testing review checklist

  • The study names a decision and a focused research question.
  • Recruitment criteria reflect actual or likely users and relevant access needs.
  • The method and sample size fit discovery or benchmarking goals.
  • Tasks describe believable goals without revealing the interface path.
  • The moderator script explains that the service is being tested and uses neutral prompts.
  • Recording consent, data handling, access arrangements, and realistic test data are planned.
  • Notes capture task outcome, behavior, context, and participant comments.
  • Analysis states limitations and distinguishes observation from interpretation.
  • Findings lead to design decisions and a plan to retest meaningful changes.

Or skip the browser setup

If you need screenshots of the flows or pages discussed in a usability study, you can capture them with a browser automation setup or use ScreenshotNeo, a website screenshot API and MCP server. It returns a screenshot or PDF from one GET request. The ScreenshotNeo API documentation covers its parameters.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks, blank pages, and failed loads are never billed, and response headers report the page verdict and billing status. Its MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Every feature is on every plan.

Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.

FAQ

How do I avoid leading questions in a usability test?

Ask open questions about the participant’s current goal or expectation, such as “What are you looking for?” or “What did you expect to happen?” Avoid naming a control or suggesting a solution.

Should I use real customer data in a session?

Only when the service can handle it securely and the data is necessary. Otherwise, use realistic dummy data that preserves the task context without exposing personal information.

Can usability testing prove that a product is accessible?

No. Sessions with people who use relevant assistive technology can uncover barriers, but they do not replace evaluation against applicable accessibility standards.

Is remote testing always more representative?

No. It can let people use their usual devices and environment, while in-person work can expose physical context and subtle cues. Choose the setting that matches the research question and document constraints.

Sources