ScreenshotNeo

BlogHow-to

How to Perform Usability Testing on a Website

Learn how to plan and run website usability tests, write realistic tasks, observe participants, and turn findings into design decisions.

By the ScreenshotNeo team4 October 202611 min read

To perform a website usability test, choose a user decision to investigate, recruit people who resemble the intended users, give them realistic tasks, and observe what they do without coaching. Record task outcomes and relevant comments, look for patterns across sessions, and turn those patterns into specific changes you can check in a later round.

Usability depends on specified users trying to achieve specified goals in a particular context. A test can examine a sketch, prototype, draft content, or working website; choose the version that can answer your question. NIST describes usability testing as testing with representative users doing representative tasks and collecting quantitative and qualitative evidence. NIST: Usability Testing

1. Define the decision your test must support

Start with a decision, not a list of pages to inspect. State what you need to learn and what you might do differently based on the answer.

  • Can a first-time visitor find the eligibility requirements for a service?
  • Can a customer compare plans and choose one that fits a stated need?
  • Can a visitor understand a return policy well enough to decide whether to buy?
  • Can a person using a keyboard complete the main application task?

Write down the scope: which flow or content is included, which device or assistive technology matters, and what falls outside the study. Test an early prototype if the question is about structure or content. Use a functioning service when the question depends on implemented interactions, loading, or validation. Testing can happen throughout design and development. NIST Handbook 161

2. Choose participants who reflect the intended audience

Describe participants by their relevant experience and circumstances: how often they perform the task, what they already know, what device or environment they use, and what access needs are relevant. Recruit actual or likely users where possible. For accessibility research, include people who use the relevant assistive technologies and recruit around functional abilities and technology use, rather than relying only on diagnostic labels. See GOV.UK user research guidance.

There is no single correct participant count for every usability test. The right number depends on whether you are exploring problems or estimating performance:

Study purpose Source guidance How to use it
Small exploratory usability test Digital.gov recommends three to five people for its described small test. Use as a practical starting point for learning where a flow causes difficulty, not as a guarantee that every issue will be found.
Qualitative usability study GOV.UK recommends five to six participants and advises recruiting more for quantitative testing. Use moderated sessions to understand what people do and why; add participants if your audience has meaningfully different groups.
Performance measurement NIST Handbook 161 notes that some organizations test eight people per user group and suggests 30 or more may be appropriate for quantitative performance testing. Plan the sample around the measure, variability, groups, and inference you need. A small exploratory sample cannot establish a precise population completion rate.

These are context-specific recommendations, not interchangeable statistical rules. Sources: Digital.gov usability testing guide, GOV.UK user research guidance, and NIST Handbook 161.

3. Write realistic, neutral tasks

Give each participant one goal at a time. Describe a situation and the outcome they need, not the route through the interface. Avoid using labels from the site when those words reveal which control to choose. Keep instructions consistent across participants.

Leading task More neutral task
“Click ‘Pricing,’ select ‘Annual,’ then open the comparison table.” “Your team is considering this service for a year. Find out what it would cost and what is included.”
“Use the search box to find the returns page.” “You bought an item last week and may need to send it back. Find out what your options are.”
“Open the menu and choose ‘Book an appointment.’” “You need to speak with someone about this service. Find a way to arrange that.”

Before sessions, check that each task can be attempted in the test environment and that its scenario does not accidentally disclose the answer. If a task has prerequisites, provide them in the scenario rather than coaching during the attempt.

4. Prepare the session and recording plan

Write a short moderator guide with the same introduction, tasks, and wrap-up questions for every participant. A moderated session may take 20 minutes to an hour; Digital.gov describes that range and also gives an example of a typical session lasting about an hour. Adjust to the scope and participant burden. Digital.gov usability test methods

  1. Explain the broad purpose, what the participant will do, and that you are evaluating the experience rather than the participant.
  2. Explain recording and observation arrangements. Obtain consent for the session and separate permission for any recording.
  3. Tell participants they can take a break or stop.
  4. Assign roles: moderator, note-taker, and observers. Give observers a shared issue log so they record evidence rather than debate solutions during the session.
  5. Check the test account, prototype state, links, device, browser, assistive technology, network, and recording setup.
  6. Prepare a reset path between sessions, including how to restore accounts, carts, forms, or other changed state.

For remote sessions, confirm that the participant can access the test environment and that the session format works with their assistive technology. For in-person sessions, make the room and equipment accessible. Choose remote or in-person based on participant access, the task, and what you need to observe; neither is always preferable. GOV.UK describes a range of possible settings, including remote arrangements. GOV.UK user research guidance

5. Run the test without teaching the interface

Invite participants to think aloud when it is useful, for example by describing what they expect or what they are looking for. Do not require a constant stream of narration if that makes the task harder. Present the task, then give the participant room to act.

  • Observe hesitation, backtracking, wrong turns, errors, workarounds, and abandonment.
  • Record whether the task was completed and whether assistance was needed.
  • Notice what the participant expected to happen, especially when the site behaved differently.
  • Ask neutral questions after a task, such as “What were you expecting there?” or “What would you do next?”
  • Do not point to a control, suggest a label, or guide the participant to a successful path.
  • If a participant is stuck, use a prepared neutral prompt such as “What are you looking for now?” Record any help you provide.

Keep task wording and moderator behavior consistent. If the participant asks for help, you can say, “Please do what you would normally do; I can’t guide you, but I’ll make a note of that.” Treat the request for help as evidence, not as a mistake by the participant.

6. Capture performance and experience evidence

Choose measures that answer the study question. Common observations include:

  • Task outcome: completed, partially completed, not completed, or completed with assistance. Define these categories before sessions.
  • Errors and recovery: missteps, invalid entries, backtracking, and whether the participant recovered.
  • Time or effort: record only when speed or effort matters; task time is affected by the scenario and participant context.
  • Comments and expectations: confusion, confidence, likes, dislikes, and what participants thought labels or controls meant.
  • Context: relevant experience, device, environment, and access setup that help explain the observation.

Quantitative measures and qualitative observations complement each other. A completion count says what happened in your sessions; participant comments and observed behavior can help explain why. If the study is small or lacks a controlled comparison, report the number of sessions and describe observed patterns. Do not present exploratory observations as population-wide rates. NIST’s guidance describes combining quantitative and qualitative evidence. NIST: Usability Testing

For a website flow, a screen recording can help the team review navigation, scrolling, and interaction sequences alongside notes. Obtain recording consent, limit access to the files, and avoid recording sensitive data unless it is necessary and appropriately handled. A clean page screenshot can also document the state or content participants saw; it does not replace observing a participant’s behavior.

7. Synthesize findings into design decisions

Debrief after sessions while observations are fresh. Group related issues, but keep each finding connected to what someone did or said and the context in which it happened. Distinguish an observed problem from a possible explanation and from a proposed solution.

A useful issue log can use these fields:

Field What to record
Task and context Which task, participant context, device, or access setup was involved.
Observed evidence Action, hesitation, error, quote paraphrase, and outcome; avoid unsupported interpretation.
Impact Whether the issue blocked a goal, caused an error, or added avoidable effort.
Pattern How many sessions showed related behavior, with the study denominator stated.
Decision Proposed change, owner, and how the team will check whether it helped.

Prioritize using your team’s judgment about task importance, severity, recurrence, and consequences. There is no universally valid severity formula in the cited guidance. Choose changes that address a clear problem, then retest meaningful revisions when appropriate.

8. Report the method and its limits

A report should let readers understand what was studied and how far the findings apply. Include:

  • Research goal and decisions in scope.
  • Participant count and relevant characteristics, without exposing identities.
  • Exact task wording and the website, prototype, or build tested.
  • Test context, procedure, device or assistive technology where relevant, and whether sessions were moderated.
  • Measures, findings tied to evidence, limitations, and design decisions.

NIST’s reporting work emphasizes clear goals, participant selection, task descriptions, test design, and procedure. NIST: Common Industry Format for Usability Test Reports

Or skip the browser setup

If you need clean captures of pages involved in a study, ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request returns a PNG, JPEG, WebP, or PDF. The DIY browser workflow above is useful for observing people; an API capture is useful for documenting page states or preparing consistent reference images.

See the ScreenshotNeo API documentation. cURL example:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));

Cookie banners, newsletter popups, and chat widgets are removed before the shot; each cleaning step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month.

Troubleshooting usability tests

Problem Likely cause What to do
Participants finish immediately without exploring The task gave away the navigation path or used the site’s own labels. Rewrite it around the user’s situation and desired outcome. Pilot the wording with someone outside the study team.
Participants interpret the same task in different ways The scenario is ambiguous, or it combines multiple goals. Separate goals, define necessary context, and use identical wording across sessions.
The moderator keeps rescuing people The guide lacks neutral prompts, or the moderator is too close to the design. Prepare neutral prompts, rehearse the introduction, and have a note-taker flag accidental coaching.
Sessions produce opinions but little actionable evidence The team asked what participants like without observing task behavior. Use realistic tasks, record outcomes and behavior, then ask about expectations after the attempt.
Findings conflict across participants Participants may have different experience, context, or access needs, or the task may expose a genuine trade-off. Review context and evidence before averaging results. Consider whether the audience contains distinct groups that need separate analysis.
Remote session cannot access the page or account Test data, permissions, network restrictions, or setup instructions are incomplete. Run a technical check, prepare a reset or backup path, and reschedule rather than turning the session into improvised support.
Recording is missing or unusable Permissions, audio, capture settings, or storage were not checked. Verify the setup before the participant joins, obtain recording permission, and have a note-taking fallback.
A page screenshot differs between captures Dynamic content, personalization, viewport, cookie state, or timing changed. Record viewport and test conditions, reset site state, and use a consistent wait condition. Treat screenshots as state records, not evidence of how a person navigated.

Performance, reliability, and cost considerations

Plan for participant time as well as team analysis time. Keep tasks focused, avoid making people repeat long flows unnecessarily, and allow breaks when sessions are lengthy. For remote work, check the environment and access needs ahead of time. Schedule enough time after sessions to review notes and agree on decisions; collecting recordings without synthesis does not complete the research.

Reliability comes from a repeatable protocol: stable task wording, a known build or prototype version, consistent account state, and a recorded procedure for assistance and failures. Note interruptions or technical failures instead of silently excluding them. If a task fails because the test environment broke, mark it separately from a usability failure.

Budget participant recruitment or compensation, moderation, note-taking, analysis, accessible session arrangements, and any recording or research tools. There is no fixed cost implied by the method; it depends on recruitment, session mode, and study scope. Keep recordings only as long as needed for the stated research purpose and follow your organization’s data-handling practices.

Frequently asked questions

How many tasks should each participant complete?

Use only the tasks needed to answer the research question within a reasonable session. A focused study may use a few tasks; broader coverage can require more sessions or separate studies. Pilot the guide to check that it fits the planned duration.

Should I tell participants to think aloud?

It can reveal expectations and decision-making, but narration can change task pace. Invite it when useful and interpret it alongside behavior rather than treating spoken explanations as a substitute for observation.

Can I usability-test a prototype?

Yes. Use a sketch or prototype to examine structure, content, and proposed interactions. Test a working service when implementation details are part of the question.

Is a usability test the same as asking users what they want?

No. A usability test observes people attempting goals with a product or representation. Interviews about needs can complement it, but stated preferences alone do not show whether a task can be completed.

What does usability mean?

NIST’s page quotes ISO 9241-11’s definition as the extent to which specified users can achieve specified goals with effectiveness, efficiency, and satisfaction in a specified context. ISO’s official page explains that the standard provides a framework and does not prescribe specific design or evaluation methods. ISO 9241-11:2018