ScreenshotNeo

BlogGuides

Usability Testing Methods for Websites: A Practical Guide

Choose a website usability testing method, plan realistic tasks, run sessions, and turn observed friction into improvements.

By the ScreenshotNeo team4 October 202611 min read

Website usability testing means observing people as they try to complete realistic tasks on a website, prototype, or service. Choose a moderated study when you need to ask follow-up questions and understand unexpected behavior; choose an unmoderated study when tasks are focused and participants can complete them alone. Decide whether the study is qualitative or quantitative based on the decision you need to make, then recruit relevant users, write neutral tasks, pilot the session, and record what participants do as well as what they say.

Usability testing answers questions about how people use an interface. It is different from functional QA, which checks whether software behaves as specified, and from expert inspection, where specialists assess an interface without observing representative users. ISO 9241-11 provides a framework for understanding usability; it does not prescribe a particular evaluation method. ISO 9241-11:2018

1. Choose the method that fits the question

Start with the uncertainty you need to resolve. If you need to learn why people hesitate or misunderstand a page, use a method that lets you observe and probe. If you need comparable outcomes for a small set of clear tasks, an unmoderated or quantitative design may fit better.

Method choice Useful when Trade-off
Moderated You need to clarify context, ask tailored follow-ups, or diagnose unexpected behavior. Requires a facilitator and scheduling; facilitation can affect what participants do.
Unmoderated Instructions and tasks are self-contained, and participants can work independently. You cannot intervene in real time, so ambiguous instructions can make results hard to interpret.
Remote Participants are distributed, or using their own environment is relevant. Technology, connectivity, and distractions can affect sessions.
In person Physical context or direct observation is important. Location and access logistics can limit who can take part.
Qualitative You want to discover problems and understand how they arise. Small samples do not establish precise population rates.
Quantitative benchmark You have defined measures, such as completion, time, or errors, and need to compare outcomes. Needs a suitable study design and generally more participants than a formative qualitative study.

Moderated remote testing keeps live interaction while participants join from different locations. Unmoderated remote testing delivers instructions and captures work without a researcher present. The latter can be convenient, but gives less opportunity to clarify what happened. NN/g guidance on remote usability testing

These choices are independent. A study can be qualitative and remote, or quantitative and moderated, for example. Avoid choosing a format only because it is familiar; choose it based on the research question, participants, access needs, and logistics.

2. Define the study before recruiting

Write down the decision the study will inform and the questions you need answered. Keep the scope small enough that each task can be observed carefully.

  1. Research question: What is uncertain? Example: “Can first-time customers find the delivery date before placing an order?”
  2. User group: Who uses or is likely to use this service? Specify relevant experience, needs, and context.
  3. Journey or pages: Which part of the website or prototype is in scope?
  4. Decision: What might the team change depending on what participants encounter?
  5. Evidence: Which behaviors or outcomes would help answer the question?

Keep observations, interpretations, and decisions distinct. “Participant selected the shipping tab, returned to the cart, and asked whether delivery was included” is an observation. “The delivery information is hard to find” is an interpretation. This distinction helps the team evaluate whether the evidence supports a proposed change.

3. Recruit participants and choose a sample size

Recruit actual or likely users and screen for characteristics that matter to the service. Avoid recruiting only colleagues who already know the interface. When accessibility or inclusive usability is in scope, include people with disabilities and prepare materials, technology, and facilitation for their needs.

For a typical qualitative study of one user group, NN/g recommends five participants as a practical way to uncover many common usability problems. It is not a universal statistical sample size, a guarantee that a fixed percentage of issues will be found, or a sufficient design for precise population estimates. Consider separate user groups, high-risk tasks, comparisons, and quantitative benchmarking when deciding whether more participants or another design is needed. NN/g on qualitative usability study sample size Digital.gov also suggests three to five people for testing a website or document as practical plain-language guidance, not a statistical guarantee. Digital.gov audience guidance

For accessibility studies, tailor participant characteristics, assistive technology, evaluation parameters, and methods to the barrier or question being investigated. Do not assume that task time or satisfaction alone will reveal the accessibility issue. W3C WAI: Involving Users in Web Projects for Better Accessibility

4. Write realistic, neutral tasks

A task should describe an outcome that matters to the participant, not disclose the interface path. If the goal is to test whether someone can locate a return policy, do not say “Click Help, then Returns.” That would test following directions rather than finding the information.

Weak: “Use the filters to select size 8 and buy the blue jacket.”
Better: “You need a jacket for a rainy weekend trip. Find an option you would consider buying in your size, and tell me what you would want to know before deciding.”

  • Make the scenario believable and give enough context to act.
  • Do not name the button, menu, or navigation label you want the participant to find.
  • Avoid implying that there is one correct answer when the task is exploratory.
  • Do not embed information that participants should discover on the site.
  • For unmoderated tests, check that each task can be understood without a live explanation.

Write down completion criteria before sessions begin. For example, decide whether a task is complete only when the participant reaches a confirmation page, and whether a hint changes full completion to partial completion. Consistent definitions make results easier to compare.

5. Prepare the session and pilot it

Prepare a discussion guide or self-guided instructions with an introduction, consent and recording explanation, tasks, neutral prompts, and a closing question. Pilot the materials and technology with someone who is not part of the study. Confirm prototype links, accounts, test data, recording, sound, and any assistive technology needed.

For moderated sessions, a common introduction is: “We are evaluating the website, not you. There are no wrong answers. Please say what you are thinking as you work, and tell me if anything is unclear.” Ask permission before recording and explain how recordings and notes will be used.

GOV.UK suggests moderated sessions commonly last 30 to 60 minutes, depending on the number and complexity of tasks. Keep the session only as long as needed to observe the relevant work. GOV.UK Service Manual: usability testing

6. Run the test without steering the participant

  1. Welcome the participant, explain the purpose and recording, and answer process questions.
  2. Remind them that the service is being evaluated, not their ability.
  3. Read each task as written. Let the participant decide what to do.
  4. Invite them to think aloud with a neutral prompt, such as “What are you looking for now?”
  5. Observe silently when possible. If you need to probe, ask open questions after the action or task rather than pointing to a control.
  6. Record what happened, including detours, errors, pauses, comments, and context.
  7. Close by asking what felt difficult or surprising, without treating preference as proof of usability.

Think-aloud can reveal expectations and interpretations, but it does not replace observing task results and actions. In unmoderated studies, participants may stop verbalizing and cannot be prompted, which can leave recordings less explanatory. GOV.UK describes the value of asking participants to think aloud as they move through a service. GOV.UK Service Manual: usability testing

7. Capture website evidence consistently

Notes are the primary record of behavior. A screen recording can help the team review navigation paths, timing, and moments of hesitation, but it should support the research question rather than become an excuse to collect everything. Obtain consent, limit access, and follow your organization’s retention and privacy practices.

When a study involves a live website, a screenshot can document the exact page state associated with an observation. It does not show what a participant understood or why they acted, so pair it with notes and session context. A captured page may also include consent banners, popups, or chat widgets that obscure the interface being discussed.

8. Analyze observations and prioritize changes

After each session, record evidence while context is fresh. Then group observations by task and issue. Separate direct evidence from explanations the team is inferring, and look for patterns without assuming that repeated behavior is the only important signal. A single severe barrier can matter even if it appears once.

Possible measures include task success, partial success, critical errors, time, navigation path, and post-task satisfaction. Define measures before testing and collect only what helps answer the question. Combine metrics with observed behavior and participant explanations; a completion rate alone rarely explains why a task failed.

Priority factor Questions to ask
Impact Did the issue block a key task, create a serious error, or cause confusion?
Evidence What did the participant do or say? Is the conclusion an observation or interpretation?
Reach Which user groups and tasks could be affected?
Confidence Did the behavior recur, or is more targeted research needed?
Next action Can the team make a change and retest the uncertain part?

Turn findings into specific changes, owners, and follow-up questions. Retest when a design decision warrants it; a new study need not repeat the full original protocol if only one important issue changed.

9. Use screenshots to document the interface

For study notes, design reviews, or a before-and-after record, a screenshot helps preserve the page appearance at a particular moment. If you capture pages programmatically, you can use a browser library or a screenshot API. Browser automation gives you control over the browser session; an API can avoid maintaining browser infrastructure. Screenshots support usability evidence, but they do not replace participant observation.

With Playwright in Node.js, install the package and its browser, then save a full-page screenshot:

npm install playwright
npx playwright install chromium
// save as capture.mjs
import { chromium } from 'playwright';

const browser = await chromium.launch({ headless: true });
const page = await browser.newPage({ viewport: { width: 1440, height: 1000 } });
try {
  await page.goto('https://example.com', { waitUntil: 'domcontentloaded', timeout: 30000 });
  await page.screenshot({ path: 'page.png', fullPage: true });
} finally {
  await browser.close();
}

Run it with node capture.mjs. Replace the example address with a page you are allowed to access. For a stable state, wait for a specific selector or application event rather than adding a long arbitrary delay. Avoid capturing participant account data or other sensitive content unless the study requires it and you have appropriate consent and controls.

10. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request returns a PNG, JPEG, WebP, or PDF; the API supports full-page captures and other capture options. See the ScreenshotNeo API documentation for configuration and response details.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com \
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
r.raise_for_status()
with open("shot.webp", "wb") as f:
    f.write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const image = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', image));

ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture; each removal step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, with page verdict and billing details in response headers. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 screenshots each month with no card, and paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

11. Troubleshooting common study problems

Problem Likely cause What to do
Participants complete tasks unusually quickly The task gives away the route, or participants already know the interface. Rewrite around a realistic outcome and recruit people closer to the intended user group.
Participants ask what a task means The scenario is ambiguous or relies on knowledge they do not have. Clarify the outcome in the script, then pilot it again. In unmoderated studies, make instructions especially self-contained.
The moderator keeps helping Silence feels uncomfortable or the team wants the participant to succeed. Use neutral prompts, wait, and record when assistance was necessary. Decide in advance how assistance affects task success.
Think-aloud produces little commentary The prompt may be forgotten, or speaking while acting may be difficult. Use a neutral reminder in moderated sessions. Treat actions and outcomes as evidence; do not infer thoughts from silence.
Results differ sharply across participants User needs or prior experience vary, or the task has multiple valid paths. Check recruiting criteria and context. Report differences instead of averaging away meaningful segments.
A recording or screenshot misses the relevant state The page had not loaded, a modal obscured it, or the capture started at the wrong moment. Verify the page state before recording or capturing, pilot the flow, and annotate the observation with context.
Small-study metrics are presented as population rates A qualitative sample is being used to make a statistical claim. Report observed outcomes for the study participants and design a larger quantitative benchmark for population estimates.
Disabled participants cannot use the setup Materials, venue, prototype, or assistive technology were not prepared for the study question. Adapt the setup and protocol with the relevant participants and accessibility needs in mind; involve users with disabilities in the research.

12. Performance, reliability, and cost

Study quality depends more on the question, participant fit, task wording, and consistent facilitation than on maximizing the number of recordings or measures. Keep sessions focused, test the setup before participants arrive, and use only the evidence needed for the decision. Recruiting, scheduling, accessible materials, and analysis all take time; include them in the plan.

Remote sessions reduce location constraints but depend on participant connectivity and the chosen setup. In-person sessions can simplify direct observation while adding venue and travel logistics. Unmoderated studies can collect consistent task outcomes without live scheduling, but unclear instructions can undermine reliability because no moderator can clarify them.

For quantitative comparisons, define success, errors, timing rules, and participant criteria in advance, and recruit a sample appropriate to the claim. For screenshots, browser automation requires installing and maintaining browser dependencies; a hosted screenshot API shifts that browser setup to a service. ScreenshotNeo offers a free tier of 1,000 shots per month and paid plans from $5 for 3,000; other listed plans are Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan. Use screenshots only as supporting documentation, not a proxy for user research.

Frequently asked questions

Is usability testing the same as A/B testing?

No. Usability testing observes people attempting tasks to understand behavior and friction. An A/B test compares variants under a defined experiment to measure differences in outcomes.

Should every session use think-aloud?

No. It is useful when expectations and interpretation matter, but participant actions and task results remain important evidence. Choose the protocol that fits the question and participants.

Can I test a prototype instead of a live site?

Yes. A prototype can answer questions about a proposed flow or content before implementation, provided its limitations do not prevent the behavior you need to observe.

How often should a website be usability tested?

Test when an important design decision is uncertain, after a meaningful change, or when evidence shows users are struggling. The cadence depends on the product’s changes and the risks of the tasks in scope.