Common Website Testing Mistakes to Avoid
Avoid late, narrow, and misleading website tests. Build a practical release process for browser coverage, accessibility, mobile behavior, and performance.
Common website testing mistakes are testing only on the developer’s device, waiting until release week, treating an automated accessibility score as proof, overlooking real mobile conditions, and reducing performance to one load-time number. Avoid them by agreeing on supported environments, testing each change early, combining automation with human evaluation, and checking loading, interaction, and smoothness from the user’s perspective.
You do not need to test every browser and device combination. You do need a clear target matrix based on the site’s audience and support commitments, repeatable checks for important tasks, and enough time to fix what those checks find. MDN’s cross-browser testing guidance likewise recommends agreeing on a supported range and expanding testing as a feature develops.
1. Testing only on your own browser, device, and network
A page that works on a developer’s laptop may break in another browser, at a narrow viewport, on slower hardware, or for someone navigating with a keyboard or screen reader. Cross-browser testing includes browsers and devices, but also assistive technology and differences in hardware capability. The goal is not pixel-identical behavior everywhere: preserve accessible core functionality across the environments you support.
Define a supported environment matrix
Write down the environments that matter for this site rather than adopting a universal list. Use audience information when available, and confirm support expectations with the site owner. A small, explicit matrix makes gaps visible and keeps “works on my machine” from becoming the release criterion.
| Dimension | What to record | Example test question |
|---|---|---|
| Browser and operating system | Supported browser families and operating systems, including any older versions the project commits to support | Can a user complete checkout in the supported desktop browsers? |
| Device and viewport | Representative phone, tablet, and desktop sizes; orientation where relevant | Does the primary action remain visible and usable at a narrow width? |
| Input and assistive technology | Keyboard-only use and the screen-reader paths important to the audience | Can a person reach, understand, and activate every primary control? |
| Hardware and connection | Whether the test uses a physical device, emulator, or virtual machine, and any throttling or network conditions | Does the page remain usable on the lower-powered devices the audience may have? |
Universal coverage is impractical. Start with representative target environments and the important user tasks. Use physical devices where possible; emulators and virtual machines can extend coverage when hardware is unavailable, but record which kind of environment produced each result. MDN describes these options and recommends testing on mobile platforms such as Android or iOS in its testing workflow.
2. Leaving all testing until the end
Late testing makes regressions harder to isolate and leaves little time to correct them. Test each small implementation phase before committing it, then broaden coverage as the feature becomes stable. This catches simple failures while the change is still fresh and the cause is easier to find.
A practical testing cadence
- While implementing: check the component or flow you just changed in a couple of stable browsers available to the team.
- Before merging: run relevant automated checks and exercise the changed feature’s main task manually.
- As the feature matures: expand to the agreed browser and device matrix, including mobile and assistive-technology paths.
- Before release: run a short regression pass over critical user journeys and review unresolved issues against the stated support commitments.
For a useful manual smoke pass, test the main journey from entry to completion, then try keyboard-only navigation and a mobile environment. Keep the test narrow enough to repeat after a fix. MDN recommends starting with stable desktop browsers, basic keyboard or screen-reader checks, and a mobile platform, then expanding toward the full target list.
3. Treating an automated accessibility score as proof
Automated accessibility tools can identify many common issues, but a score does not establish that a site conforms to an accessibility standard or that people can use it successfully. W3C says WCAG evaluation combines automated testing and human evaluation; it also recommends usability testing, including people with disabilities in test groups. See W3C’s explanation of WCAG conformance.
Combine automated checks with human review
- Run an automated accessibility scan during development or CI to catch issues the tool can detect.
- Review text and background contrast, and do not use color as the only way to convey meaning.
- Disable CSS temporarily and inspect whether the source order still makes sense.
- Use the site without a mouse: tab through links and controls, activate them, and check whether focus is visible and in a sensible order.
- Use a screen reader for key paths, such as navigation, forms, error messages, and dialogs.
- Run task-based usability sessions with people who represent the audience, including people with disabilities where feasible.
These checks answer different questions. A conformance evaluation checks criteria against a named standard; usability testing asks whether people can complete intended tasks. W3C notes that satisfying success criteria does not necessarily mean content is usable by people with a wide range of disabilities. Automated tools such as Lighthouse accessibility audits, axe, and WAVE are examples of aids, not substitutes for human evaluation.
Name the standard and target
Replace vague claims such as “accessibility tested” with a test plan that names the standard and conformance target. WCAG 2.2 is a W3C Recommendation, published in 2023 and updated in 2024; it adds nine success criteria beyond WCAG 2.1. W3C advises using the latest WCAG version when developing or updating policies, while the applicable legal or contractual requirements still depend on the project and jurisdiction. See the WCAG 2.2 Recommendation.
A useful evaluation record states the scope, pages and flows covered, standard and target, browsers and assistive technologies used, automated and manual methods, findings, and known gaps. W3C’s WCAG Evaluation Methodology overview outlines defining scope, exploring the product, selecting representative pages, evaluating them, and reporting results.
4. Assuming responsive design guarantees mobile behavior
A layout that shrinks correctly in a desktop browser’s responsive mode can still behave differently on a phone. Touch input, browser behavior, available hardware, orientation, and the actual viewport all matter. Test real mobile environments in the support matrix when possible. Emulators and virtual machines are useful alternatives or additions, but a single device does not represent every supported combination.
On each representative mobile environment, exercise the same important tasks used on desktop. Check that content remains readable, controls can be operated with touch, menus and dialogs work, and the page still provides its core function. Record the device or emulator, operating system, browser, viewport, and the specific steps that failed so another person can reproduce the issue.
5. Calling a page fast based on one stopwatch number
Performance includes how quickly content loads, how soon the page becomes usable, how promptly it responds to input, and whether scrolling or animation feels smooth. A single page-load result cannot describe all of those experiences. MDN’s overview of web performance covers loading, usability, smoothness, interactivity, and perceived performance.
Measure a repeatable user journey
- Choose a representative page and task, such as opening a product page and submitting its main form.
- Record the browser, device or emulation, viewport, network conditions, and whether the run is a first visit or a repeat visit.
- Measure loading and rendering, then check when the primary task becomes usable.
- Interact with the page and observe responsiveness, scrolling, and animations; note delays or visible layout shifts.
- Repeat under the same conditions after a change and compare like with like.
Media, JavaScript, HTML, CSS, and rendering work can all affect the result. If a page feels slow, identify which part of the journey is slow before changing code. Do not report a performance result without saying what was measured and under what conditions.
6. A repeatable release checklist
- Scope: the supported browsers, operating systems, devices, viewports, and relevant assistive-technology paths are written down.
- Timing: testing happens during implementation and before release, not only at the end.
- Tasks: the important user journeys have explicit steps and expected outcomes.
- Accessibility: automated checks are paired with manual keyboard and screen-reader checks; usability is evaluated separately.
- Standard: the accessibility target is named, with scope and limitations recorded.
- Performance: loading, readiness for interaction, responsiveness, and smoothness are considered under recorded conditions.
- Evidence: results include environment, steps, expected and actual behavior, and enough detail to reproduce failures.
- Follow-up: failures are assigned and retested after fixes.
Or skip the browser setup
For a screenshot of a public page, ScreenshotNeo can return a PNG, JPEG, WebP, or PDF from one request. A screenshot can help compare page rendering across target viewport sizes, but it does not replace interaction, accessibility, or performance testing. See the ScreenshotNeo API docs.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
with open("shot.webp", "wb") as f:
f.write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', new Uint8Array(await res.arrayBuffer()));
Cookie and consent banners are accepted like a visitor and removed along with more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are never billed, and response headers report the page verdict and billing status. An MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.
Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.
Troubleshooting common testing problems
| Symptom | Likely cause | What to do |
|---|---|---|
| “It works on my machine,” but fails for someone else | The test covered only one browser, device, or configuration. | Reproduce the issue in the target matrix; record browser, OS, device, viewport, and exact steps. |
| A screenshot looks correct, but the control does nothing | Visual review checked appearance without exercising behavior. | Run the user task manually and add an interaction check; screenshots do not prove functionality. |
| The accessibility scan passes, but keyboard users get stuck | The issue requires interaction or human evaluation beyond the scan’s checks. | Repeat the flow with keyboard-only navigation and inspect focus order, focus visibility, and operability. |
| The mobile layout looks fine in emulation, but fails on a phone | The physical browser, input method, viewport, or device capability differs. | Reproduce on a physical target device where possible; compare recorded conditions and retain the emulator result as separate evidence. |
| Performance results disagree between runs | Runs may differ in network, cache state, device load, or other conditions. | Record the environment, distinguish first from repeat visits, repeat consistently, and compare the same task under like conditions. |
| A “WCAG compliant” report is disputed | The report may not define the standard version, conformance level, scope, or human evaluation performed. | State the target and scope; document automated and manual methods and usability findings separately. Do not treat a score as a blanket legal conclusion. |
FAQ
Do websites need to look identical in every browser?
No. Define the supported range and preserve accessible core functionality. Some visual or advanced behavior may reasonably differ by browser or device.
Can automated tests replace manual testing?
No. Automation is useful for repeatable checks, but keyboard use, assistive-technology behavior, and whether people can complete tasks need human evaluation.
How many browsers and devices should a team test?
There is no universal count. Select representative environments from the audience and project support requirements, and record the resulting scope.
Does one mobile device prove a site works on mobile?
No. A physical device gives useful evidence for that environment. Combine representative devices with emulators or virtual machines as needed to cover the agreed matrix.
Does passing WCAG checks prove legal compliance?
Not by itself. Name the WCAG version and target, document evaluation scope, and check the legal or contractual requirements that apply to the project.


