How to Test for Broken Links with Selenium
Use Selenium to test whether a real user journey reaches a working page. For a site-wide broken-link inventory, use an HTTP crawler instead.
Selenium can test whether a link works in a real browser flow: open the page, click the link, and assert that the expected destination content appears. It does not include a built-in broken-link checker, and Selenium’s own guidance discourages using WebDriver to spider a whole site. For a site-wide inventory, use an HTTP crawler; for a user-facing journey or JavaScript-rendered error page, use Selenium.
1. Choose the right method
The key question is whether you need to validate what a user experiences or enumerate links across a site.
| Need | Best fit | What it tells you |
|---|---|---|
| Check a user journey or rendered destination | Selenium WebDriver | Whether browser interaction reaches expected content or an error page |
| Inventory links across many pages | HTTP crawler, such as a curl-based workflow or BeautifulSoup-based crawler | Which discovered URLs fail direct HTTP requests |
| Capture network response details during a browser flow | Proxy or supported WebDriver BiDi events | Network observations tied to the browser flow; setup and browser support vary |
Selenium drives a browser in a way intended to represent a user’s interaction. Starting browsers and traversing the DOM adds overhead for link spidering, so the Selenium project recommends tools such as curl or BeautifulSoup for that job. These are fit-for-purpose distinctions, not published speed benchmarks. Selenium: Link spidering · Selenium: WebDriver
2. Test a link as a user would with Python
This runnable example uses Selenium 4 with Chrome. It opens a page, waits for the target link, clicks it, and checks a stable page heading. Replace the URL, selector, and expected heading with values from your application. Install Selenium with python -m pip install selenium; Selenium Manager can manage browser drivers for supported setups.
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
start_url = "https://www.selenium.dev/"
link_selector = "a[href='https://www.selenium.dev/documentation/']"
expected_heading = "WebDriver"
driver = webdriver.Chrome()
try:
driver.get(start_url)
wait = WebDriverWait(driver, 10)
link = wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, link_selector)))
link.click()
wait.until(EC.url_contains("/documentation/"))
heading = wait.until(EC.visibility_of_element_located((By.TAG_NAME, "h1")))
assert expected_heading.lower() in heading.text.lower(), (
f"Unexpected destination heading: {heading.text!r}; URL: {driver.current_url}"
)
print(f"OK: {driver.current_url} — {heading.text}")
finally:
driver.quit()
A successful assertion means this particular browser path produced the expected page content. It does not prove every link on the site works, and it does not assert an HTTP status code. For a new tab, wait for the window count to increase, switch to the new window, then make the same destination assertion. For a redirect, assert the final URL or page content rather than requiring the original URL to remain unchanged.
Assert a rendered error page
When a click lands on an application or server error page, check its title or a reliable element such as an H1. Selenium’s guidance for functional tests says the user-visible failure and steps leading to it are often more useful than status-code inspection alone. Choose an error marker your application controls; generic text such as “not found” can appear in unrelated content.
title = driver.title.lower()
heading_text = driver.find_element(By.TAG_NAME, "h1").text.lower()
assert "page not found" not in title
assert "page not found" not in heading_text
This assertion is only an example: adjust it to your site’s expected success and error states. A page can return an error status while rendering a custom page, or return a successful status while showing an application-level failure. Selenium: HTTP response codes
3. Wait for links created by JavaScript
Navigation reaching the browser’s normal readiness state does not guarantee that client-side code has finished changing the page. If links are inserted after an API call, wait for the specific link or condition you need. Prefer a targeted explicit wait over a fixed sleep.
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.common.by import By
wait = WebDriverWait(driver, 15)
link = wait.until(EC.presence_of_element_located(
(By.CSS_SELECTOR, "a[data-testid='report-link']")
))
wait.until(EC.element_to_be_clickable(
(By.CSS_SELECTOR, "a[data-testid='report-link']")
)).click()
Use presence when the element only needs to exist, visibility when it must be visible, and clickability when the test is about clicking it. If the page replaces the element during rendering, locate it again after the change rather than reusing a stale element. Selenium’s waiting guidance explains the distinction between document readiness and later JavaScript updates. Selenium: Waiting Strategies
4. Check HTTP responses when that is the requirement
WebDriver does not provide a universal, built-in API for collecting every link’s HTTP status. When the test must assert status codes while still exercising a browser journey, Selenium documents using a proxy as an advanced approach; the exact setup depends on the proxy and browser. Browser support for exposing response codes varies.
WebDriver BiDi can stream browser events such as network requests, console messages, and JavaScript errors. It is an observability capability, not a one-step site crawler; availability and event handling depend on browser and client support. Do not treat either proxy capture or BiDi as a replacement for defining which URLs to discover and which failures count. HTTP response codes · WebDriver and BiDi
5. Crawl links at site scale with HTTP requests
For a link inventory, use a crawler workflow: discover pages within an allowed scope, extract links, normalize and de-duplicate URLs, request each destination, and report failures for review. Choose the URL scope, redirect policy, concurrency, retry rules, and status-code policy for your site. These are crawler design decisions, not Selenium defaults.
A direct request is not identical to a browser visit. It may miss links created only by JavaScript, require authentication or cookies, or receive bot protection. Conversely, a crawler can check many discovered URLs without starting a browser for each one. Use Selenium on a smaller set of high-value flows when the rendered experience matters.
Simple curl spot-check
For one known URL, inspect the response headers and follow redirects:
curl -L -sS -o /dev/null -w "%{http_code} %{url_effective}\n" "https://www.selenium.dev/documentation/"
This checks one URL only; it does not discover links. A crawler should define what counts as failure: for example, whether to report all non-2xx responses, how to treat redirects, and whether access-denied or rate-limited responses need a retry or manual review.
6. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Element not found | The selector is wrong, the link is in another frame, or JavaScript has not inserted it yet | Verify the selector in the rendered DOM, switch to the relevant frame if needed, and wait for the element explicitly |
| Click intercepted or element not clickable | An overlay, animation, or consent prompt covers the link | Wait for the overlay to disappear, use a real user path to dismiss it, and confirm the target is visible and enabled |
| Stale element reference | The page re-rendered after the element was located | Wait for the update and locate the element again immediately before interacting |
| Test passes despite a broken destination | The assertion checks only that navigation occurred, or the server renders an error page with ordinary content | Assert a stable success marker such as the destination heading, and add an explicit error-page assertion |
| Driver or browser startup error | Browser and driver mismatch, missing browser, or environment setup problem | Check the Selenium troubleshooting guide; isolate browser-driver issues by trying another supported browser |
| Intermittent timeout | Slow network, delayed client rendering, or an unstable condition | Wait for the specific state required, inspect the failure URL and logs, and avoid increasing timeouts without diagnosing the delay |
| HTTP crawler reports a failure for a link that works in the browser | The target needs authentication, cookies, JavaScript, or blocks automated requests | Reproduce the request context, classify the result, or verify that target with a browser flow |
Selenium troubleshooting guidance notes that synchronization and the underlying browser driver can both cause failures. Selenium: Troubleshooting Assistance
7. Reliability, runtime, and cost considerations
- Coverage: Browser tests see rendered, interactive behavior; HTTP crawlers are suited to broad URL inventories but may not see client-generated links.
- Runtime: Browser startup and page navigation carry overhead. The Selenium project recommends avoiding WebDriver for spidering; no universal speed ratio applies.
- Reliability: Keep assertions tied to stable page content, use explicit waits, and distinguish confirmed broken links from transient timeouts, access controls, and rate limits.
- Test scope: Check representative user journeys in Selenium and run a crawler for breadth. Keep crawler concurrency and retry behavior appropriate for the site.
- Cost: Selenium itself is an open-source browser automation project, but running browser sessions consumes CI or machine resources. HTTP requests also consume network and target-server resources; set a responsible crawl rate.
8. Or skip the browser setup
If the task is capturing a page rather than validating link behavior, ScreenshotNeo is a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF; its API documentation covers capture options. For example:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. This captures pages; it does not replace a Selenium link test or a site-wide link crawler.
Sign up free for 1,000 screenshots a month, with no card required.
9. FAQ
Can Selenium check whether a link is broken?
Yes, for a defined browser journey: click the link and assert the expected destination content or error state. Selenium does not provide a built-in whole-site broken-link checker.
Should I assert on the HTTP status code?
Only when status is the requirement and you have a supported way to capture it. For functional browser tests, assert the experience the user should see.
Will Selenium find links added after page load?
It can, if the test waits for the client-side change and then locates the rendered link. Initial document readiness alone may not be sufficient.
What should a broken-link report include?
At minimum, record the source page, destination URL, observed result, and whether redirects or transient failures were involved. That makes failures easier to reproduce and triage.


