How to Find All Links Using BeautifulSoup and Python
Learn how to extract every anchor link with BeautifulSoup, resolve relative URLs, handle malformed HTML, and process JavaScript-rendered pages.

How do I find all links on a web page with BeautifulSoup? Parse the HTML, select every <a> element, and read its href attribute:
from bs4 import BeautifulSoup
html = """
<a href="/about">About</a>
<a href="https://example.com/docs">Docs</a>
<a>This anchor has no href</a>
"""
soup = BeautifulSoup(html, "html.parser")
links = [a.get("href") for a in soup.find_all("a")]
print(links)
find_all("a") returns the anchor tags. get("href") safely returns the attribute value, or None when an anchor has no href. In the basic recipe, “all links” means hyperlinks represented by <a> tags. Images, canonical tags, scripts, forms and other URL-bearing elements require their own searches.
This guide builds the short example into a reusable crawler component. It covers fetching, relative URLs, filtering, duplicate handling, parser choices, malformed markup, JavaScript-rendered links, performance, security and troubleshooting.
1. Install BeautifulSoup and choose a parser
Install Beautiful Soup with pip:
python -m pip install beautifulsoup4
Beautiful Soup can use Python’s built-in html.parser, lxml, or html5lib. The parser can change the tree produced from malformed HTML, so name the parser explicitly when repeatability matters. The Beautiful Soup documentation lists the parser trade-offs: lxml is generally the fastest listed option when installed, html5lib follows browser-like HTML5 parsing, and html.parser requires no extra parser package.
# Optional parser choices
python -m pip install lxml html5lib
Use one of these constructors:
BeautifulSoup(markup, "html.parser")
BeautifulSoup(markup, "lxml")
BeautifulSoup(markup, "html5lib")
2. Extract raw href values
The raw value preserves exactly what appears in the document. This is useful when you need to reproduce source markup or distinguish relative and absolute references.
from bs4 import BeautifulSoup
html = """
<nav>
<a href="/products">Products</a>
<a href="team.html">Team</a>
<a href="https://external.example.net/page">External</a>
<a href="#pricing">Pricing section</a>
<a href="mailto:support@example.com">Email</a>
<a>Missing href</a>
</nav>
"""
soup = BeautifulSoup(html, "html.parser")
for anchor in soup.find_all("a"):
href = anchor.get("href")
print(href)
Output can include None, fragments, mail addresses and other schemes. Decide which forms your application accepts instead of assuming every value is an HTTP page.
Keep the anchor text too
records = [
{
"text": anchor.get_text(" ", strip=True),
"href": anchor.get("href"),
}
for anchor in soup.find_all("a")
]
get_text(" ", strip=True) joins nested text nodes with spaces and removes surrounding whitespace. Empty text is valid; icon-only or image-only links may have no useful visible label.
3. Fetch a page, then parse its response
Fetching and parsing are separate operations. The following complete script downloads one page, checks the response, and prints links:

import requests
from bs4 import BeautifulSoup
page_url = "https://example.com/"
response = requests.get(
page_url,
timeout=30,
headers={"User-Agent": "link-extractor/1.0"},
)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
for anchor in soup.find_all("a"):
href = anchor.get("href")
if href is not None:
print(href)
Use a timeout so a stalled origin does not hold a worker forever. Check the HTTP status before parsing; an error page can be valid HTML while containing none of the links you expected. Respect the target site’s terms, robots policy and rate limits when fetching pages.
4. Convert relative links to absolute URLs
Pages commonly use /about, team.html or ../assets/page. Resolve them against the page URL with urllib.parse.urljoin. Python’s urljoin documentation explains that an absolute or scheme-relative input can replace the base host or scheme.
from urllib.parse import urljoin
page_url = "https://example.com/docs/start.html"
absolute_links = []
for anchor in soup.find_all("a"):
href = anchor.get("href")
if href:
absolute_links.append(urljoin(page_url, href))
for url in absolute_links:
print(url)
Examples:
| href | Resolved result |
|---|---|
/about |
https://example.com/about |
team.html |
https://example.com/docs/team.html |
#install |
https://example.com/docs/start.html#install |
https://other.example/x |
https://other.example/x |
If URLs are untrusted and will later be requested, validate the result. A value such as https://attacker.example/ is allowed to replace your base host by design. Restrict schemes to http and https, and enforce an allowlist when your job must remain on one domain.
5. A robust extraction function
This function returns normalized records, optionally removes fragments, filters schemes, and deduplicates while preserving document order:
from collections import OrderedDict
from urllib.parse import urldefrag, urljoin, urlparse
from bs4 import BeautifulSoup
def extract_links(html: str, page_url: str, *, same_host: bool = False):
soup = BeautifulSoup(html, "html.parser")
seen = OrderedDict()
for anchor in soup.find_all("a"):
raw_href = anchor.get("href")
if not raw_href:
continue
absolute = urljoin(page_url, raw_href)
parsed = urlparse(absolute)
if parsed.scheme not in {"http", "https"}:
continue
clean_url, _fragment = urldefrag(absolute)
if same_host and urlparse(clean_url).netloc != urlparse(page_url).netloc:
continue
seen.setdefault(clean_url, {
"url": clean_url,
"text": anchor.get_text(" ", strip=True),
"raw_href": raw_href,
})
return list(seen.values())
html = """<a href='/about#top'>About</a>
<a href='/about#team'>About again</a>
<a href='mailto:a@example.com'>Mail</a>"""
for item in extract_links(html, "https://example.com/docs/page", same_host=True):
print(item)
Fragment removal makes /about#top and /about#team the same crawl target. Keep fragments if your application treats in-page destinations as distinct. Deduplication is also a policy decision: the same URL can have different anchor text and contexts.
6. Search for specific links
Beautiful Soup accepts filters for attributes, text and tags. To find only anchors with an href:
anchors_with_href = soup.find_all("a", href=True)
external = soup.find_all("a", href=lambda value: value and value.startswith("https://"))
To limit extraction to a navigation region, first select that region:
nav = soup.select_one("nav")
nav_links = nav.find_all("a", href=True) if nav else []
CSS selectors are useful for classes and attributes:
download_links = soup.select('a[href$=".pdf"]')
article_links = soup.select("article a[href]")
Do not assume every URL appears in an anchor. Common alternatives include <link href>, <script src>, <img src>, <form action> and Open Graph metadata. Search each tag and attribute deliberately:
assets = {
"stylesheets": [tag.get("href") for tag in soup.find_all("link", href=True)],
"scripts": [tag.get("src") for tag in soup.find_all("script", src=True)],
"images": [tag.get("src") for tag in soup.find_all("img", src=True)],
"forms": [tag.get("action") for tag in soup.find_all("form", action=True)],
}
7. JavaScript-rendered links
A static HTTP response contains the HTML returned by the server. If JavaScript adds links after load, Beautiful Soup cannot see those generated elements because it does not execute a browser runtime. An empty result can therefore mean the response is an app shell rather than the final page.
Use a browser automation tool when you need post-render DOM content. A typical workflow is:
- Open the URL in a browser context.
- Wait for a selector, a short delay or network idle.
- Read the rendered HTML.
- Pass that HTML to Beautiful Soup, or query anchors directly.
For pages that require consent, authentication, geolocation or interaction, configure the browser context accordingly. Keep static parsing for server-rendered pages because it is faster and uses fewer resources.
8. Parser differences and malformed HTML
Malformed markup can produce different trees under different parsers. If two machines return different links, check that they use the same Beautiful Soup version, parser and input bytes. Specify the parser in code and pin dependencies in your project.
| Parser | Use when | Trade-off |
|---|---|---|
html.parser |
You want zero extra parser installs. | Behavior can differ from browser HTML5 parsing. |
lxml |
You want speed and can install a native dependency. | Requires the lxml package. |
html5lib |
You need HTML5-style error recovery. | Usually slower and requires an extra package. |
9. Troubleshooting checklist
No links are returned
- Print
response.url,response.status_codeand the first 500 characters ofresponse.text; you may have received a login page, block page or error document. - Check that the document actually contains
<a>tags withhrefattributes. - If the site is client-rendered, use a browser rendering step.
- Confirm that your selector is not scoped to a missing element such as a nonexistent
nav.
KeyError: 'href'
Some anchors have no href. Replace anchor['href'] with anchor.get('href'), or search with find_all('a', href=True).
Relative URLs look wrong
Pass the final page URL, including its path, to urljoin. A redirect can change the effective base, so use response.url rather than assuming the requested URL.
Unexpected external hosts appear
urljoin intentionally honors absolute and scheme-relative href values. Parse the result and enforce an allowed hostname before queueing it for a crawler.
Parser or encoding errors
Let the HTTP library decode the response when possible, inspect response.encoding, and pass text or bytes consistently. Install and explicitly select the parser required by your deployment.
Requests hang or fail intermittently
Set connect and read timeouts, retry only idempotent requests with backoff, and limit concurrency. Treat repeated 403, 429 and 5xx responses according to the site’s policy instead of retrying without bound.
10. Performance, reliability and cost
Parsing one document is normally inexpensive; network latency and browser rendering dominate larger jobs. For a link inventory:

- Fetch each page once and cache its HTML when revisiting.
- Use
lxmlwhen parser speed matters and its dependency is acceptable. - Stream or batch your output instead of keeping an entire site graph in memory.
- Deduplicate URLs before scheduling more requests.
- Set bounded concurrency and per-request timeouts.
- Record source URL, final redirected URL, status code and parser choice for reproducibility.
Beautiful Soup itself has no service charge; your costs come from compute, bandwidth, storage and any browser automation or proxy service. Browser rendering is more expensive than downloading server HTML, so reserve it for pages whose links are created after load.
Or skip the browser setup
If you need a clean screenshot or rendered page artifact while analyzing links, ScreenshotNeo provides a website screenshot API and MCP server. It can render the page, accept cookie and consent banners, and remove more than 60 known consent platforms, newsletter popups and chat widgets before capture. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and each response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers.
One GET request is enough:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for authentication and options. The same request in Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
And Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo supports full-page captures with lazy images loaded, CSS-selector element captures, dark mode, device presets and custom viewports, retina scale, PDF output, custom CSS and JavaScript, click and wait actions, request blocking, headers, cookies, user agents, authorization, timezone and geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API and an OpenAPI specification. Existing parameter names used by other screenshot APIs also work, which can simplify migration.
The MCP server exposes take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
11. Frequently asked questions
Does Beautiful Soup crawl a whole website?
No. It parses one piece of markup. A crawler must fetch pages, resolve and filter URLs, track visited pages and enforce rate limits.
How do I include links without visible text?
Read the href independently of anchor text. Image-only anchors may require inspecting nested img alt text.
Should I remove URL fragments?
Remove them when fragments do not identify separate crawl targets. Keep them when your application needs in-page navigation destinations.
Can I parse a PDF with Beautiful Soup?
No. Beautiful Soup parses HTML and XML-like markup. Extract PDF links from HTML, then use a PDF-specific tool for the downloaded document.
Why specify a parser in production?
Different parsers recover malformed HTML differently. Naming one makes deployments more predictable and easier to debug.


