ScreenshotNeo

BlogHow-to

How to Find All Links Using BeautifulSoup and Python

Learn how to extract every anchor link with BeautifulSoup, resolve relative URLs, handle malformed HTML, and process JavaScript-rendered pages.

By the ScreenshotNeo team30 September 20269 min read

How to Find All Links Using BeautifulSoup and Python

How do I find all links on a web page with BeautifulSoup? Parse the HTML, select every <a> element, and read its href attribute:

from bs4 import BeautifulSoup

html = """
<a href="/about">About</a>
<a href="https://example.com/docs">Docs</a>
<a>This anchor has no href</a>
"""

soup = BeautifulSoup(html, "html.parser")
links = [a.get("href") for a in soup.find_all("a")]
print(links)

find_all("a") returns the anchor tags. get("href") safely returns the attribute value, or None when an anchor has no href. In the basic recipe, “all links” means hyperlinks represented by <a> tags. Images, canonical tags, scripts, forms and other URL-bearing elements require their own searches.

This guide builds the short example into a reusable crawler component. It covers fetching, relative URLs, filtering, duplicate handling, parser choices, malformed markup, JavaScript-rendered links, performance, security and troubleshooting.

1. Install BeautifulSoup and choose a parser

Install Beautiful Soup with pip:

python -m pip install beautifulsoup4

Beautiful Soup can use Python’s built-in html.parser, lxml, or html5lib. The parser can change the tree produced from malformed HTML, so name the parser explicitly when repeatability matters. The Beautiful Soup documentation lists the parser trade-offs: lxml is generally the fastest listed option when installed, html5lib follows browser-like HTML5 parsing, and html.parser requires no extra parser package.

# Optional parser choices
python -m pip install lxml html5lib

Use one of these constructors:

BeautifulSoup(markup, "html.parser")
BeautifulSoup(markup, "lxml")
BeautifulSoup(markup, "html5lib")

2. Extract raw href values

The raw value preserves exactly what appears in the document. This is useful when you need to reproduce source markup or distinguish relative and absolute references.

from bs4 import BeautifulSoup

html = """
<nav>
  <a href="/products">Products</a>
  <a href="team.html">Team</a>
  <a href="https://external.example.net/page">External</a>
  <a href="#pricing">Pricing section</a>
  <a href="mailto:support@example.com">Email</a>
  <a>Missing href</a>
</nav>
"""

soup = BeautifulSoup(html, "html.parser")
for anchor in soup.find_all("a"):
    href = anchor.get("href")
    print(href)

Output can include None, fragments, mail addresses and other schemes. Decide which forms your application accepts instead of assuming every value is an HTTP page.

Keep the anchor text too

records = [
    {
        "text": anchor.get_text(" ", strip=True),
        "href": anchor.get("href"),
    }
    for anchor in soup.find_all("a")
]

get_text(" ", strip=True) joins nested text nodes with spaces and removes surrounding whitespace. Empty text is valid; icon-only or image-only links may have no useful visible label.

3. Fetch a page, then parse its response

Fetching and parsing are separate operations. The following complete script downloads one page, checks the response, and prints links:

The extraction pipeline: fetch HTML, parse anchors, resolve URLs and filter results.
The extraction pipeline: fetch HTML, parse anchors, resolve URLs and filter results.
import requests
from bs4 import BeautifulSoup

page_url = "https://example.com/"
response = requests.get(
    page_url,
    timeout=30,
    headers={"User-Agent": "link-extractor/1.0"},
)
response.raise_for_status()

soup = BeautifulSoup(response.text, "html.parser")
for anchor in soup.find_all("a"):
    href = anchor.get("href")
    if href is not None:
        print(href)

Use a timeout so a stalled origin does not hold a worker forever. Check the HTTP status before parsing; an error page can be valid HTML while containing none of the links you expected. Respect the target site’s terms, robots policy and rate limits when fetching pages.

Pages commonly use /about, team.html or ../assets/page. Resolve them against the page URL with urllib.parse.urljoin. Python’s urljoin documentation explains that an absolute or scheme-relative input can replace the base host or scheme.

from urllib.parse import urljoin

page_url = "https://example.com/docs/start.html"
absolute_links = []

for anchor in soup.find_all("a"):
    href = anchor.get("href")
    if href:
        absolute_links.append(urljoin(page_url, href))

for url in absolute_links:
    print(url)

Examples:

href Resolved result
/about https://example.com/about
team.html https://example.com/docs/team.html
#install https://example.com/docs/start.html#install
https://other.example/x https://other.example/x

If URLs are untrusted and will later be requested, validate the result. A value such as https://attacker.example/ is allowed to replace your base host by design. Restrict schemes to http and https, and enforce an allowlist when your job must remain on one domain.

5. A robust extraction function

This function returns normalized records, optionally removes fragments, filters schemes, and deduplicates while preserving document order:

from collections import OrderedDict
from urllib.parse import urldefrag, urljoin, urlparse
from bs4 import BeautifulSoup

def extract_links(html: str, page_url: str, *, same_host: bool = False):
    soup = BeautifulSoup(html, "html.parser")
    seen = OrderedDict()

    for anchor in soup.find_all("a"):
        raw_href = anchor.get("href")
        if not raw_href:
            continue

        absolute = urljoin(page_url, raw_href)
        parsed = urlparse(absolute)
        if parsed.scheme not in {"http", "https"}:
            continue

        clean_url, _fragment = urldefrag(absolute)
        if same_host and urlparse(clean_url).netloc != urlparse(page_url).netloc:
            continue

        seen.setdefault(clean_url, {
            "url": clean_url,
            "text": anchor.get_text(" ", strip=True),
            "raw_href": raw_href,
        })

    return list(seen.values())

html = """<a href='/about#top'>About</a>
<a href='/about#team'>About again</a>
<a href='mailto:a@example.com'>Mail</a>"""

for item in extract_links(html, "https://example.com/docs/page", same_host=True):
    print(item)

Fragment removal makes /about#top and /about#team the same crawl target. Keep fragments if your application treats in-page destinations as distinct. Deduplication is also a policy decision: the same URL can have different anchor text and contexts.

Beautiful Soup accepts filters for attributes, text and tags. To find only anchors with an href:

anchors_with_href = soup.find_all("a", href=True)
external = soup.find_all("a", href=lambda value: value and value.startswith("https://"))

To limit extraction to a navigation region, first select that region:

nav = soup.select_one("nav")
nav_links = nav.find_all("a", href=True) if nav else []

CSS selectors are useful for classes and attributes:

download_links = soup.select('a[href$=".pdf"]')
article_links = soup.select("article a[href]")

Do not assume every URL appears in an anchor. Common alternatives include <link href>, <script src>, <img src>, <form action> and Open Graph metadata. Search each tag and attribute deliberately:

assets = {
    "stylesheets": [tag.get("href") for tag in soup.find_all("link", href=True)],
    "scripts": [tag.get("src") for tag in soup.find_all("script", src=True)],
    "images": [tag.get("src") for tag in soup.find_all("img", src=True)],
    "forms": [tag.get("action") for tag in soup.find_all("form", action=True)],
}

A static HTTP response contains the HTML returned by the server. If JavaScript adds links after load, Beautiful Soup cannot see those generated elements because it does not execute a browser runtime. An empty result can therefore mean the response is an app shell rather than the final page.

Use a browser automation tool when you need post-render DOM content. A typical workflow is:

  1. Open the URL in a browser context.
  2. Wait for a selector, a short delay or network idle.
  3. Read the rendered HTML.
  4. Pass that HTML to Beautiful Soup, or query anchors directly.

For pages that require consent, authentication, geolocation or interaction, configure the browser context accordingly. Keep static parsing for server-rendered pages because it is faster and uses fewer resources.

8. Parser differences and malformed HTML

Malformed markup can produce different trees under different parsers. If two machines return different links, check that they use the same Beautiful Soup version, parser and input bytes. Specify the parser in code and pin dependencies in your project.

Parser Use when Trade-off
html.parser You want zero extra parser installs. Behavior can differ from browser HTML5 parsing.
lxml You want speed and can install a native dependency. Requires the lxml package.
html5lib You need HTML5-style error recovery. Usually slower and requires an extra package.

9. Troubleshooting checklist

  • Print response.url, response.status_code and the first 500 characters of response.text; you may have received a login page, block page or error document.
  • Check that the document actually contains <a> tags with href attributes.
  • If the site is client-rendered, use a browser rendering step.
  • Confirm that your selector is not scoped to a missing element such as a nonexistent nav.

KeyError: 'href'

Some anchors have no href. Replace anchor['href'] with anchor.get('href'), or search with find_all('a', href=True).

Relative URLs look wrong

Pass the final page URL, including its path, to urljoin. A redirect can change the effective base, so use response.url rather than assuming the requested URL.

Unexpected external hosts appear

urljoin intentionally honors absolute and scheme-relative href values. Parse the result and enforce an allowed hostname before queueing it for a crawler.

Parser or encoding errors

Let the HTTP library decode the response when possible, inspect response.encoding, and pass text or bytes consistently. Install and explicitly select the parser required by your deployment.

Requests hang or fail intermittently

Set connect and read timeouts, retry only idempotent requests with backoff, and limit concurrency. Treat repeated 403, 429 and 5xx responses according to the site’s policy instead of retrying without bound.

10. Performance, reliability and cost

Parsing one document is normally inexpensive; network latency and browser rendering dominate larger jobs. For a link inventory:

Rendered capture can remove consent overlays and other widgets before producing a usable page image.
Rendered capture can remove consent overlays and other widgets before producing a usable page image.
  • Fetch each page once and cache its HTML when revisiting.
  • Use lxml when parser speed matters and its dependency is acceptable.
  • Stream or batch your output instead of keeping an entire site graph in memory.
  • Deduplicate URLs before scheduling more requests.
  • Set bounded concurrency and per-request timeouts.
  • Record source URL, final redirected URL, status code and parser choice for reproducibility.

Beautiful Soup itself has no service charge; your costs come from compute, bandwidth, storage and any browser automation or proxy service. Browser rendering is more expensive than downloading server HTML, so reserve it for pages whose links are created after load.

Or skip the browser setup

If you need a clean screenshot or rendered page artifact while analyzing links, ScreenshotNeo provides a website screenshot API and MCP server. It can render the page, accept cookie and consent banners, and remove more than 60 known consent platforms, newsletter popups and chat widgets before capture. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and each response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers.

One GET request is enough:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for authentication and options. The same request in Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

And Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo supports full-page captures with lazy images loaded, CSS-selector element captures, dark mode, device presets and custom viewports, retina scale, PDF output, custom CSS and JavaScript, click and wait actions, request blocking, headers, cookies, user agents, authorization, timezone and geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API and an OpenAPI specification. Existing parameter names used by other screenshot APIs also work, which can simplify migration.

The MCP server exposes take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

11. Frequently asked questions

Does Beautiful Soup crawl a whole website?

No. It parses one piece of markup. A crawler must fetch pages, resolve and filter URLs, track visited pages and enforce rate limits.

Read the href independently of anchor text. Image-only anchors may require inspecting nested img alt text.

Should I remove URL fragments?

Remove them when fragments do not identify separate crawl targets. Keep them when your application needs in-page navigation destinations.

Can I parse a PDF with Beautiful Soup?

No. Beautiful Soup parses HTML and XML-like markup. Extract PDF links from HTML, then use a PDF-specific tool for the downloaded document.

Why specify a parser in production?

Different parsers recover malformed HTML differently. Naming one makes deployments more predictable and easier to debug.