How to Find HTML Elements by Attribute Using BeautifulSoup
Find one or many HTML elements by id, class, data-*, aria-label, or any attribute with BeautifulSoup’s find(), find_all(), and CSS selectors.

To find HTML elements by attribute in BeautifulSoup, parse the document and pass an attribute filter to find() or find_all(). Use find() for the first match and find_all() for every match. For arbitrary, hyphenated, or reserved attribute names, use attrs={...}; for the HTML class attribute, use class_ or a CSS selector.
from bs4 import BeautifulSoup
html = '<a data-id="42">Answer</a><a data-id="43">Other</a>'
soup = BeautifulSoup(html, "html.parser")
first = soup.find("a", attrs={"data-id": "42"})
all_matches = soup.find_all("a", attrs={"data-id": "42"})
print(first.get_text(strip=True)) # Answer
print(len(all_matches)) # 1
Install the library with python -m pip install beautifulsoup4. BeautifulSoup parses HTML you already have; it does not fetch a web page. If your input comes from a URL, retrieve its HTML separately, then pass the response text to BeautifulSoup. The rest of this guide covers exact and flexible matching, classes, data attributes, combined selectors, edge cases, and debugging.
1. Install BeautifulSoup and parse the HTML
BeautifulSoup is distributed as beautifulsoup4 and imported as bs4. The standard-library html.parser is a practical parser choice for a small, self-contained script:
python -m pip install beautifulsoup4
from bs4 import BeautifulSoup
html = """<main>
<div id="main" data-state="active">Ready</div>
</main>"""
soup = BeautifulSoup(html, "html.parser")
main = soup.find("div", id="main")
print(main.get_text(strip=True))
The parser builds a navigable tree from the markup. Pick a parser explicitly so the script’s behavior is clear. BeautifulSoup also supports parsers such as lxml when installed; parser choice can affect how malformed markup is interpreted. A search can only find elements represented in the parsed input, so confirm that you are parsing the HTML you intend to inspect.
2. Choose between find() and find_all()
find() returns the first matching tag, or None if there is no match. find_all() returns a list-like result containing all matches; no matches produces an empty result. Check the first result before using tag methods, and iterate over all results when multiple elements are expected.

html = """<ul>
<li data-kind="fruit">Pear</li>
<li data-kind="fruit">Plum</li>
<li data-kind="tool">Saw</li>
</ul>"""
soup = BeautifulSoup(html, "html.parser")
one = soup.find("li", attrs={"data-kind": "fruit"})
print(one.get_text(strip=True)) # Pear: first matching element in document order
many = soup.find_all("li", attrs={"data-kind": "fruit"})
for item in many:
print(item.get_text(strip=True))
Use find() when document order is meaningful and you only need the first match. Use find_all() when you need to process every matching element or verify how many exist. Avoid treating a missing result as a tag: None.get_text() raises an error.
3. Search by id, class, and ordinary attributes
Attribute names that are valid keyword arguments can be passed directly. For example, id and type work as keyword filters. Python reserves the word class, so BeautifulSoup exposes class_ for that attribute.
by_id = soup.find("div", id="main")
emails = soup.find_all("input", type="email")
cards = soup.find_all("div", class_="card")
A class value is a list of class tokens in HTML. A search for class_="card" matches an element with class="card featured", because one token matches. If you need an element that has both card and featured classes regardless of their order, use:
featured_cards = soup.select("div.card.featured")
An exact multi-class string comparison such as class_="card featured" is order-sensitive. Prefer the selector when the requirement is “has both class tokens.”
4. Find data-*, aria-label, name, and unusual attributes
Use attrs as a dictionary for any attribute name, especially names containing hyphens or names that conflict with function parameters. This is the reliable general form for data-*, aria-*, and an HTML name attribute.
checkout = soup.find_all(attrs={"data-test-id": "checkout"})
close_buttons = soup.find_all(attrs={"aria-label": "Close"})
email_field = soup.find("input", attrs={"name": "email"})
active = soup.find_all(attrs={"data-state": "active"})
For an attribute such as name, passing name="email" can be confused with BeautifulSoup’s tag-name argument. Put it in attrs to state unambiguously that you mean the HTML attribute. Likewise, a hyphenated key such as data-test-id cannot be written as a Python keyword argument, so use the dictionary form.
To find every tag that has a given attribute regardless of its value, use True:
with_test_id = soup.find_all(attrs={"data-test-id": True})
5. Match values with strings, lists, regular expressions, and callables
Attribute filters can express more than exact equality. Use an exact string for a known value, a list for any of several values, a regular expression for a pattern, a callable for custom logic, and True to test whether the attribute is present. Use None when searching for tags that lack an attribute.
import re
# Exact value
home_links = soup.find_all("a", attrs={"href": "/home"})
# Any listed value
open_or_active = soup.find_all(attrs={"data-state": ["open", "active"]})
# Value begins with a path prefix
product_links = soup.find_all("a", href=re.compile(r"^/products/"))
# Attribute is present, even if its value is empty
with_disabled = soup.find_all("button", attrs={"disabled": True})
# Attribute is absent
without_title = soup.find_all("a", attrs={"title": None})
# A callable can apply application-specific matching rules
menus = soup.find_all(
attrs={"aria-label": lambda value: value and "menu" in value.lower()}
)
A callable receives the candidate attribute value, which may be None. Guard it before calling string methods, as the value and ... expression does above. For boolean HTML attributes such as disabled, presence is often the fact you care about; the literal value may be empty or represented differently than you expect. Check presence rather than assuming a string value.
6. Combine attributes and document structure with CSS selectors
Use select() when CSS syntax makes a condition easier to read, especially when combining multiple attributes, class tokens, or ancestor/descendant structure. It returns all matching elements; use select_one() for the first matching element.
# Exact attribute value
links = soup.select('a[href="/home"]')
# Any element with this data attribute value
cards = soup.select('[data-role="card"]')
# A heading link nested in a qualifying article
article_links = soup.select('article[data-kind="news"] h2 a')
# Attribute exists
labeled = soup.select("[aria-label]")
# Both classes, in either class order
body_strikeout = soup.select("p.body.strikeout")
# First match only
first_card = soup.select_one("[data-role='card']")
CSS attribute selectors can also express common value patterns. For example, [href^="/products/"] means the value starts with the prefix; [href$=".pdf"] means it ends with that suffix; and [href*="docs"] means it contains that text. BeautifulSoup’s select() uses SoupSieve. For a one-attribute lookup, find_all(attrs=...) is often more direct; for a compound condition, use whichever form communicates the requirement most clearly.
| Need | Good starting point |
|---|---|
| First tag with an exact attribute value | find(tag, attrs={...}) |
| Every matching tag | find_all(tag, attrs={...}) |
| Any of several values or a custom predicate | find_all() with a list, regex, or callable |
| Several classes or structural relationships | select() or select_one() |
7. A complete runnable example
This script parses a small HTML fragment, gets all links whose data-section value is one of two allowed values, and handles an absent match safely. Save it as find_elements.py and run python find_elements.py.
from bs4 import BeautifulSoup
html = """\
<nav>
<a data-section="products" href="/products/a">Product A</a>
<a data-section="support" href="/help">Help</a>
<a data-section="products" href="/products/b">Product B</a>
</nav>
"""
soup = BeautifulSoup(html, "html.parser")
links = soup.find_all(
"a",
attrs={"data-section": ["products", "support"]},
)
if not links:
print("No matching links found")
else:
for link in links:
print({
"label": link.get_text(" ", strip=True),
"href": link.get("href"),
"section": link.get("data-section"),
})
To adapt it, replace the tag name and attribute mapping. If any tag type may match, leave out the first argument and provide only attrs. To require nested structure or multiple class tokens, express the condition with a CSS selector. Keep the no-match path in production code: real pages change, and selectors that used to match may stop doing so.
8. When the HTML comes from a web page
BeautifulSoup searches the HTML you provide. For a static page, a common workflow is to request the page, check the HTTP response, then parse its body. This example uses Python’s requests package, which must be installed separately with python -m pip install requests.

import requests
from bs4 import BeautifulSoup
url = "https://example.com/"
response = requests.get(url, timeout=20)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
heading = soup.find("h1")
print(heading.get_text(" ", strip=True) if heading else "No h1 in response HTML")
This request only demonstrates acquiring HTML to parse. A server-rendered page may expose the target attributes in the response; a JavaScript-heavy page may add them later in the browser, after the initial response was delivered. In that case, parsing the response alone cannot find those later elements. You need a browser rendering step, or a page/API response that contains the relevant data. Respect the site’s access rules and rate limits when retrieving pages.
9. Troubleshooting common misses
| Symptom | Likely cause | Fix |
|---|---|---|
find() returns None |
The tag or attribute value differs, the element is absent, or it is created after initial page load. | Print or inspect the parsed HTML; confirm the tag and exact value. For dynamic content, obtain rendered HTML. |
AttributeError after a search |
Code called a tag method on the None result. |
Check if tag is not None before accessing it. |
A class_ search misses a multi-class element |
The code expects an exact class string or is combining tokens incorrectly. | For any one token, search with class_="token". For all required tokens, use .token1.token2 in CSS. |
| Hyphenated attribute raises a syntax error | Python keyword arguments cannot contain hyphens. | Use attrs={"data-test-id": "value"}. |
Searching name= seems to target the wrong thing |
name is also used as the tag-name parameter. |
Use attrs={"name": "email"}. |
| Callable filter raises an error | The attribute is absent on a candidate and the callable used a string method on None. |
Guard the value, for example lambda v: v and "menu" in v.lower(). |
| Response HTML has no target element | The page populates content using JavaScript after load, or the request returned a different page. | Inspect the actual response and status. Use rendered browser output if the element is created client-side. |
10. Performance, reliability, and cost considerations
For typical documents, the main practical cost is parsing the input and traversing matching tags. Search for the specific tag and attribute you need rather than collecting every tag and filtering afterward when a direct filter expresses the same condition. Avoid repeatedly parsing identical HTML inside a loop; parse once, then run the needed searches on the tree.
Reliability depends on the source markup and selector stability. Prefer attributes intended as stable identifiers, such as a documented data-* field, where available. Visible text and generated class names can change with localization, redesigns, or build tooling. Treat no matches as a normal result, validate required fields before using them, and log enough context to diagnose source changes without dumping sensitive page contents.
BeautifulSoup is a Python library; its use does not itself incur an API charge. If retrieving pages, the network request, proxy, or hosted browser service you choose may have separate limits or costs. The supplied research contains no benchmark figures, so performance should be measured on your own document sizes and workload rather than inferred from a universal number.
11. Or skip the browser setup
If your target content is already in HTML, BeautifulSoup is the direct tool for attribute searches. When the missing step is getting a clean screenshot of a live page, ScreenshotNeo is a website screenshot API and MCP server for developers. Its one-call API returns an image or PDF, so you do not need to set up a browser for capture. See the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie banners are accepted like a visitor would accept consent, and 60+ known consent platforms, newsletter popups, and chat widgets are removed before the shot; each step can be turned off.
- Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing. Response headers report the page verdict and whether it was billed.
- An MCP server provides
take_screenshot,get_page_info, andcapture_pdftools for Claude, Cursor, and other MCP clients. - 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000 screenshots. Every feature is on every plan.
Create a free ScreenshotNeo account to get 1,000 screenshots per month without a card.
12. FAQ
Can I search without specifying a tag name?
Yes. Call soup.find_all(attrs={"data-role": "card"}) to match across tag names, or use soup.select('[data-role="card"]').
Does BeautifulSoup execute JavaScript?
No. It parses the HTML string it receives. Use a browser-rendered source when JavaScript creates the elements you need.
How do I get the attribute value from a match?
Use tag.get("data-id") or tag["data-id"]. The get() form returns None when the attribute is missing; bracket access raises a key error if it is absent.
Can I find an element whose attribute contains a substring?
Use a regular expression with find_all(), or a CSS attribute selector such as [href*="docs"] with select().
Primary reference
See the Beautiful Soup documentation for the complete search API, attribute filters, CSS selectors, and class matching behavior.


