How to Select Elements by Class in XPath
Use a whitespace-aware XPath 1.0 expression to match a class token exactly, even when an element has multiple classes. See runnable examples, pitfalls, and fixes.

To select elements by an exact class token in XPath 1.0, use a whitespace-aware predicate:
//*[contains(concat(' ', normalize-space(@class), ' '), ' notice ')]
Replace notice with the class name you need. This matches class="notice highlighted" and class="highlighted notice", but not class="noticeable". A plain @class='notice' requires the entire attribute to equal that one value, while contains(@class, 'notice') can match unintended substrings. Parsel and Scrapy document the whitespace-aware pattern and its pitfalls. Parsel selector documentation · Scrapy selector documentation
1. The exact class-token pattern
HTML class attributes contain whitespace-separated tokens. To test whether a specific token is present, normalize the attribute’s whitespace, add a space to both ends, and search for the desired token with spaces around it:

//*[contains(concat(' ', normalize-space(@class), ' '), ' notice ')]
Each part has a job:
@classreads the class attribute on the current element.normalize-space()trims leading and trailing whitespace and reduces runs of whitespace to a single space.concat(' ', ..., ' ')pads the normalized value so a class at the beginning or end also has a boundary.contains(..., ' notice ')searches for the complete token, with its space boundaries.
For example, it matches <div class="notice highlighted"> and <div class="highlighted notice">. It does not match <div class="noticeable"> because that value has no space after notice.
This idiom is intended for XPath 1.0, the version commonly exposed by scraping and browser automation libraries. XPath syntax and evaluation rules are defined in the XPath 1.0 Recommendation.
2. Adapt the expression to your query
Limit matches to a tag
Use a specific node test instead of * when you know the element type:
//div[contains(concat(' ', normalize-space(@class), ' '), ' notice ')]
This finds matching div elements, whereas //*[...] checks every element type.
Require multiple classes
Add a predicate for every required class. Both conditions must be true on the same element:
//*[contains(concat(' ', normalize-space(@class), ' '), ' notice ') and contains(concat(' ', normalize-space(@class), ' '), ' urgent ')]
This matches elements with both tokens, in either order, and allows additional classes. In CSS, the equivalent compound selector is .notice.urgent.
Search within a selected element
When evaluating from a context node, use a relative path beginning with a dot:
.//*[contains(concat(' ', normalize-space(@class), ' '), ' notice ')]
The leading dot scopes the search to descendants of the current node. A leading // instead starts from the document root, which may escape the scope you intended. Parsel demonstrates chaining a CSS selection into a relative XPath expression in its usage documentation.
Select the first matching element
Parentheses determine whether a positional predicate applies per parent or to the complete result:
(//*[contains(concat(' ', normalize-space(@class), ' '), ' notice ')])[1]
This selects the first matching element in the document result. By contrast, //*[...][1] selects elements that are first matching children in their respective parent contexts; it can return more than one node. The related distinction is documented in Parsel’s selector guide.
3. Runnable examples in Python, Node.js, and cURL
XPath evaluates against a document tree. It does not fetch a URL or render JavaScript by itself. The examples below parse supplied HTML and evaluate the XPath locally. The HTML parsing behavior depends on the parser and the document it receives.
Python with lxml
Install the parser with python -m pip install lxml. Save this as select_class.py and run python select_class.py:
from lxml import html
source = """
<html><body>
<div class="notice highlighted">First notice</div>
<p class="noticeable">Not a notice token</p>
<section class="urgent notice">Second notice</section>
</body></html>
"""
doc = html.fromstring(source)
expr = "//*[contains(concat(' ', normalize-space(@class), ' '), ' notice ')]"
for element in doc.xpath(expr):
print(element.tag, element.get("class"), element.text_content().strip())
To require two tokens, put both predicates in the expression:
expr = ("//*[contains(concat(' ', normalize-space(@class), ' '), ' notice ') "
"and contains(concat(' ', normalize-space(@class), ' '), ' urgent ')]")
In a larger scraping program, use the XPath evaluation method provided by its selector object. Confirm whether that method returns element objects, strings, or a result wrapper before accessing attributes or text.
Node.js with an XPath-capable DOM
For a static HTML string, install @xmldom/xmldom and xpath with npm install @xmldom/xmldom xpath. Save as select-class.cjs and run node select-class.cjs:
const { DOMParser } = require('@xmldom/xmldom');
const xpath = require('xpath');
const source = `
<html><body>
<div class="notice highlighted">First notice</div>
<p class="noticeable">Not a notice token</p>
<section class="urgent notice">Second notice</section>
</body></html>
`;
const doc = new DOMParser().parseFromString(source, 'text/html');
const expr = "//*[contains(concat(' ', normalize-space(@class), ' '), ' notice ')]";
const matches = xpath.select(expr, doc);
for (const element of matches) {
console.log(element.nodeName, element.getAttribute('class'), element.textContent.trim());
}
DOM parser APIs and support for HTML parsing vary. If your parser expects XML, provide well-formed XML or use an HTML parser mode; malformed markup may be repaired or represented differently by the parser.
cURL and browser automation
cURL does not include an XPath engine. It can retrieve a response body, but you need a parser or browser automation library to evaluate an XPath. For a page whose markup is directly available as HTML, download it and pass it to a local parser:
curl -L --fail https://example.com/ -o page.html
Then load page.html with the Python or Node.js example above, adapting the code to read the file. A response fetched with cURL does not include DOM changes made later by client-side JavaScript. For rendered pages or dynamic content, evaluate the XPath in a browser automation environment after navigation and the required wait condition. Selenium’s locator documentation describes its browser element lookup context: Selenium locators.
4. XPath or CSS for class lookup?
If your tool supports CSS and the task is only to find a class, CSS is shorter:
.notice
.notice.urgent
The W3C Selectors specification defines class membership for HTML, SVG, and MathML in terms of whitespace-separated class tokens. See Selectors Level 4. Parsel recommends CSS for ordinary class lookup and shows XPath when subsequent navigation or extraction needs XPath expressions.
| Need | Use | Example |
|---|---|---|
| One class token | CSS if supported | .notice |
| Multiple required classes | CSS compound selector | .notice.urgent |
| Class matching in an XPath expression | Whitespace-aware XPath predicate | //*[contains(concat(' ', normalize-space(@class), ' '), ' notice ')] |
| Navigate by XPath axes or combine with XPath predicates | XPath | Class predicate plus parent, ancestor, or text conditions |
Use the selector language the host API supports and keep scope clear: a query may be document-wide or relative to a current element. Parsel’s documentation covers both CSS and XPath selector workflows.
5. Build the XPath safely in host code
When the class name is fixed in source code, write it directly in the quoted XPath string. When it comes from user input or scraped data, quote it safely before inserting it into an XPath literal. XPath 1.0 has no backslash escape inside string literals. A value containing an apostrophe cannot simply be wrapped in single quotes, and a value containing both quote types may need an XPath concat() expression.
Here is a Python helper that emits a safe XPath string literal for arbitrary text:
def xpath_literal(value: str) -> str:
if "'" not in value:
return "'" + value + "'"
if '"' not in value:
return '"' + value + '"'
parts = value.split("'")
pieces = []
for index, part in enumerate(parts):
if part:
pieces.append("'" + part + "'")
if index < len(parts) - 1:
pieces.append('"\'"')
return "concat(" + ", ".join(pieces) + ")"
class_name = "notice"
expr = ("//*[contains(concat(' ', normalize-space(@class), ' '), "
+ xpath_literal(" " + class_name + " ") + ")]" )
Class tokens are normally chosen by page authors and follow the host language’s class-token rules, but robust quoting matters when values are externally supplied. Validate the class name as a single token if that is the intended input. If it contains whitespace, the padded search would look for that sequence rather than one class token.
6. Edge cases and common mistakes
Empty or missing class attribute
An element with no class attribute will not match a nonempty token. An empty attribute will not match either. The normalization expression safely yields an empty string for an absent attribute in standard XPath 1.0 conversion behavior.
Whitespace in class values
Extra spaces, tabs, and newlines in the attribute are normalized before searching, so formatting whitespace does not make a normal token fail. The class attribute’s token semantics come from the document environment; XML namespaces, HTML parsing, and case behavior may differ by host and document type.
Substring matching
Avoid contains(@class, 'notice') for exact class membership. It can match noticeable or promo-notice. The padded expression checks token boundaries and avoids that common overmatch, as the Scrapy docs explain.
Exact attribute comparison
//*[@class='notice'] only matches when the entire attribute is exactly notice. It misses elements that have additional classes or different whitespace.
Wrong search scope
In a scoped query, a document-root expression such as //* may search outside the current node. Use .//* for descendants of that context node. If you want the context node itself as well as descendants, use a suitable expression such as self::*[...] | .//*[...], adjusting the context to your API.
First match versus first per parent
If a positional predicate produces too many results, parenthesize the entire result: (//li)[1] means the first list item in document order, while //li[1] means list items that are first among their siblings.
7. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| No elements returned | The XPath is evaluated against a different tree than expected, the page content is not loaded, the class token differs, or the query is scoped too narrowly. | Inspect the actual parsed HTML/DOM, verify the exact class value, use document-root or relative scope intentionally, and wait for dynamically added content before querying. |
| Too many results | A substring query such as contains(@class, 'notice') matches longer tokens, or the wildcard searches every element type. |
Use the padded token expression and narrow * to a tag such as div if appropriate. |
| XPath syntax error near quotes | The host-language string quoting and XPath literal quoting were mixed up, especially for a class name containing an apostrophe. | Use a safe XPath-literal helper or a library’s variable binding feature if available. XPath 1.0 literals do not support backslash escaping. |
| Query works in one library but not another | Different tools expose different XPath versions, context nodes, return types, namespaces, or HTML parsing behavior. | Check the host library’s XPath API and parser documentation; reduce the query to a small known document and confirm result type. |
| Static HTML has no target but browser shows it | The target is created or modified by JavaScript after the initial response. | Use a browser automation tool, wait for the element, then evaluate the XPath against the rendered DOM. |
Wrong element selected with [1] |
The positional predicate is applied at each parent context rather than after assembling the full result. | Wrap the complete location path in parentheses before applying [1]. |
8. Performance, reliability, and cost
For a small local document, the main practical cost is usually parsing the document and the number of nodes the expression must inspect. A broad //* query checks all elements; a tag-specific query such as //div[...] narrows candidates and communicates intent. Avoid repeatedly parsing the same HTML when several selectors can share one parsed tree. Measure in the actual host environment if selector evaluation is on a high-volume path; this article does not claim a benchmark.
Reliability depends on querying the intended document state. In scraping, ensure you have the response body you expect. In a browser, wait for the content you need rather than assuming navigation completion means every asynchronous widget or component is ready. Handle empty result sets, parser errors, network failures, and changed markup explicitly. Class names generated by a site’s build system may change between releases, so prefer stable attributes or structural context where available.
XPath itself has no service fee. Costs come from the surrounding work: network requests, browser instances, compute, and any third-party capture or automation service. Static HTML parsing is generally simpler than starting a full browser when the needed markup is already present; JavaScript-rendered content requires a browser-capable approach if the target does not exist in the response source.
9. Or skip the browser setup
If the reason you need a browser is to capture a page screenshot, ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. For the capture step, one GET request returns an image or PDF. The XPath expression above remains useful when you need to query a DOM in your own scraper or automation code.

Example request using cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo API documentation for request options. ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000.
Sign up for ScreenshotNeo’s free plan and get 1,000 screenshots a month with no card.
10. FAQ
Does XPath have a class selector like CSS?
No dedicated XPath class shorthand exists in XPath 1.0. Use the whitespace-aware predicate, or use CSS .className where the host API supports CSS selectors.
Can I match a class that contains a hyphen?
Yes. A class such as promo-card is still one token; put it between the padded spaces in the XPath search string.
Does this select the element’s children too?
No. The predicate identifies elements whose own class attribute contains the token. Select descendants separately with a path such as //section[... ]//a.
Why does the XPath find nothing when I can see the element?
Check that the XPath is running against the rendered DOM rather than an initial response, and verify context scope, exact class spelling, and whether the selector engine is querying the expected document.


