ScreenshotNeo

BlogHow-to

How to Use XPath in Python

Use Python’s built-in ElementTree for simple XPath-style lookups, or lxml for full XPath 1.0. Runnable examples, namespace handling, troubleshooting, and practical guidance.

By the ScreenshotNeo team4 October 20268 min read

Python offers two practical ways to query XML with XPath expressions: the standard-library xml.etree.ElementTree supports a limited XPath subset, while lxml.etree evaluates full XPath 1.0 expressions. Use ElementTree for straightforward paths when you want no extra dependency; choose lxml when you need richer predicates, functions, namespaces, or parameterized queries.

1. Choose ElementTree or lxml

Need Use Why
Simple child or descendant lookups xml.etree.ElementTree It ships with Python and supports a limited set of XPath-style paths.
Full XPath 1.0 expressions, functions, or richer predicates lxml.etree It provides XPath 1.0 evaluation through .xpath().
XPath against namespaced XML Either, depending on the query ElementTree has namespace-aware path syntax; lxml accepts explicit prefix-to-URI mappings in .xpath().
Reuse a query with changing values lxml.etree XPath variables keep changing values separate from the expression.

The Python documentation describes ElementTree as providing “limited support for XPath expressions for locating elements in a tree.” It is not a complete XPath engine; do not assume that every standard XPath function, axis, or expression is supported. See the official ElementTree documentation and the lxml XPath guide.

2. Use XPath-style paths with ElementTree

For uncomplicated paths, ElementTree’s find(), findall(), and iterfind() methods are often sufficient. This runnable example parses a small document, finds direct book children, then locates all titles below the root:

import xml.etree.ElementTree as ET

xml_text = """<catalog>
  <book id="b1"><title>XPath Basics</title></book>
  <book id="b2"><title>Python XML</title></book>
</catalog>"""

root = ET.fromstring(xml_text)
books = root.findall("./book")
matching_titles = root.findall(".//book/title")

for title in matching_titles:
    print(title.text)

Output:

XPath Basics
Python XML

Common ElementTree path forms

Path Meaning
./book Direct book children of the current element.
.//title Descendant title elements at any depth.
[@id='b2'] An attribute predicate; use it as part of a path, such as ./book[@id='b2'].
.. Parent step in supported path contexts.
[1] A positional predicate in supported contexts; check the documented subset for its exact semantics.

These examples are illustrative of the supported subset, not a promise that arbitrary XPath syntax will work. If a path fails because it uses an unsupported function or axis, either rewrite it using supported lookups or use lxml.

3. Use full XPath 1.0 with lxml

Install lxml in the active environment, then call .xpath() on an element or tree. This complete example selects the book with ID b2 and prints its title:

python -m pip install lxml
from lxml import etree

xml_text = """<catalog>
  <book id="b1"><title>XPath Basics</title></book>
  <book id="b2"><title>Python XML</title></book>
</catalog>"""

root = etree.fromstring(xml_text.encode("utf-8"))
books = root.xpath("//book[@id='b2']")

if books:
    print(books[0].findtext("title"))

Output:

Python XML

Unlike ElementTree’s limited lookup syntax, lxml’s XPath evaluator supports XPath 1.0. An XPath result is not always a list of elements: depending on the expression, it can be a scalar such as a string, number, or boolean. For example, string(//book[1]/title) returns a string. Write callers to expect the result type implied by the expression.

4. Handle variables and namespaces

Pass changing values as lxml variables

Do not build an XPath expression by inserting untrusted or changing text into the expression string. Pass it as a variable instead:

from lxml import etree

xml_text = b"<catalog><book id='b1'/><book id='b2'/></catalog>"
root = etree.fromstring(xml_text)

book_id = "b2"
results = root.xpath("//book[@id=$book_id]", book_id=book_id)
print(len(results))  # 1

Variables help avoid quoting mistakes and keep data separate from the XPath expression. They are especially useful when values contain quotes or come from user input.

Query namespaced XML

In XPath, an unprefixed element name does not match an element in a default XML namespace. Bind a convenient prefix to the namespace URI and use that prefix in the XPath expression. The chosen XPath prefix need not be the same prefix, or lack of prefix, used in the source document:

from lxml import etree

xml_data = b'<root xmlns="urn:catalog"><item>Example</item></root>'
root = etree.fromstring(xml_data)

items = root.xpath("//c:item", namespaces={"c": "urn:catalog"})
print(items[0].text if items else "No item found")

For ElementTree, provide a prefix-to-URI mapping to its path lookup method too:

import xml.etree.ElementTree as ET

xml_data = b'<root xmlns="urn:catalog"><item>Example</item></root>'
root = ET.fromstring(xml_data)
items = root.findall(".//c:item", {"c": "urn:catalog"})
print(items[0].text if items else "No item found")

ElementTree still supports only its documented path subset in this example. A namespace mapping does not turn it into a full XPath implementation.

5. Parse XML from a file safely

For a local XML file you control, parse it and run the same lookup methods on the resulting root or tree:

import xml.etree.ElementTree as ET

tree = ET.parse("catalog.xml")
root = tree.getroot()
for book in root.findall("./book"):
    print(book.get("id"), book.findtext("title"))

Use lxml.etree.parse("catalog.xml") and tree.xpath(...) when the query needs XPath 1.0. For untrusted XML, review the XML parser’s security guidance and configure parsing for the threat model; avoid casually enabling external entity or network access. Also distinguish malformed XML from a query that found no matching nodes: parsing raises a parse error, while a valid query with no matches usually returns an empty result.

6. Common errors and fixes

Symptom Likely cause Fix
ElementTree raises a syntax or path error The expression uses syntax outside ElementTree’s supported subset. Check the official supported-path documentation; simplify the lookup or use lxml.etree.xpath().
Query returns an empty list Wrong path context, capitalization, namespace, or element hierarchy. Inspect the parsed tree and test a simple path such as .//*; bind the namespace URI and query with a prefix where needed.
Namespaced elements never match The XML elements belong to a namespace, but the XPath uses unqualified names. Use a prefix mapped to the exact namespace URI in the XPath call.
IndexError after selecting results The code assumes a match exists. Check whether the result list is nonempty before indexing, or handle the missing case explicitly.
ModuleNotFoundError: No module named 'lxml' lxml is not installed in the Python environment running the script. Run python -m pip install lxml with the same interpreter or environment used to execute the program.
Unexpected string instead of elements The XPath expression returns a scalar value, for example through string() or count(). Match the code to the expression’s result type rather than treating every result as an element.
Parse error before XPath runs The input is malformed XML or encoded differently than expected. Validate the source document and its encoding; fix parsing first, then debug the query.

7. Performance, reliability, and dependency trade-offs

  • Choose by capability first. ElementTree avoids an added dependency and is adequate for simple paths. lxml adds an installable library in exchange for full XPath 1.0 support.
  • Measure your workload. There is no single useful speed claim for all documents and expressions. Parsing cost, tree size, query shape, library version, and how often the query runs all matter. Benchmark representative inputs in your own environment before changing libraries for speed.
  • Reuse parsed trees when appropriate. If several queries operate on the same unchanged document, parse once and run the queries against that tree rather than reparsing for each lookup.
  • Handle missing data explicitly. Empty matches are normal; check results before indexing and define what the application should do when expected nodes are absent.
  • Keep expressions maintainable. Store complex XPath expressions in named constants, test them against representative XML fixtures, and pass variable values separately when using lxml.
  • Account for input trust. XML parsing and XPath evaluation are separate stages. Apply suitable parser security settings to untrusted input and avoid constructing expressions from untrusted strings.

8. Or skip the browser setup

XPath is for selecting nodes from XML. If your actual goal is to capture a rendered web page, you need a browser-rendering step rather than an XML XPath query. ScreenshotNeo is a website screenshot API and MCP server: one GET request returns a PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. Every feature is on every plan.

Sign up free for 1,000 screenshots a month with no card.

9. FAQ

Is ElementTree a full XPath engine?

No. It implements a documented subset of XPath-style syntax for locating elements.

Does lxml support XPath 2.0 or 3.0?

The cited lxml guide documents XPath 1.0. This guide does not rely on support for later XPath versions.

Should I switch to lxml just to make queries faster?

Do not assume it will be faster for your workload. Choose based on required XPath features, then benchmark representative documents if performance is the reason for switching.

Can I use XPath to select elements from a live web page?

Not directly from a URL with these XML APIs. You must first obtain and parse suitable markup, and browser-rendered content may differ from the original response. For a screenshot of the rendered page, use a browser capture service such as ScreenshotNeo.

Further reading

For historical background, O’Reilly’s Python & XML: XML Processing with Python covers Python, XML, and XPath queries. It was published in 2001, so use current Python and lxml documentation for contemporary APIs and behavior.