How to Get an Element’s Attribute by XPath in Pyppeteer
Use Pyppeteer’s page.xpath() and page.evaluate() to read any DOM attribute safely, with complete examples, edge cases, and fixes.
Use await page.xpath() to find matching elements, then pass an ElementHandle to page.evaluate() and call the browser DOM method getAttribute():
matches = await page.xpath("//a[@class='download']")
if not matches:
attribute_value = None
else:
attribute_value = await page.evaluate(
'(element) => element.getAttribute("href")',
matches[0],
)
print(attribute_value)
page.xpath() returns a list. Check whether it is empty before reading an index. A matching element can still return None when the requested attribute does not exist.
How XPath attribute extraction works
- Build an XPath expression for the element you need.
- Call
await page.xpath(expression). - Choose one returned
ElementHandle, or iterate over all handles. - Run JavaScript in the page with
page.evaluate(). - Call the DOM element’s
getAttribute(name)method.
The Pyppeteer API reference defines Page.xpath() as returning a list of element handles and returning an empty list when there is no match. It also documents passing an ElementHandle as an argument to Page.evaluate(). See the Pyppeteer API reference.
Complete runnable example
Install Pyppeteer and run this script against a page containing links:
python -m pip install pyppeteer
import asyncio
from pyppeteer import launch
async def main():
browser = await launch(headless=True)
page = await browser.newPage()
await page.goto("https://example.com", {"waitUntil": "networkidle2"})
matches = await page.xpath("//a")
if not matches:
print("No matching element")
else:
href = await page.evaluate(
'(element) => element.getAttribute("href")',
matches[0],
)
print("href:", href)
await browser.close()
asyncio.run(main())
Replace the URL, XPath, and attribute name with your own values. The first launch may download a Chromium revision, so allow network access and disk space for that setup step.
Read an attribute from every XPath match
Because XPath can match multiple nodes, evaluate each handle separately:
matches = await page.xpath("//img")
values = [
await page.evaluate(
'(element) => element.getAttribute("src")',
element,
)
for element in matches
]
print(values)
The resulting list preserves the order returned by XPath. It can contain None for elements without the attribute.
Handle missing elements and missing attributes
These are separate cases:
| Situation | Result | Recommended handling |
|---|---|---|
| No element matches the XPath | page.xpath() returns [] |
Check the list before indexing or iterating |
| An element matches but lacks the attribute | getAttribute() returns None |
Handle None as an absent value |
| The attribute exists with an empty value | Returns an empty string | Do not confuse "" with None |
matches = await page.xpath("//button[@data-id]")
if not matches:
value = None
else:
value = await page.evaluate(
'(element) => element.getAttribute("data-id")',
matches[0],
)
if value is None:
print("The element has no data-id attribute")
else:
print("data-id:", value)
Useful XPath patterns
# Exact attribute value
//a[@class='download']
# Any element with an attribute
//*[@data-testid]
# Attribute containing text
//a[contains(@href, '/download')]
# Element text
//button[normalize-space(.)='Continue']
# Descendant relationship
//main//img[@alt]
# First matching element
(//article)[1]
Use stable attributes such as data-testid when a site provides them. Avoid selectors that depend on generated class names or a fragile position in the document.
Use Pyppeteer’s Python method names
JavaScript Puppeteer examples often show page.$x(). In Pyppeteer, use page.xpath() or its shorthand page.Jx(); Python cannot call a method named with the JavaScript $ syntax. The Pyppeteer documentation and the project’s README describe this naming difference.
Evaluate expressions safely
The callback form is the clearest approach:
value = await page.evaluate(
'(element) => element.getAttribute("aria-label")',
handle,
)
Pyppeteer also accepts JavaScript expressions as strings. If an expression is detected incorrectly, the documentation provides force_expr=True:
value = await page.evaluate(
'element => element.getAttribute("href")',
handle,
)
# Use force_expr=True only for an expression that your installed
# Pyppeteer version misclassifies.
value = await page.evaluate(
'document.title',
force_expr=True,
)
The arrow-function callback above is a function, so it normally does not need force_expr. Check the behavior of the Pyppeteer version installed in your environment before relying on less common evaluation forms.
Wait for dynamic content before querying
Run the XPath query after navigation and after the page has rendered the target element:
await page.goto(url, {"waitUntil": "networkidle2"})
await page.waitForSelector("a.download")
matches = await page.xpath("//a[contains(@class, 'download')]")
If the site does not expose a reliable CSS selector, wait for a short, bounded delay and still handle an empty XPath result:
await page.goto(url, {"waitUntil": "domcontentloaded"})
await asyncio.sleep(1)
matches = await page.xpath("//a[@data-file]")
Or skip the browser setup
If your goal is obtaining a clean page image rather than inspecting the DOM, ScreenshotNeo returns a screenshot or PDF from one GET request. See the ScreenshotNeo API documentation for all options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and billing result. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card, and paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
Troubleshooting
| Error or symptom | Cause | Fix |
|---|---|---|
IndexError: list index out of range |
The XPath matched nothing | Check if matches before using matches[0]; verify the XPath and wait for content |
Returned value is None |
The element exists but lacks that attribute | Confirm the attribute name and treat None as absent |
ElementHandleError during evaluation |
The handle became detached after the DOM re-rendered | Run the XPath query again immediately before evaluation |
| No results on a JavaScript-rendered page | Query ran before the app inserted the element | Wait for navigation, a selector, or a bounded delay |
page.$x raises an attribute error |
That is Puppeteer’s JavaScript method name | Use Pyppeteer’s page.xpath() or page.Jx() |
| Evaluation syntax error | JavaScript string was malformed or misdetected | Use the arrow-function callback; try force_expr=True for a true expression |
| Browser fails to launch | Chromium download, sandbox, or dependency problem | Complete the Pyppeteer Chromium setup, inspect launch logs, and use the launch flags required by your host |
Performance and reliability tips
- Reuse one browser process and create pages as needed instead of launching Chromium for every attribute.
- Keep XPath expressions specific to reduce matching and evaluation work.
- Evaluate one handle at a time when correctness matters; each call is a browser round trip.
- For many values on a stable page, consider one page-side evaluation that queries the DOM directly, but verify that approach against your installed Pyppeteer version.
- Set navigation and application-level timeouts so a stalled site does not hold a worker indefinitely.
- Close pages and browsers in a
finallyblock in long-running services. - Do not trust an attribute as safe HTML or a URL without validating it for the next system that consumes it.
FAQ
Can I get an attribute without clicking the element?
Yes. XPath selection and getAttribute() read the current DOM without a click.
What does getAttribute() return for a missing attribute?
It returns None through Pyppeteer’s JavaScript-to-Python conversion.
Why is my XPath result a list?
XPath can match zero, one, or many elements, so Pyppeteer always returns a list of handles.
Should I use CSS selectors instead?
Use whichever locator expresses the page structure reliably. XPath is useful for text conditions, ancestor relationships, and attribute tests.
Does the API reference describe the newest Pyppeteer release?
The cited API reference is versioned Pyppeteer 0.0.25. Treat it as the documented behavior for that version, not as a release tracker.


