ScreenshotNeo

BlogHow-to

How to Get an Element’s Attribute by XPath in Pyppeteer

Use Pyppeteer’s page.xpath() and page.evaluate() to read any DOM attribute safely, with complete examples, edge cases, and fixes.

By the ScreenshotNeo team1 October 20266 min read

Use await page.xpath() to find matching elements, then pass an ElementHandle to page.evaluate() and call the browser DOM method getAttribute():

matches = await page.xpath("//a[@class='download']")
if not matches:
    attribute_value = None
else:
    attribute_value = await page.evaluate(
        '(element) => element.getAttribute("href")',
        matches[0],
    )

print(attribute_value)

page.xpath() returns a list. Check whether it is empty before reading an index. A matching element can still return None when the requested attribute does not exist.

How XPath attribute extraction works

  1. Build an XPath expression for the element you need.
  2. Call await page.xpath(expression).
  3. Choose one returned ElementHandle, or iterate over all handles.
  4. Run JavaScript in the page with page.evaluate().
  5. Call the DOM element’s getAttribute(name) method.

The Pyppeteer API reference defines Page.xpath() as returning a list of element handles and returning an empty list when there is no match. It also documents passing an ElementHandle as an argument to Page.evaluate(). See the Pyppeteer API reference.

Complete runnable example

Install Pyppeteer and run this script against a page containing links:

python -m pip install pyppeteer
import asyncio
from pyppeteer import launch


async def main():
    browser = await launch(headless=True)
    page = await browser.newPage()
    await page.goto("https://example.com", {"waitUntil": "networkidle2"})

    matches = await page.xpath("//a")
    if not matches:
        print("No matching element")
    else:
        href = await page.evaluate(
            '(element) => element.getAttribute("href")',
            matches[0],
        )
        print("href:", href)

    await browser.close()


asyncio.run(main())

Replace the URL, XPath, and attribute name with your own values. The first launch may download a Chromium revision, so allow network access and disk space for that setup step.

Read an attribute from every XPath match

Because XPath can match multiple nodes, evaluate each handle separately:

matches = await page.xpath("//img")
values = [
    await page.evaluate(
        '(element) => element.getAttribute("src")',
        element,
    )
    for element in matches
]

print(values)

The resulting list preserves the order returned by XPath. It can contain None for elements without the attribute.

Handle missing elements and missing attributes

These are separate cases:

Situation Result Recommended handling
No element matches the XPath page.xpath() returns [] Check the list before indexing or iterating
An element matches but lacks the attribute getAttribute() returns None Handle None as an absent value
The attribute exists with an empty value Returns an empty string Do not confuse "" with None
matches = await page.xpath("//button[@data-id]")
if not matches:
    value = None
else:
    value = await page.evaluate(
        '(element) => element.getAttribute("data-id")',
        matches[0],
    )

if value is None:
    print("The element has no data-id attribute")
else:
    print("data-id:", value)

Useful XPath patterns

# Exact attribute value
//a[@class='download']

# Any element with an attribute
//*[@data-testid]

# Attribute containing text
//a[contains(@href, '/download')]

# Element text
//button[normalize-space(.)='Continue']

# Descendant relationship
//main//img[@alt]

# First matching element
(//article)[1]

Use stable attributes such as data-testid when a site provides them. Avoid selectors that depend on generated class names or a fragile position in the document.

Use Pyppeteer’s Python method names

JavaScript Puppeteer examples often show page.$x(). In Pyppeteer, use page.xpath() or its shorthand page.Jx(); Python cannot call a method named with the JavaScript $ syntax. The Pyppeteer documentation and the project’s README describe this naming difference.

Evaluate expressions safely

The callback form is the clearest approach:

value = await page.evaluate(
    '(element) => element.getAttribute("aria-label")',
    handle,
)

Pyppeteer also accepts JavaScript expressions as strings. If an expression is detected incorrectly, the documentation provides force_expr=True:

value = await page.evaluate(
    'element => element.getAttribute("href")',
    handle,
)

# Use force_expr=True only for an expression that your installed
# Pyppeteer version misclassifies.
value = await page.evaluate(
    'document.title',
    force_expr=True,
)

The arrow-function callback above is a function, so it normally does not need force_expr. Check the behavior of the Pyppeteer version installed in your environment before relying on less common evaluation forms.

Wait for dynamic content before querying

Run the XPath query after navigation and after the page has rendered the target element:

await page.goto(url, {"waitUntil": "networkidle2"})
await page.waitForSelector("a.download")
matches = await page.xpath("//a[contains(@class, 'download')]")

If the site does not expose a reliable CSS selector, wait for a short, bounded delay and still handle an empty XPath result:

await page.goto(url, {"waitUntil": "domcontentloaded"})
await asyncio.sleep(1)
matches = await page.xpath("//a[@data-file]")

Or skip the browser setup

If your goal is obtaining a clean page image rather than inspecting the DOM, ScreenshotNeo returns a screenshot or PDF from one GET request. See the ScreenshotNeo API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and billing result. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card, and paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

Troubleshooting

Error or symptom Cause Fix
IndexError: list index out of range The XPath matched nothing Check if matches before using matches[0]; verify the XPath and wait for content
Returned value is None The element exists but lacks that attribute Confirm the attribute name and treat None as absent
ElementHandleError during evaluation The handle became detached after the DOM re-rendered Run the XPath query again immediately before evaluation
No results on a JavaScript-rendered page Query ran before the app inserted the element Wait for navigation, a selector, or a bounded delay
page.$x raises an attribute error That is Puppeteer’s JavaScript method name Use Pyppeteer’s page.xpath() or page.Jx()
Evaluation syntax error JavaScript string was malformed or misdetected Use the arrow-function callback; try force_expr=True for a true expression
Browser fails to launch Chromium download, sandbox, or dependency problem Complete the Pyppeteer Chromium setup, inspect launch logs, and use the launch flags required by your host

Performance and reliability tips

  • Reuse one browser process and create pages as needed instead of launching Chromium for every attribute.
  • Keep XPath expressions specific to reduce matching and evaluation work.
  • Evaluate one handle at a time when correctness matters; each call is a browser round trip.
  • For many values on a stable page, consider one page-side evaluation that queries the DOM directly, but verify that approach against your installed Pyppeteer version.
  • Set navigation and application-level timeouts so a stalled site does not hold a worker indefinitely.
  • Close pages and browsers in a finally block in long-running services.
  • Do not trust an attribute as safe HTML or a URL without validating it for the next system that consumes it.

FAQ

Can I get an attribute without clicking the element?

Yes. XPath selection and getAttribute() read the current DOM without a click.

What does getAttribute() return for a missing attribute?

It returns None through Pyppeteer’s JavaScript-to-Python conversion.

Why is my XPath result a list?

XPath can match zero, one, or many elements, so Pyppeteer always returns a list of handles.

Should I use CSS selectors instead?

Use whichever locator expresses the page structure reliably. XPath is useful for text conditions, ancestor relationships, and attribute tests.

Does the API reference describe the newest Pyppeteer release?

The cited API reference is versioned Pyppeteer 0.0.25. Treat it as the documented behavior for that version, not as a release tracker.