ScreenshotNeo

BlogHow-to

How to Use CSS Selectors in Python

Learn CSS selectors in Python with Beautiful Soup, lxml, and cssselect, including runnable examples, debugging, performance, and production tips.

By the ScreenshotNeo team29 September 202611 min read

How to Use CSS Selectors in Python

Short answer: parse the HTML first, then run a CSS selector against the parsed tree. For most Python projects, Beautiful Soup provides the simplest API: use select() for every match and select_one() for the first match. If your project already uses lxml or XPath, use lxml.cssselect.CSSSelector instead.

A CSS selector is a query. It does not download a page, execute JavaScript, or create a document by itself. Your program must obtain HTML from a file, HTTP response, browser session, or another source before parsing and selecting it. Python’s standard-library html.parser can parse markup through callbacks, but it does not include a CSS selector query method; the official documentation describes an HTMLParser instance that receives data and calls handler methods for tags, text, comments, and other markup events. See the Python html.parser reference.

1. The basic Beautiful Soup workflow

Install Beautiful Soup with:

CSS selectors query the tree created by an HTML parser.
CSS selectors query the tree created by an HTML parser.
python -m pip install beautifulsoup4

Beautiful Soup’s current documentation says its CSS selector implementation is Soup Sieve, which is installed with Beautiful Soup through pip. This complete example parses an HTML string, selects all matching articles, selects one heading, reads an attribute, and handles a missing result safely.

from bs4 import BeautifulSoup

html = '''
<main>
  <article class='story' data-kind='guide'>
    <h2>Selectors</h2>
    <a href='/learn'>Read more</a>
  </article>
  <article class='story' data-kind='reference'>
    <h2>Reference</h2>
    <a href='/reference'>Open reference</a>
  </article>
</main>
'''

soup = BeautifulSoup(html, 'html.parser')

# All matching elements: returns a list of Tag objects.
articles = soup.select("article.story[data-kind='guide']")

# First matching element, or None when there is no match.
heading = soup.select_one('article.story h2')

print([article.get_text(' ', strip=True) for article in articles])
print(heading.get_text(strip=True) if heading else 'No heading found')

# Read an attribute from the first matching link.
link = soup.select_one('article.story a[href]')
print(link.get('href') if link else 'No link found')

The selector combines a type selector (article), a class selector (.story), an attribute equality test ([data-kind='guide']), and a descendant combinator (the space before h2). Beautiful Soup documents these forms along with child combinators, positional selectors, and attribute prefix, suffix, and substring matching.

select() versus select_one()

Method Result Use it when
soup.select(selector) List of all matching tags You need every card, link, row, or repeated field
soup.select_one(selector) First matching tag, or None A field is expected once or is optional
tag.select(selector) Matches below a specific tag You want to restrict a query to one component

Always check optional results before calling methods on them. Calling get_text() on None raises an AttributeError. For repeated data, iterate over the returned list and use tag.get('attribute') rather than indexing an attribute directly.

2. CSS selector patterns you will use most

Selector Meaning Example
article Every element of a type soup.select('article')
.story Any element with a class soup.select('.story')
#content The element with an ID soup.select_one('#content')
main h1 An h1 anywhere inside main soup.select_one('main h1')
main > h1 An immediate child soup.select_one('main > h1')
a[href] A link that has an href soup.select('a[href]')
a[href^='/docs'] href starts with a value Internal documentation links
img[src$='.webp'] src ends with a value WebP image URLs
div[data-id*='user'] Attribute contains a value Data attributes with a known fragment
ul li:nth-of-type(2) The second li of its type Position within a list
h1, h2 Either selector Collecting multiple heading levels

Prefer stable attributes such as semantic tags, data-testid, or documented IDs. Long chains of generated classes are fragile because a template redesign can change them without changing the content you need.

Reading text and attributes

for card in soup.select('article.card'):
    title = card.select_one('h2, h3')
    price = card.select_one('[data-price]')

    record = {
        'title': title.get_text(' ', strip=True) if title else None,
        'price': price.get('data-price') if price else None,
    }
    print(record)

get_text(' ', strip=True) joins nested text with spaces and removes surrounding whitespace. Use tag.get('href') for an optional attribute; it returns None when the attribute is absent. If you need all values, collect them explicitly:

links = [
    a.get('href')
    for a in soup.select('nav a[href]')
    if a.get('href')
]

3. Selecting from an HTTP response

Fetching and selecting are separate operations. The following example uses requests to obtain HTML and Beautiful Soup to query it.

import requests
from bs4 import BeautifulSoup

url = 'https://example.com/'
response = requests.get(url, timeout=30)
response.raise_for_status()

soup = BeautifulSoup(response.text, 'html.parser')
title = soup.select_one('title')
print(title.get_text(strip=True) if title else 'No title')

A response can be successful while containing an error page, a login page, or only a client-side application shell. Inspect response.status_code, response.url, the Content-Type header, and a short prefix of response.text when a selector unexpectedly returns no results. Keep site policies, authentication requirements, and permission to retrieve content separate from selector syntax.

4. lxml and cssselect

Use lxml when your project already relies on lxml’s tree and XPath APIs, or when you want to translate CSS selectors into XPath. The lxml.cssselect documentation describes CSSSelector as a convenience API that translates a CSS selector for lxml’s XPath engine. The standalone cssselect project translates CSS3 selectors to XPath 1.0 expressions.

from lxml import html
from lxml.cssselect import CSSSelector

markup = '''
<main>
  <article class='story' data-kind='guide'>
    <h2>Selectors</h2>
    <a href='/learn'>Read more</a>
  </article>
</main>
'''

tree = html.fromstring(markup)
select_guides = CSSSelector("article.story[data-kind='guide']")

for article in select_guides(tree):
    heading = article.cssselect('h2')
    print(heading[0].text_content().strip() if heading else 'No heading')

# The same selection expressed directly as XPath.
print(tree.xpath("//article[contains(concat(' ', normalize-space(@class), ' '), ' story ') and @data-kind='guide']"))

lxml and Beautiful Soup differ in parser behavior, tree APIs, and supported selector details. Choose based on the rest of your application, the malformed markup you expect, and the selector features supported by your installed versions. Beautiful Soup’s documentation recommends lxml for selector-only workflows and describes it as faster; that is qualitative project guidance, not a universal benchmark. Measure your own workload if throughput matters.

5. Python’s built-in html.parser

html.parser is useful when you want a dependency-free, event-driven parser or need to build a custom data structure while the document is read. It is not a CSS selector engine. A handler can record tags and text, but implementing reliable CSS queries yourself means creating a tree, parsing selector syntax, and matching combinators and attributes.

from html.parser import HTMLParser

class HeadingParser(HTMLParser):
    def __init__(self):
        super().__init__()
        self.in_h1 = False
        self.parts = []

    def handle_starttag(self, tag, attrs):
        if tag == 'h1':
            self.in_h1 = True

    def handle_endtag(self, tag):
        if tag == 'h1':
            self.in_h1 = False

    def handle_data(self, data):
        if self.in_h1:
            self.parts.append(data)

parser = HeadingParser()
parser.feed('<h1>A heading</h1>')
print(''.join(parser.parts).strip())

Use Beautiful Soup or lxml when the requirement is “find elements with this CSS selector.” Use html.parser when callback processing and a small custom parser are more appropriate.

6. Version and compatibility details

Beautiful Soup’s documentation records Soup Sieve integration beginning in Beautiful Soup 4.7.0 and the .css property arriving in 4.12.0. Confirm the API against the version installed in your project, especially when deploying from a lockfile or supporting multiple environments.

from bs4 import BeautifulSoup

soup = BeautifulSoup('<main><h1>Hello</h1></main>', 'html.parser')

# Equivalent styles in versions that expose the css interface.
print(soup.select_one('main h1').get_text(strip=True))
print(soup.css.select_one('main h1').get_text(strip=True))

Selector support varies between implementations and versions. A selector that works in a browser is not automatically supported by every Python library. Check the documentation for your installed Beautiful Soup/Soup Sieve, lxml, or cssselect version before relying on advanced pseudo-classes.

7. Debugging selectors

  1. Print the actual markup. Save or print the response body around the expected element. You may be parsing a redirect, an error page, or a different template.
  2. Start with a broad query. Try soup.select('article'), then add the class, attribute, and relationship one condition at a time.
  3. Check nesting. A descendant selector such as main h2 allows intervening elements; main > h2 requires a direct child.
  4. Check attribute spelling and whitespace. HTML class values are space-separated. Use .card rather than comparing the entire class attribute.
  5. Check parser input. Pass the HTML string, not a requests.Response object, to Beautiful Soup.
  6. Check rendering. If the HTML source lacks content that appears in an interactive browser, a JavaScript-capable browser capture or API is required before parsing.

Common errors and fixes

Symptom Likely cause Fix
select_one() returns None No matching element in the parsed HTML Inspect the source and simplify the selector
AttributeError: 'NoneType' object has no attribute ... Code used a missing optional match as though it existed Check the result before reading text or attributes
Empty list from select() Wrong nesting, class, attribute, or response body Print response.url and a markup sample
Browser shows content, parser does not Content is inserted after JavaScript runs Obtain rendered HTML with a permitted browser workflow or use an appropriate capture service
Selector raises a syntax error Unsupported or malformed selector Check the library’s selector documentation and reduce the expression
Unexpected duplicate text Nested elements are both selected or text is repeated Select the smallest required node and inspect descendants

8. Performance, reliability, and cost

For one document, selector matching is usually a small part of total work compared with downloading, parsing, and (if applicable) rendering the page. Avoid repeatedly parsing the same HTML. Parse once, scope queries to a containing element, and reuse a compiled CSSSelector in lxml when processing many documents with the same selector.

Reliability comes from validating inputs and outputs: set network timeouts, call raise_for_status(), record the final URL after redirects, and treat missing elements as a normal data-quality case when pages vary. Store the response or a sanitized fixture when a production selector fails so you can reproduce the exact markup.

Library parsing has no per-request service fee, but your application still pays for network traffic, compute, and browser infrastructure. A browser is necessary when you need JavaScript-rendered content, interactions, authentication flows, or a visual screenshot. Keep browser concurrency bounded, cache stable results where policy permits, and use retries only for transient failures.

9. Or skip the browser setup

If your goal is to inspect or capture a rendered page before selecting content, ScreenshotNeo provides a one-call website capture API. It accepts a URL and returns a PNG, JPEG, WebP, or PDF. Before capture it accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off.

A rendered capture can remove obstructing overlays before you inspect or save the page.
A rendered capture can remove obstructing overlays before you inspect or save the page.

Only clean shots are billed. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and each response reports the result through X-Page-Verdict and X-Billed headers.

cURL

curl -G 'https://api.screenshotneo.com/v1/shot' \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp

Python

import requests

r = requests.get(
    'https://api.screenshotneo.com/v1/shot',
    params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'},
    timeout=90,
)
r.raise_for_status()
open('shot.webp', 'wb').write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const bytes = Buffer.from(await res.arrayBuffer());
require('node:fs').writeFileSync('shot.webp', bytes);

See the ScreenshotNeo API documentation for the complete option list. Relevant capture controls include full-page screenshots with lazy images loaded, one-element capture by CSS selector, dark mode, 12 device presets or any viewport, retina scale, PDF paper size and margins, custom CSS and JavaScript, click and wait actions, network-idle or selector waits, blocked ads and resource types, custom headers, cookies, user agents and Authorization, timezone and geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, usage reporting, and an OpenAPI specification. Parameter names used by other screenshot APIs also work to ease migration.

ScreenshotNeo also includes an MCP server for AI agents, with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. Pricing is Free for 1,000 shots per month with no card, then Starter is $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000; yearly billing gives two months free and every feature is on every plan.

Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.

10. Practical checklist

  • Obtain the HTML and confirm it is the document you expect.
  • Parse once with Beautiful Soup or lxml.
  • Use select() for collections and select_one() for an optional single result.
  • Read text with get_text(' ', strip=True) and attributes with get().
  • Prefer stable semantic selectors and documented data attributes.
  • Test selectors against fixtures representing missing fields and template variations.
  • Use a rendered browser capture when required content is absent from the raw HTML.
  • Log the selector, URL, parser version, and a safe markup sample when extraction fails.

FAQ

Can CSS selectors fetch a web page?

No. They query a parsed document. Fetch the HTML separately with an HTTP client, read a file, or obtain rendered markup from a browser workflow.

Which Python library should a beginner choose?

Beautiful Soup is usually the most approachable when you want CSS selectors and straightforward tree navigation. Choose lxml when XPath, lxml trees, or an existing lxml codebase are central to the project.

Can I use browser-only selectors unchanged in Python?

Not always. Selector support depends on the implementation and version. Check Soup Sieve, lxml, or cssselect documentation for the syntax you plan to use.

How do I select an element by class in Beautiful Soup?

Use a class selector such as soup.select_one('.story') for one element or soup.select('.story') for all matching elements.

How do I select an element by attribute?

Use an attribute selector such as soup.select('a[href]') or soup.select("[data-kind='guide']"), then read the value with tag.get('href').

When should I use XPath instead?

Use XPath when your lxml application already uses XPath expressions or needs XPath-specific relationships and functions. lxml’s CSSSelector lets you keep CSS syntax while using the XPath engine.