CSS Selectors: A Cheatsheet for Web Scraping and HTML Parsing
Learn CSS selector syntax for scraping and parsing HTML, with runnable browser and Python examples, troubleshooting tips, and guidance on static versus rendered pages.

A CSS selector is a pattern for matching elements in a document tree. For scraping, use it with the tree your tool actually has: a browser can query its current DOM, while a static parser such as Beautiful Soup, Scrapy, or lxml queries the tree built from fetched HTML. A selector cannot fetch a page, execute JavaScript, or retrieve rendered pseudo-elements by itself.
For example, article.product a[href] matches links with an href inside elements that are both article and product. Start with a narrow selector, inspect the matched elements, and then extract the attributes or text you need. This guide covers the syntax, runnable examples, parser differences, and ways to debug empty or incorrect results.
1. What are CSS selectors?
Selectors describe patterns for matching elements in a document tree. The Selectors Level 4 specification covers HTML and XML trees and defines type, class, ID, attribute, combinator, and pseudo-class selectors. Selectors are not parsers: your scraping library first needs to build a tree, and your code then queries that tree. W3C Selectors Level 4
A selector’s job is to say which elements. Your code decides what to do with the matches, such as read text, collect a link, or parse a price. Use the simplest selector that identifies the intended elements; deeply nested selectors can break when unrelated page structure changes.
2. CSS selector cheatsheet
| Goal | Selector | What it matches |
|---|---|---|
| Every paragraph | p |
All p elements |
| Element with an ID | #main |
The element whose ID is main |
| Element with a class | .product |
Elements whose class list includes product |
| Type and class together | article.product |
article elements with class product |
| Descendant at any depth | article p |
Paragraphs anywhere inside an article |
| Direct child | ul > li |
li elements directly inside a ul |
| Next sibling | h2 + p |
A paragraph immediately after an h2 sibling |
| Later sibling | h2 ~ p |
Paragraph siblings occurring after an h2 |
| Attribute exists | a[href] |
Links with an href attribute |
| Exact attribute value | input[type="email"] |
Email inputs |
| Attribute starts with | a[href^="https"] |
Links whose href starts with https |
| Attribute ends with | a[href$=".pdf"] |
Links whose href ends in .pdf |
| Attribute contains | [data-id*="item"] |
Elements whose data-id contains item |
| Any listed alternative | h1, h2, h3 |
Elements matching any branch |
| First among siblings | li:first-child |
An li that is first among its siblings |
| Logical alternatives | button:is(.primary, .submit) |
Buttons matching either argument |
| Relational condition | article:has(img) |
Articles containing a matching image descendant |
Commas form a selector list: each branch is an alternative, not a sequence. Thus h1, h2 selects both heading types. By contrast, h1 h2 selects an h2 nested inside an h1, which is usually not what a scraper means.
3. Select by class, ID, attribute, and structure
Class and ID
Use .class-name for a class and #element-id for an ID. A class can appear on many elements; IDs are intended to identify one element in a document, but real markup can be imperfect. Combine selectors to narrow the match: section.results article.product means a product article inside a results section.
Multiple classes are written without spaces: .product.featured matches an element carrying both classes. A space means a descendant, so .product .featured instead looks for a descendant with the second class.
Attribute tests
Attribute selectors can test presence, exact value, whitespace-separated tokens, a hyphen-separated prefix, or substring patterns. The common substring operators are ^= (starts with), $= (ends with), and *= (contains). For example, [class~="featured"] checks for a whitespace-separated token, while [lang|="en"] matches the value en or a value beginning en-. See the MDN attribute selector reference for syntax details.
Be precise about what an attribute contains. An href may be relative, and a data-id can be generated or change between page loads. If you need stable extraction, prefer meaningful semantic attributes when the page provides them, and validate the resulting values in your code.
Combinators and pseudo-classes
The four common relationships are descendant (space), child (>), next sibling (+), and subsequent sibling (~). Pseudo-classes add conditions, such as structural position (:first-child) or logical matching (:is(), :where()). The relational :has() tests for a related matching element. Although these appear in the standard, support depends on the browser or parser implementation. Check the documentation for your actual runtime, especially for newer pseudo-classes. MDN CSS selectors guide
4. How do I use CSS selectors for web scraping?
- Fetch or load the page. Decide whether your source is the response HTML or a browser-rendered DOM.
- Inspect the tree. Find the element and a stable class, ID, or attribute that identifies it.
- Write a focused selector. Prefer a short selector tied to useful structure over a long chain of incidental wrappers.
- Count and inspect matches. Check that the number and content make sense before extracting at scale.
- Extract and validate. Handle missing attributes, blank text, duplicates, and relative links explicitly.
Use browser developer tools to inspect the live DOM when working in a browser. For static parsing, inspect the downloaded HTML or print a small portion of the parsed tree. Test the selector in the same environment that will run the scraper: identical syntax can have different support across browser engines and parser libraries.
5. Runnable examples
JavaScript in a browser
This example selects product links from the current document, extracts text and href values, and handles missing attributes. querySelector() returns the first match or null; querySelectorAll() returns all matches as a static NodeList. An invalid selector throws a SyntaxError. MDN querySelector()
const cards = document.querySelectorAll("article.product");
const products = Array.from(cards, card => {
const link = card.querySelector("a[href]");
return {
title: link?.textContent?.trim() ?? "",
href: link?.getAttribute("href") ?? null
};
});
console.log(products);
If a selector value comes from untrusted or arbitrary data, do not concatenate it after # or . without escaping. HTML IDs and class values are not guaranteed to be valid CSS identifiers. In a browser, use CSS.escape() when building an identifier selector:
const idFromData = "item:42";
const element = document.querySelector(`#${CSS.escape(idFromData)}`);
For values that are not identifiers, such as a full attribute value, use a correctly quoted and escaped CSS string or query broadly and compare the attribute in JavaScript.
Python with Beautiful Soup
Install the packages with python -m pip install requests beautifulsoup4. This complete example fetches a page, parses its response HTML, selects product cards, and tolerates missing links. Beautiful Soup documents select() and select_one(); its parser and version determine the tree and supported selector behavior. Beautiful Soup documentation
import requests
from bs4 import BeautifulSoup
url = "https://example.com"
response = requests.get(url, timeout=20)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
products = []
for card in soup.select("article.product"):
link = card.select_one("a[href]")
products.append({
"title": link.get_text(" ", strip=True) if link else "",
"href": link.get("href") if link else None,
})
for product in products:
print(product)
example.com is a placeholder; replace it with a page you are permitted to access. This code parses the returned response body. It does not execute page JavaScript, so content inserted only after rendering will not be present.
Scrapy and lxml
In Scrapy, a response selector exposes CSS queries alongside XPath. For example, inside a spider callback, response.css("article.product a::attr(href)").getall() retrieves matching href values; consult the Scrapy selector documentation for its syntax and extraction methods. For lxml, CSS selectors are provided through lxml.cssselect, which translates CSS selector expressions to XPath; check the lxml.cssselect documentation for dependencies and supported constructs.
Choose a library based on parser/tree construction, selector subset, and extraction workflow. Beautiful Soup integrates selectors with its tree API; Scrapy provides selectors as part of its crawling framework; lxml offers CSS support through XPath translation. Measure performance in your own workload instead of assuming a universal speed ranking. Beautiful Soup’s documentation itself notes lxml is faster and supports more selectors when CSS alone is needed, but that is not an independent benchmark for your page or setup.
6. Static HTML versus rendered browser DOM
A selector only sees the tree presented to it. A static parser selects from the HTML response it parsed. Browser APIs select from the browser document, which page scripts may have changed. A product, menu, or article visible in a browser may therefore be missing from the original response HTML. Before rewriting a selector that returns no matches, check whether the element exists in the source tree you are querying.
Pseudo-elements such as ::before are rendered abstractions, not ordinary document nodes. They are generally not a way to retrieve an HTML element’s text from a parser. If the information is represented only in generated styling content, inspect the page’s data source or computed style in a browser context instead of expecting a static parser to find a node.
If rendered content is required, use a browser automation or browser capture workflow and query after the page reaches the needed state. A screenshot can help you see what a rendered page looks like, but it does not turn image pixels into parsed HTML or make an absent source element match a selector.
7. Why does my CSS selector return no results?
| Symptom | Likely cause | What to do |
|---|---|---|
| Zero matches, no error | The target is absent from the tree, selector scope is wrong, or spelling/case differs. | Inspect the exact response or browser DOM, then test a broad selector such as article before narrowing it. |
| SyntaxError in browser | The selector string is malformed or uses unsupported syntax. | Check brackets, quotes, commas, and pseudo-class support; try it in the same browser context. |
| Some pages match, others do not | Markup varies, content is conditional, or classes are generated dynamically. | Use stable attributes, handle optional elements, and validate per-page results. |
| Wrong elements are selected | A space was used where a child combinator was needed, or a selector list was misunderstood. | Check the relationship: space means any descendant, > direct child, and comma means either branch. |
| Browser finds it, Python does not | Scripts changed the browser DOM after the original response was parsed. | Confirm the element exists in response HTML; use a rendering workflow if it is added client-side. |
| Dynamic ID causes an exception or mismatch | The value is not a valid CSS identifier or has selector-significant characters. | Escape identifier values with CSS.escape() in browsers, or avoid string construction where possible. |
| Text is empty despite a match | The node has no text, text is in a descendant, or content is rendered elsewhere. | Inspect child nodes and attributes separately; select the specific text-bearing descendant. |
:has() works in a browser but not a parser |
The parser implements a smaller selector subset. | Consult the parser’s own docs, or express the relationship with supported selectors and code. |
Browser selector APIs raise a SyntaxError DOMException for an invalid selector string. They return null or an empty collection for a valid selector with no matches; treat those as different debugging cases. MDN querySelectorAll()
8. Reliability, performance, and operating costs
- Keep selectors maintainable. A short selector using stable semantic attributes is easier to review and less sensitive to unrelated wrapper changes.
- Validate output. Record match counts and check required fields. A successful query can still return the wrong repeated element or stale page content.
- Bound network work. Set request timeouts, check HTTP status, and use an appropriate crawl policy for the site. Selector syntax does not provide these protections.
- Reuse page data. Fetch once, parse once, and run the needed selectors against that tree instead of repeatedly downloading the same page.
- Benchmark the actual task. Parsing cost depends on response size, parser, selector pattern, and workload. The dossier contains no comparative benchmark; measure on representative pages.
- Budget for rendering separately. Static requests and browser rendering have different setup and resource needs. Use rendering only when the needed content requires it, and account for its execution time and infrastructure.
For reliability, make extraction tolerant of optional elements and fail clearly when required fields are missing. Save representative HTML fixtures where appropriate so selector changes can be checked against known page structures. Do not treat a non-empty match as proof that the page is complete or current.
9. Or skip the browser setup
If the task is to capture a page image or PDF rather than extract DOM data, ScreenshotNeo is a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. Its screenshot API is separate from CSS-selector-based HTML extraction. For CSS selection during a capture, its element-capture option targets one element by selector. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits cost nothing, with response headers indicating the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page info, and PDF capture. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Every feature is on every plan. Use it when you need a clean screenshot, not as a replacement for parsing HTML.
Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.
10. Frequently asked questions
What is the difference between querySelector() and querySelectorAll()?
The first returns one matching element or null; the second returns a static NodeList containing all matches. Use the one that fits whether you expect a single result or a collection.
Can CSS selectors scrape a website on their own?
No. They match nodes in a tree supplied by a browser or parser. Fetching, rendering, extraction, and network policy are separate parts of a scraper.
Do CSS selectors work in every HTML parser?
No. Libraries can support different selector subsets and construct trees differently. Check the documentation for the specific library and version you run.
Can a CSS selector retrieve text from ::before?
Not as an ordinary HTML node. Pseudo-elements are rendered abstractions; a parser’s document tree generally does not contain them.
Should I use CSS or XPath?
Use whichever is supported and clearest for the relationship you need. Scrapy documents both. Compare maintainability and workload behavior in your own project rather than relying on a universal performance claim.


