Frequently Asked Questions About Web Scraping and CSS Selectors
Learn how CSS selectors work in web scraping, when to use XPath, how to extract text and attributes, and how to fix empty results.
CSS selectors are patterns that identify elements in parsed HTML. A scraper uses them to find headings, links, prices, attributes, and other page data. For example, article.product h2 finds an h2 inside an element with the product class.
This guide explains selector syntax, CSS versus XPath, extraction in Scrapy and Beautiful Soup, JavaScript-rendered pages, debugging, performance, responsible crawling, and practical code examples.
What is a CSS selector in web scraping?
A CSS selector is an expression that selects nodes from an HTML document. Scrapy describes selectors as selecting parts of an HTML document with CSS or XPath expressions (Scrapy selectors documentation).
| Selector | Matches | Example |
|---|---|---|
article |
Every article element |
response.css("article") |
.product-card |
Elements with that class | response.css(".product-card") |
#main-content |
The element with that ID | response.css("#main-content") |
a[href] |
Links that have an href |
response.css("a[href]") |
[data-testid="price"] |
An exact attribute value | response.css('[data-testid="price"]') |
.product-card a.title |
A title link inside a product card | response.css(".product-card a.title") |
.menu > li |
Direct child list items | response.css(".menu > li") |
h2, h3 |
Either heading type | response.css("h2, h3") |
Use a type selector when the element name is distinctive, a stable semantic class when one exists, an ID only when it is stable and unique, and a published data attribute when the site provides one. Keep the path short. A selector such as .product-card [data-testid="price"] is generally easier to maintain than a long chain of generated classes.
How do CSS selectors work in Scrapy?
Scrapy exposes parallel CSS and XPath APIs. Both return selector lists; .get() returns the first serialized result and .getall() returns every result.
import scrapy
class ProductsSpider(scrapy.Spider):
name = "products"
start_urls = ["https://example.com/products"]
def parse(self, response):
for card in response.css("article.product-card"):
title = card.css("h2::text").get()
price = card.css("[data-testid='price']::text").get()
href = card.css("a.title::attr(href)").get()
yield {
"title": title.strip() if title else None,
"price": price.strip() if price else None,
"url": response.urljoin(href) if href else None,
}
Scrapy’s ::text and ::attr(name) forms are scraping-library extensions. They are convenient in Scrapy and Parsel but are not portable CSS syntax; the older Scrapy documentation warns that they may not work in lxml or PyQuery (Scrapy CSS extensions).
Equivalent XPath expressions
titles = response.xpath("//article[contains(@class, 'product-card')]//h2/text()").getall()
links = response.xpath("//article[contains(@class, 'product-card')]//a/@href").getall()
Use XPath when predicates, ancestor navigation, or node relationships express the condition more clearly. Scrapy translates CSS queries into XPath through its selector implementation (Scrapy selector details).
How do I extract text and attributes?
Text in Scrapy
headline = response.css("h1::text").get()
all_paragraph_text = response.css("article p::text").getall()
clean_headline = headline.strip() if headline else None
Text can be split across nested elements. To collect all descendant text, select the element and use its XPath string value:
body_text = response.css("article").xpath("string(.)").get()
body_text = " ".join(body_text.split()) if body_text else ""
Attributes in Scrapy
hrefs = response.css("a::attr(href)").getall()
image_sources = response.css("img::attr(src)").getall()
canonical = response.xpath("//link[@rel='canonical']/@href").get()
XPath attribute syntax such as //a/@href and the selector .attrib property are alternatives when you need explicit attribute handling.
Text and attributes with Beautiful Soup
from bs4 import BeautifulSoup
html = """<article class='product-card'>
<h2>Keyboard</h2>
<a class='title' href='/keyboard'>Details</a>
<span data-testid='price'>$49</span>
</article>"""
soup = BeautifulSoup(html, "html.parser")
prices = [node.get_text(strip=True) for node in soup.select(".product-card [data-testid='price']")]
first_link = soup.select_one(".product-card a.title")
url = first_link.get("href") if first_link else None
print(prices, url)
soup.select() returns all matches, while soup.select_one() returns the first match. If CSS selection is all you need, Beautiful Soup’s documentation recommends parsing with lxml directly for better speed (Beautiful Soup CSS selectors).
Should I use CSS selectors or XPath?
| Choose CSS when | Choose XPath when |
|---|---|
| You need readable type, class, ID, attribute, child, sibling, or descendant matching. | You need complex predicates or navigation to ancestors and related nodes. |
| Your team uses the same selectors in browser tools and scraping libraries. | The condition is naturally expressed as an XPath relationship. |
| You want a short selector that is easy to review. | You need functions such as contains(), starts-with(), or string(). |
Neither syntax automatically makes a scraper reliable. Stability comes from selecting meaningful containers and attributes, validating results, and monitoring changes.
Why does my selector return no results?
- The HTML is rendered by JavaScript. A normal HTTP response may contain only an empty shell. Save and inspect the response before changing the selector. Use a browser renderer when the data appears only after scripts run.
- The selector targets a generated class. Frameworks often produce classes that change between deployments. Prefer semantic classes, IDs, or
data-*attributes. - The scope is wrong. A selector run against a card, row, or response fragment cannot find nodes outside that container.
- The page structure differs by variant. Logged-in, mobile, locale, consent, and empty states can use different markup. Handle each expected state.
- The content is inside an iframe. The parent HTML does not contain the iframe document. Fetch the frame URL separately when permitted.
- You expected one result but received zero or many. Treat
None, an empty list, and multiple matches as normal cases and validate cardinality. - Pagination was missed. The first page may be correct while later pages use another URL or cursor.
results = response.css("article.product-card")
if not results:
self.logger.warning("No products at %s (status=%s)", response.url, response.status)
with open("debug-response.html", "wb") as f:
f.write(response.body)
else:
self.logger.info("Found %d products at %s", len(results), response.url)
How do I make selectors resilient?
- Start with a meaningful container such as
article.product-card. - Prefer stable semantic attributes over generated class names.
- Use the shallowest selector that uniquely identifies the data.
- Keep extraction rules tolerant of missing optional fields.
- Record URL, status, selector version, and match counts.
- Write fixture tests against saved HTML for important selectors.
- Alert when a required selector changes from one match to zero or unexpectedly many matches.
What are the performance considerations?
CSS and XPath performance depends on the parser, document size, selector complexity, and number of pages. Restrict selection to a container before extracting fields, avoid repeatedly parsing the same response, and compile or reuse parsing setup where your library supports it. Network time, JavaScript rendering, retries, and rate limits usually dominate total crawl time.
For Beautiful Soup workloads that only need CSS selection, lxml can be faster according to the Beautiful Soup documentation. Benchmark with your actual HTML and extraction volume rather than assuming one parser wins every workload.
How should I handle JavaScript pages and screenshots?
Use a browser-capable capture or rendering step when the required DOM appears only after JavaScript runs. A screenshot can also verify what a visitor sees, but image pixels are not a substitute for structured extraction. Wait for a meaningful selector or network idle state, and account for consent banners, delayed images, and infinite scrolling.
Or skip the browser setup
ScreenshotNeo provides a website screenshot API and MCP server. The one-call request below returns a WebP image; see the ScreenshotNeo API documentation for options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and the response identifies the page verdict and billing status in headers. Its MCP server lets Claude, Cursor, and other MCP clients use take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots each month without a card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
Does robots.txt make scraping legal?
robots.txt is a public file at a site’s root that communicates crawler preferences and can reduce load. It is optional and does not protect private data; malicious robots may ignore it (MDN robots.txt reference).
Treat it as one operational signal. Also review the site’s terms, authentication boundaries, applicable law, and rate limits. Identify your crawler where appropriate, cache responses, obey reasonable limits, and collect only the data you need.
Common errors and fixes
| Error or symptom | Likely cause | Fix |
|---|---|---|
None from .get() |
No match or missing optional field | Check saved HTML and handle None explicitly. |
Empty list from .getall() |
Wrong selector, scope, or JavaScript-rendered content | Inspect the fetched document and render the page if needed. |
| Too many matches | Selector is too broad | Scope it to a semantic container and validate expected counts. |
| Text is incomplete | Text is nested in child elements | Use XPath string(.) or collect descendant text. |
| Links are relative | HTML contains paths such as /pricing |
Resolve with response.urljoin(href). |
| Works locally, fails in production | Different headers, cookies, locale, rate, or response variant | Log status and response samples; make request context explicit. |
FAQ
Are CSS selectors the same as browser selectors?
The core syntax is shared, but scraping libraries add APIs and extensions. Scrapy’s ::text and ::attr() forms are examples of library-specific behavior.
Can a CSS selector extract data from JSON?
No. CSS selectors operate on parsed HTML or XML nodes. Parse JSON with a JSON parser, or locate the script element first and then decode its contents.
What is the safest selector to start with?
Use a short selector anchored to a stable semantic container or published data attribute, then verify it against saved responses and expected match counts.
Should I scrape an API instead of HTML?
If a documented, permitted API provides the required data, it is often more structured and stable than HTML. Follow its authentication, usage, and rate rules.
How do I know whether a page needs a browser?
Compare the raw response HTML with the DOM after scripts run. If the target nodes exist only after rendering, use a browser or rendering service and wait for a deterministic condition.


