ScreenshotNeo

BlogHow-to

How to Parse Dynamic CSS Classes When Web Scraping

Learn how to handle changing CSS classes in scraped pages, distinguish unstable selectors from JavaScript-rendered content, and extract data reliably with Python and Playwright.

By the ScreenshotNeo team29 September 202610 min read

How to Parse Dynamic CSS Classes When Web Scraping

When CSS class names change during web scraping, avoid relying on generated styling classes as your only selector. First check whether the target data exists in the original HTML or appears only after JavaScript runs. Then prefer a stable semantic signal, such as an accessible role and name, label, ID, or explicit data-* attribute. Use a class only when you have checked that it is stable across representative pages and renders.

These are two separate problems: an unstable class name makes an element hard to identify, while JavaScript-rendered content may not exist in the HTML you downloaded. A better selector cannot find content that has not been rendered. This guide shows how to diagnose both cases, parse static HTML with Python, inspect rendered pages with Playwright, and validate your extraction so changes fail visibly.

1. Diagnose what “dynamic” means

A page may use classes that look generated, such as card_x7f3a, or classes whose names vary between builds. That is a selector-stability problem. Separately, a page may initially return a shell and fill in the data using client-side JavaScript. That is an availability problem. It can have stable classes and still be absent from a plain HTTP response.

Check whether the target data is already in the response or appears only after the browser runs the page.
Check whether the target data is already in the response or appears only after the browser runs the page.

Start with a representative URL and compare the original response HTML with the browser-rendered DOM. In your browser’s developer tools, inspect the target element. You can also fetch the URL and search the response for a distinctive piece of its expected content. If the data is in the response, a static parser may be sufficient. If it appears only after scripts run, use browser automation or another permitted way to obtain the rendered data.

Observation Likely issue Next step
Data is in the response, but its class token changes Unstable selector Find a semantic or explicit identifier; otherwise constrain class matching and validate it
Data is missing from the response but visible in the browser Client-side rendering Wait for a meaningful rendered element, then inspect the DOM
Data is in the response and a meaningful attribute is present Neither, or both are manageable Parse the document using that attribute

2. Choose a locator that describes the target

Use a selector as a contract with the page, not just a transcription of whatever class happens to be visible today. Prefer, in order, a user-facing role and accessible name, a label, a meaningful ID, or an explicit data attribute intended for stable identification. Playwright recommends prioritizing user-facing attributes and explicit contracts such as page.getByRole() [Playwright locator guidance](https://playwright.dev/docs/locators).

Prefer a locator that expresses the target's meaning over a class used only for styling.
Prefer a locator that expresses the target's meaning over a class used only for styling.

A locator based on article, a button’s accessible name, or data-product-id says more about the target than a generated class token. When CSS or XPath is necessary, keep it short and focused. Long chains that depend on exact nesting or positional structure are more likely to break when the DOM changes. Playwright supports CSS and XPath, but its locator guidance explains why structural selectors can be brittle [Playwright locator guidance](https://playwright.dev/docs/locators).

Signal Example When to use
Accessible role and name getByRole('link', { name: 'Details' }) The element has user-facing semantics and a useful accessible name
Label getByLabel('Price') A form control is associated with a meaningful label
Explicit data attribute [data-product-id] The page exposes an identifier as a stable data or test contract
ID or semantic HTML #main, article The identifier or element meaning matches the extraction target
Class .product-card No better signal exists and the class has been checked across examples

Do not assume every class that looks random is guaranteed to change, or that every readable class is stable. Check several representative pages and, where relevant, more than one render. If the page provides only changing classes, identify the useful node with surrounding semantics or attributes, and match only the part of the class that is known to be meaningful. Avoid treating a short prefix as durable unless observation supports it.

3. Parse static HTML with Python

For an HTML response that already contains the target, Python’s Requests and Beautiful Soup provide a straightforward workflow. Install the dependencies with python -m pip install requests beautifulsoup4. This example targets a deliberately generic product card carrying an explicit data attribute; replace the URL and attribute with ones actually present in the page.

import requests
from bs4 import BeautifulSoup

url = "https://example.com/catalog"
response = requests.get(url, timeout=30)
response.raise_for_status()

soup = BeautifulSoup(response.text, "html.parser")

for card in soup.select("[data-product-id]"):
    title = card.select_one("[data-field='title']")
    price = card.select_one("[data-field='price']")
    if title is None or price is None:
        continue
    print({
        "id": card.get("data-product-id"),
        "title": title.get_text(" ", strip=True),
        "price": price.get_text(" ", strip=True),
    })

Beautiful Soup supports class lookup through class_ and CSS selection through Tag.select() [Beautiful Soup documentation](https://www.crummy.com/software/BeautifulSoup/bs4/doc/). For example, a known class can be queried like this:

cards = soup.find_all(class_="product-card")
# Equivalent CSS-style selection for a single class:
cards = soup.select(".product-card")

Multiple class matching needs care. In CSS, .item.featured means an element has both classes, while .item .featured means a descendant with the second class. Beautiful Soup’s CSS selector support follows CSS-style selection. If you need to inspect a class attribute without matching a generated token exactly, inspect the values and then apply a narrowly justified predicate:

for node in soup.find_all(True):
    classes = node.get("class", [])
    if any(name.startswith("product-card-") for name in classes):
        print(node.get_text(" ", strip=True))

This prefix example is only appropriate if you confirmed that the prefix is a meaningful, recurring part of the page’s markup. Otherwise it merely replaces one fragile assumption with another. Prefer an explicit attribute or semantic relationship whenever possible.

4. Inspect JavaScript-rendered pages with Playwright

If the response lacks the target, load the page in a browser and wait for a meaningful element or state. Do not rely on a fixed delay as the default: it may waste time on fast pages and still be too short on slow ones. Playwright locators can wait for their target to become available, and support user-facing locators as well as CSS and XPath [Playwright locators](https://playwright.dev/docs/locators), [page API](https://playwright.dev/docs/api/class-page).

Install Playwright for Python and its browser with python -m pip install playwright followed by python -m playwright install chromium. This runnable example waits for an explicit data marker and extracts visible text:

from playwright.sync_api import sync_playwright

url = "https://example.com/catalog"

with sync_playwright() as p:
    browser = p.chromium.launch()
    page = browser.new_page()
    page.goto(url, wait_until="domcontentloaded", timeout=30000)
    page.locator("[data-product-id]").first.wait_for(state="visible", timeout=15000)

    cards = page.locator("[data-product-id]").all()
    for card in cards:
        product_id = card.get_attribute("data-product-id")
        title = card.locator("[data-field='title']").inner_text()
        price = card.locator("[data-field='price']").inner_text()
        print({"id": product_id, "title": title, "price": price})

    browser.close()

Replace the data selectors with the stable identifiers available on the target. If the page has no explicit marker, a role locator may express the intent better; for example, page.get_by_role("link", name="Product details"). If you must use a class, use the smallest selector that identifies the intended item, then assert that the count and extracted fields make sense.

Wait for the right state

  • Navigation: choose a load condition appropriate to the page. domcontentloaded waits for the document to be parsed, not for every application request or delayed widget.
  • Target readiness: wait for the actual data-bearing element, a known result count, or another meaningful state. This is stronger than assuming that navigation means the data is ready.
  • Lazy content: if records load only after scrolling or interaction, reproduce the required interaction and wait for the next batch before extraction.
  • Multiple matches: check the number of cards and handle zero or duplicate matches explicitly; do not silently accept the first result when the page should contain exactly one.

5. Validate and make changes visible

Selectors can stop matching after a site update, and a selector can also match the wrong content without raising an error. Validate your assumptions rather than allowing an empty or misleading dataset to pass unnoticed. This is implementation guidance based on the brittleness of selectors and DOM structure noted in Playwright’s documentation [Playwright locator guidance](https://playwright.dev/docs/locators).

expected = 1
count = page.locator("[data-product-id='sku-123']").count()
if count != expected:
    raise RuntimeError(f"Expected {expected} product, found {count}")

price = page.locator("[data-product-id='sku-123'] [data-field='price']")
if price.count() != 1:
    raise RuntimeError("Product price is missing or ambiguous")

For a repeated listing, validate a sensible range or known invariant rather than expecting a fixed count if the catalog naturally changes. Record enough context to diagnose failures, such as the URL, response status, match count, and whether the expected text was present. Avoid logging credentials, private cookies, or unnecessary personal data.

6. Troubleshooting

Symptom Likely cause Fix
Static parser finds zero target elements Content is rendered by JavaScript, selector is wrong, or the request returned another page Check status and response HTML; inspect rendered DOM if needed; verify the target selector against the actual markup
Selector matched yesterday but not today Generated class changed or DOM structure was redesigned Replace the selector with a semantic or explicit identifier; reduce structural dependencies
Selector returns too many elements It matches a generic class shared by unrelated nodes Scope it to a semantic container or stable data attribute and assert the expected count
Playwright times out waiting for a locator The target never appeared, the page is slower than the timeout, or the selector is invalid Inspect the live DOM and console/network behavior; wait for the correct state; increase the timeout only when the expected rendering legitimately takes longer
Text is empty or incomplete Extraction ran before content was filled, or text is in a different node Wait for the data field itself, then inspect its parent and child nodes
Records are duplicated The selector includes hidden templates, repeated responsive markup, or nested cards Scope to visible result containers and check each record’s unique identifier
Request returns access denied or a challenge The site blocks or challenges the request Respect the site’s rules and access controls; do not attempt to bypass a challenge. Confirm permitted access or use an authorized data interface

7. Performance, reliability, and operating cost

Static HTTP parsing generally avoids launching a browser, so it has less setup and fewer moving parts. Use it when the needed data is present in the response. Browser automation is appropriate when the data requires client-side rendering or interaction, but a browser carries more resource and timing overhead. These are architectural tradeoffs, not benchmark claims.

Improve reliability by reusing a browser process for a batch of pages where practical, waiting for the target state rather than sleeping, and limiting concurrency to what the machine and target site can handle. Retry transient navigation failures with a bounded policy, but do not retry a permanently invalid selector as if it were a network problem. Distinguish a timeout from a successful page that contains zero results, and preserve enough diagnostics to inspect failures.

Keep request volume reasonable and check the target site’s terms, robots directives, rate limits, and access requirements. They vary by site; this technical guide is not permission to scrape any particular site. If you pay for compute, account for browser runtime and retries as well as request volume. A direct parser has fewer runtime needs, while a rendered-browser workflow may need more memory and longer worker time. No universal cost or speed figure applies without measuring your own pages and environment.

8. Or skip the browser setup

If you need a rendered screenshot to inspect what the page looks like after loading, ScreenshotNeo can return an image or PDF from one request. It is a website screenshot API and MCP server; its capture options include waiting for a selector or network idle, custom CSS and JavaScript, and element capture. A screenshot helps inspect the rendered result, though it does not replace DOM extraction when your output needs structured fields.

See the ScreenshotNeo API documentation for request options. Example request for a page image:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));

Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. The MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up free for 1,000 screenshots a month, no card required.

9. Frequently asked questions

Can I use a regular expression to match a changing class?

You can inspect class attributes and apply a prefix or pattern, but that is only reliable when the matched portion is a known, stable convention. Prefer an explicit data attribute or semantic locator if the page offers one.

Should I parse the screenshot to get the data?

Usually no. A screenshot is useful for visual inspection, while structured extraction should use the document or rendered DOM. Image-based extraction is a separate approach and may lose structure and precision.

Does waiting for network idle guarantee the data is ready?

No. A page may keep connections open, or it may update after network activity settles. Waiting for the data-bearing element or state is usually more directly tied to the result you need.

What if the site provides no stable identifier?

Use the shortest selector supported by the surrounding semantics, validate its match count and extracted values, and monitor it across representative pages. Expect to revisit it if the site’s markup changes.

References