ScreenshotNeo

BlogHow-to

How to Scrape JavaScript-Rendered Websites with Python

Learn when to use Requests or a browser, then scrape rendered pages with Playwright or Selenium using reliable waits, network inspection, and validation.

By the ScreenshotNeo team4 October 202610 min read

To scrape a JavaScript-rendered website with Python, first check whether the data is already present in the server’s HTTP response. If it is, use an HTTP client such as Requests and parse the response. If JavaScript creates the data in the browser, use browser automation such as Playwright or Selenium, wait for the specific content or network response you need, then extract and validate it.

Requests is an HTTP library; it does not execute a page’s JavaScript. Browser automation runs the page in a browser, which lets client-side code render content and supports interactions such as clicking a button. The right choice depends on where the data comes from and what the page requires.

1. Diagnose whether the page needs a browser

Start with the simplest method. Fetch the page once and inspect the response for your target data. If the text or structured data is already there, parse that response directly. If the response contains only an app shell and scripts, and the content appears only after those scripts run, use browser automation.

import requests

url = "https://example.com"
response = requests.get(url, timeout=30)
response.raise_for_status()
html = response.text

needle = "Target text"
print("Target is in initial response:", needle in html)
print(html[:1000])

This check is a useful first diagnosis, not proof that every record is present. Content may be split across pages, loaded after an interaction, or returned by a separate data request. Requests is designed for HTTP requests, not browser rendering: Requests documentation.

2. Scrape rendered content with Playwright

Playwright’s Python API can automate a browser, query elements with locators, evaluate JavaScript in the page context, and observe network requests. Install the package and its browser binaries in the environment where the script will run:

python -m pip install playwright
python -m playwright install chromium

Save this as scrape.py and run python scrape.py. Replace the example URL and selector with the page and stable element that contain your data.

from playwright.sync_api import sync_playwright

URL = "https://example.com"
SELECTOR = "main article"

with sync_playwright() as playwright:
    browser = playwright.chromium.launch(headless=True)
    page = browser.new_page()
    page.goto(URL, wait_until="domcontentloaded", timeout=60_000)

    # Wait for the actual content, rather than assuming navigation means
    # the application has finished rendering.
    target = page.locator(SELECTOR)
    target.wait_for(state="visible", timeout=30_000)
    rendered_text = target.inner_text()

    if not rendered_text.strip():
        raise RuntimeError(f"The target element {SELECTOR!r} was empty")

    print(rendered_text)
    browser.close()

The selector is site-specific; inspect the page’s DOM and choose a selector that identifies the intended content. A navigation completing does not establish that a client-side application has finished loading its data. Playwright notes that modern pages can fetch data lazily and run scripts after the load event. Wait for evidence of the state you need: Playwright navigation guide.

Wait for a specific element, text, or URL

Prefer condition-based waits over a fixed sleep. A fixed delay can waste time on fast responses and still be too short on slow ones.

# Element appears
page.locator(".results-table tr").first.wait_for(state="visible", timeout=30_000)

# A locator's text matches a condition
page.get_by_text("Search results", exact=True).wait_for(timeout=30_000)

# Navigation caused by an action
with page.expect_navigation(url="**/results**"):
    page.get_by_role("button", name="Search").click()

For an in-place update, there may be no navigation. Wait for the updated locator or the response triggered by the action instead. Locators provide Playwright’s element query and interaction layer: Playwright locators.

Wait for the data response triggered by an interaction

If clicking a control loads results, watch for the expected response while performing the click. Match a distinctive part of the endpoint and, when useful, check the response status:

with page.expect_response(
    lambda response: "/api/search" in response.url and response.status == 200,
    timeout=30_000,
) as response_info:
    page.get_by_role("button", name="Search").click()

response = response_info.value
print("Data response:", response.url)
print(response.text())

Use this route when the response itself contains the records you need. If you need browser-rendered text or application behavior, extract from the DOM. Playwright can monitor HTTP and HTTPS traffic, including XHR and fetch requests, and wait for a response: Playwright network documentation.

Evaluate JavaScript in the page

For a DOM shape that is awkward to express with locators, evaluate a small expression in the page context. Python and page JavaScript run in separate environments, so pass values as explicit arguments:

titles = page.evaluate("""(selector) =>
    Array.from(document.querySelectorAll(selector), el => el.textContent.trim())
""", "main article h2")
print(titles)

For larger jobs, prefer returning only the fields you need rather than serializing an entire page. See Playwright’s JavaScript evaluation guide.

3. Use Selenium when its WebDriver model fits

Selenium is another reasonable browser automation choice. It may fit an existing project that already uses WebDriver or an environment where its browser setup and Remote protocol are useful. The Selenium Python API documentation currently describes Python 3.10+ support and Selenium Manager setup for modern Selenium on most supported platforms; verify runtime requirements against the version you install because these details can change: Selenium Python API documentation.

python -m pip install selenium

Example script, saved as scrape_selenium.py:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait

URL = "https://example.com"

options = webdriver.ChromeOptions()
options.add_argument("--headless")

driver = webdriver.Chrome(options=options)
try:
    driver.get(URL)
    article = WebDriverWait(driver, 30).until(
        EC.visibility_of_element_located((By.CSS_SELECTOR, "main article"))
    )
    text = article.text
    if not text.strip():
        raise RuntimeError("The article element was empty")
    print(text)
finally:
    driver.quit()

Selenium’s explicit wait checks for a condition, such as an element becoming visible. Avoid mixing implicit and explicit waits without understanding their interaction; use a clear wait strategy for each workflow. See Selenium waiting strategies.

4. Inspect network traffic for a direct data source

When the page’s DOM is difficult to interpret, inspect its network activity in browser developer tools or use Playwright’s request and response events. Look for the request made when the page loads or when you perform the action that reveals the data. If a normal endpoint returns the records and the site’s rules permit using it, calling it directly with Requests may be simpler than rendering the entire page each time.

  1. Open the page and inspect its XHR or fetch traffic.
  2. Trigger the action that loads the records, such as submitting a search.
  3. Identify the request URL, method, query or body parameters, and response shape.
  4. Check whether results are paginated or require additional requests.
  5. Reproduce only the request you need, and validate its output against the page.

Do not assume that an endpoint is stable, public, or permitted for your use simply because a browser can call it. Check applicable site terms, access controls, and relevant law. Network inspection explains how the page obtains data; it does not establish permission to collect it.

5. Extract, validate, and handle edge cases

Successful navigation is not successful extraction. Confirm that the result belongs to the requested page or query and includes the fields you expect.

  • Check for empty output: raise or record a clear error if the target locator is empty.
  • Check record counts: compare the number of extracted records with a known page indicator or expected range.
  • Check pagination: determine whether more pages exist and follow the site’s pagination controls or data parameters.
  • Handle optional fields: a selector may be absent for some records; treat that as a missing value rather than assuming every record has the same shape.
  • Handle lazy content: scroll or interact only when needed, then wait for a specific new element or response.
  • Separate navigation from updates: an in-place search may update the DOM without changing the URL.
  • Close the browser: put shutdown in a finally block or a context manager so failures do not leave browser processes running.

Prefer stable attributes, roles, or page structure over brittle positional selectors. Store enough context with each result to detect accidental duplicates or results from a previous query.

6. Choose between Requests, Playwright, and Selenium

Approach Use it when Consider
Requests and an HTML parser The needed content is in the HTTP response or an accessible data endpoint. It does not run page JavaScript or perform browser interactions.
Playwright You need browser rendering, locators, page-context evaluation, or response observation. Install and manage the browser runtime; wait for the target state explicitly.
Selenium Your project or environment fits WebDriver, its supported browser setup, or Remote protocol. Choose explicit waits and manage browser shutdown on errors.

Both Playwright and Selenium can automate browser behavior. Pick based on the APIs, browser environment, remote execution needs, and existing team code that fit the task. The available documentation does not establish that one is always faster or more reliable.

7. Troubleshooting

Symptom Likely cause Fix
Requests returns no target data The page creates the content in the browser, or data comes from another request. Inspect the response and network traffic; use browser automation or an appropriate data request.
Timeout waiting for a locator The selector is wrong, the content is not present for this query, or the page is still waiting on data. Inspect the rendered DOM, confirm the query and selector, then wait for the actual element or expected response.
Navigation finishes but extracted text is empty The application renders after navigation or updates in place. Wait for a content condition or data response instead of treating navigation as readiness.
A click seems to do nothing The application may not yet be hydrated, or the control selector may match the wrong element. Wait for the control to be actionable, verify its locator, and wait for a response or visible state change after the click.
Selenium cannot start the browser The browser or runtime setup is missing or incompatible. Check the installed Selenium and browser versions, runtime requirements, and Selenium Manager setup for the platform.
Results are incomplete or duplicated Pagination, lazy loading, or stale results were not handled. Track page/query state, wait for new records, and validate counts and unique keys.
Fixed sleeps work intermittently Load time varies and the delay is not tied to page state. Replace the sleep with a locator, URL, text, or response condition.

8. Performance, reliability, and cost

An HTTP request is generally the simpler path when it contains the needed data: it avoids starting a browser and executing the page. Browser rendering adds setup and runtime work, but it is necessary when the page creates or reveals the target content in the browser. No universal speed ratio follows from the documentation; measure your own target and workload.

  • Reuse the browser: for multiple pages in one process, reuse a browser instance where practical and create isolated pages or contexts for separate sessions.
  • Wait narrowly: wait for the specific content or response, not an arbitrary long delay or every network request to stop.
  • Extract only what you need: reduce data transferred from the page context and avoid unnecessary full-page work.
  • Bound resources: set navigation and condition timeouts, close pages and browsers, and cap concurrency to what the machine can support.
  • Make retries selective: retry transient navigation or network failures with a limit; do not retry a bad selector or empty result forever.
  • Track failure types: record URL, query, wait condition, and error so a changed page is distinguishable from a slow response.

Browser scraping has compute and maintenance costs: browser processes use more resources than a simple HTTP request, and page redesigns can invalidate selectors. A direct endpoint can reduce rendering work, but it may change and still needs pagination, response validation, and permission checks.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. If your task is to inspect a rendered page visually or capture it for an agent, one API request returns an image or PDF. It does not return scraped DOM data, so use the Python browser workflow above when you need structured page content.

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo API documentation for request options. Cookie banners, newsletter popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents use screenshot and PDF tools. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.

Get 1,000 free screenshots a month with no card.

FAQ

Can Requests scrape a JavaScript-rendered page?

Requests can fetch the server’s HTTP response, but it does not run browser JavaScript. Use it when that response or a data endpoint contains what you need; otherwise use browser automation.

Should I use Playwright or Selenium?

Either can be appropriate. Choose according to the browser environment, interaction and waiting APIs, remote execution needs, and existing project setup.

Is it enough to wait for the page load event?

No. Modern pages can fetch or render data after that event. Wait for the element, text, URL, or response that demonstrates the specific data is ready.

Can I scrape any endpoint I see in browser traffic?

Finding an endpoint does not establish that its use is allowed. Check the site’s rules, access controls, and applicable law before collecting data.

Does ScreenshotNeo return the page’s text or DOM?

No. ScreenshotNeo returns screenshots or PDFs. Use Playwright, Selenium, or an appropriate HTTP request when you need text or structured data.