How to Scrape AJAX Websites with Python
Learn when to call an AJAX endpoint directly, when to use Playwright, and how to wait for and validate JavaScript-loaded data in Python.

Direct answer: inspect the browser’s Network panel first. If the data comes from a reproducible JSON or HTML request, call that endpoint with Python’s requests library. If JavaScript execution, scrolling, clicks, authentication state, or client-side rendering is required, use Playwright for Python. In either case, wait for the specific response or DOM condition that means your data is ready, then validate the status and payload before extracting fields.
An AJAX page can finish its initial navigation while the content you need is still being fetched. A completed load event is therefore not a reliable signal that the page is ready. Playwright’s navigation guide puts it plainly: “There is no way to tell that the page is loaded, it depends on the page, framework, etc.” Playwright navigation guidance explains why readiness must be defined for the page and data you are collecting.
1. What AJAX scraping means
AJAX is the common name for browser code that requests data after the initial document arrives. The page may use fetch(), XMLHttpRequest, GraphQL, or another client-side transport. JavaScript then inserts the response into the DOM. If you fetch the URL with requests.get() and inspect only the initial HTML, you may see an empty table, a loading placeholder, or no records at all.
Your first decision is architectural:
| Situation | Recommended approach | What you wait for |
|---|---|---|
| A stable endpoint returns the needed records | Direct HTTP with requests or Playwright’s API request client |
HTTP response, status, and expected JSON shape |
| Data appears only after JavaScript runs | Playwright browser automation | Matching network response or a specific locator |
| A click, scroll, filter, or login triggers the request | Playwright page actions | Response associated with that action |
| Content is handled by a service worker | Playwright with service workers blocked when routing is required | Observed request or rendered content |
This is a workflow recommendation based on the documented capabilities of Python HTTP clients and Playwright. It does not establish that any discovered endpoint is public, stable, or permitted for high-volume collection. Check the target site’s terms, access controls, and applicable rules before running a scraper.
2. Inspect the page before writing code
- Open the page in a normal browser.
- Open Developer Tools and select the Network panel.
- Reload the page with the panel open.
- Trigger the action that reveals the data: click a tab, submit a search, change a filter, or scroll.
- Filter requests by
Fetch/XHRand inspect candidates. - Record the method, URL, query parameters, request body, relevant headers, cookies, and response format.
Playwright can monitor browser network activity, including XHR and fetch, and can wait for a response after an action. The official Network documentation shows this response-wait pattern. Treat the inspection as diagnosis rather than proof that an endpoint is intended for unrestricted automation.

3. Directly call the AJAX endpoint with Python
Use a direct request when the endpoint supplies everything you need and you can reproduce its required inputs. This is normally faster and simpler than launching a browser because it avoids rendering, layout, and JavaScript execution.
import requests
url = "https://example.com/api/items"
params = {
"page": 1,
"limit": 50,
"query": "python",
}
headers = {
"Accept": "application/json",
"User-Agent": "ajax-scraper/1.0",
}
response = requests.get(url, params=params, headers=headers, timeout=30)
response.raise_for_status()
payload = response.json()
items = payload.get("items", [])
for item in items:
print(item)
Replace the URL, parameters, and JSON field names with values observed on the target. Keep the timeout finite. Use response.raise_for_status() so a 404 or 503 cannot silently become an empty result.
Send a POST request
import requests
response = requests.post(
"https://example.com/api/search",
json={"query": "python", "page": 1},
headers={"Accept": "application/json"},
timeout=30,
)
response.raise_for_status()
result = response.json()
print(result)
Sessions, cookies, and authentication
Use a Session when several calls share cookies or connection state. Never hard-code credentials in source control; load them from environment variables or a secret manager.
import os
import requests
with requests.Session() as session:
session.headers.update({
"Accept": "application/json",
"User-Agent": "ajax-scraper/1.0",
})
session.cookies.set("consent", "accepted", domain="example.com")
response = session.get(
"https://example.com/api/items",
headers={"Authorization": f"Bearer {os.environ['API_TOKEN']}"},
timeout=30,
)
response.raise_for_status()
data = response.json()
A browser’s request may depend on a CSRF token, a short-lived cookie, a signed parameter, or a request body generated by JavaScript. In those cases, copying one request from DevTools may work briefly and then fail. Either reproduce the complete flow or use browser automation.
4. Use Playwright when JavaScript or interaction is required
Install the Python package and a browser:
python -m pip install playwright
python -m playwright install chromium
The most reliable pattern is to pair the action with page.expect_response(). The context manager starts waiting before the click, preventing a fast response from being missed.
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
page.goto("https://example.com", wait_until="domcontentloaded", timeout=60_000)
with page.expect_response(
lambda response: "/api/data" in response.url and response.request.method == "GET",
timeout=30_000,
) as response_info:
page.get_by_text("Load data").click()
response = response_info.value
if not response.ok:
raise RuntimeError(f"Unexpected HTTP status: {response.status}")
payload = response.json()
print(payload)
browser.close()
The domain, endpoint, and button above are placeholders. Replace them with values found during inspection. Use a narrow predicate: matching only a stable URL fragment can accidentally capture an unrelated request when a page makes several similar calls.
Wait for rendered content
When the useful result is visible in the DOM rather than convenient JSON, wait for a stable locator or condition.
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch()
page = browser.new_page()
page.goto("https://example.com/search", wait_until="domcontentloaded")
page.get_by_role("button", name="Search").click()
results = page.locator("[data-testid='result']")
results.first.wait_for(state="visible", timeout=30_000)
for index in range(results.count()):
print(results.nth(index).inner_text())
browser.close()
A fixed sleep() is a poor default: it is either too short during a slow response or wasteful during a fast one. Wait for the response or content condition that represents readiness.
Navigation and network-idle caveats
page.goto() completing means navigation completed, not that every later AJAX task completed. Some pages keep analytics, streams, or polling requests open, so a universal “network idle” rule can be slow or never reached. Prefer the one response or locator your extraction needs. Playwright’s navigation documentation describes these page-lifecycle distinctions.
5. Validate every response before extracting
HTTP completion and HTTP success are different. Playwright documents that responses such as 404 and 503 still complete as HTTP responses. Check all of the following:
- The response was received before the timeout.
- The status is acceptable, usually 2xx.
- The content type is what you expect.
- The JSON or HTML has the expected top-level structure.
- The result is not an error object, login page, consent page, or empty placeholder.
def require_items(payload):
if not isinstance(payload, dict):
raise ValueError("Expected a JSON object")
items = payload.get("items")
if not isinstance(items, list):
raise ValueError("Expected an items array")
return items
Fail loudly with the URL, status, and a short response excerpt. Avoid logging tokens, cookies, or personal data.
6. Pagination, lazy loading, and scrolling
AJAX data is often paginated. Prefer the endpoint’s explicit page, cursor, or offset parameters when available. Stop when the API returns no records or no next cursor, and put a maximum page count in place to prevent an accidental infinite loop.
import requests
all_items = []
cursor = None
for _ in range(100):
params = {"limit": 100}
if cursor:
params["cursor"] = cursor
response = requests.get("https://example.com/api/items", params=params, timeout=30)
response.raise_for_status()
payload = response.json()
batch = payload.get("items", [])
all_items.extend(batch)
cursor = payload.get("next_cursor")
if not batch or not cursor:
break
else:
raise RuntimeError("Pagination limit reached")
If records appear only after scrolling, automate bounded scrolls and wait for the count to increase. Do not assume that reaching the bottom once loads everything; virtualized lists may remove earlier DOM nodes while retaining data elsewhere.
7. Service workers, threads, and browser lifecycle
Browser routing can miss requests intercepted by service workers. The Playwright Page reference recommends blocking service workers when request interception must observe those requests. Configure this deliberately because it can change page behavior.
Playwright’s Python API is not thread-safe. The Python library documentation advises creating an independent Playwright instance per thread if a multi-threaded design is necessary. A safer default is one worker process or browser context per job, with bounded concurrency and explicit cleanup.
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch()
context = browser.new_context(service_workers="block")
page = context.new_page()
page.goto("https://example.com", wait_until="domcontentloaded")
# Observe or extract the required data here.
context.close()
browser.close()
8. Reliability and performance checklist
- Use direct HTTP when possible: it usually consumes fewer CPU and memory resources than a browser.
- Reuse connections: use a
requests.Sessionfor related calls. - Set bounded timeouts: separate connect and read timeouts when your client supports them.
- Retry selectively: retry transient network failures and selected 5xx responses with backoff; do not blindly retry authentication or validation errors.
- Respect rate limits: pace requests and honor explicit site instructions.
- Cache carefully: cache immutable pages or responses, but do not serve stale data when freshness matters.
- Keep browser contexts short-lived: close pages, contexts, and browsers even after exceptions.
- Record provenance: store the request URL, retrieval time, status, and parser version with extracted records.
- Monitor shape changes: alert when expected fields disappear or result counts suddenly become zero.
There is no universal speed or success percentage for AJAX scraping. Performance depends on the target, endpoint, payload size, browser work, network path, and access policy. Measure your own workload rather than assuming a benchmark.
9. Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| Initial HTML has no records | Records are inserted after navigation | Find the XHR/fetch request or use Playwright and wait for the result. |
expect_response() times out |
Wrong URL predicate, action did not fire, or request was served by a service worker | Inspect the exact request, verify the locator, and consider blocking service workers for routing. |
| Response arrives but JSON parsing fails | HTML login page, error document, or incorrect content type | Check status and Content-Type; log a redacted body excerpt. |
| Empty results after a successful status | Missing cursor, filter, cookie, CSRF token, or required header | Compare the complete browser request with your Python request. |
| Works manually but fails in automation | Authentication state, consent, timing, or browser-only JavaScript dependency | Use a persistent authenticated context where permitted, wait on a condition, and reproduce required state. |
| Intermittent failures under concurrency | Shared Playwright objects or excessive parallelism | Use one instance per thread/process, bound concurrency, and close resources. |
| 404 or 503 treated as valid data | Completion was checked but status was not | Check response.ok or call raise_for_status() before parsing. |
10. Or skip the browser setup
If your goal is a clean screenshot of a JavaScript-rendered page rather than structured record extraction, ScreenshotNeo provides a one-request website screenshot API. It runs the capture in a browser and accepts the cookie or consent banner like a visitor, then removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot. Each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers.

See the ScreenshotNeo API documentation for the full option list, including full-page capture, element selectors, device presets, retina scale, custom CSS and JavaScript, waits, request blocking, headers, cookies, geolocation, caching, PDFs, async jobs, bulk capture, signed links, and usage reporting.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also has an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
11. Cost and operational planning
Direct HTTP requests generally have lower infrastructure overhead because they do not require a browser process. Browser automation costs more CPU and memory, especially with multiple contexts, but it is often the only reliable option for client-rendered workflows. Keep browser concurrency bounded, reuse contexts where state can safely be shared, and collect timing and failure data before tuning.
For ScreenshotNeo, only clean shots are billed. Cache hits, bot checks or CAPTCHAs, blank pages, timeouts, and failed loads are not billed, and the response headers tell you which case occurred. The available plans are Free: 1,000 shots/month; Starter: $5 for 3,000; Growth: $15 for 15,000; Pro: $39 for 60,000; Scale: $99 for 250,000; and Business: $249 for 1,000,000. Yearly billing gives two months free, and every feature is available on every plan.
12. FAQ
How do I know whether a page uses AJAX?
Reload it with Developer Tools open and watch Fetch/XHR requests. If the visible data arrives after the document, it is being loaded dynamically.
Should I copy an AJAX URL directly?
Only when you can reproduce its required method, parameters, headers, cookies, and authentication state, and your use complies with the site’s rules.
Is Selenium required?
No. Playwright provides a Python browser automation API and explicit network-response waiting. Selenium is another option, but the patterns here use Playwright.
Why did my scraper return an empty list without an exception?
It may have parsed a valid error object, login page, or placeholder. Validate status, content type, and the expected response shape before extraction.
Can I run Playwright from multiple threads?
Playwright’s Python API is not thread-safe. Use a separate instance per thread or isolate work in processes.
When should I choose ScreenshotNeo?
Choose it when the output you need is a screenshot or PDF of a JavaScript-rendered page and you want browser setup, consent cleanup, and billing signals handled by an API.
Conclusion
Reliable AJAX scraping starts with observation. Find the request that supplies the data, use direct HTTP when it is reproducible, and switch to Playwright when JavaScript or interaction is part of the workflow. Wait for the specific response or content condition, inspect status and payload shape, and design bounded retries, pagination, concurrency, and cleanup from the beginning.


