How to Capture Dynamic Content in a Web App with Infinite Scroll and Filters
Capture filtered infinite-scroll results reliably with Playwright: inspect requests, synchronize each load, paginate safely, and validate completeness.
To capture all results from a web app with infinite scroll and filters, first inspect how its filter and scroll actions affect network requests. Then automate those same actions with Playwright, wait for a matching response or visible result change after each action, extract and deduplicate records, and stop only when the app exposes a credible end-of-results signal. The pagination method, response shape, and filter behavior are specific to each app.
This guide uses Python and Playwright for the browser workflow. It also shows how to inspect requests in browser developer tools, how to reason about cursors and filter state, and how to validate what you collect. No target app was provided, so the example uses placeholders that you must adapt to the app’s selectors and response format.
1. Inspect the app before automating it
Start with the results page in a normal browser session. Learn what the app actually does before choosing whether to extract rendered elements, capture structured network responses, or use both.
- Identify the scrollable area. The whole document may scroll, or results may live in an inner panel with its own scrollbar.
- Note how each filter is represented: a select, checkbox, button, text field, or custom widget. Check whether changing a filter changes the URL.
- Open the browser’s developer tools and select the Network panel. Filter to Fetch/XHR, apply one filter, and inspect the resulting request and response.
- Trigger one more batch by scrolling. Compare the new request: look for a page number, offset, cursor, query parameter, or request body, and check the response for a next-page or completion field.
- Record which fields identify a result and whether the response contains the same records shown in the interface.
Infinite-scroll apps commonly request the next page, batch, cursor, or search results when scrolling. That is a common pattern, not a guarantee: some apps load everything in the page, virtualize rendered rows, or use a different mechanism. The Playwright Automation Practice Infinite Scroll Lab demonstrates the loading pattern, but its book list is a client-side simulation without a real backend endpoint. Treat it as a demonstration, not evidence of how your target service works.
Network inspection can expose useful structured data, but the request may depend on the current session, headers, cookies, or a POST body. Preserve only the context the app requires, and never put credentials in shared logs.
2. Choose what to collect: responses, rendered content, or both
| Method | Use it when | Watch for |
|---|---|---|
| Capture matching network responses | The app returns the result records in a parseable response and you need structured fields. | Requests and response formats are app-specific. A response may include metadata or fields that the UI does not display. |
| Read rendered DOM elements | The visible results are the reliable source, or the response is not useful for your task. | Selectors can change, and virtualized lists may remove off-screen rows from the DOM. |
| Use both | You want structured records while checking that the interface reflects the expected batch. | Correlate the response with the filter and pagination state that produced it. |
Playwright can monitor browser network activity, including XHR and fetch. Use the browser to discover the app’s behavior first. Replaying a captured request directly can be simpler later, but it may not reproduce client-side behavior and will fail if required session context, request headers, or body are missing.
3. Install Playwright and prepare the page
Install the Python package and browser binaries in your project environment:
python -m pip install playwright
python -m playwright install chromium
Replace the example URL, filter selector, result selector, and response predicate below with values observed in your target app. The example demonstrates a whole-document scroll and DOM extraction. If the app uses an inner scroll panel, scroll that element instead, as shown in the next section.
4. Automate filters and capture each batch in Python
Register a response wait before the action that should trigger the request. This avoids missing a fast response. If the request is not predictable, wait for a visible result change instead. The code deliberately leaves response fields and selectors as placeholders because no universal schema exists.
import asyncio
import json
from urllib.parse import urlparse
from playwright.async_api import async_playwright, TimeoutError as PlaywrightTimeoutError
URL = "https://example.com/results"
FILTER_SELECTOR = "select[name='category']"
RESULT_SELECTOR = "article.result-card"
TITLE_SELECTOR = "h2"
ID_ATTRIBUTE = "data-id"
async def main():
records_by_id = {}
seen_without_id = set()
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True)
page = await browser.new_page()
await page.goto(URL, wait_until="domcontentloaded")
# Apply the filter using the same control a user operates.
# For a custom control, use its observed role/label or click sequence.
await page.locator(FILTER_SELECTOR).select_option(label="Example category")
# Replace this URL test with a predicate matching the request observed
# in DevTools. Register the wait before triggering a load action.
try:
async with page.expect_response(
lambda response: "/api/search" in response.url
and response.request.method in ("GET", "POST"),
timeout=15000,
) as response_info:
# If the filter itself causes the request, move the filter action
# inside this expect_response block.
await page.locator(RESULT_SELECTOR).first.wait_for()
response = await response_info.value
print("Initial response:", response.status, response.url)
except PlaywrightTimeoutError:
# Some apps render from cached state or use a different endpoint.
# Replace this with a known visible-state condition for that app.
await page.locator(RESULT_SELECTOR).first.wait_for(timeout=10000)
previous_count = -1
stable_rounds = 0
max_batches = 100 # Safety bound; choose an app-appropriate limit.
for batch_number in range(max_batches):
cards = page.locator(RESULT_SELECTOR)
count = await cards.count()
for i in range(count):
card = cards.nth(i)
record_id = await card.get_attribute(ID_ATTRIBUTE)
title = (await card.locator(TITLE_SELECTOR).inner_text()).strip()
record = {"id": record_id, "title": title}
if record_id:
records_by_id[record_id] = record
else:
# A title-only key can merge distinct records with the same
# title. Prefer a stable ID or a composite key when available.
seen_without_id.add(title)
# If the app exposes an explicit end marker, check it here and stop
# only when it is associated with the active filtered result set.
end_marker = page.get_by_text("No more results", exact=True)
if await end_marker.count():
break
# A whole-document scroll example. For a nested panel, use the
# scroll-container example below instead.
before = await page.locator(RESULT_SELECTOR).count()
await page.evaluate("window.scrollTo(0, document.body.scrollHeight)")
try:
# Prefer a matching response wait when the request is known.
async with page.expect_response(
lambda response: "/api/search" in response.url,
timeout=15000,
) as response_info:
# Trigger the scroll after the waiter is registered.
await page.evaluate("window.scrollTo(0, document.body.scrollHeight)")
response = await response_info.value
if not response.ok:
raise RuntimeError(
f"Batch request failed: HTTP {response.status} {response.url}"
)
# Optionally parse response.json() here if the endpoint's schema
# was inspected and its records are more reliable than the DOM.
except PlaywrightTimeoutError:
# Fallback: wait for the rendered list to grow. If the app
# replaces rows instead of appending, wait for a known batch
# marker or compare stable record IDs instead.
try:
await page.wait_for_function(
"([selector, oldCount]) => "
"document.querySelectorAll(selector).length > oldCount",
arg=[RESULT_SELECTOR, before],
timeout=10000,
)
except PlaywrightTimeoutError:
current_count = await page.locator(RESULT_SELECTOR).count()
if current_count == before:
stable_rounds += 1
else:
stable_rounds = 0
current_count = await page.locator(RESULT_SELECTOR).count()
if current_count == previous_count:
stable_rounds += 1
else:
stable_rounds = 0
previous_count = current_count
# This is only a bounded fallback, not proof of completion. Prefer
# server pagination metadata or an explicit end-of-results signal.
if stable_rounds >= 3:
print("No visible growth for three rounds; completion is uncertain.")
break
await browser.close()
output = {
"records": list(records_by_id.values()),
"records_without_stable_id": sorted(seen_without_id),
}
with open("results.json", "w", encoding="utf-8") as f:
json.dump(output, f, ensure_ascii=False, indent=2)
print(f"Saved {len(records_by_id)} records with stable IDs")
asyncio.run(main())
The sample has two deliberate adaptation points: the filter may itself be the action that triggers the first request, and scrolling may trigger a later request. Put each action inside an expect_response block when you know its request signature. Avoid using a response predicate so broad that it matches analytics, images, or an unrelated background request.
5. Handle nested scroll containers and virtualized lists
A scroll command on the document does not move an inner panel. Inspect the app to find the element whose scroll position changes, then target it explicitly. Playwright’s Page API documents page interaction and scrolling behavior, including bringing elements into view.
# Example: results live in an inner panel.
PANEL_SELECTOR = "div.results-panel"
panel = page.locator(PANEL_SELECTOR)
await panel.evaluate("el => { el.scrollTop = el.scrollHeight }")
For a virtualized list, the DOM may contain only the visible rows. In that case, repeated extraction of the current DOM will miss records that have been removed after scrolling. Capture the structured response batches when available, or extract and persist each visible batch before moving on. Deduplicate using a stable record ID where possible.
6. Treat each filter combination as its own query
Think of filter state as defining a query, while the pagination token advances within that query. This is a useful operating model, not a guarantee about every app. When a filter changes, check whether the app replaces results, resets a page number or cursor, or appends to the existing feed. Treat an old cursor as stale unless inspection shows that the app intentionally keeps it.
- Store the selected filter values alongside every captured batch.
- Record the request URL, method, relevant query parameters, and pagination state without logging secrets.
- Keep output grouped by filter combination when each combination represents a separate result set.
- Deduplicate across batches by a stable ID if one exists. If it does not, document your composite key and the possibility of collisions.
- Do not assume that every result visible in multiple filter states is a duplicate; that depends on what the app’s records represent.
Pagination may use a page number, offset, cursor token, or POST search body. These are examples to look for during inspection, not fields every app provides. If the app returns hasMore, a nextCursor, or a total count, use the observed field and its actual semantics to decide when to stop.
7. Synchronize on the action, then verify completion
A scroll command only moves the viewport; it does not prove the next batch has loaded. For each filter or scroll action, wait for one of these conditions:
- A response matching the request triggered by the action.
- A known result card or batch marker to appear or change.
- A visible count, cursor marker, or end-of-results indicator to update.
Playwright’s documentation exposes response-waiting and network monitoring APIs. Its Page documentation cautions against using networkidle as a generic readiness check: “Don’t use this method for testing, rely on web assertions to assess readiness instead.” A fixed sleep can also race with slow requests or waste time when requests finish quickly. Use bounded timeouts around an observable condition and report a timeout as a failed or incomplete batch.
Prefer a server-provided completion field or a visible finished-state marker tied to the active query. A page height that stops changing once is weak evidence: delayed requests, nested scroll areas, and virtualized rows can all make height misleading. If the app exposes no trustworthy completion signal, report that completeness is uncertain and validate the capture against the interface.
8. Validate and make the capture resumable
Validation is part of the capture, not an optional cleanup step. For representative filter states:
- Compare a sample of saved records with visible results in the app.
- Compare counts with the UI or observed pagination totals when those values are available and have matching meanings.
- Check for duplicate IDs, missing fields, and batches with unexpectedly few or zero records.
- Store the filter values and page/cursor for each batch so you can diagnose failures and resume when the app permits it.
- Repeat a filter state to see whether results or ordering changed. Record that difference rather than assuming the data is static.
Keep enough diagnostic detail to tell which query and batch failed, but redact access tokens, session cookies, and personal data from logs. Observe the service’s access rules and rate limits; the app’s network behavior does not itself grant permission to collect or reuse its data.
9. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| The first filter selection produces no matching response. | The predicate does not match the real endpoint, the control did not submit, or the response happened before the waiter was registered. | Inspect the request in DevTools, register the waiter before the filter action, and match a stable endpoint substring plus method. |
| Scrolling does not load more results. | The wrong element was scrolled, the app needs a different trigger, or the current filter has no more results. | Check which element’s scroll position changes. Scroll the inner panel if present, and inspect whether a new request is sent. |
| The script times out waiting for a response. | The app uses another endpoint or transport, the action did not fire, or the request is blocked or delayed. | Confirm the request manually, narrow or correct the response predicate, and use a visible state change if the load has no relevant response. |
| New rows appear but the response waiter times out. | The response predicate is wrong, or the app renders from cached or client-side data. | Fix the predicate based on the observed request, or wait for an action-specific locator/count change instead. |
| Some records are missing after scrolling. | The list is virtualized and old DOM rows disappear, or the script scrolls again before extraction finishes. | Persist each batch before advancing. Prefer structured response records when available, and deduplicate by stable IDs. |
| The capture stops too early. | A temporary stall or unchanged page height was mistaken for completion. | Use observed pagination metadata or an explicit end marker. Treat repeated no-growth as uncertain and validate counts or samples. |
| Changing a filter returns unexpected or repeated records. | Pagination state may have carried over, or results may append instead of replace. | Inspect the filter request and result transition. Start a fresh query state where appropriate and group data by filter combination. |
| The captured request works in the browser but fails when replayed. | It may require session cookies, authorization, headers, a POST body, or other app-specific state. | Inspect the full request shape and preserve required context securely. If replay remains unreliable, keep the browser workflow. |
| A field or selector breaks after an app update. | The UI or undocumented response schema changed. | Log the failing query and batch, inspect the current DOM/request, update selectors or parsing, and rerun representative validation. |
10. Performance, reliability, and cost
- Performance: Response extraction can avoid repeatedly parsing a growing DOM when the app exposes usable structured batches. DOM extraction is straightforward but may do more work as the feed grows. Measure on the target app rather than assuming one method is always faster.
- Reliability: Action-specific waits, bounded timeouts, resumable batch records, and a credible completion signal help distinguish a complete capture from a stalled one. UI selectors and undocumented request formats can both change.
- Completeness: Virtualization, changing data, stale cursors, and ambiguous end markers can all affect results. Preserve query state and report uncertainty when the app provides no trustworthy completion signal.
- Cost: A local Playwright script has no per-screenshot API charge, but it uses compute and requires browser maintenance and engineering time. Direct request replay may use fewer browser resources after it is understood, but adds the burden of maintaining app-specific request context.
- Access: Respect the target service’s permissions, terms, and rate limits. Do not treat a discovered endpoint as authorization to collect its data.
11. ScreenshotNeo for a clean view of the page
ScreenshotNeo is a website screenshot API and MCP server for developers. A screenshot can help you inspect how a filtered or scrolled state looks, but it does not collect every result or replace the Playwright workflow above. Its one-request API returns a PNG, JPEG, WebP, or PDF, and its options include full-page capture with lazy images loaded, element capture, custom waits, and custom JavaScript. See the ScreenshotNeo API documentation for request options.
Or skip the browser setup
For a screenshot of a page state, make one GET request:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/results -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com/results"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/results' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, and failed loads are never billed, and cache hits cost nothing. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots. Every feature is on every plan.
Sign up free for 1,000 screenshots a month, with no card required.
12. FAQ
Can a screenshot capture every item in an infinite list?
A screenshot captures a visual page state; it is not a complete record export. Use browser automation and app-specific pagination signals to collect records.
Should I extract data from the API response or the page?
Use the response when inspection shows it contains the records you need. Use the DOM when rendered content is the relevant source, and account for virtualization. You can also compare both.
Is there a universal end condition for infinite scroll?
No. Use the target app’s observed pagination metadata or end marker. If neither is available, validate carefully and state that completeness cannot be confirmed from the interface alone.
Does ScreenshotNeo replace Playwright for filtered data collection?
No. ScreenshotNeo captures page images or PDFs. Playwright remains the appropriate approach in this guide for operating filters, following pagination, and collecting records.


