How to Scrape Multiple Pages on a Dynamic Website
Learn how to find a site's data source, follow pagination, and use Playwright when dynamic content requires a browser.
To scrape multiple pages on a dynamic website, first find out where the page’s records come from. If a network request returns them as JSON or HTML, fetch that response directly and follow its page links or cursor. Use browser automation only when the records depend on browser rendering, cookies, scrolling, or interaction. In either approach, define a clear stopping condition, pace requests conservatively, and validate the collected records.
This guide shows both approaches with Python, including runnable examples for a sample site you control. The site’s actual endpoint, selectors, pagination, and access rules must be inspected before adapting the examples.
1. Inspect the site before choosing a scraper
A page that looks dynamic in a browser may still load its records from a simple HTTP endpoint. Reproducing that request is usually lighter than rendering a whole browser: it can return structured data and avoid transferring page assets. Scrapy’s dynamic content guidance recommends looking for the original data request when practical.
- Open the listing page in a browser and compare its visible records with the raw HTML response. If the records are already present in the HTML, a regular HTTP client and parser may be enough.
- Open developer tools and inspect the Network panel. Filter for Fetch/XHR requests, then reload the page.
- Record which requests occur when you click Next, change a filter, scroll to the bottom, or load more results. Look at the URL, query parameters, request headers, response type, and response body.
- Check whether the response contains the records and a next-page URL, page number, or cursor. Identify the request’s required cookies or headers if any.
- Check the site’s documented API, export options, robots.txt, terms, and applicable access rules before collecting data.
Prefer a documented API or export when it provides the required data. If a JSON request is visible, reproduce it and parse the response. If the data is only available after browser-visible interaction, use Playwright and wait for evidence that new records appeared.
2. Follow pagination with direct HTTP requests
For ordinary link pagination, fetch a page, extract its records, resolve the next link against the current URL, and stop when no next link exists. For numbered pages or known cursors, generate those requests directly when it is safe to do so. Scrapy’s tutorial demonstrates following discovered links and scheduling requests.
The following Python example works with a listing whose HTML has article.product records, an h2 title, and an a.next pagination link. Change these selectors to match the inspected site. It uses a same-host check, a page cap, a timeout, and a visited set to avoid loops.
from urllib.parse import urljoin, urlparse
import json
import time
import requests
from bs4 import BeautifulSoup
START_URL = "https://example.com/catalog"
MAX_PAGES = 100
DELAY_SECONDS = 1.0
ALLOWED_HOST = urlparse(START_URL).netloc
session = requests.Session()
session.headers.update({"User-Agent": "Research scraper contact: you@example.com"})
seen_pages = set()
items_by_id = {}
url = START_URL
while url and url not in seen_pages and len(seen_pages) < MAX_PAGES:
if urlparse(url).netloc != ALLOWED_HOST:
raise ValueError(f"Refusing to follow off-site pagination URL: {url}")
seen_pages.add(url)
response = session.get(url, timeout=(5, 30))
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
for card in soup.select("article.product"):
title_node = card.select_one("h2")
link_node = card.select_one("a[href]")
if not title_node or not link_node:
continue
item_url = urljoin(response.url, link_node["href"])
# Prefer a stable site-provided identifier when one is available.
item_id = card.get("data-id") or item_url
items_by_id[item_id] = {
"id": item_id,
"title": title_node.get_text(" ", strip=True),
"url": item_url,
"source_page": response.url,
}
next_node = soup.select_one("a.next[href]")
next_url = urljoin(response.url, next_node["href"]) if next_node else None
if next_url == url:
break
url = next_url
if url:
time.sleep(DELAY_SECONDS)
with open("items.json", "w", encoding="utf-8") as output:
json.dump(list(items_by_id.values()), output, ensure_ascii=False, indent=2)
print(f"Fetched {len(seen_pages)} pages; saved {len(items_by_id)} unique items")
Install the dependencies with python -m pip install requests beautifulsoup4. Replace the sample selectors and start URL after inspecting the target. If a page returns an API response rather than HTML, parse JSON instead of using Beautiful Soup.
When the endpoint returns JSON
Use the request found in the Network panel. The names and pagination fields below are examples; inspect the actual response and adapt them.
import json
import time
import requests
BASE_URL = "https://example.com/api/products"
params = {"limit": 50}
seen_cursors = set()
items_by_id = {}
with requests.Session() as session:
while True:
response = session.get(BASE_URL, params=params, timeout=(5, 30))
response.raise_for_status()
payload = response.json()
for item in payload["items"]:
item_id = str(item["id"])
items_by_id[item_id] = item
next_cursor = payload.get("next_cursor")
if not next_cursor or next_cursor in seen_cursors:
break
seen_cursors.add(next_cursor)
params["cursor"] = next_cursor
time.sleep(1.0)
with open("items.json", "w", encoding="utf-8") as output:
json.dump(list(items_by_id.values()), output, ensure_ascii=False, indent=2)
print(f"Saved {len(items_by_id)} unique items")
Some APIs use page numbers, offset/limit parameters, Link headers, or a boolean such as has_more. Follow the endpoint’s actual contract. Avoid assuming a cursor is numeric or that the final page is short; use the server’s explicit continuation signal where available.
Numbered pages and known page counts
If the site documents a stable page URL pattern and the number of pages is known, create the URLs directly. For a production crawl, still cap the total, record failed pages, and verify that page numbering does not skip or repeat records. A sequential crawl is simpler and gentler; parallel requests can reduce elapsed time but should be limited by the target’s published policy and observed response behavior.
3. Use Playwright when the content needs a browser
Use a browser when records appear only after JavaScript runs, a user action is required, or browser state is necessary. Playwright’s Python documentation covers installation and browser automation. Prefer a condition tied to content changing over a fixed sleep: network delays vary, and a sleep can be both wasteful and too short.
This example assumes the page has a Next button and cards with article.product. It waits for the first card’s text to change after clicking Next. Adapt the selectors and change condition to the site. It saves a record snapshot per page and stops when Next is disabled or the cap is reached.
import asyncio
import json
from playwright.async_api import async_playwright
START_URL = "https://example.com/catalog"
MAX_PAGES = 100
async def main():
all_items = {}
visited_page_signatures = set()
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True)
page = await browser.new_page()
await page.goto(START_URL, wait_until="domcontentloaded", timeout=45_000)
await page.locator("article.product").first.wait_for(state="visible", timeout=20_000)
for page_number in range(1, MAX_PAGES + 1):
cards = page.locator("article.product")
count = await cards.count()
page_items = []
for index in range(count):
card = cards.nth(index)
title = (await card.locator("h2").inner_text()).strip()
href = await card.locator("a[href]").first.get_attribute("href")
item_url = page.url if not href else page.url.rstrip("/") + "/" + href.lstrip("/")
item_id = await card.get_attribute("data-id") or item_url
item = {"id": item_id, "title": title, "url": item_url, "source_page": page.url}
all_items[item_id] = item
page_items.append(item)
signature = tuple(item["id"] for item in page_items)
if not signature or signature in visited_page_signatures:
break
visited_page_signatures.add(signature)
next_button = page.get_by_role("button", name="Next")
if await next_button.count() == 0 or not await next_button.is_enabled():
break
first_title = await cards.first.locator("h2").inner_text()
await next_button.click()
await page.wait_for_function(
"previous => { const el = document.querySelector('article.product h2'); "
"return el && el.textContent.trim() !== previous; }",
arg=first_title.strip(),
timeout=20_000,
)
await browser.close()
with open("items.json", "w", encoding="utf-8") as output:
json.dump(list(all_items.values()), output, ensure_ascii=False, indent=2)
print(f"Saved {len(all_items)} unique items")
asyncio.run(main())
Install Playwright with python -m pip install playwright and install a browser with python -m playwright install chromium. In a real crawler, resolve relative links with urllib.parse.urljoin, store page URLs and statuses, and ensure your change condition is strong enough. If the first record can remain the same across pages, wait for a changed page number, changed URL, changed cursor, or a new stable item ID instead.
Infinite scroll and load-more controls
For infinite scrolling, inspect the request that loads another batch first. If direct access is impractical, scroll the relevant container or click Load more, then wait until the item count or last item ID changes. Stop when the control disappears, becomes disabled, or repeated actions add no new IDs. Set a maximum item count and action count so a broken stop condition cannot run forever. Avoid assuming that scrolling the window works when the page uses a nested scroll container.
4. Choose the right approach
| Approach | Use it when | Trade-off |
|---|---|---|
| Documented API or export | The site provides the needed records and pagination contract. | Usually the clearest interface; access and limits still apply. |
| Direct HTTP plus parser | The records are available in HTML or a reproducible JSON request. | Low browser overhead; site changes can break selectors or request assumptions. |
| Scrapy | You need a crawl scheduler, request handling, parsing, and crawl controls. | More structure for larger crawls; learn its settings and item pipeline. |
| Playwright | Browser rendering, interaction, or browser state is required. | Higher memory and runtime cost than fetching the data endpoint directly. |
Scrapy’s dynamic-content guide explains why the underlying request is often preferable and when browser automation is useful. Its robots settings documentation also notes that robots directives and crawl pacing need explicit configuration: Scrapy does not automatically apply robots.txt Crawl-delay or Request-rate directives. Translate applicable directives into delay and concurrency settings.
5. Make pagination safe and complete
- Use an explicit end condition. Stop when the next link is absent, a cursor is exhausted, the Next control is disabled, or no new items arrive after a load action.
- Prevent loops. Keep a set of visited page URLs, cursors, or page signatures. Stop if the next value repeats.
- Set hard limits. Add a maximum page, cursor, item, and elapsed-time limit appropriate to the job.
- Deduplicate by stable identity. Prefer a site record ID. A canonical item URL can be a fallback, but URLs may change or represent variants.
- Preserve provenance. Store source page or cursor, fetch time, response status, and item count so gaps can be investigated.
- Checkpoint progress. Save each page or batch as it completes. For long jobs, persist the next cursor and processed IDs so a restart does not require starting over.
- Handle updates and shifting pages. If the listing changes during a crawl, offset pagination can skip or repeat entries. Prefer a stable cursor or date boundary if the site provides one; deduplicate and reconcile on a later run.
6. Pace requests and verify the result
Start conservatively. Watch latency, timeouts, retries, and HTTP errors as concurrency increases. Scrapy’s AutoThrottle documentation describes adjusting request rates based on observed response latency. Rising 429 or 503 responses, ban pages, growing retry counts, or rising latency are signs to reduce request pressure. Follow documented limits and access rules; do not attempt to bypass access controls.
After each run, check:
- Did the crawl reach its intended end condition, or stop only because it hit a safety cap?
- Are page or cursor values continuous, with no unexplained gaps or repeats?
- Did the number of records per page look plausible, including on the final page?
- Are stable IDs unique, and do repeated runs produce expected overlap?
- Did any response contain a login page, challenge, error page, or empty shell instead of records?
- Can you resume from the saved checkpoint without duplicating output?
7. Common errors and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| HTML contains no records, but the browser shows them | JavaScript fetches the data after page load. | Inspect Fetch/XHR requests and parse the data response; use Playwright if the request depends on browser state or interaction. |
| Only the first page is saved | The scraper never follows the next link, or it treats a button as a link. | Inspect the pagination action and response. Follow the URL or cursor it produces; automate the button only when necessary. |
| Every page contains the same records | The page action did not complete, or the scraper read before the page updated. | Wait for a changed URL, page marker, cursor, or item ID before extracting the next batch. |
| Pagination loops forever | The next link points to the current page, a cursor repeats, or a disabled control still matches. | Track visited values, verify disabled state, and use a maximum page or action count. |
| Some records are missing or duplicated | Offset pages shifted during collection, cards were skipped, or identity handling is weak. | Use stable IDs and cursor pagination if available; log per-page counts and reconcile the final set. |
| HTTP 403 or challenge page | The request is not permitted, lacks expected session context, or the site applies access controls. | Check the site’s documented access route and rules. Use an authorized API or request permission; do not evade the control. |
| HTTP 429 or 503 | Request pressure is too high or the service is temporarily unavailable. | Reduce concurrency, increase delay, honor published limits, and retry transient failures with bounded backoff. |
| Playwright times out waiting for content | The selector is wrong, the page has not reached the expected state, or a consent/login step blocks access. | Inspect the rendered DOM and page state, use a meaningful site-specific wait, and handle permitted prerequisites explicitly. |
| Relative links resolve to the wrong address | Links are joined by string concatenation or against the wrong base URL. | Resolve with a URL parser such as Python’s urljoin(response.url, href). |
8. Performance, reliability, and cost
Direct HTTP requests are generally cheaper in compute and faster per page than launching a browser, especially when the endpoint returns compact structured data. Browser automation consumes more memory and CPU, and pages may load extra scripts, images, and fonts. For a browser crawl, block unneeded resources only if doing so does not prevent the page from producing the records. Reuse a session or browser context where appropriate, and close resources cleanly.
Reliability comes from bounded retries, timeouts, checkpoints, idempotent output, and observable progress. Retry transient network failures and server errors with a maximum attempt count and backoff; do not retry indefinitely or hammer a site returning rate-limit responses. Save raw responses or a small diagnostic sample when permitted, so parser changes can be distinguished from missing upstream data.
Self-hosted scraping costs include development time, compute, bandwidth, storage, and maintenance when the site changes. A managed browser service can reduce infrastructure work, but compare its current price, limits, output format, retention, and data handling with your own requirements before choosing one. Costs and capabilities change, so check the provider’s current documentation and pricing. No target-specific performance or cost can be estimated without knowing the site, volume, and environment.
Or skip the browser setup
For page screenshots during debugging, documentation, or visual checks, ScreenshotNeo is a website screenshot API and MCP server. It is not a replacement for extracting a site’s records: it returns a screenshot or PDF. One GET request captures a URL as PNG, JPEG, or WebP, or as a PDF. See the ScreenshotNeo API documentation for its parameters and formats.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Cookie and consent banners are accepted and removed before the shot, along with known newsletter popups and chat widgets. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for AI agents. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.
Sign up for ScreenshotNeo’s free plan: 1,000 screenshots a month, no card required.
FAQ
Should I scrape a page URL or the API request behind it?
Use the documented API when available. Otherwise, if the page’s own request returns the records and is permitted to reproduce, fetching that request is often simpler than rendering the page.
How do I know when an infinite-scroll list is complete?
Use the site’s continuation signal when available. Otherwise, stop when the load control ends or repeated actions produce no new stable IDs, with a hard action and item cap.
Can I scrape pages in parallel?
Only when the site’s rules and observed responses allow it. Begin with low concurrency, then monitor latency and errors; reduce pressure if rate limits or server errors rise.
What should I do if the target changes its markup?
Keep selectors in one place, log per-page extraction counts, and alert on sudden empty or unusually small batches. Reinspect the page and update selectors or request parsing after confirming the data is still available.


