How to Create a PDF of an Infinite-Scroll Web Page Using Playwright Python
Load an infinite-scroll page with Playwright Python, wait for its content, and export a PDF with print or screen styling.
To create a PDF of an infinite-scroll page with Playwright Python, scroll the page or its scroll container until the content you need has loaded, wait for a site-specific completion signal, then call page.pdf(). Playwright does not automatically load every item in an infinite list just because page.goto() completed. The example below uses document-height stability as a bounded fallback; for reliable results, replace that heuristic with the page’s item selector, loading indicator, or end marker.
PDF generation requires Chromium in headless mode. By default, page.pdf() uses print CSS. Call page.emulate_media(media="screen") first if you want screen styling. See the official Playwright input documentation, Page API, and navigation guide.
1. Install Playwright and Chromium
python -m pip install playwright
python -m playwright install chromium
Save the script below as infinite_pdf.py. Run it with python infinite_pdf.py https://example.com output.pdf. Use a page you are authorized to access; if it requires login, provide an authenticated browser context rather than bypassing access controls.
2. Runnable Python example
import asyncio
import sys
from playwright.async_api import async_playwright
async def save_infinite_page_as_pdf(url: str, output_path: str) -> None:
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True)
page = await browser.new_page()
try:
response = await page.goto(url, wait_until="load", timeout=60_000)
if response is not None and response.status >= 400:
raise RuntimeError(f"Navigation returned HTTP {response.status}")
previous_height = -1
stable_rounds = 0
max_rounds = 40
for _ in range(max_rounds):
height = await page.locator("body").evaluate("el => el.scrollHeight")
if height == previous_height:
stable_rounds += 1
else:
stable_rounds = 0
if stable_rounds >= 3:
break
previous_height = height
await page.mouse.wheel(0, 1000)
await page.wait_for_timeout(800)
# Uncomment to use screen CSS instead of the default print CSS.
# await page.emulate_media(media="screen")
await page.pdf(
path=output_path,
format="A4",
print_background=True,
prefer_css_page_size=True,
)
finally:
await browser.close()
if __name__ == "__main__":
if len(sys.argv) != 3:
raise SystemExit("Usage: python infinite_pdf.py URL OUTPUT.pdf")
asyncio.run(save_infinite_page_as_pdf(sys.argv[1], sys.argv[2]))
The loop scrolls in bounded increments, waits briefly for updates, and stops after three rounds without a change in body.scrollHeight. The 40-round cap prevents a runaway loop. Both the delay and limits are starting values, not guarantees: tune them to the site and the amount of content you need.
3. Make the stopping condition reliable
Height stability is only a heuristic. A page can append items inside a fixed-height scrolling element without changing the body height, or can take longer than the delay to fetch and render new entries. A stable item count, explicit end marker, or loading-state transition is usually a better completion condition.
Wait for a new item or an end marker
Adapt these selectors to the target page. Scroll, then wait for either a new item or the site’s explicit end marker. If the site exposes neither, retain a timeout and a maximum number of scrolls.
items = page.locator(".feed-item")
end_marker = page.locator(".feed-end")
previous_count = await items.count()
for _ in range(40):
await page.mouse.wheel(0, 900)
try:
await page.wait_for_function(
"({selector, previous}) => "
"document.querySelectorAll(selector).length > previous || "
"document.querySelector('.feed-end') !== null",
arg={"selector": ".feed-item", "previous": previous_count},
timeout=5_000,
)
except Exception:
# No new item or end marker within the wait; inspect before deciding
# whether to retry or treat this as the end.
pass
current_count = await items.count()
if await end_marker.count():
break
if current_count <= previous_count:
# Site-specific policy: retry, report incomplete content, or stop.
break
previous_count = current_count
await page.pdf(path="page.pdf", format="A4", print_background=True)
For production scripts, distinguish a confirmed end marker from a timeout: a timeout may mean slow loading rather than completion. Record the item count and report when a configured scroll limit is reached so an incomplete export is visible.
Scroll the actual container
Some feeds scroll inside a nested element. Find that element in the page’s DOM and advance its scrollTop, checking its own height and item count. Playwright documents scrolling an element with locator.evaluate(); it can also scroll an element such as a footer into view to trigger more content. See the input documentation.
feed = page.locator(".feed-scroll-container")
await feed.evaluate("el => el.scrollTop += 900")
await page.wait_for_timeout(800)
For a virtualized list, old rows may be removed from the DOM as new rows appear. A PDF made from the final DOM can therefore omit items that scrolled out of view. In that case, collect each item’s data or rendered markup as it appears, or use a page-specific export route. A single final page.pdf() call cannot print elements that the application has removed.
4. Choose PDF layout and page options
- Print CSS (default): suited to a readable document. Sites may hide navigation or reflow columns for print.
- Screen CSS: call
await page.emulate_media(media="screen")beforepage.pdf()when the screen layout is preferred. - Page size: use
format="A4"orformat="Letter", or specify dimensions and margins with CSS units. When the site defines its own page size,prefer_css_page_size=Truelets that size take precedence. - Backgrounds and colors: set
print_background=Trueto include background graphics. Print color handling can alter colors; CSS-webkit-print-color-adjust: exactcan request exact colors where supported. See the Playwright Page API. - Margins and orientation: use the PDF API’s margin settings and
landscape=Truewhen the output needs them. Check for site print styles that already define page dimensions.
Example with explicit margins:
await page.pdf(
path="page.pdf",
format="A4",
landscape=False,
print_background=True,
margin={"top": "12mm", "right": "12mm", "bottom": "12mm", "left": "12mm"},
)
PDF options and details are documented in the Python Page API. page.screenshot(full_page=True) captures a full-page image; it does not produce a PDF. Choose PDF for paginated, printable output and an image when visual fidelity as a single raster capture matters more. See the screenshot guide.
5. cURL, Python requests, and Node.js alternatives
These are useful when the source already provides a finite, printable page or when you want to call a screenshot service. A plain request does not execute the site’s browser-side infinite-scroll behavior; it cannot substitute for the Playwright loading loop on a client-rendered feed.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
Python requests
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));
These calls return a screenshot image, not a PDF. For ScreenshotNeo’s PDF options, formats, and other API parameters, use the ScreenshotNeo API documentation. The query parameter names used by other screenshot APIs also work, which can ease migration.
6. Or skip the browser setup
For a direct screenshot request, ScreenshotNeo returns an image or PDF from one GET request. Here is its documented image call:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for PDF parameters. ScreenshotNeo accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits cost nothing, with verdict and billing information in response headers. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. This is a straightforward option when a site is already ready to capture; pages that require scrolling through a client-side infinite feed may still need the Playwright workflow above.
Sign up for 1,000 free screenshots a month, with no card required.
7. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| PDF contains only the first items | The script printed before more items loaded, or the site uses a nested scroller. | Wait on an item-count or end-marker condition and scroll the element that actually scrolls. |
| Some rows are missing even after scrolling | The app virtualizes its list and removes off-screen rows. | Collect each row as it loads or use the site’s export endpoint; the final DOM may not contain the full history. |
| Loop stops too early | Body height stayed constant, a delay was too short, or content loads into a fixed-height container. | Use the container’s scroll metrics, item count, loading state, or explicit end marker; increase the wait only after checking the page behavior. |
| Loop never ends | Height changes continuously, ads or animations keep changing layout, or no end condition is checked. | Keep a maximum scroll count, use a specific completion signal, and report when the cap is reached. |
| PDF styling differs from the browser | page.pdf() uses print media by default. |
Use print CSS for document output or emulate screen media before PDF generation. |
| Colors or backgrounds are absent | Print rendering omits backgrounds or adjusts colors. | Set print_background=True and, if needed, use -webkit-print-color-adjust: exact in the page CSS. |
| PDF call fails in another browser | PDF generation is supported in headless Chromium. | Launch Chromium in headless mode for this workflow. |
| Navigation timeout or incomplete page | Network activity or the site’s load event does not correspond to content readiness. | Check the navigation response and then wait for the content-specific selector; goto() completing does not prove all dynamic data has loaded. |
8. Performance, reliability, and cost
Every scroll, wait, and rendered item adds time and browser work. Set a maximum number of rounds, avoid excessively small wheel increments, and wait for the specific next item instead of using long fixed sleeps when the site exposes a reliable signal. Large feeds can produce very large PDFs and consume substantial memory; limit the requested range when possible and inspect the resulting file before distributing it.
Reliability depends on the site’s loading rules, network responses, authentication, lazy images, and DOM strategy. Use bounded retries and make incomplete output visible in logs rather than treating every timeout as end-of-list. Always close the browser in a finally block, as in the runnable example. Playwright and Chromium have no per-shot API charge in this local workflow, but you pay for the compute and storage used to run and retain the output.
FAQ
Does wait_until="load" wait for every infinite-scroll item?
No. It waits for navigation’s load event, not for future items that the application fetches as you scroll. Wait for the page’s own content signal.
Can I use page.screenshot(full_page=True) to make the PDF?
No. It returns a full-page image. Use page.pdf() for a paginated PDF.
Will page.pdf() include items that were loaded earlier but removed from the DOM?
No. It prints the current rendered document. Capture virtualized items as they appear or use a page-specific export method.
Can I produce a PDF with selectable text?
Browser-generated PDFs generally retain rendered text where the page exposes it as text, while image-only content remains image content. Check the target page’s output if text selection is a requirement.


