ScreenshotNeo

BlogHow-to

How to Archive an Infinite Scroll Page as a Complete PDF

Load the full range of an infinite-scroll page before saving it as a PDF, then inspect the result. Here’s how to handle missing content and when to use a web archive instead.

By the ScreenshotNeo team4 October 20267 min read

To save an infinite-scroll page as a complete PDF, load the full section you want in the browser first, then print or save the page as a PDF and inspect the saved file from beginning to end. Scrolling is necessary because some pages fetch content only as you move down. It is not a guarantee: a site may discard older items, defer images, or render differently for printing.

A PDF is a static document for reading, sharing, and filing. If you need to preserve a browser-loaded experience for replay, use a web archive format such as WARC or WACZ instead; those are not PDFs.

1. Load the content you want to preserve

  1. Open the page in a desktop browser and wait for its initial content to settle.
  2. Scroll down through the intended range. Pause when new items appear to give text, images, and other content time to load.
  3. Continue until you reach the intended final item or end point. If you need the whole feed, look for an obvious end state, such as a message that there is no more content.
  4. Scroll back upward and check whether the earlier content is still present. Some sites remove older items as newer ones load to save memory.
  5. Use the browser’s print or save-to-PDF function. Labels and behavior vary by browser and version, so follow the controls in your browser rather than assuming a specific menu path.
  6. Open the resulting PDF. Check the first page, a middle section, and the final page. Confirm the expected beginning and end are present and that important images or embedded content are readable.

A page that looks complete in the browser—or a print preview that appears complete—is not proof that every item made it into the saved file. Inspect the PDF itself before relying on it.

2. Troubleshoot missing or incomplete content

Symptom Likely cause What to try
The PDF stops before the page’s apparent end Later items had not loaded when printing began. Return to the page, scroll farther, pause for loading, and save again. Confirm the last expected item appears in the PDF.
Earlier items are missing The site may replace or remove older entries while loading newer ones. Check the page after scrolling back up. If the site does not retain the full range at once, try a capture tool that interacts with the live page or look for a direct download or print view offered by the site.
Images are blank or incomplete Images may be lazy-loaded, still loading, or handled differently in print view. Pause near image-heavy sections before continuing. Check the saved PDF at several points; if images remain absent, look for a site-provided download or another capture method.
Print preview differs from the page The site or browser may use a separate print layout. Save a PDF and inspect it rather than relying on preview alone. If important material is omitted, try the site’s own print or download option.
The page keeps loading indefinitely The site may have no clear end, or the feed may continue fetching items. Define the range you actually need, stop at its final item, and record that scope. Do not describe a partial range as the complete feed.

3. Choose PDF or an interactive web archive

Choose based on what you need to do with the saved result:

Need Suitable output
Read, share, or file a fixed document PDF, after checking that the desired range was captured.
Replay a captured page and its related resources later An interactive web archive. ArchiveWeb.page can export WARC and WACZ; its documentation says captured data remains local unless shared and can be viewed offline.

ArchiveWeb.page’s Autopilot can scroll or interact with certain complex pages. Its guide describes a single-page workflow aimed at sites such as social-media pages and infinite-scroll pages. It does not promise that every site will capture successfully, and the resulting WARC or WACZ is an archive, not a PDF. See the ArchiveWeb.page capture guide and official project page.

4. Automate a static page with a browser

For a page whose content is already available in the document, a browser automation script can open it and print it to PDF. The following is a minimal Playwright example in Python. It does not implement infinite scrolling: you must add page-specific scrolling and loading checks before calling page.pdf(). There is no universal scroll loop that can guarantee completeness because pages load and discard content in different ways.

from pathlib import Path
from playwright.sync_api import sync_playwright

URL = "https://example.com"
OUTPUT = "page.pdf"

with sync_playwright() as p:
    browser = p.chromium.launch()
    page = browser.new_page()
    page.goto(URL, wait_until="domcontentloaded", timeout=60_000)
    page.pdf(path=OUTPUT, format="A4", print_background=True)
    browser.close()

assert Path(OUTPUT).is_file()

Install Playwright and its browser before running the script using the official Playwright Python installation guide. Replace the example URL. For an infinite-scroll page, first add site-specific logic to scroll through the desired range, wait for new content, and verify that earlier items remain available. Then inspect the output PDF.

Manual loading loop: only when the page’s behavior is understood

A simple scroll-to-bottom loop can help on pages that append items and retain them in the document. It cannot detect every loading pattern or prove that all content has been captured. Set a maximum number of passes, watch for an end condition, and verify the final item and earlier content before printing.

from playwright.sync_api import sync_playwright

URL = "https://example.com/feed"
MAX_PASSES = 50

with sync_playwright() as p:
    browser = p.chromium.launch()
    page = browser.new_page()
    page.goto(URL, wait_until="domcontentloaded", timeout=60_000)

    previous_height = 0
    for _ in range(MAX_PASSES):
        page.evaluate("window.scrollTo(0, document.body.scrollHeight)")
        page.wait_for_timeout(1200)  # Allow this page's requests to settle.
        height = page.evaluate("document.body.scrollHeight")
        if height == previous_height:
            break
        previous_height = height

    # Add a page-specific check here for the expected final item and
    # confirm that earlier items have not been removed from the document.
    page.pdf(path="feed.pdf", format="A4", print_background=True)
    browser.close()

The delay is an example, not a reliable universal wait value. A site may load after a longer delay, require a user action, use a nested scroll container, or virtualize older items out of the page. Replace the generic height check with a selector or end condition specific to the page where possible.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. One GET request can capture a URL as PNG, JPEG, WebP, or PDF. The API can produce a PDF, but a screenshot service does not automatically walk an infinite feed: for a page that reveals content only as you scroll, first confirm that the desired content is available to the capture, or use the browser workflow above. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

Replace YOUR_API_KEY with your key and the example URL with the page to capture. The Node.js example uses Bun’s file-writing helper; with Node.js, save the response body using fs/promises. Check the API documentation for output format and capture options.

ScreenshotNeo accepts cookie or consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots; yearly billing gives two months free, and every feature is on every plan. Sign up for 1,000 free screenshots a month, with no card.

Performance, reliability, and cost

  • Browser workflow: time depends on how much content the page loads and how long its requests take. Scroll in measured steps and wait for content to settle rather than repeatedly triggering loads before the previous batch appears.
  • Completeness: infinite scroll has no universal end signal. Set a clear range, check the expected first and last items, and inspect the PDF at the start, middle, and end.
  • Memory and page behavior: long feeds can become heavy, and some sites remove older entries. If earlier content disappears, printing the current page may not preserve it.
  • Cost: browser printing and ArchiveWeb.page are separate workflows; consult their current official information for any applicable costs. ScreenshotNeo’s listed plans start with 1,000 free shots monthly, then $5 for 3,000; choose based on the number of captures you need. A screenshot call is not a substitute for scrolling a feed that has not loaded its full range.

FAQ

Can I save an unlimited feed as one complete PDF?

Only if you define what “complete” means. A feed with no end can continue indefinitely, and a site may not keep every item loaded at once. Specify a date, item count, or other endpoint and verify it in the saved file.

Does an interactive web archive count as a PDF?

No. WARC and WACZ preserve captured web material for replay; PDF is a static document. Use the format that matches how you need to access the result.

Will a screenshot API automatically capture everything revealed by scrolling?

Do not assume so. A screenshot captures the page state it can access. Load and verify the intended range first, or use a page-specific browser automation workflow.