How to Capture Every Comment on an Infinite Scroll Discussion Page
Learn how to load and expand an entire discussion, handle nested replies, and verify what your capture includes.
Short answer: Load comments in a loop: start at the top, scroll to trigger more content, open every “Load more” and reply control, wait for the page to update, then repeat until the content stops growing and no more controls remain. Treat that as a practical stopping check, not proof that you captured every comment the platform stores. Sorting, hidden replies, access restrictions, time limits, and changing page layouts can leave items out.
If the platform offers a documented export or pagination interface, use it where practical. Native pagination gives you a clearer stopping condition than visual scrolling. For a visual record of the page, use a browser capture or archive tool that supports the page’s infinite-scroll behavior and expands replies where needed.
1. Check the discussion and its sort order
Open the exact discussion and note the selected sort order, date, and whether you are signed in. A relevance-based or filtered view may not show the same items as an “All comments” view. This behavior is platform-specific: for example, Piazza’s project documentation describes selecting “All comments” for its supported Facebook and Instagram pages. Do not assume that instruction applies to other sites.
Before capturing, check whether the site has a documented export, API, or pagination option. If it does, prefer that for a structured record. The capture process can only collect content that the account and page can access.
2. Load comments and expand replies in a loop
- Begin at the top of the discussion and note any visible comment count or reply controls.
- Scroll down far enough to trigger the next batch. Wait for the page to finish adding comments before continuing.
- Activate visible controls such as “Load more comments,” “Show more,” or “View replies.” Wording varies by site.
- Inspect each newly revealed comment for truncated text, more-replies controls, and nested reply branches. Expand those too.
- Repeat the scroll, wait, and expand steps until no new comments appear and no load or reply controls remain.
- Record whether the process ended because the page appeared exhausted or because it stalled or timed out.
Replies may be nested behind separate controls, so expanding only top-level comments is not enough. A capture tool can automate scrolling and opening discovered reply controls, but its behavior depends on the site’s current layout and supported controls. Spool’s listing describes a “Load all” action for scrolling and opening discovered “show more replies” controls; check its current supported-site behavior because layouts can change.
WebsiteArchiver documents scrolling to reveal scroll-triggered content and an option to open all Reddit comments before saving. ArchiveWeb.page describes automated page behavior and scrolling for sites that include social media and infinite-scroll pages. These are descriptions from the respective projects, not guarantees of completeness across sites.
3. Prefer documented pagination when available
For a platform with native pagination, continue until there is no next page or cursor. This provides a more explicit way to determine whether the available pages have been exhausted.
GitHub Discussions with GitHub CLI
GitHub CLI documents options to show discussion comments, retrieve a full reply thread, and continue through comment pages with an --after cursor. Use the current GitHub CLI discussion view documentation for the command syntax and authentication requirements. Continue through the returned pages until there is no next cursor. This example applies to GitHub Discussions; it is not a generic interface for other sites.
WordPress comments
Some WordPress sites use paginated comment navigation. WordPress documents previous and next comment links and paginated links in its comment pagination reference. Follow the site’s actual pagination until no next page is available. Whether pagination is enabled and what comments are visible depends on the site.
4. Save a format that suits the task
Choose an output that retains the information you need to inspect. Spool’s listing describes Markdown, JSON, CSV, plain-text, and HTML exports, with nested replies represented in its formats. Check the current tool behavior before relying on a particular export structure.
A web archive can preserve a visual page view, but it should not be treated as a structured dataset of every comment unless the capture method explicitly expands the relevant controls. Keep the original page URL and capture date with your saved result.
5. Verify the stopping point and document gaps
- Check that no load-more, show-more, or reply controls remain visible in the accessible discussion.
- Check that the comment count in your output has stopped growing after another scroll-and-wait cycle.
- Compare your count with the platform’s displayed count when available. Counts may differ because of hidden, filtered, deleted, or inaccessible items.
- Record the platform, capture date, selected sort order, whether replies were expanded, and whether the process ended at apparent exhaustion or a time limit.
- Describe the result as a capture of the content visible to your account and method, rather than a guarantee that every stored comment was collected.
The Piazza collector’s documentation is a useful model for honest reporting: it says its collector does not promise completeness and records whether it stopped at exhaustion or a time limit. Any generic browser method has similar limits. Page markup can change, and access restrictions or site-imposed limits can hide content.
6. Troubleshoot missing or incomplete comments
| Symptom | Likely cause | What to do |
|---|---|---|
| The capture stops after the first screen | The page has not triggered another load, or the capture method does not handle this site’s scroll behavior. | Scroll farther, wait for new content, and repeat. Check whether the site has a native export or pagination method. |
| Top-level comments appear, but replies are missing | Replies are behind separate “View replies” or equivalent controls. | Expand each reply control, including controls inside newly revealed replies, then check again for nested branches. |
| The result is smaller than the displayed count | Some comments may be filtered, hidden, deleted, inaccessible, or omitted by a timeout. | Check the sort and filter settings, account access, reply branches, and whether the process stopped at a time limit. Treat the displayed count as a comparison, not proof of which items should be accessible. |
| The count keeps changing or the page never seems to finish | More content may be loading, or the page may be stalled or continually updating. | Wait for the current batch to settle. Record a time limit if you use one, and report that the capture stopped at that limit rather than claiming exhaustion. |
| A browser tool no longer finds the controls | The site may have changed its layout or labels. | Use the visible page controls manually or look for documented native pagination. Recheck whether the tool currently supports the site. |
| You need to diagnose what loads in the background | The visible interaction may be unclear, or requests may fail. | Use Chrome DevTools’ Network panel: preserve the log across page loads, trigger another comment batch, and inspect the resulting requests. A HAR can help with analysis, but it is not a complete comment export. Chrome offers sanitized HAR export by default and a separate sensitive-data export option; handle exported request data carefully. |
7. Performance, reliability, and privacy
Infinite-scroll capture takes time because each batch may depend on a page update, and nested replies add more interactions. Waiting for the page to settle after each trigger reduces the chance of racing ahead before comments appear. A fixed time limit helps bound a run, but it also means the result may be partial; record that limit and outcome.
Reliability depends on the page, account access, current markup, selected sort, and the capture method’s ability to trigger loads and find reply controls. No generic visual technique can establish that inaccessible or filtered comments were collected. When completeness matters, use a documented platform-native export or pagination method where available, and report the method’s boundaries.
Network inspection is an advanced troubleshooting step. A HAR records network activity and may contain sensitive request data depending on export settings. Chrome documents sanitized export as the default and a separate option for including sensitive data. WebKit Web Inspector also documents HAR export. Store and share those files with care.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. A screenshot records the rendered page; it is not a structured export of every comment, and a single capture does not itself expand every nested reply. For a visual snapshot after you have loaded the discussion, make one GET request. See the ScreenshotNeo API documentation for options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/discussion -o discussion.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com/discussion"},
timeout=90,
)
open("discussion.webp", "wb").write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://example.com/discussion'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('discussion.webp', Buffer.from(await res.arrayBuffer()));
ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month with no card.
FAQ
Does reaching the bottom mean I have captured every comment?
No. On an infinite-scroll page, the current bottom may just trigger another batch. Also check for nested reply controls and any platform filters or access limits.
Can a screenshot contain every comment in the discussion?
Only if the relevant comments have been loaded into the rendered page and the capture includes them. A screenshot is a visual record, not a structured export or a completeness guarantee.
Is a HAR file the same as a comment export?
No. A HAR records network activity that can help diagnose loading behavior. It does not by itself provide a verified, structured list of all comments.


