How to Capture Selenium Screenshots of Websites That Use Infinite Scrolling
Load an infinite-scroll page before capturing it. This guide shows a bounded Selenium workflow for scrolling, waiting for new content, and saving reliable screenshots.
Load the content you want before taking the screenshot. Selenium’s screenshot command captures the browser’s current state; it does not automatically scroll an infinite feed or make the page fetch later items. Scroll the element that owns the feed, wait for a meaningful change, and stop when you reach your target or the page stops making progress. Then capture the viewport or, in supported Firefox Python setups, the full document.
This guide uses Python and Selenium. The examples use explicit waits and bounded loops so a page that keeps changing cannot leave the script scrolling forever. Adjust the selectors and stopping rule to match the site you are allowed to access.
1. Install Selenium and open the page
Install Selenium in your Python environment:
python -m pip install selenium
The example below uses Selenium Manager to obtain a compatible browser driver when needed. Install a supported browser such as Chrome or Firefox on the machine running the script.
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
URL = "https://example.com/feed"
options = Options()
options.add_argument("--window-size=1440,1000")
# Uncomment to run without opening a visible browser window:
# options.add_argument("--headless=new")
driver = webdriver.Chrome(options=options)
try:
driver.get(URL)
WebDriverWait(driver, 20).until(
lambda d: d.execute_script("return document.readyState") == "complete"
)
# Continue with one of the bounded scrolling workflows below.
finally:
driver.quit()
document.readyState == "complete" only indicates that the document load event has completed. It does not guarantee that a JavaScript feed, images, or asynchronous requests have finished. Wait for a page-specific element or content condition when you know one.
2. Find the element that actually scrolls
Many pages scroll the document, but feeds inside a panel may use a nested scroll container. Scrolling window will not advance a separately scrolling panel. Inspect the page’s DOM and identify the element whose scrollHeight exceeds its clientHeight and whose scroll position changes as you scroll the feed.
For a document-scrolling page, use the document element:
SCROLLER = "document.scrollingElement"
metrics = driver.execute_script("""
const el = document.scrollingElement;
return {top: el.scrollTop, height: el.scrollHeight, client: el.clientHeight};
""")
print(metrics)
For a nested feed, substitute a selector found in the page, for example div.feed:
SCROLLER = "document.querySelector('div.feed')"
metrics = driver.execute_script("""
const el = document.querySelector('div.feed');
if (!el) return null;
return {top: el.scrollTop, height: el.scrollHeight, client: el.clientHeight};
""")
print(metrics)
If the result is null, the selector does not match. If the element’s height does not exceed its visible client height, it may not be the scrolling container. Some sites use a shadow root or an iframe; those require locating the element in the correct DOM context before scrolling.
3. Scroll, wait for progress, and stop safely
A reliable loop records a progress signal, scrolls, waits for a change, and applies a finite stopping rule. Useful signals include feed item count, the text or identity of the last item, a loading indicator disappearing, or the scroller’s height increasing. The best signal depends on the site.
Document-scrolling example: wait for item count
Replace .feed-item with a selector that matches loaded entries. This version stops when the target count is reached, an end marker appears, no new entries load for several rounds, or the maximum rounds are exhausted.
from selenium.webdriver.support.ui import WebDriverWait
ITEMS = ".feed-item" # Replace with the site's item selector.
END_MARKER = ".feed-end" # Optional; replace or set to None.
TARGET_ITEMS = 100
MAX_ROUNDS = 30
NO_PROGRESS_LIMIT = 3
WAIT_SECONDS = 8
wait = WebDriverWait(driver, WAIT_SECONDS)
no_progress = 0
for round_number in range(MAX_ROUNDS):
before = driver.find_elements("css selector", ITEMS)
before_count = len(before)
if before_count >= TARGET_ITEMS:
break
if END_MARKER and driver.find_elements("css selector", END_MARKER):
break
driver.execute_script("window.scrollTo(0, document.body.scrollHeight)")
try:
wait.until(lambda d: (
len(d.find_elements("css selector", ITEMS)) > before_count
or (END_MARKER and d.find_elements("css selector", END_MARKER))
))
except Exception:
# A timeout is not proof the page is finished; count it as a no-progress round.
pass
after_count = len(driver.find_elements("css selector", ITEMS))
if after_count > before_count:
no_progress = 0
else:
no_progress += 1
print(f"round={round_number + 1}, items={after_count}")
if END_MARKER and driver.find_elements("css selector", END_MARKER):
break
if no_progress >= NO_PROGRESS_LIMIT:
break
print("Loaded items:", len(driver.find_elements("css selector", ITEMS)))
For production code, catch TimeoutException specifically rather than catching every exception. A broad catch is shown only to keep the snippet focused; unexpected WebDriver errors should be allowed to surface and be investigated.
Nested feed example
Scroll the feed element itself and wait for its item count to increase. This assumes the item selector is scoped to the feed and the feed is in the top-level document.
from selenium.webdriver.support.ui import WebDriverWait
FEED = "div.feed"
ITEMS_IN_FEED = "div.feed .feed-item"
MAX_ROUNDS = 30
NO_PROGRESS_LIMIT = 3
wait = WebDriverWait(driver, 8)
no_progress = 0
for round_number in range(MAX_ROUNDS):
feed = driver.find_element("css selector", FEED)
before_count = len(driver.find_elements("css selector", ITEMS_IN_FEED))
driver.execute_script(
"arguments[0].scrollTop = arguments[0].scrollHeight", feed
)
try:
wait.until(lambda d: len(
d.find_elements("css selector", ITEMS_IN_FEED)
) > before_count)
except Exception:
pass
after_count = len(driver.find_elements("css selector", ITEMS_IN_FEED))
if after_count > before_count:
no_progress = 0
else:
no_progress += 1
if no_progress >= NO_PROGRESS_LIMIT:
break
print("Loaded items:", len(driver.find_elements("css selector", ITEMS_IN_FEED)))
If scrolling to the current bottom fails to trigger loading, scroll in smaller increments instead. Some sites trigger requests only when the feed approaches the viewport boundary:
driver.execute_script("""
const el = arguments[0];
el.scrollTop = Math.min(el.scrollTop + Math.floor(el.clientHeight * 0.8), el.scrollHeight);
""", feed)
For document scrolling, use the same approach with window.scrollBy(0, Math.floor(window.innerHeight * 0.8)). Repeat it inside the bounded loop and use the same progress wait.
4. Capture after the desired content has loaded
Capture the current viewport
Selenium’s regular screenshot API captures the current browsing context or an element. It does not turn the entire infinite feed into one tall image. Save a viewport screenshot with:
driver.save_screenshot("feed-viewport.png")
Or capture a particular element:
feed = driver.find_element("css selector", "div.feed")
feed.screenshot("feed-panel.png")
The element screenshot contains that element’s visible rendered area; it does not guarantee that all off-screen feed entries are included.
Firefox Python full-document screenshot
The Firefox Python WebDriver reference documents full-document screenshot methods. They change the capture extent; they do not trigger lazy loading or fetch more entries. Load the content first, then use the Firefox-specific method:
from selenium import webdriver
from selenium.webdriver.firefox.options import Options
options = Options()
# options.add_argument("-headless")
driver = webdriver.Firefox(options=options)
try:
driver.get("https://example.com/feed")
# Run the appropriate scrolling and waiting loop before capture.
driver.save_full_page_screenshot("feed-full.png")
finally:
driver.quit()
Confirm that the method is available in the Selenium Python and Firefox versions in your environment. Selenium’s general screenshot API and Firefox’s full-document API are distinct capabilities. See the Firefox WebDriver Python reference.
Capture long or virtualized feeds in sections
Some feeds recycle DOM nodes: as new entries appear, earlier off-screen entries are removed. In that case, a final full-document screenshot can omit items that were previously visible. Capture each viewport as the feed advances and keep a record of which items or scroll positions each image covers. If you later stitch images, account for overlap and changing sticky headers; inspect the combined output for gaps or duplicated rows.
5. Make captures repeatable
- Set a fixed browser window size so the viewport and responsive layout are consistent.
- Use a page-specific readiness condition, such as the first feed item appearing, rather than relying only on document readiness.
- Wait for meaningful progress after each scroll. A fixed sleep can be a fallback for sites with no observable signal, but it can be both slower and less reliable than a condition-based wait.
- Use a finite maximum number of scroll rounds, a target item count, or an end marker. Record which condition stopped the loop.
- Wait for late images or layout shifts to settle before capture if they affect the result. For image-heavy pages, verify that the intended images have loaded.
- When reproducibility matters, record the browser and driver versions, viewport dimensions, URL, item selector, and stopping condition.
Selenium can execute JavaScript through WebDriver, which is useful for inspecting dimensions and scrolling. Its screenshot and JavaScript execution APIs are documented in the Selenium WebDriver documentation and the Selenium Python WebDriver reference.
6. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| The screenshot contains only the first screen. | The screenshot command captures the current state; the page was not scrolled far enough or the full-page API was not used. | Run the bounded scroll-and-wait loop before capture. Use a supported full-document method only after loading content. |
| Scrolling the window loads nothing. | The feed uses a nested scrolling element, or the page is inside a different frame context. | Inspect the DOM and scroll the element that owns the feed. Switch into the relevant iframe before locating its content. |
| The loop stops after one scroll despite more items appearing later. | The wait condition watches the wrong selector, or the site loads after a delay longer than the timeout. | Verify the item selector and loading indicator in the live DOM. Increase the explicit wait modestly or wait for a better page-specific signal. |
| The loop never terminates. | The feed continually appends content, or the script has no finite stop rule. | Set a maximum round count and target item count. Stop after repeated no-progress checks or when the site’s end marker appears. |
| Earlier entries are missing from the final image. | The site virtualizes the feed and removes off-screen nodes. | Capture successive viewport sections while scrolling, or use an appropriate permitted data/export route if one is available. |
| Images are blank or elements overlap. | Lazy images, fonts, animations, sticky elements, or layout shifts were still settling. | Wait for relevant image elements or a stable page-specific condition, then inspect the capture. Consider section captures for very long pages. |
| Firefox full-page method is unavailable. | The chosen browser, Selenium binding, or installed version does not expose that Firefox-specific method. | Check the installed Selenium Firefox API reference and versions. Use viewport or section captures when the method is unsupported. |
| A selector lookup fails. | The selector is incorrect, the feed is in an iframe or shadow DOM, or the content has not appeared yet. | Inspect the page structure, enter the correct frame context, handle shadow roots as needed, and wait for the matching element. |
7. Performance, reliability, and cost
Infinite feeds can require many network requests and browser rendering cycles. A smaller scroll increment can trigger lazy loading more reliably but usually takes more iterations; a jump to the bottom can be faster but may skip a threshold-based trigger. Tune the increment and wait condition for the target page, and set a maximum item count or time budget for large feeds.
For repeatable work, keep the browser and viewport fixed and save captures in sections when one very tall image is unwieldy or the DOM is virtualized. Screenshot quality depends on the page state and browser rendering as well as the capture call, so inspect the resulting files. The research sources provide no benchmark or universal runtime figure; measure the specific page and environment if performance matters.
Running Selenium requires a browser and driver environment that you operate. Account for browser execution time and infrastructure in your own workload. If your task is simply to obtain a clean screenshot without maintaining that browser setup, the option below uses ScreenshotNeo’s API.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. For a normal page capture, one GET request returns an image or PDF; it does not expose a Selenium scrolling loop for controlling an infinite feed, so use Selenium when you need to load and verify feed entries before capture. See the ScreenshotNeo API documentation for options and parameters.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie banners are accepted and removed before the shot, along with supported newsletter popups and chat widgets; each step can be turned off.
- Bot checks, blank pages, failed loads, timeouts, and cache hits cost nothing. Response headers report the page verdict and billing status.
- An MCP server provides
take_screenshot,get_page_info, andcapture_pdftools for AI agents and MCP clients. - The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots; every feature is on every plan.
Sign up for ScreenshotNeo’s free plan and get 1,000 screenshots a month with no card.
FAQ
Does Selenium automatically load every item in an infinite feed?
No. Your script must scroll the page’s actual feed container and wait for the site to load more content.
Can I use a full-page screenshot to load lazy content?
No. A full-document screenshot changes how much of the current page is captured; it does not itself cause the page to fetch every later item.
What should I do if the site never signals that the feed is finished?
Choose a finite target, such as an item count or maximum number of scroll rounds, and stop after repeated checks show no progress.
Will one tall screenshot include a virtualized feed’s earlier entries?
Not necessarily. If the page removes off-screen entries from the DOM, capture sections as you scroll and verify that the output covers the full range.


