How to Screenshot Infinite Scroll Pages with Scrapy and Headless Chrome
Load an infinite feed before capturing it. This guide shows Scrapy with scrapy-playwright, bounded scrolling, full-page screenshots, and common fixes.
To screenshot an infinite-scroll page with Scrapy and headless Chrome, use scrapy-playwright to render the page, scroll it until the feed reaches a known end or stops growing, wait for each batch of content to appear, and then call Playwright’s full-page screenshot method. full_page=True captures the page’s current scrollable extent; it does not load items that have not been triggered by scrolling.
First inspect how the page gets its content. If its network requests expose a data endpoint that you can reproduce, ordinary Scrapy requests are often simpler. Use a browser when you need the rendered DOM or an actual screenshot. Scrapy recommends scrapy-playwright for browser integration in a Scrapy project. Scrapy’s dynamic-content guide
1. Choose between reproducing requests and rendering a browser
| Page or requirement | Approach | Trade-off |
|---|---|---|
| The feed data comes from a repeatable API request, embedded JSON, or another request you can inspect | Reproduce the request with Scrapy and parse its response | Avoids browser rendering and lets you work with structured data. Inspect the real method, URL, body, pagination token, and required headers. |
| You need a screenshot or the content only appears in the rendered page | Render with a headless browser | Matches browser-visible state, with browser startup, memory, and page-lifecycle costs. |
| You already have a Scrapy crawler and need browser interactions per request | Use scrapy-playwright |
Keeps work in Scrapy’s request workflow while giving the callback access to a Playwright Page. |
| You need one screenshot and do not need Scrapy’s crawling components | A direct Playwright script may be sufficient | Scrapy notes that direct Playwright use inside a spider can bypass Scrapy components such as middleware and duplicate filtering. See Scrapy’s guidance. |
There is no universal scroll count, wait duration, or end-of-feed signal. Use a page-specific condition where possible; the bounded stability loop below is a fallback, not a guarantee for every site.
2. Install Scrapy, Playwright, and scrapy-playwright
The integration project documents minimum requirements of Python 3.10, Scrapy 2.7, and Playwright 1.40. Requirements can change, so check the current scrapy-playwright README against the versions in your environment.
python -m venv .venv
source .venv/bin/activate
python -m pip install Scrapy scrapy-playwright
playwright install chromium
On Windows PowerShell, activate the virtual environment with .venv\\Scripts\\Activate.ps1. Install the browser binary in the same environment that runs the spider. The integration README’s general installation command is playwright install; installing Chromium explicitly keeps the example aligned with its default browser type.
In settings.py, configure the browser download handler and asyncio reactor:
DOWNLOAD_HANDLERS = {
"https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
}
TWISTED_REACTOR = "twisted.internet.asyncioreactor.AsyncioSelectorReactor"
# Tune this to the memory available and cost of pages in your crawl.
CONCURRENT_REQUESTS = 4
The documented minimal setup registers the HTTPS handler, which is usually enough for modern sites. Check your Scrapy version and existing project settings before copying the reactor configuration; new Scrapy projects have used the asyncio-based reactor by default since Scrapy 2.7. Refer to the integration README for configuration details.
3. Build a spider that scrolls, waits, and captures
Save this as infinite_screenshot.py. The spider scrolls in bounded steps, waits for the document height or item count to change, and stops after repeated checks without growth or at a hard maximum. Replace the URL, item selector, and end-marker selector with values for the target page.
import asyncio
from pathlib import Path
import scrapy
from playwright.async_api import Page
async def scroll_until_loaded(
page: Page,
*,
item_selector: str,
end_selector: str | None = None,
max_steps: int = 40,
stable_checks_required: int = 3,
) -> dict[str, int | bool]:
"""Scroll a feed with a hard cap; tailor its signals to the target site."""
stable_checks = 0
previous_height = 0
previous_count = 0
steps = 0
for steps in range(1, max_steps + 1):
if end_selector and await page.locator(end_selector).count():
return {"steps": steps - 1, "items": await page.locator(item_selector).count(), "end_marker": True}
before_count = await page.locator(item_selector).count()
before_height = await page.evaluate("document.documentElement.scrollHeight")
await page.evaluate("window.scrollTo(0, document.documentElement.scrollHeight)")
# Wait for observable growth; the timeout permits a no-growth check,
# so this loop can reach its stopping condition when the feed is done.
try:
await page.wait_for_function(
"args => document.documentElement.scrollHeight > args.height "
"|| document.querySelectorAll(args.selector).length > args.count",
arg={"height": before_height, "selector": item_selector, "count": before_count},
timeout=5000,
)
except Exception:
pass
await page.wait_for_timeout(300)
after_height = await page.evaluate("document.documentElement.scrollHeight")
after_count = await page.locator(item_selector).count()
if after_height > previous_height or after_count > previous_count:
stable_checks = 0
else:
stable_checks += 1
previous_height = after_height
previous_count = after_count
if end_selector and await page.locator(end_selector).count():
return {"steps": steps, "items": after_count, "end_marker": True}
if stable_checks >= stable_checks_required:
break
return {
"steps": steps,
"items": await page.locator(item_selector).count(),
"end_marker": bool(end_selector and await page.locator(end_selector).count()),
}
class InfiniteScreenshotSpider(scrapy.Spider):
name = "infinite_screenshot"
start_urls = ["https://example.org/long-feed"]
async def start(self):
for url in self.start_urls:
yield scrapy.Request(
url,
callback=self.capture,
errback=self.close_failed_page,
meta={
"playwright": True,
"playwright_include_page": True,
},
)
async def capture(self, response):
page: Page = response.meta["playwright_page"]
try:
result = await scroll_until_loaded(
page,
item_selector="article.feed-item", # Change for the target site.
end_selector=".end-of-feed", # Set to None if unavailable.
max_steps=40,
stable_checks_required=3,
)
self.logger.info("Feed loading result: %s", result)
# Give in-view lazy assets a chance to finish loading.
await page.evaluate("window.scrollTo(0, 0)")
await page.wait_for_timeout(500)
await page.screenshot(path="feed.png", full_page=True)
yield {"url": response.url, "screenshot": str(Path("feed.png")), **result}
finally:
await page.close()
async def close_failed_page(self, failure):
page = failure.request.meta.get("playwright_page")
if page and not page.is_closed():
await page.close()
self.logger.error("Request failed: %s", failure.value)
Run the spider from the project directory:
scrapy runspider infinite_screenshot.py
The selectors in this example are placeholders. Choose an item selector that matches each feed entry and, if available, a specific end marker. The script’s brief post-scroll wait is only a settling interval; the main signal is observable DOM growth. The maximum step count prevents an unbounded crawl if the page continually appends content.
The loop compares both document height and item count. A site may append or replace items without changing height, so count is useful; virtualized lists may reuse a fixed number of DOM nodes, in which case neither count nor height reliably indicates progress. Prefer waiting for the site’s actual response, cursor change, item identity, or end marker for such pages.
async def start is available in current Scrapy versions. In an older project, use its supported start-request pattern, such as start_requests, and confirm compatibility with the installed Scrapy and scrapy-playwright versions. The integration documents marking a request with meta={"playwright": True}, including the Page with playwright_include_page=True, and accessing it as response.meta["playwright_page"] in an async callback. Included pages should be closed after use. scrapy-playwright README
4. Make the loading condition fit the page
Prefer an explicit end condition
If the page displays an end-of-feed marker, check it after each scroll and stop when it appears. If the page exposes a known total, compare the number of unique loaded items with that total. These signals are stronger than guessing from a fixed number of scrolls.
Wait for the page’s observable change
A fixed sleep alone is fragile: slow responses may need longer, while fast responses waste time. Better signals include a new item appearing, the item count increasing, a loading indicator disappearing, or the specific feed request completing. Playwright provides locator waits and page-level event handling; choose a condition the page actually exposes. If no signal is available, use bounded retries and treat the result as potentially incomplete.
Handle nested scrollers and “load more” controls
Some feeds scroll inside a panel instead of the window. Identify the scroll container and change its scrollTop or use locator scrolling. Other pages require clicking a “Load more” button. In both cases, wait for the item content to change after each action. Scrolling the window will not load a feed whose own container has the scroll event listener.
Account for virtualized feeds
A virtualized list may remove old cards as new ones enter the viewport. Its document height or DOM item count may stay constant even though the feed is advancing. Track stable item identifiers or the last visible item’s text or link, and stop on a known final identifier, explicit end marker, or a bounded period with no new identifiers. A full-page screenshot cannot restore content that the page has already removed from the DOM; a site-specific capture strategy may need to save viewport segments as the feed advances.
5. Capture the full page after loading
The key call is await page.screenshot(path="feed.png", full_page=True). Playwright defines the full-page option as capturing the full scrollable page rather than just the current viewport; the option’s default is false. Capture only after the feed-loading loop ends. Playwright Page API
For other output formats, Playwright’s screenshot API supports options such as type="jpeg" and a quality value for JPEG, as well as transparency for supported formats. Check the current Page API for the complete option set and constraints. Large full-page images can consume substantial memory; consider viewport-by-viewport captures when a page is exceptionally long or the browser cannot produce one image.
Loading the feed and loading its images are separate concerns. Scrolling entries into view often triggers lazy loading, but does not guarantee every image has finished loading. If complete images matter, wait for relevant image elements to finish or verify their natural dimensions before capture. Avoid assuming full_page=True itself triggers every lazy-load mechanism.
6. Troubleshoot common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| Screenshot contains only the first viewport | The scroll loop did not run, or the screenshot happened before feed loading completed. | Check callback execution and log scroll steps and item counts. Capture after the loading loop. Confirm the page’s scroll event is being triggered. |
| Only one batch loads | The site needs a longer response wait, a different scroll increment, a nested container, or a button click. | Inspect the browser network requests and page structure. Wait for the actual response or item change and interact with the correct element. |
| The loop never stops | The page keeps changing height, or the feed has no end signal. | Keep the hard maximum. Add an explicit end condition or stop after several checks with no new unique items. Consider a maximum elapsed time as well. |
| The loop stops although more content exists | A brief network delay or a transient no-growth check was mistaken for completion. | Increase the number of stable checks, wait for a target-specific signal, and inspect whether the next request is still pending. Do not treat the fallback stability condition as proof of completeness. |
| Height and item count do not change | The page may use a nested scroller or virtualized list that reuses DOM nodes. | Scroll the correct container and track unique item identifiers or the last visible item instead of global height or count. |
| Cards appear but images are blank | Images are lazy-loaded or still downloading when capture begins. | Scroll them into view, then wait for required image loads or verify image dimensions before taking the screenshot. |
| Scrapy middleware or duplicate filtering seems absent | Browser work was performed through a separately created Playwright browser instead of the Scrapy integration. | Route browser requests through scrapy-playwright when Scrapy’s scheduling and components matter. Scrapy dynamic-content guide |
| Memory usage grows during a crawl | Included Page objects were not closed, or too many heavy pages are open concurrently. | Close each included Page in a finally block, close pages on error, and lower browser concurrency to fit available memory. |
| Browser executable or handler error | Browser binaries are missing from the active environment, or the download handler/reactor is misconfigured. | Run playwright install chromium in the crawler environment; check Scrapy settings and the installed integration requirements. |
7. Performance, reliability, and cost considerations
- Request replay can avoid rendering work. When the page’s data request is reproducible, Scrapy can request and parse that data directly. Browser rendering is appropriate when browser behavior or a screenshot is required. The reviewed documentation publishes no fixed performance ratio, so benchmark against the target workload rather than relying on a universal speed claim.
- Bound every loop. Set a maximum scroll count and, for production crawls, a total time budget. Log final item count, end-marker state, and whether the loop exited through stability or the cap. A screenshot can look complete while omitting feed items.
- Control concurrency. Browser pages use more resources than plain HTTP requests. Begin with conservative concurrency, monitor memory, and increase only while the machine and target site remain stable. The example’s concurrency value is a starting configuration, not a measured recommendation.
- Expect page-specific behavior. Network delay, rate limits, bot checks, authentication, and client-side errors can affect completeness. Handle request failures and close pages on every path. Follow the target site’s access rules.
- Plan for image size. Very long full-page images may be large and resource-intensive. Use JPEG when lossy compression is acceptable, or capture viewport segments if one full image is impractical. Verify the output before treating it as an archival record.
- Browser cost is operational. Account for browser CPU and memory, crawl duration, and storage for screenshots. Direct data extraction may reduce those costs when it satisfies the task. No benchmark or universal cost figure is established by the sources.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. One GET request returns an image or PDF, so you do not have to install Chromium or write a scroll loop for a one-off screenshot. Its screenshot API captures the rendered page, while full-page capture loads lazy images. This API call is a direct alternative for capturing a page; it does not expose a Scrapy feed-extraction loop or guarantee that every infinite feed has loaded before capture.
See the ScreenshotNeo API documentation for parameters and configuration.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.org/long-feed \
-o feed.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.org/long-feed"},
timeout=90,
)
r.raise_for_status()
open("feed.webp", "wb").write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://example.org/long-feed',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('feed.webp', new Uint8Array(await res.arrayBuffer()));
ScreenshotNeo removes supported cookie and consent banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers report the page verdict and billing status. Its MCP server provides screenshot tools for AI agents. The free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 screenshots.
Sign up free for 1,000 screenshots a month, with no card required.
Frequently asked questions
Does full_page=True load the entire infinite feed?
No. It captures the page’s current scrollable extent. Trigger feed loading first, then take the full-page screenshot.
Should I use Scrapy or Playwright?
Use ordinary Scrapy requests when you can reproduce the data request and do not need browser rendering. Use scrapy-playwright when browser behavior or a rendered screenshot is required inside a Scrapy project.
Can I know that every item was captured?
Only if the page exposes a reliable completion signal, such as a total count or end marker, or you validate the loaded item identities against a known result. A bounded no-growth loop is a practical fallback, not proof.
Can I use this pattern on a site that requires authentication?
Potentially, if you configure the browser request with the required session or cookies and have permission to access the page. Authentication and anti-automation behavior are site-specific; test the resulting rendered state before saving the screenshot.


