How to Take a Screenshot of an Infinite-Scroll News Website with Playwright
Load the stories you need by scrolling the right feed and waiting for new cards, then capture the page with Playwright’s full-page screenshot option.
To capture an infinite-scroll news page with Playwright, first scroll the feed to make the site load the stories you want, wait for evidence that each batch arrived, and then call page.screenshot({ fullPage: true }). The fullPage option captures the page’s current scrollable extent; it does not fetch every item from an infinite data source by itself.
The exact story selector, scroll container, loading signal, and stopping condition depend on the news site. The example below is runnable after you replace the URL and selectors with ones from the target page.
1. Install Playwright and identify the feed
This example uses Playwright’s Node.js library. Install it and its browser:
npm init -y
npm install playwright
npx playwright install chromium
Inspect the page in a regular browser or with Playwright’s locator tools to identify:
- A locator matching each story card, such as
articleor a site-specific card class. - Whether the document scrolls, or whether stories sit inside a nested element with its own scrollbar.
- A reliable signal that a batch finished loading: a story count increase, a loading indicator disappearing, or an end-of-feed marker.
Do not assume a universal selector or that network inactivity means the feed is finished. Sites may poll, load ads, or fetch stories only after a specific element enters view.
2. Scroll, wait for new stories, and capture
Save this as capture-news.js. Set NEWS_URL, STORY_SELECTOR, and optionally END_SELECTOR for the page you are capturing. The loop is bounded by both a story target and a maximum number of scrolls, so a feed that never grows cannot run forever.
const { chromium } = require('playwright');
const url = process.env.NEWS_URL || 'https://example.com/news';
const storySelector = process.env.STORY_SELECTOR || 'article';
const endSelector = process.env.END_SELECTOR; // Optional explicit end marker
const targetStories = Number(process.env.TARGET_STORIES || 30);
const maxScrolls = Number(process.env.MAX_SCROLLS || 40);
(async () => {
const browser = await chromium.launch({ headless: true });
const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
try {
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 45000 });
await page.locator(storySelector).first().waitFor({ state: 'visible', timeout: 15000 });
let stableRounds = 0;
for (let i = 0; i < maxScrolls; i++) {
const countBefore = await page.locator(storySelector).count();
if (countBefore >= targetStories) break;
if (endSelector && await page.locator(endSelector).count() > 0) break;
// Bring the last currently loaded story into view to trigger another batch.
await page.locator(storySelector).last().scrollIntoViewIfNeeded();
try {
await page.waitForFunction(
({ selector, previous }) => document.querySelectorAll(selector).length > previous,
{ selector: storySelector, previous: countBefore },
{ timeout: 8000 }
);
stableRounds = 0;
} catch {
// A timeout can mean the feed ended, stalled, or needs a different trigger.
stableRounds++;
}
if (endSelector && await page.locator(endSelector).count() > 0) break;
const countAfter = await page.locator(storySelector).count();
if (countAfter >= targetStories) break;
if (stableRounds >= 2) break;
}
await page.screenshot({ path: 'news.png', fullPage: true });
console.log(`Saved news.png with ${await page.locator(storySelector).count()} story cards in the DOM.`);
} finally {
await browser.close();
}
})().catch(error => {
console.error(error);
process.exitCode = 1;
});
Run it like this:
NEWS_URL='https://example.com/news' \
STORY_SELECTOR='article.story-card' \
END_SELECTOR='.end-of-feed' \
TARGET_STORIES=50 \
node capture-news.js
If the site has no end marker, omit END_SELECTOR. The script stops at the requested number of cards, the maximum scroll count, or after two consecutive scrolls produce no increase. That no-growth rule is a practical fallback, not proof that the site has no more stories. Increase the wait, alter the scroll trigger, or use a site-specific loading signal if batches take longer than eight seconds.
3. Choose the right scroll target and capture scope
Document-scrolling feeds
When the browser document is the scroller, bringing the last loaded story into view often triggers the next batch. Another option is a wheel scroll:
await page.mouse.wheel(0, 700);
await page.waitForTimeout(500); // Replace with a content or loading signal when possible
Scrolling in smaller steps can help when a site triggers loading near the viewport boundary. Prefer waiting for a new card or a known loading signal over relying only on a fixed delay.
Feeds inside a nested scroll container
If the feed has its own scrollbar, scroll that element. A locator-based approach is useful when the container is known:
const feed = page.locator('.news-feed');
await feed.evaluate(element => { element.scrollTop = element.scrollHeight; });
await page.waitForFunction(() => {
const feed = document.querySelector('.news-feed');
return feed && feed.scrollTop + feed.clientHeight >= feed.scrollHeight - 2;
});
For repeated loading, record the story count, scroll the container, and wait for the count or loading state to change before scrolling again. A full-page screenshot captures the document’s scrollable extent; it may not represent a nested feed as one continuous long page. If the desired output is the feed region itself, capture that element after loading the desired stories:
await page.locator('.news-feed').screenshot({ path: 'feed.png' });
Full page, viewport, and element screenshots
await page.screenshot({ path: 'news.png', fullPage: true })captures the full current scrollable page.await page.screenshot({ path: 'viewport.png' })captures only the current viewport.await page.locator('.news-feed').screenshot({ path: 'feed.png' })captures one selected element.
Take the screenshot only after the desired stories are present. A full-page capture can be very tall, and it does not guarantee that images or content which load only on scroll have finished rendering.
4. Make the stop condition reliable
Pick a stopping condition that matches the task:
| Condition | Use it when | Watch out for |
|---|---|---|
| Target story count | You need a known minimum number of cards. | Duplicate cards or non-story articles can inflate the count; use a specific locator. |
| End marker | The site exposes a stable “end” or “no more stories” element. | Some sites omit it or render it only after another scroll. |
| Stable no-growth rounds | No explicit end signal exists. | A slow network or failed request can look like the end. Use a bounded wait and report the observed count. |
| Known story identifier | You need a specific article in the feed. | Check for that story’s unique link or ID rather than relying on total count. |
For production capture, combine a hard iteration limit with a meaningful condition, such as a target count or end marker. If the page exposes a loading element, wait for it to appear and then disappear, or wait for the card count to increase. The right selectors and signals must be confirmed for the specific site.
5. Python and cURL alternatives
Playwright’s documented scroll and screenshot methods are available in multiple language bindings. This Python example uses the synchronous API. Install it with pip install playwright and playwright install chromium, then adapt the URL and card selector:
from playwright.sync_api import sync_playwright, TimeoutError as PlaywrightTimeoutError
url = 'https://example.com/news'
story_selector = 'article.story-card'
target_stories = 30
max_scrolls = 40
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page(viewport={"width": 1440, "height": 900})
try:
page.goto(url, wait_until='domcontentloaded', timeout=45000)
page.locator(story_selector).first.wait_for(state='visible', timeout=15000)
stable_rounds = 0
for _ in range(max_scrolls):
before = page.locator(story_selector).count()
if before >= target_stories:
break
page.locator(story_selector).last.scroll_into_view_if_needed()
try:
page.wait_for_function(
"({selector, previous}) => document.querySelectorAll(selector).length > previous",
{"selector": story_selector, "previous": before},
timeout=8000,
)
stable_rounds = 0
except PlaywrightTimeoutError:
stable_rounds += 1
if page.locator(story_selector).count() >= target_stories or stable_rounds >= 2:
break
page.screenshot(path='news.png', full_page=True)
print(f"Saved news.png with {page.locator(story_selector).count()} story cards in the DOM")
finally:
browser.close()
Plain cURL cannot trigger browser scrolling, execute page JavaScript, or wait for client-side infinite-scroll batches. It can fetch an HTTP response, but that is not equivalent to loading and capturing the rendered feed. Use Playwright when you need the site’s browser behavior. For a hosted screenshot API that captures a URL with one request, see the ScreenshotNeo option below.
6. Performance, reliability, and cost
- Performance: Every scroll and wait adds time. Stop as soon as the required stories or a trustworthy end condition is reached. A target count avoids unnecessary traversal of a long feed.
- Reliability: Use bounded navigation and content waits, a maximum iteration count, and a site-specific signal. Record the final story count so downstream steps can detect a short capture. The general Playwright guidance does not prescribe a universal selector or end condition.
- Memory and image size: Full-page screenshots of long feeds can create large images and use substantial browser memory. Capture only the required region or viewport if the whole document is unnecessary.
- Cost: Self-hosted Playwright has no per-screenshot API charge, but browser compute, storage, and engineering time have costs. A hosted API trades browser maintenance for service usage charges; compare its billing rules and supported behavior against your needs.
7. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Screenshot contains only the first few stories | fullPage captures the loaded document extent but did not trigger more feed requests. |
Scroll the correct target in a loop, wait for new cards, and confirm the final count before capture. |
| Story count never increases | Wrong card selector, wrong scroll container, a blocked request, or a feed trigger other than scroll. | Inspect the DOM and network behavior; scroll the nested feed if present and use the site’s loading signal. |
| Loop stops before the next batch arrives | The wait window is shorter than the site’s response time. | Increase the bounded timeout or wait for a specific loading indicator or card change. Keep a maximum iteration limit. |
| Loop runs until the maximum | The feed keeps loading, the selector counts unrelated elements, or no end signal exists. | Narrow the story selector and add a target count or explicit end marker. |
| Nested feed stays at the top | The document was scrolled while the feed has its own scrolling element. | Set scrollTop or use wheel input over the feed container, then wait for story growth. |
| Images are blank or incomplete | Images load lazily or need more time after the final story appears. | Scroll content into view and wait for relevant images to complete before capture. A card count alone only proves the cards exist. |
| Navigation times out | The page keeps background connections open or the site is slow. | Use domcontentloaded when appropriate, then wait explicitly for the feed. Avoid treating a global network-idle state as a universal completion signal. |
| Screenshot is too tall or memory use spikes | The feed loaded more stories than needed. | Use a lower target, an element screenshot, or a viewport screenshot. |
8. Or skip the browser setup
If you need a hosted capture of a URL rather than a scrolling archive of a feed, ScreenshotNeo returns an image or PDF from one GET request. See the ScreenshotNeo API documentation for parameters and response details.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
The service can load lazy images and supports full-page capture, but a one-call URL capture should not be treated as a guarantee that an infinite feed has been exhausted; use the browser loop above when you need a specific number of stories loaded. ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot, with each cleanup step configurable. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page info, and PDF capture. The free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 screenshots.
Sign up for 1,000 free screenshots a month, no card required.
FAQ
Does fullPage: true load every story?
No. It captures the current full scrollable extent. Trigger the feed’s loading behavior first, then capture.
How do I know when to stop scrolling?
Use a site-specific target such as a story count, end marker, or the appearance of a particular article. Bound the loop in case the feed stalls or never ends.
Can I capture only the news feed?
Yes. Use a locator’s screenshot method after loading the desired content. For a nested scroller, verify that the element capture includes the content and range you need.
Can I use cURL alone for an infinite-scroll page?
Not to reproduce browser scrolling and client-side loading. cURL does not execute the page’s JavaScript or interact with its feed.


