Capture a Hindi Website That Loads on Scroll Using Puppeteer
Scroll a Hindi page in bounded steps, wait for its content to appear, then capture the full page with Puppeteer.
To capture content that appears only after scrolling, make Puppeteer scroll the page in bounded steps, wait for a condition tied to the page’s actual content, and then take a full-page screenshot. fullPage: true expands the screenshot area; it does not itself trigger scroll-based loading. Puppeteer documents Page.screenshot(), page evaluation, and predicate waits as separate capabilities, so treat scrolling, readiness, and capture as separate steps. Puppeteer’s screenshot guide shows the screenshot API, while its Page.evaluate() and Page.waitForFunction() references cover page-side actions and waits.
1. Install Puppeteer and choose the page condition
Use a selector that identifies content which is added as the page scrolls, such as .article-card. If the page updates an existing container instead of adding nodes, wait for a known text, item count, or explicit end marker instead. The selectors below are examples: inspect the target page and replace them with selectors and conditions that match its markup.
mkdir scroll-capture
cd scroll-capture
npm init -y
npm install puppeteer
Save the following script as capture.cjs. Set TARGET_URL to the Hindi website you are allowed to capture, and set ITEM_SELECTOR to its repeating content selector. The script scrolls down by less than a viewport at a time, gives the page a short bounded pause to react, checks for page growth, and stops at a stable bottom or after a maximum number of steps. Its height check is a fallback signal; a site-specific condition is more reliable.
2. Scroll, wait for content, and capture
const puppeteer = require('puppeteer');
const TARGET_URL = 'https://example.com/hindi-news';
const ITEM_SELECTOR = '.article-card'; // Replace with the site's repeating item selector.
const OUTPUT = 'capture.png';
const MAX_STEPS = 40;
const STEP_PAUSE_MS = 500;
const SELECTOR_TIMEOUT_MS = 3000;
(async () => {
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
await page.setViewport({ width: 1365, height: 900, deviceScaleFactor: 1 });
await page.goto(TARGET_URL, {
waitUntil: 'networkidle2',
timeout: 45000,
});
// Confirm that the initial content exists before starting the scroll loop.
await page.waitForSelector(ITEM_SELECTOR, { timeout: 15000 });
let unchangedBottomRounds = 0;
let previousHeight = await page.evaluate(() =>
document.documentElement.scrollHeight
);
for (let step = 0; step < MAX_STEPS; step++) {
await page.evaluate(() => {
window.scrollBy(0, Math.max(300, Math.floor(window.innerHeight * 0.8)));
});
// A bounded pause lets scroll handlers and lazy loaders begin their work.
await new Promise(resolve => setTimeout(resolve, STEP_PAUSE_MS));
// This observes height growth only. Replace or supplement it with a
// selector/count/text condition when the page has a known loading signal.
await page.waitForFunction(
oldHeight => document.documentElement.scrollHeight > oldHeight,
{ timeout: SELECTOR_TIMEOUT_MS },
previousHeight
).catch(() => {});
const state = await page.evaluate(() => ({
height: document.documentElement.scrollHeight,
atBottom: window.innerHeight + window.scrollY >=
document.documentElement.scrollHeight - 2,
}));
if (state.atBottom && state.height === previousHeight) {
unchangedBottomRounds++;
if (unchangedBottomRounds >= 2) break;
} else {
unchangedBottomRounds = 0;
}
previousHeight = state.height;
}
// Optional final verification: choose a condition meaningful to this site.
const itemCount = await page.$$eval(ITEM_SELECTOR, items => items.length);
if (itemCount === 0) {
throw new Error(`No elements matched ${ITEM_SELECTOR}; update the selector.`);
}
console.log(`Capturing ${itemCount} matching items from ${page.url()}`);
await page.screenshot({ path: OUTPUT, fullPage: true, type: 'png' });
console.log(`Saved ${OUTPUT}`);
} finally {
await browser.close();
}
})().catch(error => {
console.error(error);
process.exitCode = 1;
});
Run it with node capture.cjs. This pattern is an adaptable implementation, not a universal infinite-scroll algorithm: some pages load content without changing document height, use an inner scrolling panel, or require a click before they fetch more items. The loop is deliberately bounded so a page that keeps extending cannot run forever.
Use a content-specific wait when possible
A short delay can allow asynchronous work to start, but elapsed time alone does not prove that the required content is ready. If you know the next item count, wait for that count. If the page exposes a loading indicator, wait for it to disappear. If it appends a known last-item marker, wait for that marker. For example, after a scroll that should add at least one item:
const oldCount = await page.$$eval('.article-card', items => items.length);
await page.evaluate(() => window.scrollBy(0, window.innerHeight * 0.8));
await page.waitForFunction(
(selector, count) => document.querySelectorAll(selector).length > count,
{ timeout: 8000 },
'.article-card',
oldCount
);
If the page can legitimately return no more items at the end, catch that timeout as an end-of-feed signal only when other evidence confirms you reached the end. Do not silently treat every timeout as success in a workflow where a complete capture is required.
3. Tune the loading and capture behavior
| Need | Adjustment | Trade-off |
|---|---|---|
| Wait for initial navigation | waitUntil: 'networkidle2' is a useful initial navigation choice; domcontentloaded can be quicker when you then wait explicitly for the application. |
Network idle concerns navigation settling. It does not show that later scroll-triggered content loaded. Some sites keep network requests open. |
| Wait for an element | page.waitForSelector('.article-card', { timeout: 15000 }) |
Good when the desired node appears; choose a selector specific enough to avoid matching a placeholder. |
| Wait for a state or count | page.waitForFunction(predicate, options, ...args) |
Supports page-specific checks such as count increase, visible end marker, or loading state removal. Define a timeout. |
| Capture everything in the document | page.screenshot({ path: 'capture.png', fullPage: true }) |
Captures the current document extent, not future content that was never loaded. |
| Capture one element | Find the element and call its screenshot() method. |
Puppeteer’s guide notes that an element screenshot scrolls a hidden element into view. It will not collect every feed item as one document. |
| Capture a specific viewport | Set width, height, and device scale factor with page.setViewport(); omit fullPage for viewport-only output. |
Different viewport sizes can change responsive layout and which elements are visible. |
For lazy-loaded images as well as text, scrolling past each image’s threshold is often necessary. Before the final screenshot, wait for the images you care about to finish loading; for example, check that visible images have complete and a nonzero naturalWidth. Some sites replace image URLs only when an element approaches the viewport, so include image readiness in the per-step condition if image completeness matters.
4. Hindi text, fonts, and page variants
Puppeteer captures the browser’s rendered page, so Hindi content is captured as rendered text and pixels; there is no separate Hindi screenshot mode. If Devanagari glyphs appear as empty boxes or fall back to an unexpected font, confirm the browser environment has a font with Devanagari coverage and wait for fonts before taking the screenshot:
await page.evaluate(async () => {
if (document.fonts?.ready) await document.fonts.ready;
});
Set the viewport to the layout you need before navigation when responsive behavior matters. If the page selects language or region through cookies, headers, or a language switch, configure those before loading and verify the resulting page content; do not infer the selected language solely from the URL. A fixed viewport, user agent, and any required session state improve repeatability, but the page can still change because of its own content and timing.
5. Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| Only the first viewport appears | The code captured before scrolling, or omitted fullPage. |
Run the scroll loop first, verify later items exist, then use fullPage: true. |
| The lower page is blank or incomplete | Full-page capture did not trigger the site’s scroll loader, or the wait condition observed height but not the content. | Scroll in bounded steps and wait for the new item, text, count, or end marker that matters. |
| The script stops too early | Height remains constant while content is inserted into a fixed-height container, or the chosen selector matches only initial items. | Check item count or content state instead of document height; inspect whether the page uses an inner scroll container. |
| The loop never finishes | The page appends content indefinitely, ads keep changing the height, or bottom detection is never stable. | Keep a maximum step count and overall deadline; stop on a known last item, explicit end state, or repeated stable bottom. |
waitForSelector times out |
The selector is wrong, appears only after a user action, or the page failed to load. | Inspect the rendered DOM, update the selector, handle any required navigation or click, and log the final URL and page errors. |
networkidle2 times out or waits too long |
The site maintains polling or long-lived requests. | Use domcontentloaded for navigation and then wait for a specific page condition with an explicit timeout. |
| Images are missing | Images are lazy loaded, still decoding, or blocked by the page/session. | Scroll them into their loading threshold and wait for the relevant image load state before capture. |
| Hindi glyphs look wrong | Required Devanagari fonts are unavailable or have not loaded yet. | Provide a suitable system font in the browser environment and await document.fonts.ready. |
| Navigation fails or displays an error page | DNS, TLS, access controls, redirects, or a site-side failure prevented the expected document from loading. | Record the final URL and response status, verify access in a normal browser, and avoid treating an error page as a successful capture. |
6. Reliability, performance, and cost
Scrolling and checking in steps takes longer than a single screenshot because the browser must run page scripts and fetch content at each stage. Use the smallest step that reliably crosses the page’s load threshold, and prefer a specific selector or count over repeated long sleeps. Keep both a maximum step count and an overall job timeout. Close the browser in a finally block, as in the example, so failures do not leave the browser process running.
For repeatable captures, fix the viewport and relevant session state, wait for fonts and required images, log the final URL and item count, and fail the job if the completion condition was not met. A maximum scroll limit protects against infinite feeds, but it can produce a partial screenshot; expose that limit in your job result or logs so downstream code can distinguish partial output from a complete capture. Browser execution cost depends on your own runtime and hosting setup; this method has no ScreenshotNeo API charge, but it requires you to provision and operate the browser environment.
7. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. Its one-call API captures a URL without requiring you to install and operate Puppeteer. The API accepts screenshot and PDF options; see the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/hindi-news -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com/hindi-news"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://example.com/hindi-news',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
ScreenshotNeo accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Every feature is on every plan. For pages requiring actual scroll-triggered app interaction or a custom stopping rule, use the Puppeteer workflow above and verify that the needed content loaded.
Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.
8. FAQ
Does fullPage: true scroll the page for me?
No. It captures the full document extent available at screenshot time. Trigger scroll-dependent loading before capture.
Should I use a mouse wheel instead of window.scrollBy()?
For pages that listen to normal scroll events, a page-side scroll is usually enough. If the page responds only to pointer or wheel input, use Puppeteer’s input APIs and check the same content-specific condition afterward.
How can I tell the screenshot is complete?
Use a site-specific completion signal such as a known last item, expected item count, or explicit end marker. Document height alone is only a useful signal on pages that grow as content loads.
Can I screenshot just one newly loaded card?
Yes. Select the card after it appears and call its element screenshot method. Puppeteer documents element screenshots separately from full-page screenshots.


