How to Take a Full-Page Screenshot of Dynamic Articles on a Hindi News Website
Use Playwright to load lazy content before capturing a Hindi news article as one full-page image. This guide covers setup, checks, and alternatives.
Short answer: use Playwright’s full-page screenshot option, but first scroll through the article so lazy-loaded text and images have a chance to load. In Node.js, the capture call is await page.screenshot({ path: 'article.png', fullPage: true }); in Python it is page.screenshot(path='article.png', full_page=True). A full-page capture covers the page’s scrollable height, but it does not guarantee that every item loaded only on scroll has finished loading. Inspect the saved image, especially its bottom and image areas.
This workflow is useful when a Hindi news article is long or fills content as the reader scrolls. It captures a rendered page rather than merely saving the article’s HTML. The output is still a visual record: it does not make the page text searchable or preserve the page’s interactive behavior.
1. What full-page capture does—and what it cannot do
Playwright describes a full-page screenshot as capturing the full scrollable page as if it were displayed on a very tall screen. The fullPage option requests that capture; the screenshot API also supports image-format, quality, and clipping options. See the Playwright screenshot guide and Page screenshot API.
That is different from loading dynamic content. A site may defer images or article blocks until they approach the viewport. A screenshot of the full document does not itself prove those items were requested and rendered. Scroll through the page before capture, wait for the content that matters, then inspect the result.
| Need | Approach | Trade-off |
|---|---|---|
| One-off manual image | Chrome DevTools full-size screenshot command | UI wording can vary by Chrome version; less repeatable for many URLs. |
| Repeatable capture | Playwright plus a scroll-and-wait routine | Requires Node.js or Python and a browser installation. |
| Article in a nested scrolling panel | Scroll that panel explicitly, then capture the relevant element or page | Document-level full-page capture may not expand the panel’s own scroll area. |
2. Playwright in Node.js: runnable full-page capture
Install Playwright and its Chromium browser in a project directory:
npm init -y
npm install playwright
npx playwright install chromium
Save this as capture.mjs. Pass the article URL as the first command-line argument. The script visits the page, scrolls down in viewport-sized increments to trigger common lazy-loading behavior, waits briefly at each step, returns to the top, waits for image elements to finish or fail, and saves a PNG.
import { chromium } from 'playwright';
const url = process.argv[2];
if (!url) {
console.error('Usage: node capture.mjs https://example.com/article');
process.exit(2);
}
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage({ viewport: { width: 1365, height: 900 }, deviceScaleFactor: 1 });
page.setDefaultNavigationTimeout(60_000);
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 60_000 });
// Scroll in steps so viewport-triggered content has an opportunity to load.
await page.evaluate(async () => {
const pause = ms => new Promise(resolve => setTimeout(resolve, ms));
let previousHeight = 0;
let stablePasses = 0;
for (let pass = 0; pass < 40 && stablePasses < 3; pass++) {
const height = document.documentElement.scrollHeight;
for (let y = 0; y < height; y += Math.max(400, window.innerHeight * 0.8)) {
window.scrollTo(0, y);
await pause(350);
}
await pause(700);
const newHeight = document.documentElement.scrollHeight;
stablePasses = newHeight === previousHeight ? stablePasses + 1 : 0;
previousHeight = newHeight;
}
window.scrollTo(0, 0);
});
// Wait for currently present image elements to complete. A failed image is
// complete too, so inspect the screenshot for missing image content.
await page.evaluate(async () => {
await Promise.all(Array.from(document.images, image => {
if (image.complete) return Promise.resolve();
return new Promise(resolve => {
image.addEventListener('load', resolve, { once: true });
image.addEventListener('error', resolve, { once: true });
setTimeout(resolve, 10_000);
});
}));
});
await page.screenshot({ path: 'article.png', fullPage: true, animations: 'disabled' });
console.log('Saved article.png');
} finally {
await browser.close();
}
Run it with:
node capture.mjs 'https://example.com/path/to/article'
Replace the example URL with the article you are authorized to access. The scroll loop caps its passes so pages that continuously append content cannot run forever; adjust the cap and pause for a specific site after observing how it loads.
Why this waits for DOM content instead of network idle
News pages may keep analytics, ads, live updates, or other network connections active. Waiting for every connection to become idle can stall even when the article itself is ready. This script waits for the initial DOM, scrolls, pauses, and checks images. If the article has a known selector such as article or a site-specific content container, add a targeted wait such as await page.locator('article').waitFor({ state: 'visible' }) after navigation. Choose a selector that actually exists on the target page.
3. Python alternative
Install Playwright’s Python package and browser:
python -m pip install playwright
python -m playwright install chromium
Save as capture.py:
import asyncio
import sys
from playwright.async_api import async_playwright
async def main():
if len(sys.argv) < 2:
raise SystemExit('Usage: python capture.py https://example.com/article')
url = sys.argv[1]
async with async_playwright() as p:
browser = await p.chromium.launch()
try:
page = await browser.new_page(viewport={"width": 1365, "height": 900}, device_scale_factor=1)
page.set_default_navigation_timeout(60_000)
await page.goto(url, wait_until='domcontentloaded', timeout=60_000)
await page.evaluate("""async () => {
const pause = ms => new Promise(resolve => setTimeout(resolve, ms));
let previousHeight = 0;
let stablePasses = 0;
for (let pass = 0; pass < 40 && stablePasses < 3; pass++) {
const height = document.documentElement.scrollHeight;
for (let y = 0; y < height; y += Math.max(400, window.innerHeight * 0.8)) {
window.scrollTo(0, y);
await pause(350);
}
await pause(700);
const newHeight = document.documentElement.scrollHeight;
stablePasses = newHeight === previousHeight ? stablePasses + 1 : 0;
previousHeight = newHeight;
}
window.scrollTo(0, 0);
}""")
await page.evaluate("""async () => {
await Promise.all(Array.from(document.images, image => {
if (image.complete) return Promise.resolve();
return new Promise(resolve => {
image.addEventListener('load', resolve, { once: true });
image.addEventListener('error', resolve, { once: true });
setTimeout(resolve, 10_000);
});
}));
}""")
await page.screenshot(path='article.png', full_page=True, animations='disabled')
print('Saved article.png')
finally:
await browser.close()
asyncio.run(main())
Run python capture.py 'https://example.com/path/to/article'. In Python, the option is spelled full_page; in JavaScript it is fullPage.
4. Manual capture with Chrome DevTools
- Open the article in desktop Chrome and let its initial content settle.
- Scroll down through the article, pausing where text or images appear on demand.
- Open DevTools and use its command menu to find the full-size screenshot action. The exact label and menu location can vary with the installed Chrome version, so search the available commands rather than relying on a fixed label.
- Save and inspect the resulting image, including the final section and any floating or fixed elements.
This is convenient for a single article when you can visually check each step. Use Playwright when you need a repeatable script or want to capture multiple pages. The research for this guide did not verify a specific Hindi news site or current Chrome menu wording.
5. Options and adjustments
| Option or adjustment | Use it for | What to consider |
|---|---|---|
fullPage: true / full_page=True |
Capture the full scrollable document | It does not trigger every site’s lazy loader by itself. |
path |
Save directly to a file | Use a unique, predictable filename if processing a batch. |
type: 'jpeg' or 'png' |
Choose output format | PNG is lossless; JPEG uses lossy compression. Playwright infers format from the file extension unless the type is specified. |
quality |
Set JPEG compression quality | Applies to JPEG, not PNG; lower quality can reduce file size while adding visible artifacts. |
scale: 'css' or 'device' |
Choose CSS-pixel or device-pixel output | Device scale can produce larger images. Check the screenshot API docs for supported values and defaults. |
animations: 'disabled' |
Reduce animation-related inconsistency | Disabling animations can change a page’s animated state during capture. |
clip |
Capture a defined rectangle instead of a full page | Not appropriate when the requirement is the whole article. |
For example, to write a JPEG instead of PNG, use await page.screenshot({ path: 'article.jpg', type: 'jpeg', quality: 85, fullPage: true }). Confirm option names and behavior in the current Playwright API reference if you need a less common setting.
6. Edge cases to check
Lazy images and content appended while scrolling
Scroll the page before capturing. The sample script repeats its pass when document height changes, but stable height alone is not proof that every article component is present. Check expected section headings, image count, and the bottom of the output against the page.
Nested scrolling containers
Some sites put the article inside an element with its own scrollbar. The document’s height may remain short while that container holds the text. Identify the actual scrollable article container and scroll it explicitly; then capture the article element with a locator screenshot if that produces the desired frame. The cited sources do not establish a universal selector or site-specific fix.
Repeated fixed headers and overlays
Full-page screenshot implementations may resize or scroll the page to produce a tall image. Fixed headers, sticky bars, or floating controls can therefore appear in unexpected positions or repeat in scroll-and-stitch methods. Inspect the output; if needed, hide a nonessential element with a targeted CSS rule before capture, while keeping a copy of the original view if fidelity matters.
Endless feeds and very tall documents
A page that keeps appending related stories may never reach a natural end. Set a maximum scroll distance or pass count, and define what counts as the article’s end. Very tall screenshots also consume more memory and can be awkward to preview or share; capture the article container or split the job into intentional sections when a single image is impractical.
Frames and blocked content
Content embedded in an iframe may load independently, and cross-origin restrictions can limit what page scripts can inspect. A screenshot can render visible frame content, but scrolling the top-level document may not trigger content inside an independently scrolling frame. If text or images remain missing, inspect the frame and its scroll behavior manually.
7. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Only the visible viewport was saved | The full-page option was omitted or misspelled. | Use fullPage: true in JS or full_page=True in Python. |
| Lower images are blank | They had not been requested or decoded before capture, or their requests failed. | Scroll in steps, wait for relevant images or a known article selector, then inspect failed resources and recapture. |
| Article stops partway down | Content loads only after another scroll, or the page has a nested scroller. | Repeat scroll passes until height settles; if the document height does not change, scroll the article’s own container. |
| Navigation timeout | The site or some long-running resources did not reach the chosen navigation condition in time. | Use domcontentloaded for initial navigation, then wait for a known article element and its content. Raise the timeout only when the page legitimately needs more time. |
| Screenshot operation times out or browser runs out of memory | The page is extremely tall, complex, or resource-heavy. | Reduce viewport width only if acceptable, target the article element, or capture sections. Do not silently discard needed content. |
| Fixed banner appears several times | The page uses sticky/fixed content and the capture method involves scrolling or stitching. | Inspect the image and hide or dismiss that element only if it is not part of the record you need. |
| Browser executable is missing | The Playwright package is installed without its browser binary. | Run npx playwright install chromium or python -m playwright install chromium. |
| Access denied, CAPTCHA, or blank response | The site may be restricting automation, blocking a resource, or returning a page without the article. | Check the URL and access in a normal browser, follow the site’s access rules, and avoid trying to bypass access controls. |
8. Reliability, performance, and cost
There is no universal speed or completeness winner between a manual capture and Playwright: the outcome depends on the page, network, browser, and how much content the page loads on demand. Manual capture is easy to inspect for one page. Automation is repeatable, but needs explicit waits, sensible limits, and output checks.
Scrolling and waiting adds time but reduces the chance of capturing before lazy content appears. Long waits on every step can make a batch slow; tune the pause after observing the site and wait on specific article elements when possible. A full-page image can be large, especially at high device scale. Choose an output format and scale that suit the archive or sharing use case.
Playwright itself is an open-source browser automation library; this local workflow has no ScreenshotNeo API charge. It still uses your compute, browser installation, and network, and the target site’s own access terms apply. No measured speed, accuracy, or cost comparison is established by the sources used for this guide.
Or skip the browser setup
ScreenshotNeo is a website screenshot API: one request can return a screenshot or PDF. Its API accepts the target URL and supports full-page capture with lazy images loaded; see the ScreenshotNeo API documentation. For an article URL, replace the sample URL below and save the response as an image:
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/article -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com/article"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/article' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
Before the shot, ScreenshotNeo accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each of these steps can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. An MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Every feature is available on every plan. These details do not guarantee that a particular news site will load or permit capture; inspect the response and result.
Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.
FAQ
Will the image contain selectable Hindi text?
No. A screenshot is a raster image of the rendered page. Keep the page or use a separate text extraction workflow if you need searchable or selectable text.
Can I capture a page that requires login?
Playwright can use a browser context configured for an authorized session, but handle credentials and session data securely. Do not use automation to bypass access restrictions.
Is a full-page screenshot an archival copy of the article?
It preserves a visual appearance at capture time. It does not preserve links, scripts, underlying source data, or future versions of the article.
What should I do if a single image becomes unwieldy?
Capture the article element or produce several deliberate sections, then verify that no paragraphs or images fall between them.


