ScreenshotNeo

BlogHow-to

How to Capture a Lazy-Loaded Hindi Webpage with a Website Screenshot API

Load below-the-fold content, wait for Hindi fonts and images, then capture the full page. Includes Playwright, Puppeteer, API, Python, and Node.js approaches.

By the ScreenshotNeo team4 October 202612 min read

Short answer: A full-page screenshot option controls how much of the page is captured; it does not necessarily trigger content that loads only when scrolled into view. For a reliable capture, open the Hindi page in a real browser, scroll through the content you need so the page can load lazy images and sections, wait for the relevant images and used fonts, then take a full-page screenshot. With a hosted screenshot API, confirm it supports the scroll or page interaction and readiness controls your page needs.

For DIY capture, the runnable examples below use Playwright and Puppeteer. If you want to avoid managing a browser, ScreenshotNeo provides a screenshot API and MCP server; see its API documentation for available parameters.

Why lazy-loaded Hindi pages need preparation

Lazy loading defers fetching offscreen images until they are near the viewport. A browser’s load event therefore does not prove that every image lower on a long page has arrived. A full-page screenshot captures the document’s extent, but the site may not have run its viewport-triggered loading behavior for every section.

Hindi pages add a font readiness concern. Web fonts may load after initial navigation, affecting Devanagari glyph shaping, matras, line breaks, and page layout. Waiting for document.fonts.ready allows used fonts and associated layout work to finish, but it does not prove the intended font loaded successfully. Inspect the final image for glyph and layout correctness.

Choose the capture method

Method Use it when Check before relying on it
Screenshot API You want an HTTP request to a managed capture service. Verify that it supports scrolling or pre-capture interactions, full-page output, viewport settings, and the waits or selectors your page requires. Locale support and provider-specific limits vary.
Playwright You need browser control, custom scrolling, readiness checks, and a full-page screenshot in one script. Install the browser runtime and set finite navigation and readiness timeouts.
Puppeteer You want browser page control through Puppeteer’s screenshot API. Install a compatible browser runtime and implement the same scroll and readiness work yourself.

Playwright and Puppeteer both document full-page screenshot options. Neither automatically guarantees that a site’s particular lazy-loading behavior has run. The examples below use Playwright’s documented APIs; adapt selectors and scroll behavior to the target page.

Capture with Playwright

Create a project and install Playwright and its Chromium browser:

npm init -y
npm install playwright
npx playwright install chromium

Save this as capture.mjs. It scrolls through the document in increments, waits for used fonts, gives image requests a bounded time to settle, attempts image decoding, returns to the top, and saves the full page. It is a starting point: pages with infinite scroll, consent gates, or app-specific loading triggers need tailored selectors and logic.

import { chromium } from 'playwright';

const url = process.argv[2];
if (!url) throw new Error('Usage: node capture.mjs <url>');

const browser = await chromium.launch({ headless: true });
const page = await browser.newPage({
  viewport: { width: 1365, height: 900 },
  deviceScaleFactor: 1,
  locale: 'hi-IN'
});

try {
  await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 45000 });

  // Trigger viewport-based lazy loading throughout the current document.
  await page.evaluate(async () => {
    const pause = ms => new Promise(resolve => setTimeout(resolve, ms));
    let previousHeight = 0;
    for (let pass = 0; pass < 5; pass++) {
      const height = document.documentElement.scrollHeight;
      const step = Math.max(300, Math.floor(window.innerHeight * 0.8));
      for (let y = 0; y < height; y += step) {
        window.scrollTo(0, y);
        await pause(150);
      }
      await pause(300);
      const nextHeight = document.documentElement.scrollHeight;
      if (nextHeight === height && nextHeight === previousHeight) break;
      previousHeight = nextHeight;
    }
    window.scrollTo(0, 0);
  });

  // Wait for used fonts and give images a bounded period to load and decode.
  const readiness = await page.evaluate(async () => {
    if (document.fonts?.ready) await document.fonts.ready;
    const images = [...document.images];
    const results = await Promise.all(images.map(async (img, index) => {
      if (!img.complete) {
        await new Promise(resolve => {
          const timer = setTimeout(resolve, 10000);
          const done = () => { clearTimeout(timer); resolve(); };
          img.addEventListener('load', done, { once: true });
          img.addEventListener('error', done, { once: true });
        });
      }
      if (img.complete && img.naturalWidth > 0 && img.decode) {
        try { await img.decode(); } catch { /* Report below as a failed image if dimensions are empty. */ }
      }
      return { index, src: img.currentSrc || img.src, complete: img.complete,
        width: img.naturalWidth, height: img.naturalHeight };
    }));
    return { images: results, failed: results.filter(x => !x.complete || x.width === 0 || x.height === 0) };
  });

  if (readiness.failed.length) {
    console.warn('Images still missing or failed:', readiness.failed);
  }
  await page.screenshot({ path: 'page.png', fullPage: true });
  console.log('Saved page.png');
} finally {
  await browser.close();
}

Run it with node capture.mjs https://example.com/page. Replace the example URL with the page to capture. The 150 ms scroll pause and five passes are bounded starting values, not universal guarantees. If the page appends content slowly, increase the pass limit or, preferably, stop based on an expected item or a page-specific condition. If the page never changes height but lazy images remain missing, inspect the image requests and site behavior rather than extending the pause indefinitely.

Capture with Puppeteer

Install Puppeteer and save the following as capture-puppeteer.mjs. This uses the same preparation sequence and Puppeteer’s fullPage screenshot option.

npm init -y
npm install puppeteer
import puppeteer from 'puppeteer';

const url = process.argv[2];
if (!url) throw new Error('Usage: node capture-puppeteer.mjs <url>');

const browser = await puppeteer.launch({ headless: true });
const page = await browser.newPage();
await page.setViewport({ width: 1365, height: 900, deviceScaleFactor: 1 });
await page.setExtraHTTPHeaders({ 'Accept-Language': 'hi-IN,hi;q=0.9,en;q=0.8' });

try {
  await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 45000 });
  await page.evaluate(async () => {
    const pause = ms => new Promise(resolve => setTimeout(resolve, ms));
    let previousHeight = 0;
    for (let pass = 0; pass < 5; pass++) {
      const height = document.documentElement.scrollHeight;
      const step = Math.max(300, Math.floor(window.innerHeight * 0.8));
      for (let y = 0; y < height; y += step) {
        window.scrollTo(0, y);
        await pause(150);
      }
      await pause(300);
      const nextHeight = document.documentElement.scrollHeight;
      if (nextHeight === height && nextHeight === previousHeight) break;
      previousHeight = nextHeight;
    }
    window.scrollTo(0, 0);
  });
  const report = await page.evaluate(async () => {
    if (document.fonts?.ready) await document.fonts.ready;
    const images = [...document.images];
    await Promise.all(images.map(img => new Promise(resolve => {
      if (img.complete) return resolve();
      const timer = setTimeout(resolve, 10000);
      const done = () => { clearTimeout(timer); resolve(); };
      img.addEventListener('load', done, { once: true });
      img.addEventListener('error', done, { once: true });
    })));
    return images.filter(img => !img.complete || img.naturalWidth === 0)
      .map(img => img.currentSrc || img.src);
  });
  if (report.length) console.warn('Images still missing or failed:', report);
  await page.screenshot({ path: 'page.png', fullPage: true });
  console.log('Saved page.png');
} finally {
  await browser.close();
}

Run node capture-puppeteer.mjs https://example.com/page. For Puppeteer’s documented screenshot options and Playwright’s full-page behavior, see Puppeteer screenshot options and Playwright screenshots.

Hosted API, Python, and cURL options

A hosted API can simplify browser installation and deployment, but the request must expose the controls your page needs. Before choosing one, check whether it can scroll or execute interactions before capture, wait for a selector or other page condition, set viewport and device scale, and return a full-page image. Do not assume a generic full-page flag triggers lazy loading. Also check locale controls if the site changes language based on browser locale; the API facts below do not establish a locale parameter.

ScreenshotNeo’s one-call endpoint returns an image or PDF for a URL. Its available controls include full-page capture with lazy images loaded, viewport and device presets, waits, custom CSS and JavaScript, and more; consult the ScreenshotNeo documentation for exact parameter names and current configuration details. A basic call looks like this:

cURL

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com/hindi-page \
  -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={
        "access_key": "YOUR_API_KEY",
        "url": "https://example.com/hindi-page",
    },
    timeout=90,
)
r.raise_for_status()
with open("shot.webp", "wb") as f:
    f.write(r.content)

Node.js

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.com/hindi-page'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));

These basic requests demonstrate URL capture. For lazy content that requires custom scrolling or a selector-specific wait, configure the relevant documented capture options and confirm the behavior for the target site. The API also supports custom headers, cookies, user agent, authorization, caching with a chosen TTL, async jobs with signed webhooks, bulk capture of up to 100 URLs per call, and a usage API. Do not expose an API key in browser-side JavaScript; call the API from a trusted server or use signed links for public image embedding as documented.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. Its clean-shot flow accepts cookie and consent banners like a visitor, then removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Only clean shots are billed: bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses include X-Page-Verdict and X-Billed headers. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com/hindi-page \
  -o shot.webp

It includes full-page capture with lazy images loaded, waits, custom JavaScript and CSS, and viewport controls; use the docs to select the options needed for a specific page. The Free plan includes 1,000 screenshots a month with no card. Paid plans start at $5 for 3,000 shots; every feature is on every plan.

Sign up free for 1,000 screenshots a month with no card.

Readiness options and edge cases

domcontentloaded is a useful point to begin preparation, but it does not mean images or application data are ready. The load event also does not guarantee offscreen lazy resources have been requested. Network idle can help as one signal, but dynamic pages may keep connections open or keep changing after network activity pauses. Prefer a page-specific condition, with a timeout as a safety bound.

Scroll strategy

Scroll in increments smaller than the viewport so content enters the visible area. Pause briefly for each step, then check whether document height or expected content changes. For infinite-scroll pages, continue until a known item appears, a “load more” control disappears, or a defined no-growth rule is met. A fixed delay alone cannot guarantee readiness on slow, variable, or blocked requests.

Image checks

For each relevant image, check complete and nonzero naturalWidth/naturalHeight. If available, await img.decode(); it can reject for corrupt or failed image data, so catch the rejection and report the URL rather than waiting without a bound. Background images and canvas content are not included in document.images; inspect their specific elements or network requests when those are missing.

Hindi fonts and locale

Set a browser locale or language header only when the tool supports it and the page uses that signal. A locale setting does not guarantee that a specific font is installed or that the site serves Hindi. Wait for used fonts, then inspect glyphs and line wrapping. If Devanagari appears as boxes or broken shaping, check font loading and browser availability. Fix viewport dimensions and device scale between runs if you need comparable output.

Full-page limits and overlays

Very long documents can consume substantial memory and produce large image files. Sticky headers may appear repeatedly or cover content depending on browser behavior; validate the top, middle, and bottom. If the output is clipped or scaled unexpectedly, capture at a smaller device scale or split the page into sections if your workflow allows. Consent banners and other overlays can obscure content, so determine whether the goal is a faithful visitor view or a clean content capture.

Troubleshooting

Symptom Likely cause Fix
Images are blank below the fold The page was captured without triggering viewport-based loading, or the image request failed. Scroll through the needed content before capture; inspect image completion, dimensions, and failed network requests. Use a page-specific wait for the missing section.
Some list items are absent The page uses infinite scroll or a “load more” interaction. Repeat scrolling or activate the page’s load control, then stop on an expected item or defined no-growth condition.
Capture times out waiting for network idle Analytics, streaming connections, or dynamic requests keep the network active. Use a navigation milestone such as DOM content loaded, then wait for the actual selector, image state, or application condition needed.
Hindi glyphs are boxes or look malformed The intended font did not load, is unavailable in the browser environment, or the page has a font/rendering issue. Wait for document.fonts.ready, inspect font requests and computed font family, and verify the resulting image. Locale alone does not install a font.
Text wraps differently across captures Viewport, device scale, font readiness, or responsive layout changed. Fix viewport dimensions and device scale; wait for fonts before capture.
Images remain incomplete indefinitely A request is blocked, broken, or never settles. Use bounded waits, log failed URLs, and capture with a report of missing assets rather than allowing one image to hang the job.
Hosted API returns an error or an unexpected result Invalid credentials, unsupported options, target access restrictions, or a capture failure. Check the provider’s current docs and response status/headers. For ScreenshotNeo, inspect X-Page-Verdict and X-Billed to see the page outcome and billing status.

Performance, reliability, and cost

  • Keep waits purposeful. Scrolling and checking a specific selector or image state is usually more reliable than increasing a global sleep. Every extra wait adds latency; use bounded timeouts.
  • Control image scale. Higher device scale and very tall full-page output increase processing, memory, and file size. Use the dimensions needed for the reader or downstream system.
  • Retry selectively. A timeout or temporary failed load may merit a bounded retry, but retry only after recording the URL and failure state. Do not retry indefinitely or treat a bot check as a successful screenshot.
  • Cache stable pages. If content changes infrequently, a chosen cache TTL can reduce repeated work where supported. Do not use a cached result when freshness is required.
  • Compare the actual billing model. Hosted API prices, quotas, retention, and limits differ and must be checked in the provider’s current documentation. ScreenshotNeo’s Free plan is 1,000 shots/month with no card; Starter is $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan. ScreenshotNeo bills only clean shots; failed loads, blank pages, bot checks/CAPTCHAs, timeouts, and cache hits are not billed.

Validation checklist

  1. Confirm the target URL, viewport, device scale, and desired capture extent.
  2. Confirm scrolling triggered the sections and images you need.
  3. Check important images for completion and usable dimensions; record failures.
  4. Wait for used fonts and inspect Devanagari glyphs, matras, punctuation, and wrapping.
  5. Inspect the top, middle, and bottom for missing content, repeated sticky elements, overlays, or clipping.
  6. For an API, record the response status and relevant verdict or billing headers.

FAQ

Does full-page capture automatically load every lazy image?

No. It specifies the capture extent. The page’s own viewport-triggered loading behavior may need to be activated before capture.

Should I use PNG, JPEG, or WebP?

Choose based on the downstream use: PNG is useful when preserving crisp text without lossy compression matters; JPEG or WebP can reduce file size depending on content and settings. Confirm which formats the chosen capture tool supports.

Can I capture a page behind a login?

Only if your capture environment can access it with an authorized session, cookies, or headers. Keep credentials on the server side and follow the target site’s access rules.

Will setting the locale make the page render in Hindi?

Not necessarily. The site must respond to that locale, and the required fonts must load in the browser. Inspect the captured result.

Sources