How to Capture Hindi Web Pages Correctly with Chrome Headless Screenshots
Capture Hindi pages with Chrome Headless by setting a predictable viewport, providing Devanagari font coverage, and waiting for fonts and content to render.
To capture Hindi web pages correctly with Chrome Headless, set the intended viewport, ensure the page has a Devanagari-capable font, wait for the page content and fonts to render, then inspect the screenshot for missing glyphs, shaping problems, or clipping. For a one-off capture, Chrome’s --screenshot flag is enough; for repeatable automation and page-specific readiness checks, use Puppeteer.
Hindi is written in Devanagari. A screenshot can have the right dimensions and still be wrong if the browser captures before a web font loads or if the available font lacks the needed glyphs. There is no universal Chrome flag that guarantees correct Hindi rendering on every site, so treat font availability, page readiness, browser version, and output inspection as part of the capture.
1. Check Chrome and choose a viewport
Record the Chrome binary and version before debugging or comparing captures. Headless behavior has changed: since Chrome 112, modern Headless uses the regular Chrome implementation. Since Chrome 132.0.6793.0, the old Headless mode is available only as the separate chrome-headless-shell binary. Older instructions using --headless=old may not apply to a current Chrome installation. See the Chrome Headless documentation.
which google-chrome || which chromium || which chromium-browser
google-chrome --version
Use the executable name installed on your system. Choose dimensions for the viewport you need rather than relying on an implicit default. Chrome’s command-line reference documents --window-size and uses 412,892 as an example. --timeout sets a maximum wait before capture; it does not confirm that a particular font or application component is ready. Refer to the official command-line reference for current flag details.
2. Capture a page from the command line
For a basic viewport screenshot, run this from a shell with Chrome installed:
google-chrome \
--headless \
--no-sandbox \
--disable-gpu \
--window-size=412,892 \
--timeout=10000 \
--screenshot=hindi-page.png \
'https://example.com/hindi-page'
Replace the URL and output filename. Use --no-sandbox only when your container or environment requires it and you understand its security implications; omit it in a normal local setup where Chrome’s sandbox is available. The timeout is a ceiling, not a readiness guarantee. If the site loads fonts or text asynchronously, use Puppeteer and wait for a page-specific condition as shown below.
The command-line route is convenient for a one-off or simple repeatable capture. It offers fewer page-specific controls than Puppeteer: it cannot conveniently wait for a known text element, check font readiness in page context, or capture a selected element with application-specific logic.
3. Ensure the page has Devanagari font coverage
If you control the page, specify a Devanagari-capable font early in the CSS font stack. Noto’s web font guidance recommends a script-specific family such as Noto Sans Devanagari for Hindi, followed by a general-script font and a generic fallback. The matching general Noto Sans family can cover punctuation and digits that are not part of the primary script family.
body {
font-family: "Noto Sans Devanagari", "Noto Sans", sans-serif;
}
Make sure the font is actually available to the page: self-host it, load it through the site’s existing font service, or install it in the capture environment when the page relies on system fonts. A CSS family name does not install or fetch a font by itself. For a site you do not control, inspect its computed font and network requests, and make sure the capture environment can reach the font resources.
System defaults vary by platform and Chromium version. Chromium source lists platform-specific Devanagari mappings, including Noto Sans Devanagari on Linux and Nirmala UI on Windows, but those mappings can change and are not a substitute for explicit font configuration when you control the page. See Chromium’s Linux font fallback source and check the corresponding platform behavior for the environment you actually run.
4. Use Puppeteer for readiness checks and repeatable captures
Puppeteer gives you control over viewport, navigation, page lifecycle, font readiness, and element screenshots. Install it in a Node.js project using the official installation guide. This example launches the bundled browser, sets a viewport, waits for navigation and fonts, checks for a page-specific marker, and saves a full-page screenshot:
import puppeteer from 'puppeteer';
const url = 'https://example.com/hindi-page';
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
await page.setViewport({ width: 412, height: 892, deviceScaleFactor: 1 });
await page.goto(url, {
waitUntil: 'networkidle2',
timeout: 30000,
});
// Replace this selector with a stable element that means the Hindi content is ready.
await page.waitForSelector('main', { timeout: 10000 });
await page.evaluate(() => document.fonts.ready);
await page.screenshot({
path: 'hindi-page.png',
fullPage: true,
});
} finally {
await browser.close();
}
The script uses networkidle2, a condition shown in Puppeteer’s screenshot guidance. It can be useful, but it does not prove that all application-specific work is complete: pages may keep connections open, render content after network activity settles, or load content only after interaction. Wait for the selector, text, or application state that signifies the content you need. document.fonts.ready waits for the document’s font loading set to resolve; it cannot fix a font that failed to load or lacks Devanagari glyphs.
Puppeteer screenshots can also target an element. For example, after the same navigation and readiness steps, replace the page screenshot call with:
const content = await page.$('main');
if (!content) throw new Error('Expected main element was not found');
await content.screenshot({ path: 'hindi-main.png' });
Element capture is useful when the page has unrelated navigation or large sidebars. Full-page capture includes content beyond the initial viewport, but very long pages may be resource-intensive and can expose lazy-loading behavior. If important images or text appear only while scrolling, scroll the page in a controlled way and wait for those resources before capturing.
For reproducibility, keep the Chrome or Puppeteer version, viewport dimensions, device scale factor, URL, and wait conditions with the capture job. A change in browser generation, operating-system fonts, or site assets can change the result.
5. Verify the screenshot
Open the output image and check the actual Hindi content. Look for:
- Missing characters or tofu boxes, which often indicate absent glyph coverage or a font load failure.
- Incorrect conjuncts, matras, or shaping, which can point to a font or rendering issue that needs inspection in the source page as well as the screenshot.
- Unexpected fallback type, such as Hindi appearing in a visibly different font from the intended design.
- Clipped lines or wrapped text caused by the viewport, page layout, or font metrics.
- Blank regions where the site uses a web font that had not loaded when capture began.
This inspection is a practical check, not a guarantee that every Hindi page will render identically across browsers and platforms. The result depends on the site’s CSS and font delivery, Chrome version, operating system, and dynamic rendering.
6. Troubleshooting common problems
| Symptom | Likely cause | What to try |
|---|---|---|
| Hindi characters appear as boxes or are missing | The selected or fallback font does not contain the required Devanagari glyphs, or the web font failed to load. | Check font network requests and computed styles. Provide a Devanagari-capable font such as Noto Sans Devanagari, then wait for font readiness and capture again. |
| Text is blank in the screenshot | The page was captured before its web font or content finished loading. | Wait for a stable content selector and document.fonts.ready. Use a page-specific readiness condition; increasing a generic timeout alone may not solve it. |
| Hindi looks different on a server than on a laptop | The environments have different operating-system fonts, browser versions, or font fallback mappings. | Use the same pinned browser and environment where possible. Explicitly load the intended web font and record the browser version and platform. |
| Some text is cut off | The viewport is too narrow or short, or full-page and lazy-loaded content were not handled as expected. | Set an explicit viewport. For long or lazy pages, trigger the needed content to load, wait for it, and choose full-page or element capture deliberately. |
| Puppeteer times out at navigation | The page keeps network connections active or does not reach the selected network idle condition. | Try a less strict navigation condition such as domcontentloaded, then separately wait for the content and font conditions you need. |
| Command-line capture happens too early | --timeout elapsed before the site finished rendering; the flag does not target a specific readiness condition. |
Use Puppeteer for selector and font checks, or increase the command timeout as a best effort for simple pages. |
| Old Headless flag or instructions fail | Headless mode changed across Chrome versions; old Headless is now a separate shell binary in newer Chrome releases. | Check the installed Chrome version and use current Chrome Headless documentation; use the standalone shell only when that is specifically the environment you intend to run. |
7. Performance, reliability, and cost
A command-line screenshot has little setup overhead for a single URL. Puppeteer adds browser startup and scripting work but gives you reusable control over readiness and capture scope. For repeated jobs, reuse a browser process where appropriate, bound navigation and selector waits, and close pages or browsers when jobs finish so a stuck page does not accumulate resources.
Reliability comes from making the rendering inputs repeatable: pin or record Chrome/Puppeteer versions, use an explicit viewport and device scale factor, ensure fonts are reachable, define a meaningful readiness condition, and retain the output for visual checks when rendering correctness matters. Network idleness is a useful signal, not proof that every page’s text, font, or lazy content is ready.
For DIY capture, direct software cost depends on your infrastructure and the page’s resource needs; the cited browser documentation does not provide a universal per-screenshot cost or performance benchmark. Do not infer a capture speed from a timeout setting.
Or skip the browser setup
If you want a screenshot API instead of managing Chrome, fonts, and capture code, ScreenshotNeo accepts a URL in one GET request and returns an image or PDF. Its clean-shot flow accepts cookie and consent banners like a visitor, then removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers say the page verdict and whether the request was billed. An MCP server provides take_screenshot, get_page_info, and capture_pdf for AI agents and other MCP clients. Free includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.com/hindi-page \
-o hindi-page.webp
See the ScreenshotNeo API documentation for request options and sign up for 1,000 free screenshots a month with no card.
FAQ
Does Chrome Headless automatically install a Hindi font?
No. Rendering depends on fonts supplied by the page and fonts available in the capture environment. When you control the page, declare a Devanagari-capable font and make sure it can load.
Is networkidle2 enough to guarantee a correct Hindi screenshot?
No. It is a navigation readiness signal, not proof that the target content or font rendered correctly. Wait for page-specific content and font readiness, then inspect the image.
Should I use modern Headless or chrome-headless-shell?
Use the browser mode your environment supports and record its version. Modern Headless uses the regular Chrome implementation; the old Headless implementation is distributed as a separate shell binary in Chrome 132.0.6793.0 and later.
Can a screenshot API guarantee that every Hindi site will render correctly?
No universal guarantee follows from the available browser and font guidance. Site font delivery, rendering code, browser version, and capture environment still matter; verify output when script rendering is important.


