ScreenshotNeo

BlogAI agents

How to Capture AI-Agent Screenshots of a Hindi News Website Without Broken Devanagari Text

Wait for the Hindi page and its used fonts to settle, check Devanagari font loading, then capture in a stable browser environment.

By the ScreenshotNeo team4 October 20268 min read

To capture a Hindi news page without broken Devanagari text, first wait until the intended article content is visible, then wait for the fonts used by the page to finish loading and inspect the expected Devanagari font face. Capture only after those checks. document.fonts.ready waits for used fonts and their layout operations; it does not prove that every declared font loaded or that the intended font contains every needed glyph. Keep the browser, operating system, installed fonts, and viewport consistent when repeatability matters.

This Playwright-oriented workflow adapts to an AI agent or browser-automation operator. The target website and agent platform are unspecified, so the example does not assume any particular page structure or claim that a site was tested.

1. Set up a repeatable capture environment

Use a pinned Playwright package and a compatible browser installation. For containerized capture, align the Playwright Docker image version with the package version and pin the image where possible. Record these details with each baseline or production run:

  • Playwright version and browser engine/version
  • Operating system or container image, including its version
  • Headless or headed mode
  • Installed fonts, especially the Devanagari fonts available to the browser
  • Viewport dimensions, device scale factor, and whether the capture is viewport-only or full page

Playwright cautions that screenshots can vary with the operating system, browser version, settings, hardware, and headless mode. Use the same environment for visual comparisons. See Playwright visual comparisons and Playwright Docker guidance.

2. Navigate to the state you intend to capture

A navigation event alone does not guarantee that an application-rendered article, consent state, menu, or other dynamic content is ready. Wait for a page-specific signal that represents the state you want. Prefer a stable article selector or another meaningful condition over an arbitrary delay. If the screenshot should include a consent dialog or menu, do not dismiss it; if the capture should show the page after that state is handled, perform the intended interaction before the font check.

3. Wait for and inspect Devanagari fonts

Once the intended content is present, evaluate document.fonts.ready. MDN describes its scope precisely: “The promise fulfills when loading and layout operations of all used fonts are done.” Unused declared faces can still be unloaded, so also inspect the font set and the status of the face you expect. A fulfilled promise by itself does not establish that the requested Devanagari face succeeded, is installed, or has all required glyphs.

Here is a complete Node.js Playwright example. Install Playwright in your project and install its Chromium browser using the commands in the Playwright getting started guide. Set TARGET_URL to the article URL and adjust CONTENT_SELECTOR to match the target site.

const { chromium } = require('playwright');

const url = process.env.TARGET_URL;
const contentSelector = process.env.CONTENT_SELECTOR || 'article';
const expectedFamily = process.env.DEVANAGARI_FONT || 'Noto Sans Devanagari';

if (!url) throw new Error('Set TARGET_URL to the Hindi news article URL');

(async () => {
  const browser = await chromium.launch({ headless: true });
  const page = await browser.newPage({
    viewport: { width: 1365, height: 900 },
    deviceScaleFactor: 1
  });

  try {
    await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 45000 });
    await page.locator(contentSelector).waitFor({ state: 'visible', timeout: 30000 });

    // Wait for the page's used fonts after its intended content is present.
    await page.evaluate(() => document.fonts.ready);

    const fontReport = await page.evaluate((family) => {
      const faces = Array.from(document.fonts).map((face) => ({
        family: face.family,
        status: face.status,
        weight: face.weight,
        style: face.style
      }));
      const matching = faces.filter((face) =>
        face.family.replaceAll('"', '').toLowerCase().includes(family.toLowerCase())
      );
      return {
        documentFontStatus: document.fonts.status,
        expectedFamily: family,
        matchingFaces: matching,
        faces
      };
    }, expectedFamily);

    console.log(JSON.stringify(fontReport, null, 2));
    if (fontReport.matchingFaces.length === 0 ||
        fontReport.matchingFaces.some((face) => face.status === 'error')) {
      throw new Error(`Expected font face not confirmed: ${expectedFamily}`);
    }

    await page.screenshot({ path: 'hindi-news.png', fullPage: true });
  } finally {
    await browser.close();
  }
})().catch((error) => {
  console.error(error);
  process.exitCode = 1;
});

The inspection is a useful warning, not a universal font-quality test. CSS family names may be quoted or include variations; a site can use a different valid Devanagari family, or use separate faces for weights and styles. Set DEVANAGARI_FONT to the family the page actually requests. If you need to verify the rendered appearance rather than just face state, inspect the saved image and the browser’s computed styles and network/font loading behavior.

4. Choose capture dimensions and validate the image

Playwright’s page.screenshot() captures the current viewport by default. Set fullPage: true to capture the whole scrollable page. Viewport captures are useful for a headline or above-the-fold check; full-page captures help review an entire article, though very long pages produce larger images and can expose layout differences as the page is expanded for capture. See the Playwright Page API.

Review the headline, body, captions, navigation, and text mixing Latin and Devanagari. Check that glyphs appear, vowel marks and conjuncts are legible, line breaks are plausible, and text is not clipped or unexpectedly rendered in a fallback face. For exact text, compare DOM text or an accessibility/text snapshot alongside the image. A screenshot is visual evidence; it is not a reliable way to establish every word.

5. If you control the site’s CSS

Put a script-appropriate family early in the font stack so the browser has a suitable face for Devanagari before broad fallbacks. For example, adapt this stack to the fonts your site actually serves and supports:

body {
  font-family: "Noto Sans Devanagari", "Noto Sans", sans-serif;
}

/* For a serif design, use an appropriate Devanagari serif face first. */
.article--serif {
  font-family: "Noto Serif Devanagari", "Noto Serif", serif;
}

Do not assume these font files are installed on every machine. If the site serves webfonts, check that the font request succeeds and that the selected face covers the needed characters. Noto’s guidance recommends putting the most important script’s family first and using a script-specific serif family for serif designs. See Noto’s font-use guidance.

6. Diagnose why Devanagari looks broken

Symptom Likely cause What to do
Capture has tofu boxes, missing glyphs, or an unrelated typeface The expected webfont failed, was blocked, is unavailable in the environment, or lacks the glyph. Inspect failed font requests and FontFace.status; confirm the font file and character coverage. Add or serve an appropriate Devanagari face if you control the site.
Text is readable in a manual browser session but broken in automation The automated browser uses a different OS, container, engine, headless mode, or installed font set. Record and pin the capture environment. Install the needed font in the image or rely on a successfully loaded site webfont; compare using the same browser image.
Text is missing or shifts in the screenshot although navigation completed The application had not rendered the article or the webfont/layout work had not finished. Wait for the intended content selector, then await document.fonts.ready and inspect the target face before capturing.
document.fonts.ready resolves but the intended face is absent The face may be declared but unused, not requested for the visible text, or not present in the document’s font set. Check computed styles on a Hindi text element and inspect font requests. Do not treat the promise as proof that all declared faces loaded.
Font readiness wait hangs or times out A specific browser/version/page combination may be stuck, or font loading may be unsettled. Log browser and Playwright versions, inspect face statuses and network activity, and reproduce with the same page and image. A Playwright issue report describes one Linux WebKit version-specific timeout; it is an anecdote, not evidence of a general failure. Avoid replacing the wait with an unbounded sleep.
Different line breaks or glyph shapes across machines Font availability, browser defaults, version, rendering settings, or viewport differ. Pin those inputs and compare captures only within the same environment. Validate the actual target image.

Fallback behavior is platform-dependent. A Chromium source change dated March 4, 2026 documents Devanagari defaults for Linux, Windows, and macOS, but that source change does not prove the mappings or font files exist in every deployed browser build or container. See the Chromium source change.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request returns an image or PDF; the API supports PNG, JPEG, and WebP. For this workflow, it can remove cookie banners, newsletter popups, and chat widgets before the shot. Its response identifies page verdict and billing status in headers. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf.

cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/hindi-article -o shot.webp

Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com/hindi-article"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.com/hindi-article'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await (await import('node:fs/promises')).writeFile('shot.webp', bytes);

Replace the example URL with the page you are authorized to capture. See the ScreenshotNeo API documentation for authentication and capture options. The service offers 1,000 screenshots a month free with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, with no card.

Performance, reliability, and cost

For a Playwright capture, large pages, client-side rendering, slow font requests, and full-page output can add time and memory use. Wait for the particular content and used fonts your screenshot needs, and use a bounded navigation and selector timeout. Record timeouts and browser details so a slow site can be distinguished from a font issue. A fixed sleep can waste time on fast pages and still be too short on slow ones.

For reliable visual comparisons, control the browser engine/version, operating system or container, font installation, headless setting, viewport, and device scale factor. Keep a text or accessibility snapshot when exact article content matters. No universal Devanagari screenshot-quality check or performance benchmark is established by the cited documentation.

In a self-hosted workflow, cost depends on the compute and maintenance for browsers, containers, and any fonts you provision; the dossier provides no benchmark for estimating that cost. ScreenshotNeo pricing is free for 1,000 shots/month, then Starter is $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan. For billing-sensitive automation, read the response’s X-Page-Verdict and X-Billed headers.

Frequently asked questions

Does a successful screenshot prove the article text is correct?

No. Pair the image with DOM text or an accessibility/text snapshot when exact wording matters.

Should I always use Noto Sans Devanagari?

No. Use the site’s intended Devanagari face when available. The key is a suitable script-specific font and a stable environment, not one mandatory family.

Will full-page capture fix missing glyphs?

No. Full-page changes the captured area; it does not repair font loading, font coverage, or platform fallback.

Can a screenshot API guarantee the same rendering as my pinned Playwright browser?

Do not assume so. When exact visual baselines matter, validate the actual output and keep the rendering environment consistent.