ScreenshotNeo

BlogHow-to

How to Make Website Thumbnails for a Hindi-Language News Links Directory

Build consistent page thumbnails for a Hindi news directory with Playwright, reliable Devanagari font rendering, and a practical storage workflow.

By the ScreenshotNeo team4 October 20269 min read

A website thumbnail for a Hindi-language news directory is a browser-rendered image of the linked page, or a selected part of it, saved and displayed in a directory card. You can generate them with a self-managed browser such as Playwright or a hosted screenshot API. For Hindi text, wait for the page and its used web fonts to finish rendering, then inspect the result at the small size readers will see.

For most directory cards, start with a consistent viewport capture. Use an element capture when one page region is the meaningful preview, and full-page capture only when a long page is useful and legible in your design. Store each image with its source URL, capture time, and status so editors can refresh stale or failed previews.

1. Choose the preview that fits the directory card

Capture type Good fit Tradeoff
Viewport A consistent snapshot of the top of each news page. Content below the visible area is omitted.
Element A headline, article hero, or other identifiable region selected by CSS. Selectors vary across sites and may stop matching after redesigns.
Full page A page overview when the whole article layout matters. Long pages shrink heavily when fitted into a small card, and the resulting image may be larger.

There is no universal thumbnail dimension established by the sources for this workflow. Choose dimensions based on your card layout, then crop or resize downstream to that display size. Keep a consistent aspect ratio where cards share a grid. Test the actual page mix: news sites differ in layout, font, overlays, and loading behavior.

JPEG or WebP can help reduce delivery size; PNG is an option when lossless image output matters. Compare visual quality and file size on representative pages rather than assuming one format will look best for every page.

2. Prepare for Hindi and mixed-script content

Test with real headlines that include Devanagari, punctuation, numerals, and Latin text mixed with Hindi. The page must have a font capable of rendering the characters it uses. Noto Sans Devanagari UI is one documented interface typeface that supports Hindi; the source site may use a different font.

For browser captures, wait until the content and relevant fonts are ready. The browser’s document.fonts.ready promise fulfills after loading and layout operations for fonts used by the document complete. A useful readiness sequence is: navigate, wait for the page content you need, wait for fonts, and capture. Do not assume that a fixed delay alone guarantees the final visual state.

3. Generate thumbnails with Playwright

Playwright can navigate to a page and save a screenshot. This Python example captures a consistent viewport, waits for a headline selector and used fonts, and writes a WebP file. Install Playwright and its browser as described in the Playwright screenshot documentation.

from pathlib import Path
from playwright.sync_api import sync_playwright

url = "https://example.com/news-article"
output = Path("thumbnails")
output.mkdir(exist_ok=True)

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page(
        viewport={"width": 1280, "height": 800},
        device_scale_factor=1,
    )
    try:
        page.goto(url, wait_until="domcontentloaded", timeout=45000)
        # Replace this with a selector that indicates useful article content.
        page.locator("h1").first.wait_for(state="visible", timeout=15000)
        page.evaluate("document.fonts.ready")
        page.screenshot(
            path=str(output / "news-article.webp"),
            type="webp",
            quality=82,
            full_page=False,
        )
    finally:
        browser.close()

Replace the example URL and output naming with values from your directory records. A real worker should catch navigation, selector, and screenshot errors per URL so one broken destination does not stop the rest of a batch. Check the resulting files before publishing them.

Capture an element or the full page

For an element thumbnail, wait for a stable selector and capture that element:

card = page.locator("article").first
card.wait_for(state="visible", timeout=15000)
page.evaluate("document.fonts.ready")
card.screenshot(path="article-element.png")

For a full-page image, change the screenshot call to full_page=True. Full-page captures can be very tall; resize or crop them to match the card dimensions and verify that important content remains recognizable. Playwright documents viewport, element, and full-page screenshots and PNG, JPEG, and WebP output in its screenshot guide.

Readiness and browser settings

  • wait_until="domcontentloaded" waits for the document to be parsed, but not necessarily for all images or asynchronous article content.
  • Wait for a meaningful selector, such as the headline or article container, when available.
  • Use document.fonts.ready when web fonts affect the capture.
  • For pages with lazy-loaded images, scroll the relevant region into view before capturing, then wait for the image to load. Full-page screenshot behavior and lazy content can vary by page.
  • Set a consistent viewport and device scale factor for repeatable card geometry. Increase the scale factor only if the higher-resolution image is useful in your display.
  • Use a per-navigation timeout and per-URL error handling. Sites can be slow, unavailable, or difficult to automate; the sources do not promise success on every destination.

4. Store, refresh, and serve the images

Keep the original destination URL, capture timestamp, image location, and capture status with each directory entry. This makes it possible to find stale images, diagnose failures, and retry selectively. The sources do not prescribe a database schema or refresh interval, so choose one that fits how often the directory’s links and source pages change.

  1. Normalize and validate each destination URL before sending it to a browser worker.
  2. Capture the page or selected region using the directory’s chosen viewport and format.
  3. Save the image to durable storage your application can serve, and record its path and capture metadata.
  4. Show a fallback card if capture fails instead of blocking the directory page.
  5. Refresh based on your editorial needs, and avoid recapturing unchanged entries unnecessarily.

For a hosted screenshot route, OpenGraph.io documents a Screenshot API with format, quality, viewport preset, full-page, selector, and capture-delay parameters. Its documentation lists link-preview thumbnail generation as a use case. The example screenshot URLs expire after 24 hours, so download or cache the image if you need to keep it longer. The cited sources do not establish comparative cost, throughput, or accuracy, so evaluate the service against your own pages and storage requirements.

5. Make thumbnail crops useful at small sizes

Inspect images at the actual card size, not only at full resolution. A complete page capture may contain more information but become unreadable when reduced. An element capture may improve focus but can fail when the selector is absent or changes. A viewport capture is predictable but may miss an article image or headline positioned farther down the page.

Research on thumbnails for data stories discusses resizing, cropping, simplifying, and embellishing as design choices, while noting that empirical consensus on effective choices was lacking. Treat crop and design decisions as choices to evaluate in context: preserve the association with the linked article, and avoid a crop that changes the apparent meaning of a photograph or chart. See the study on thumbnails for data stories.

6. Troubleshoot common capture problems

Problem Likely cause Fix
Hindi characters appear as boxes or missing glyphs The source page or capture environment lacks a font with the needed Devanagari glyphs, or the font has not loaded yet. Check the page’s font stack and network loading, wait for document.fonts.ready, and test a Devanagari-capable font such as Noto Sans Devanagari UI where you control the template.
Thumbnail shows a blank or incomplete page The page was captured before meaningful content rendered, or scripts/content failed to load. Wait for a content selector, inspect the page’s load behavior, and capture failures separately for retry rather than treating every navigation completion as success.
Headline or article image is missing The selected region is below the fold, lazy-loaded, or not present on that site. Scroll the relevant area into view, wait for it to appear, and use a site-appropriate selector or viewport.
Element capture times out The CSS selector does not match, the page layout changed, or the content is hidden. Check the selector against the current page, set a deliberate timeout, and fall back to viewport capture when the element is unavailable.
Consent banner or popup covers the content The destination presents an overlay before capture. Use a permitted interaction or a capture option that handles overlays; otherwise record the limitation and avoid publishing a misleading preview.
Images differ in size or framing Viewport dimensions, device scale, page layout, or capture mode varies. Standardize viewport and scale settings, then resize or crop to the directory’s output dimensions.
Hosted screenshot disappears later The returned example URL may be temporary. Download or cache the image in storage you control; OpenGraph.io documents 24-hour expiry for its example screenshot URLs.

7. Performance, reliability, and cost

Browser automation gives you control over browser behavior and lets you integrate capture, retries, and storage into your own workflow. You also operate the browser workers and decide how to handle failed URLs. A hosted API avoids running that browser infrastructure, while its capture controls and persistence behavior depend on the service. The available research does not provide a measured cost, throughput, or accuracy comparison between these approaches.

For either route, keep capture work off the directory’s reader-facing request path: generate or refresh previews in a background workflow and serve already stored images. Cache results for unchanged URLs, track failures, and retry selectively. This is operational guidance for a directory workflow, not a benchmark or a source-backed guarantee.

ScreenshotNeo is a website screenshot API and MCP server for developers. Its one-call endpoint returns an image or PDF, and its clean-capture flow accepts cookie/consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. It bills only clean shots: bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status.

Or skip the browser setup

Use the API to capture a directory destination as a WebP image:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/news-article -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com/news-article"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/news-article' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));

See the ScreenshotNeo API documentation for request options. Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents use the screenshot tools. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. ScreenshotNeo supports viewport and full-page capture, element selection, WebP/JPEG/PNG, custom CSS and JavaScript, wait conditions, caching, bulk capture, and more. It is made by Yorker Media. Sign up for 1,000 free screenshots a month, with no card required.

FAQ

Should every directory card use a full-page screenshot?

No. Choose the capture mode based on what readers need to recognize in the card. Full-page images can become too small to read when reduced.

Can I use the same font on every news site?

Usually you do not control the source site’s typography. Ensure your own directory template supports Devanagari, and wait for each source page’s used fonts before capture.

How often should thumbnails be refreshed?

There is no source-established interval. Base refreshes on how often your directory links or the usefulness of their previews change, and track the last capture time.

Can I rely on one capture succeeding for every destination?

No. Pages can change, load slowly, or block automation. Keep a fallback image and a way to identify and retry failed captures.