Generate Hindi News Article Thumbnails from URLs with PHP and Headless Chrome
Build a PHP workflow that captures Hindi news pages in headless Chrome, waits for Devanagari fonts, and saves a consistently framed thumbnail.
To generate a thumbnail from a Hindi news article URL in PHP, launch Chrome or Chromium with the chrome-php/chrome library, set a fixed viewport, navigate to the validated URL, wait for the page and Devanagari fonts to render, then save a screenshot. Choose a deliberate crop and output format; there is no single correct thumbnail size unless you know the destination platform.
This approach captures what the browser renders. If you need a designed news card with a headline, image, and branding, build that composition in HTML first and capture the composed page rather than using an arbitrary crop of the full article.
1. Install PHP and Chrome
Install the PHP library with Composer:
composer require chrome-php/chrome
Provide a Chrome or Chromium executable in the runtime environment. The package README documents PHP 7.4–8.5 and Chrome/Chromium 65 or later, but these requirements can change; check the current project README and your installed versions before deployment.
On servers, ensure the PHP process can launch the browser and write to the output directory. Keep the output directory outside public uploads unless you intend generated files to be public.
2. Capture a URL with PHP
The following example uses a fixed viewport and waits for navigation, the article content, and browser font readiness before capturing a WebP image. Replace the example URL with a trusted article URL. For a public service, add strict URL and network controls before accepting arbitrary URLs.
<?php
require __DIR__ . '/vendor/autoload.php';
use HeadlessChromium\BrowserFactory;
$url = 'https://example.com/hindi-article';
$output = __DIR__ . '/output/article-thumbnail.webp';
// Use a real URL allowlist and private-network blocking in a public service.
$parts = parse_url($url);
if (!$parts || ($parts['scheme'] ?? '') !== 'https' || empty($parts['host'])) {
throw new InvalidArgumentException('Expected an HTTPS article URL.');
}
if (!is_dir(dirname($output)) && !mkdir(dirname($output), 0750, true)) {
throw new RuntimeException('Could not create the output directory.');
}
$browserFactory = new BrowserFactory();
$browser = $browserFactory->createBrowser([
'headless' => true,
'noSandbox' => true, // Use only with process/container isolation and network controls.
]);
try {
$page = $browser->createPage();
$page->setViewport(1200, 675);
$page->navigate($url)->waitForNavigation();
// Wait for an article landmark when the site exposes one. Replace the selector
// with a site-specific selector when necessary.
$page->waitForElement('article', 10000);
// Ask the page to wait for document fonts and image elements to settle.
// This returns a promise resolved when the checks finish.
$page->evaluate(
'document.fonts.ready.then(() => Promise.all(Array.from(document.images).map(img => {
if (img.complete) return Promise.resolve();
return new Promise(resolve => { img.addEventListener("load", resolve, { once: true }); img.addEventListener("error", resolve, { once: true }); });
})))'
)->waitForResponse();
$page->screenshot([
'format' => 'webp',
'quality' => 82,
])->saveToFile($output);
} finally {
$browser->close();
}
if (!is_file($output) || filesize($output) === 0) {
throw new RuntimeException('Screenshot output was not created.');
}
echo $output . PHP_EOL;
The library API can vary by release. Confirm method signatures against the installed package version if your editor or runtime reports an undefined method. The essential flow is to create a browser and page, set the viewport, navigate, wait for the required content, capture, and close the browser even if an operation fails.
Wait for a site-specific article element
Some pages do not use an <article> element. Replace it with a selector that identifies the main story, such as a site-specific content container. Waiting for an element helps avoid capturing a page shell before client-rendered content appears.
$page->waitForElement('.story-content', 10000);
Do not use a very broad selector such as body as a readiness check: it usually exists before the article content has loaded.
3. Make Hindi text render reliably
Use a font that includes Devanagari glyphs. Noto documentation recommends Noto Sans Devanagari for Hindi sans-serif text, with a corresponding Noto Sans style for punctuation, digits, and other characters. A CSS fallback stack can provide additional glyph coverage:
body {
font-family: "Noto Sans Devanagari", "Noto Sans", sans-serif;
}
If the article depends on a remote web font, the browser may paint the page while that font is still loading, leaving its text blank temporarily. Wait for document.fonts.ready before the screenshot, or install/serve the required font in the browser environment. Check the result at the actual thumbnail scale: legible article text at desktop size may become too small after cropping.
For consistent output, control the browser’s installed fonts and avoid relying on whichever fallback happens to be present on a server. Font fallback can also affect punctuation and numerals, so inspect representative Hindi headlines, digits, and punctuation.
4. Choose framing, dimensions, and format
| Choice | Use it when | Trade-off |
|---|---|---|
| Viewport screenshot | You want a predictable, fixed composition. | Content outside the viewport is omitted. |
| Clipped region | You know the page coordinates or have prepared a fixed thumbnail layout. | Coordinates can shift with responsive layout, fonts, and content. |
| Full-page screenshot | You need the entire article for review or archival capture. | It is usually not a useful thumbnail composition and can create a very tall image. |
The example’s 1200 × 675 viewport is only an illustrative 16:9 composition, not a universal platform requirement. Set the size to match the destination’s documented dimensions. For an article screenshot, the page’s top viewport may include navigation, ads, and consent UI; a purpose-built HTML thumbnail gives more control over headline, image, and crop.
- PNG: useful when lossless output or crisp interface details matter; files may be larger.
- JPEG: useful for photographic content where a smaller file matters; quality is adjustable.
- WebP: supported by the documented library/protocol interfaces and often useful for web delivery; verify that the destination accepts it.
The PHP library supports PNG, JPEG, and WebP, JPEG/WebP quality options, a clip rectangle, and full-page capture through a full-page clip with captureBeyondViewport. Chrome’s screenshot interface also documents PNG, JPEG, and WebP and clipping controls. Choose the format based on destination support and inspect the output rather than assuming the smallest file will look best.
5. Validate URLs before opening them
A service that opens user-supplied URLs is an internet-facing fetcher, so URL validation alone is not enough. At minimum:
- Accept only the URL schemes you intend to support, usually HTTPS.
- Resolve the hostname and reject loopback, private, link-local, and other internal network destinations.
- Re-check destinations after redirects; a public URL can redirect to an internal address.
- Restrict browser egress at the network or container layer, including DNS rebinding scenarios.
- Run Chrome in an isolated, least-privilege process/container, keep it updated, and impose CPU, memory, time, and output-size limits.
- Use an allowlist where the workflow permits known news sites only.
Do not expose a browser endpoint that can reach cloud metadata, internal services, local files, or administrative networks. Headless mode does not make arbitrary URL fetching safe. The research sources do not define a complete server-side URL-fetch policy; have the controls reviewed against current security guidance for your hosting environment.
6. cURL, Python, and Node.js options
These examples show direct headless Chrome invocation or browser automation in other common languages. The PHP library remains the path above for a PHP application. The command-line screenshot flag writes an image in the current working directory; include an explicit working directory and output path in a service.
Chrome command line
chrome --headless --window-size=1200,675 --screenshot=article.png https://example.com/hindi-article
The exact executable name and flags can depend on the installed Chrome/Chromium build and environment. Run it as an unprivileged isolated process, apply a timeout, and validate the output file.
Python with Selenium
from pathlib import Path
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
url = "https://example.com/hindi-article"
out = Path("output/article-thumbnail.png")
out.parent.mkdir(parents=True, exist_ok=True)
options = Options()
options.add_argument("--headless")
options.add_argument("--window-size=1200,675")
# Configure sandboxing and process isolation for your environment.
driver = webdriver.Chrome(options=options)
try:
driver.set_page_load_timeout(30)
driver.get(url)
WebDriverWait(driver, 10).until(
lambda d: d.execute_script("return document.fonts.status") == "loaded"
)
driver.save_screenshot(str(out))
finally:
driver.quit()
if not out.exists() or out.stat().st_size == 0:
raise RuntimeError("Screenshot was not created")
This Python sample assumes Selenium and a compatible Chrome driver are installed. Set the versions and browser binary explicitly in a managed runtime; the PHP package requirements do not establish Python package or driver requirements.
Node.js with Playwright
const { chromium } = require('playwright');
const fs = require('node:fs/promises');
(async () => {
const url = 'https://example.com/hindi-article';
const out = 'output/article-thumbnail.webp';
await fs.mkdir('output', { recursive: true });
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage({ viewport: { width: 1200, height: 675 } });
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30000 });
await page.locator('article').waitFor({ state: 'visible', timeout: 10000 });
await page.evaluate(() => document.fonts.ready);
await page.screenshot({ path: out, type: 'webp', quality: 82 });
} finally {
await browser.close();
}
})().catch(error => {
console.error(error);
process.exitCode = 1;
});
Install Playwright and its browser runtime according to its official documentation. Replace article when the page uses another content selector. For content that loads below the fold, scrolling or a site-specific readiness condition may be needed before capture.
7. Handle dynamic pages and failure cases
- Client-rendered article: wait for a meaningful story selector, not merely initial navigation.
- Lazy-loaded images: scroll the relevant image into view or use a deliberate page-specific wait before capturing. A fixed viewport does not guarantee below-the-fold assets load.
- Consent banners and popups: dismiss them where permitted and appropriate, or use a dedicated composition instead of capturing the live page chrome.
- Navigation never settles: pages with long-running requests may never reach a network-idle condition. Prefer a bounded navigation event followed by explicit checks for the article and fonts.
- Redirects: record or validate the final URL, and enforce URL/network policy at every redirect.
- Browser cleanup: close the browser in a
finallyblock so failed navigation does not leave Chrome processes consuming resources. - Output validation: check that the file exists, is nonempty, has the expected dimensions, and can be decoded before publishing it.
8. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Hindi characters show as boxes or incorrect glyphs | The runtime lacks a Devanagari font or font fallback is incomplete. | Install or load Noto Sans Devanagari, define a fallback stack, wait for font readiness, and verify punctuation and digits too. |
| Headline text is blank | A remote web font has not loaded at capture time. | Wait for document.fonts.ready or serve the font locally. |
| Screenshot is blank or only shows a loader | The capture happened before client-side rendering or the site blocked/failed the request. | Wait for a meaningful article selector, inspect the final navigation state, and report a bounded failure. |
| Image or headline is cut off | The viewport or clip does not match the intended composition. | Set the target dimensions deliberately, use a stable crop, and inspect the image at the destination size. |
| Navigation timeout | The page is slow, blocked, or keeps network connections open. | Use a bounded timeout and a suitable navigation event, then wait specifically for required content and fonts. |
| Chrome fails to start in a container | Missing browser dependencies, permissions, executable path, or unsuitable sandbox setup. | Install runtime dependencies, set the binary path if needed, run with least privilege, and isolate the process. Do not disable browser protections without compensating isolation. |
| Screenshot save fails | The directory is missing or unwritable, or the process ran out of space. | Create a controlled output directory, check permissions and free space, and validate the result after saving. |
| Unexpected internal page appears | A supplied URL or redirect reached a private destination. | Block private and internal networks at the network layer, validate redirects, and restrict allowed destinations. |
9. Performance, reliability, and cost
Each capture needs a browser process or an available browser session, page navigation, rendering, and image encoding. Reusing browser infrastructure can reduce startup overhead, but isolate pages and cap concurrent work so slow or hostile sites cannot exhaust the service. No comparative performance measurements are available here, so benchmark with representative pages in your own deployment.
Bound navigation and total job time, limit output dimensions and file size, and make cleanup unconditional. For a queue-based workflow, record the input URL, final URL, completion status, and failure reason; retry only transient failures and cap retries. Avoid retrying deterministic failures indefinitely.
Self-hosted cost includes compute, memory, storage, bandwidth, and maintenance of PHP, Chrome/Chromium, fonts, and dependencies. Keep generated files only as long as needed and avoid logging sensitive URL query strings. The appropriate dimensions and format depend on the publishing destination and its upload constraints.
10. Frequently asked questions
Should I capture the whole article page for a thumbnail?
Usually not. A thumbnail benefits from a chosen crop or a purpose-built card. Full-page capture is more appropriate for review or archival use.
Why does the same URL produce different images?
Page content, responsive layout, remote assets, font availability, consent state, and timing can change between captures. Fix the viewport and readiness checks, and control the browser’s fonts and runtime where repeatability matters.
Which thumbnail dimensions should I use?
Use the dimensions documented by the destination where you will publish. The title does not identify a platform, so no one size can be prescribed.
Or skip the browser setup
ScreenshotNeo captures a URL with one API request and returns an image or PDF. Its clean-shot flow accepts cookie/consent banners like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the page verdict and billing status in headers. An MCP server provides screenshot tools for AI agents. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000.
See the ScreenshotNeo API documentation for the available capture options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/hindi-article -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com/hindi-article"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://example.com/hindi-article'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await require('node:fs/promises').writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
For recurring captures, use the free plan to start and choose a paid allowance based on your volume: Starter is $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000; yearly billing gives two months free. Sign up for 1,000 free screenshots a month with no card.


