ScreenshotNeo

BlogHow-to

Convert a Hindi News Webpage to PDF with Puppeteer

Use Puppeteer to print a Hindi news page to PDF, wait for fonts and content, and troubleshoot Devanagari rendering and pagination.

By the ScreenshotNeo team4 October 20269 min read

Use Puppeteer to open the news page in Chromium and call page.pdf(). The PDF method uses print CSS by default. Wait for the article content and fonts to be ready, then inspect the generated file for Devanagari glyphs, images, and page breaks. The browser, page markup, and available fonts determine the result; this script cannot guarantee that every publisher permits automation or renders Hindi identically in every environment.

Install Puppeteer and choose a browser

The examples below use the puppeteer package, which downloads a compatible Chrome for Testing browser during installation. Use puppeteer-core if you will connect to a separately managed or remote browser; it does not download one, so browser installation and version compatibility become your responsibility. Puppeteer guarantees compatibility with its bundled browser, not an arbitrary executable.

mkdir hindi-page-pdf
cd hindi-page-pdf
npm init -y
npm install puppeteer

If your package manager blocks install scripts and Chrome is missing, install Puppeteer’s browser explicitly with npx puppeteer browsers install chrome. The downloaded browser is substantial: Puppeteer’s installation guide estimates about 170 MB on macOS, 282 MB on Linux, and 280 MB on Windows; these estimates can change.

Complete Node.js example

Save as save-hindi-pdf.cjs. Pass the article URL and optional output path on the command line. The script waits for navigation, then for a page-specific article selector when supplied, waits for fonts, and writes an A4 PDF. It uses bounded timeouts so a page that keeps making requests does not hang forever.

const puppeteer = require('puppeteer');

async function main() {
  const url = process.argv[2];
  const output = process.argv[3] || 'article.pdf';
  const articleSelector = process.env.ARTICLE_SELECTOR;
  if (!url) throw new Error('Usage: node save-hindi-pdf.cjs <url> [output.pdf]');

  const browser = await puppeteer.launch({ headless: true });
  try {
    const page = await browser.newPage();
    page.setDefaultNavigationTimeout(45000);
    const response = await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 45000 });
    if (!response) throw new Error('Navigation returned no response; the page may have redirected or closed.');
    if (!response.ok()) throw new Error(`Navigation failed: HTTP ${response.status()}`);

    if (articleSelector) {
      await page.waitForSelector(articleSelector, { visible: true, timeout: 20000 });
    } else {
      // Useful for many pages, but not proof that all lazy content is present.
      await page.waitForNetworkIdle({ idleTime: 800, timeout: 15000 }).catch(() => {});
    }

    await page.evaluate(async () => {
      if (document.fonts) await document.fonts.ready;
      const images = Array.from(document.images);
      await Promise.all(images.map(img => {
        if (img.complete) return Promise.resolve();
        return new Promise(resolve => {
          img.addEventListener('load', resolve, { once: true });
          img.addEventListener('error', resolve, { once: true });
        });
      }));
    });

    await page.pdf({
      path: output,
      format: 'A4',
      printBackground: true,
      preferCSSPageSize: true,
      margin: { top: '15mm', right: '14mm', bottom: '15mm', left: '14mm' },
      // page.pdf waits for fonts by default; the explicit wait above makes readiness visible.
    });
    console.log(`Saved ${output}`);
  } finally {
    await browser.close();
  }
}

main().catch(error => {
  console.error(error);
  process.exitCode = 1;
});

Run it with:

node save-hindi-pdf.cjs 'https://example.com/hindi-news-article' article.pdf

Replace the example URL with the article you are authorized to access and save. For a site with a stable article container, provide its selector so the script waits for actual content instead of relying on network quiet:

ARTICLE_SELECTOR='article' node save-hindi-pdf.cjs 'https://example.com/hindi-news-article' article.pdf

Make Hindi text render correctly

Puppeteer waits for fonts when producing a PDF, and document.fonts.ready resolves after fonts used by the document have loaded and layout has completed. That does not ensure the page has a suitable Devanagari font or that a particular publisher’s font request succeeds. Open the PDF and check matras, conjuncts, punctuation, and line wrapping. If characters are missing or incorrectly shaped, inspect the page’s computed fonts and font-loading requests; the appropriate fix depends on the page and runtime.

A page can use a remote @font-face source or a font installed locally. Ensure the browser process can reach the remote font host, or that the required font is installed in the runtime. Do not assume a Latin-only fallback will cover Hindi glyphs. Avoid injecting a font without checking its license and whether it is appropriate for the page.

Choose print or screen styling and page layout

page.pdf() uses print media by default. This is usually suitable for paper-like output, but a site’s print stylesheet may hide navigation, change colors, or rearrange content. To render screen CSS instead, emulate screen media before creating the PDF:

await page.emulateMediaType('screen');
await page.pdf({ path: 'article-screen-style.pdf', format: 'A4', printBackground: true });

Choose A4 for a common international page size or Letter when that is the intended paper. Puppeteer also supports named paper formats such as Legal and custom width and height values. CSS page rules can take precedence when preferCSSPageSize is enabled. Check the output because the page’s own print CSS may set margins, page size, or page breaks.

Setting Use Consideration
format Choose A4, Letter, or another supported paper preset. Page width changes line wrapping and page count.
printBackground Include background colors and images. Printing without backgrounds can reduce ink, but may remove visual cues.
preferCSSPageSize Honor CSS @page dimensions. Use when the publisher defines its intended paper size.
margin Set top, right, bottom, and left margins. Large margins reduce text width and may increase page count.
pageRanges Export only selected pages, for example '1-3'. Ranges apply to the laid-out PDF, not article sections.
landscape Use a wider page orientation. Usually unnecessary for a text article.

Chromium adjusts colors for printing by default. If exact screen colors matter, the page can use -webkit-print-color-adjust: exact in print CSS. Use this selectively: background-heavy pages can make large PDFs and consume more printer ink.

Wait for dynamic and lazy-loaded content

waitForNetworkIdle() is one readiness signal, not a guarantee that an article is complete. Analytics or streaming requests can prevent idleness; client-rendered pages may become quiet before the article appears; images may load only after scrolling. Prefer waiting for a stable article selector or a page-specific condition. If content appears only after scrolling, scroll in steps before printing and then wait for images and fonts again.

await page.waitForSelector('article h1', { visible: true, timeout: 20000 });
await page.evaluate(async () => {
  for (let y = 0; y < document.body.scrollHeight; y += window.innerHeight) {
    window.scrollTo(0, y);
    await new Promise(resolve => setTimeout(resolve, 150));
  }
  window.scrollTo(0, 0);
  if (document.fonts) await document.fonts.ready;
});

Use a selector that matches the actual article and verify the final PDF. The scroll delay is a starting point, not a universal timing guarantee. Some sites append content only after interaction or serve different content to automated browsers.

Alternative implementations

cURL

cURL does not render a webpage or create a PDF by itself. It can call a service that performs the browser capture and returns a PDF. For local Puppeteer conversion, run the Node.js script above; for a one-request hosted option, see the ScreenshotNeo block below.

Python with a locally running Puppeteer script

Puppeteer is a Node.js library. Python can invoke the complete Node script as a subprocess when an existing Python workflow needs to orchestrate it:

import subprocess

subprocess.run(
    ["node", "save-hindi-pdf.cjs", "https://example.com/hindi-news-article", "article.pdf"],
    check=True,
    timeout=90,
)

This requires Node.js, the installed Puppeteer package, and its compatible browser. For Python-native browser automation, a different library would be needed; this is not Puppeteer code.

Direct Puppeteer PDF options

To export a page range or landscape PDF, pass the corresponding options to page.pdf():

await page.pdf({
  path: 'selected-pages.pdf',
  format: 'A4',
  landscape: true,
  pageRanges: '1-2',
  printBackground: true,
  margin: { top: '12mm', right: '12mm', bottom: '12mm', left: '12mm' }
});

Or skip the browser setup

ScreenshotNeo accepts a URL and can return a PDF. See the API documentation for PDF parameters and response handling.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/hindi-news-article -d format=pdf -o article.pdf

Cookie banners are accepted like a visitor and removed before capture, along with known consent platforms, newsletter popups, and chat widgets; each of those steps can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing status. An MCP server lets AI agents use screenshot, page-info, and PDF-capture tools. The free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000 shots. ScreenshotNeo is the publisher’s website screenshot API and MCP server: ScreenshotNeo. Its PDF output still depends on what the target page serves, so inspect Hindi glyphs and pagination.

Sign up for ScreenshotNeo’s free plan: 1,000 screenshots a month, no card required.

Troubleshooting

Symptom Likely cause Fix
Chrome executable not found Browser download was skipped or install scripts were blocked. Run npx puppeteer browsers install chrome, or configure the separately managed browser when using puppeteer-core.
Navigation timeout The page is slow or has persistent requests. Use domcontentloaded plus a content-specific selector, increase a bounded timeout, and avoid requiring network idle on pages that never settle.
PDF is missing part of the story Client rendering or lazy loading had not completed. Wait for an article-specific selector, scroll to trigger lazy content, then inspect the PDF.
Hindi letters appear as boxes or are malformed Required Devanagari font did not load or available fallback lacks glyphs. Check font requests and computed fonts, wait for document.fonts.ready, and ensure an appropriate font is available in the browser environment.
PDF differs from the visible page PDF uses print CSS and may alter colors or hide elements. Use emulateMediaType('screen') if screen styling is desired; check paper size, margins, backgrounds, and the site’s @media print rules.
Images are absent Images are lazy-loaded, blocked, or failed to fetch. Scroll through the page, wait for image completion, and check network access and image errors.
Output has unexpected breaks Paper width, CSS page rules, or print styles cause repagination. Try A4 or Letter deliberately, review margins and @page, and inspect the resulting page boundaries.
Access denied or CAPTCHA appears The publisher may restrict automated access. Respect the site’s access rules; do not attempt to bypass an access control. A browser script cannot guarantee access to content unavailable to it.

Performance, reliability, and cost

Launching a browser has setup and memory costs, so for repeated conversions, keep a browser process alive and create a fresh page per job, while closing pages and the browser cleanly. Bound navigation, selector, and overall job times. Reuse a compatible bundled browser for predictable behavior; with an external browser, pin and maintain the version yourself. For parallel jobs, limit concurrency according to the memory and CPU available to the host because each active page consumes resources.

The article’s images, fonts, scripts, and other remote resources affect load time and output. Blocking resources can speed capture but may remove article content or alter layout, so measure the tradeoff against the document you need. Puppeteer itself is open source; operational cost comes from the machine, browser storage/download, network traffic, and any separately hosted browser infrastructure. Validate representative pages when changing browser versions, fonts, or runtime images. No conversion timing or output quality is guaranteed for an unspecified news site.

FAQ

Does Puppeteer translate the Hindi article?

No. It prints the rendered webpage; it does not translate its content.

Will every Hindi website produce a searchable PDF?

Not necessarily. Text-based page content is generally rendered as text, but the specific site and PDF output should be checked; text embedded in images remains image content.

Can I save a page as a PDF without installing Chrome myself?

The puppeteer package downloads its compatible Chrome during installation. For managed browser infrastructure, use puppeteer-core and provide a compatible browser.

Can I create a PDF that looks exactly like the screen?

Emulating screen media helps select screen CSS, but paper dimensions and browser printing still affect pagination. Inspect the file and adjust styles or page settings as needed.

Official references