ScreenshotNeo

BlogHow-to

How to Take a Puppeteer Screenshot of an Urdu HTML Page with Right-to-Left Text

Set Urdu text direction correctly, wait for fonts and page content, then capture a reliable Puppeteer screenshot with runnable code and troubleshooting tips.

By the ScreenshotNeo team4 October 20268 min read

To take a Puppeteer screenshot of an Urdu HTML page with right-to-left text, set lang="ur" dir="rtl" on the document, make the intended font available, wait for fonts and any application-specific content to finish loading, then call page.screenshot(). Use fullPage: true when you need the whole document rather than the visible viewport.

Direction controls text flow and bidirectional layout; it does not install or select a font. Urdu and English fragments can coexist: mark a distinct English phrase, identifier, or URL as left-to-right when its ordering needs to stay LTR. The examples below use documented Puppeteer and browser APIs. [MDN: direction] [MDN: writing mode systems]

1. Set Urdu language and direction in the HTML

For a primarily Urdu page, declare the language and direction on the html element. Include a doctype, UTF-8 metadata, and a viewport declaration. If only part of the page is Urdu, put lang="ur" dir="rtl" on that section instead.

<!doctype html>
<html lang="ur" dir="rtl">
<head>
  <meta charset="utf-8">
  <meta name="viewport" content="width=device-width, initial-scale=1">
  <title>Urdu document</title>
  <style>
    body {
      font-family: "Your Urdu Font", sans-serif;
      line-height: 1.8;
      margin: 2rem;
    }
    .english {
      direction: ltr;
      text-align: left;
    }
  </style>
</head>
<body>
  <main>
    <h1>اردو متن</h1>
    <p>یہ دائیں سے بائیں لکھی جانے والی مثال ہے۔</p>
    <p>English fragment:
      <span class="english" lang="en" dir="ltr">Puppeteer screenshot</span>
    </p>
  </main>
</body>
</html>

HTML’s dir attribute is the preferred way to declare base direction where possible. It affects inline direction and alignment as well as flow in relevant layout contexts. A CSS font stack only names candidate fonts: the selected Urdu-capable font must also be available to Chromium, either through a stylesheet-loaded web font or the runtime environment. [MDN: direction]

2. Capture an Urdu HTML string with Puppeteer

Install Puppeteer in a Node.js project with npm install puppeteer. The following ES module creates a browser, loads a complete HTML document with page.setContent(), waits for used fonts, captures the full page as a PNG, and closes the browser even if capture fails.

import puppeteer from 'puppeteer';

const html = `<!doctype html>
<html lang="ur" dir="rtl">
<head>
  <meta charset="utf-8">
  <meta name="viewport" content="width=device-width, initial-scale=1">
  <style>
    body { font-family: sans-serif; margin: 2rem; line-height: 1.8; }
    .english { direction: ltr; text-align: left; }
  </style>
</head>
<body>
  <main>
    <h1>اردو متن</h1>
    <p>یہ دائیں سے بائیں لکھی جانے والی مثال ہے۔</p>
    <p>English fragment:
      <span class="english" lang="en" dir="ltr">Puppeteer screenshot</span>
    </p>
  </main>
</body>
</html>`;

const browser = await puppeteer.launch();
try {
  const page = await browser.newPage();
  await page.setContent(html);
  await page.evaluate(() => document.fonts.ready);
  await page.screenshot({ path: 'urdu.png', fullPage: true });
} finally {
  await browser.close();
}

page.setContent() assigns markup to the page. document.fonts.ready fulfills when used fonts have finished loading and associated layout operations are complete. It does not wait for your application to fetch and render data after the initial document is set. Add a separate wait for the condition that means your own content is ready. [Puppeteer: setContent()] [MDN: Document.fonts]

3. Capture a live Urdu page by URL

For a website, use page.goto(), then wait for fonts and a page-specific readiness signal before capturing. The selector in this example is illustrative: replace [data-capture-ready] with a selector your page sets only after the intended content is rendered.

import puppeteer from 'puppeteer';

const browser = await puppeteer.launch();
try {
  const page = await browser.newPage();
  await page.setViewport({ width: 1365, height: 900, deviceScaleFactor: 1 });
  await page.goto('https://example.com/urdu-page', { waitUntil: 'domcontentloaded' });
  await page.waitForSelector('[data-capture-ready]', { timeout: 15000 });
  await page.evaluate(() => document.fonts.ready);
  await page.screenshot({ path: 'urdu-page.png', fullPage: true });
} finally {
  await browser.close();
}

Choose the navigation event based on the site. domcontentloaded means the initial document was parsed, not that fonts, images, or client-rendered data are ready. networkidle can be useful for some pages, but ongoing requests or delayed application updates mean it is not proof that the exact visual state you want has arrived. Prefer a known selector or other application-specific readiness condition, then wait for fonts. Puppeteer’s screenshot method captures the rendered state at the time it is called. [Puppeteer: screenshot()]

4. Choose viewport or full-page capture

  • Viewport screenshot: omit fullPage or set it to false. The image covers the current viewport.
  • Full-page screenshot: set fullPage: true to request a capture of the full document.
  • Specific region: pass a clip rectangle with the desired coordinates and dimensions. Ensure the target is within the rendered page and account for the viewport and device scale.
// Viewport PNG
await page.screenshot({ path: 'viewport.png' });

// Full-document PNG
await page.screenshot({ path: 'full-page.png', fullPage: true });

// A region in CSS pixel coordinates
await page.screenshot({
  path: 'region.png',
  clip: { x: 0, y: 0, width: 900, height: 600 }
});

When a path is provided, Puppeteer infers the image type from its extension. Screenshot options include full-page capture, clipping, and capture beyond the viewport; consult the current interface documentation for the exact options supported by the installed Puppeteer version. [Puppeteer: ScreenshotOptions]

5. Keep mixed Urdu and English text readable

Urdu is written horizontally right to left. In mixed-direction paragraphs, Unicode bidirectional handling determines how directional runs appear. Give an isolated English phrase, code sample, account identifier, or URL an LTR direction when its characters or punctuation need to remain ordered as expected.

<p dir="rtl" lang="ur">
  رابطہ:
  <bdi dir="ltr">dev@example.com</bdi>
</p>

<p dir="rtl" lang="ur">
  کمانڈ:
  <code dir="ltr">npm install puppeteer</code>
</p>

Do not apply unicode-bidi: bidi-override to the whole document as a general RTL repair. An override forces ordering by sequence and can defeat the normal implicit bidi rules. Use semantic direction markup on the smallest relevant fragment instead. [MDN: unicode-bidi]

6. Troubleshoot common screenshot problems

Symptom Likely cause Fix
Urdu text starts or aligns on the wrong side The document or intended container lacks RTL direction, or a later style resets it. Set dir="rtl" on the html element or Urdu container and inspect computed styles for overrides.
English punctuation, URLs, or identifiers look reordered A mixed-direction run is relying on surrounding context. Mark the fragment dir="ltr"; avoid a document-wide bidi override.
Urdu glyphs look different from the design The intended font is unavailable to Chromium, or font loading has not finished. Make the font available through CSS or the runtime, check that it loaded, and await document.fonts.ready.
The bottom of the page is missing The default capture covers only the viewport. Set fullPage: true, or use a clip for a deliberate region.
Page is captured before data or images appear Navigation completed before the application reached its final state. Wait for an application-specific selector or readiness signal, and wait for fonts before capture.
Navigation or selector wait times out The URL is unreachable, the expected state never occurs, or the timeout is too short for the page. Check URL access and selector spelling; wait on the correct state and set a realistic timeout for the environment.
Capture fails in a container The Chromium process may not launch correctly in that runtime or may lack required system dependencies. Review Puppeteer’s launch error and container setup, and use the supported browser installation and runtime configuration for that environment.

7. Improve repeatability, performance, and cost control

  • Wait for what matters: a precise readiness selector avoids both premature captures and unnecessary arbitrary delays. Await fonts separately.
  • Keep capture dimensions bounded: very long full-page documents create larger images and more rendering work. Use viewport or clipped captures when the use case permits.
  • Reuse browser processes in services: launching Chromium for every request adds startup work. A long-lived browser with isolated pages can reduce that overhead; always close each page and handle browser disconnections.
  • Bound timeouts and concurrency: set navigation and readiness timeouts appropriate to your environment, and cap simultaneous renders to manage CPU and memory.
  • Make output reproducible: keep viewport, device scale, font assets, page data, and readiness criteria consistent. Browser versions and installed fonts can change rendering, so inspect output in the target runtime when pixel-level fidelity matters.
  • Account for infrastructure: self-hosted Puppeteer has no per-screenshot API charge, but uses compute, memory, browser maintenance, and engineering time. Remote font fetching and page assets also affect latency and reliability.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from ScreenshotNeo. Its API takes a URL in one GET request and returns an image or PDF. See the ScreenshotNeo API docs for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/urdu-page -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com/urdu-page"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.com/urdu-page'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const image = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', image));

Cookie and consent banners are accepted or removed before capture, along with known newsletter popups and chat widgets. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers identify the page verdict and billing status. AI agents can use the MCP server’s take_screenshot, get_page_info, and capture_pdf tools. The free plan includes 1,000 screenshots monthly with no card; paid plans start at $5 for 3,000 screenshots.

Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.

FAQ

Does Puppeteer need a special Urdu screenshot option?

No. Set the document’s language and direction in HTML, make the desired font available, and capture the rendered page normally.

Will dir="rtl" choose an Urdu font?

No. Direction controls writing flow. Font selection and availability are separate.

Should I use dir="rtl" or CSS direction: rtl?

Prefer the HTML dir attribute for document or element direction where possible. Use CSS for styling needs that cannot be expressed by the markup.

Why does waiting for network idle sometimes produce the wrong image?

A network state does not necessarily mean the page’s application data, fonts, or desired visual state is complete. Wait for the condition your page uses to indicate capture readiness.

References