How to Capture a Hindi Website as a PDF with BrowserCat
Use BrowserCat with Playwright to save a Hindi webpage as a PDF. Configure print options, wait for fonts and content, and check the finished file.
Use Playwright’s page.pdf() method in a BrowserCat Chromium session. It creates a browser-printed PDF document; it is not a screenshot image embedded in a PDF. Navigate to the Hindi page, wait for its content and fonts to load, set print options such as A4 and background printing, then save the PDF. BrowserCat documents Playwright as its recommended way to get started, and its current sessions run on Chromium-based browsers. BrowserCat quick start · Playwright PDF and screenshot options · Browser configuration.
Hindi rendering depends on the page’s loaded fonts and assets. The reviewed BrowserCat documentation does not promise Hindi-specific font coverage or glyph shaping for every site, so inspect the resulting PDF before relying on it.
1. Create a BrowserCat PDF with Playwright
The following Node.js script connects Playwright to BrowserCat, visits a Hindi page, waits for the page and its fonts, and writes an A4 PDF. Replace the example URL with the page you need. Set BROWSERCAT_API_KEY to an API key from your BrowserCat account.
- Install Node.js and Playwright:
npm install playwright. - Create a BrowserCat account and API key using its quick-start instructions.
- Set the key in your environment, then save the script below as
save-hindi-pdf.mjs. - Run
node save-hindi-pdf.mjs.
import { chromium } from 'playwright';
const apiKey = process.env.BROWSERCAT_API_KEY;
if (!apiKey) {
throw new Error('Set BROWSERCAT_API_KEY before running this script.');
}
const url = process.env.TARGET_URL ?? 'https://example.com';
const browser = await chromium.connect('wss://api.browsercat.com/connect', {
headers: { 'Api-Key': apiKey },
});
try {
const page = await browser.newPage({
viewport: { width: 1280, height: 900 },
});
const response = await page.goto(url, {
waitUntil: 'load',
timeout: 60000,
});
if (response && !response.ok()) {
throw new Error(`Page returned HTTP ${response.status()}`);
}
// Wait for web fonts to finish loading when the page uses them.
await page.evaluate(() => document.fonts.ready);
// Optional: wait for a known page-specific element, for example:
// await page.locator('main').waitFor({ state: 'visible', timeout: 15000 });
await page.pdf({
path: 'hindi-page.pdf',
format: 'A4',
printBackground: true,
margin: {
top: '12mm',
right: '12mm',
bottom: '12mm',
left: '12mm',
},
});
console.log('Saved hindi-page.pdf');
} finally {
await browser.close();
}
For example, run it with an explicit target URL by setting TARGET_URL in the environment as well as the API key. The BrowserCat connection endpoint and API key header follow its Playwright quick start. Since this is a cloud browser session, the script connects with chromium.connect() instead of starting a local browser with chromium.launch().
2. Tune the PDF for the page
Playwright’s page.pdf() returns a PDF buffer if no path is supplied; provide path to save it directly. Choose the options that suit the document:
| Option | When to use it |
|---|---|
format |
Choose a standard size such as A4, A3, Letter, Legal, or Tabloid. The documented default is Letter, so set A4 explicitly when that is what you need. |
landscape |
Set to true for wide tables or layouts. The documented default is portrait. |
printBackground |
Set to true when page background colors or images matter. The documented default is off. |
margin |
Set top, right, bottom, and left margins with CSS lengths such as 12mm, 1cm, or 0.5in. The documented defaults are zero. |
width and height |
Set custom page dimensions using units such as px, in, cm, or mm. |
preferCSSPageSize |
Set to true to prefer a page size declared by the page’s print CSS, when present. |
pageRanges |
Export only selected pages when a large PDF does not need every page. |
displayHeaderFooter, headerTemplate, footerTemplate |
Add browser-generated headers or footers when you need page labels or numbering. |
tagged |
Request tagged PDF output when your downstream workflow needs it. |
Options and defaults are documented in the BrowserCat Playwright cheatsheet. To use the page’s CSS-defined print size, you can add a print stylesheet before exporting:
await page.addStyleTag({
content: '@page { size: A4 landscape; margin: 12mm; }',
});
await page.pdf({ path: 'hindi-page.pdf', preferCSSPageSize: true });
Print layout can differ from the normal browser view: sites may apply print-specific CSS, hide navigation, or reflow columns. Check the output for clipped content, missing backgrounds, unexpected page breaks, or headers and footers that the website itself adds.
3. Wait for Hindi text, fonts, and dynamic content
A successful navigation does not guarantee that every font, image, or client-rendered section is ready. Use a readiness condition based on the page you are capturing:
- Wait for navigation:
page.goto()supports load-state choices such asloadanddomcontentloaded. The sample waits forload; a page that keeps network connections open may need a different condition. - Wait for a meaningful element: if the main article appears after JavaScript runs, wait for a stable selector such as
mainor the article’s content container. - Wait for fonts:
await page.evaluate(() => document.fonts.ready)waits for the document’s font loading set. It cannot make a missing font available; check that the page actually loads its intended Devanagari font. - Wait for images if needed: for pages where images matter, wait for visible images to complete loading and decode before printing. Lazy-loaded images may require scrolling through the page first.
- Use a bounded delay only as a fallback: a fixed wait can help with a known short animation or delayed widget, but it is less reliable than waiting for the actual content.
There is no universal wait condition for every website. A page can report that it loaded while a particular article, translation, or font is still pending. Verify a sample PDF for the target site and add an explicit selector wait if needed.
4. PDF document or screenshot image?
Use page.pdf() when you want a paginated, printable document. Use page.screenshot() when you want a visual image of a viewport, full page, or selected region. BrowserCat documents screenshot controls for viewport versus full-page output, PNG or JPEG, clipping, masking, scaling, and screenshot styling in its Playwright cheatsheet.
To save an image, for example, use await page.screenshot({ path: 'page.png', fullPage: true }). That creates a PNG image, not a browser-printed PDF. If your requirement is specifically a picture placed inside a PDF, capture the image first and create a PDF from that image as a separate conversion step.
5. Troubleshooting
| Problem | Likely cause | Fix |
|---|---|---|
| Connection fails before the page opens | The API key is missing or invalid, or the connection options do not include the expected header. | Check that BROWSERCAT_API_KEY is set and passed as Api-Key. Follow the current BrowserCat connection example. |
| Navigation times out | The site is slow, blocked, or keeps requests active after its useful content appears. | Check the URL and site access. Consider waiting for domcontentloaded and then for a page-specific selector rather than waiting for every resource. |
| PDF is blank or missing article text | The page may render content after navigation, require interaction, or show an error page. | Wait for the relevant content selector, check the response status, and inspect the rendered page before exporting. |
| Hindi characters appear as boxes, missing marks, or incorrect glyphs | The intended font may not have loaded, or the page/browser combination may not render the glyphs as expected. | Wait for document.fonts.ready, confirm the page’s font requests succeed, and inspect the PDF. BrowserCat’s reviewed docs do not make a Hindi-specific rendering guarantee. |
| Background colors or images are missing | PDF background printing is off by default. | Set printBackground: true. |
| Wide content is clipped or hard to read | Portrait pages, narrow paper, or large margins do not fit the layout. | Try A4 landscape, reduce margins, or use a custom paper size. Inspect print CSS and page breaks. |
| The result is an image when a document was expected | page.screenshot() was used. |
Call page.pdf() for a browser-printed PDF. |
| PDF file is not written | The script failed before the PDF call completed or the path is not writable in the process environment. | Check the thrown error and write to a path your script can access. The example writes to its current working directory. |
6. Performance, reliability, and cost considerations
- Keep waits specific: waiting for all network activity can be unsuitable for sites with analytics, streaming, or long-lived requests. A selector or font readiness check often expresses the capture requirement more clearly.
- Reuse the browser when capturing multiple pages: connect once, create pages as needed, and close the browser in a
finallyblock so the session is released after success or failure. - Bound slow operations: set navigation and selector timeouts, handle failures, and retry only transient errors. Avoid blindly retrying a page that consistently returns an error or requires authentication.
- Inspect representative outputs: test pages with long Devanagari passages, mixed Latin and Hindi, tables, and web fonts. Print CSS and content vary by site, so one successful page does not establish correct output for every page.
- Plan around BrowserCat usage and billing: consult BrowserCat’s account and billing information for current costs and limits. The reviewed research does not establish a specific price, quota, or PDF-specific charge, so this guide does not state one.
7. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. One GET request can return PNG, JPEG, WebP, or a PDF. The example below requests a PDF; see the ScreenshotNeo API documentation for supported parameters.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.com \
-d format=pdf \
-o hindi-page.pdf
Cookie banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents use screenshot tools. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. ScreenshotNeo also supports PDF output, but inspect the result to confirm the target page renders Hindi correctly.
Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.
FAQ
Does saving a Hindi webpage as PDF preserve selectable text?
A browser-printed PDF may preserve text, but the result depends on the page and its rendering. Inspect the generated file to confirm that text selection and search work as needed.
Does BrowserCat run this workflow in Firefox?
The reviewed BrowserCat browser configuration says current sessions run on Chromium-based browsers; Firefox and WebKit are listed as roadmap items.
Can I create a PDF without writing a script?
Playwright’s CLI includes a PDF command for a URL. BrowserCat’s documented cloud connection flow, however, is shown with a Playwright script connected to its browser endpoint.


