How to Take a Playwright Screenshot of a Webpage Hosted on an Indian-Language Domain
Use Playwright to capture a webpage on a Unicode Indian-language domain. Handle IDNs correctly, choose the capture scope, and diagnose navigation errors.
Use the site’s complete URL, including its scheme, in page.goto(), then call page.screenshot(). Playwright’s browser-compatible URL handling supports internationalized domain names (IDNs), including hostnames written in Indian scripts. A navigation can return an HTTP error page such as 404 or 500 without throwing, so check the response status separately.
The hostname उदाहरण.भारत below is illustrative, not a real destination. Replace it with the exact address of the site you want to capture. Playwright documents the screenshot workflow and Page API.
1. Install Playwright and its browser
For a new Node.js project, install Playwright and its Chromium browser:
npm init -y
npm install playwright
npx playwright install chromium
Save the following as screenshot.js. Run it with node screenshot.js.
2. Navigate to the Unicode domain and capture the page
const { chromium } = require('playwright');
(async () => {
const browser = await chromium.launch();
try {
const page = await browser.newPage();
const response = await page.goto('https://उदाहरण.भारत', {
waitUntil: 'domcontentloaded',
});
// fullPage: true captures the full scrollable page.
// Remove it or set it to false for the visible viewport only.
await page.screenshot({ path: 'page.png', fullPage: true });
console.log('HTTP status:', response?.status());
} finally {
await browser.close();
}
})();
Use https:// for an HTTPS site or the site’s actual scheme. The script writes page.png in the current directory. The finally block closes the browser even if navigation or capture fails.
3. Choose what to capture
Visible viewport or full page
By default, page.screenshot() captures the visible viewport. Set fullPage: true to capture the page’s full scrollable area. A full-page image can be much taller and larger than a viewport image, so use it only when the extra content is needed.
A single element
To capture one component, such as a banner or article, take a locator screenshot instead:
const target = page.locator('main article');
await target.screenshot({ path: 'article.png' });
Choose a selector that identifies the intended element. If the locator matches no element, the capture cannot proceed; if it matches multiple elements, make the target unambiguous.
Format, path, and scale
- Path: set
pathto the output filename, such aspage.pngorpage.jpg. - Format: Playwright supports PNG, JPEG, and WebP. The format can be inferred from the file extension or set with the screenshot
typeoption. - JPEG quality: use the
qualityoption for JPEG or WebP when a smaller, lossy image is acceptable. It does not apply to PNG. - Scale: choose CSS-pixel output or device-pixel output with
scale: 'css'orscale: 'device'. CSS scale keeps the image dimensions closer to the page’s CSS dimensions; device scale preserves the device-pixel ratio and can produce a larger image. - Clip: use the
clipoption to capture a specified rectangle when neither the viewport nor a whole element is the right scope.
For details on the supported screenshot options and their constraints, see the Page API.
4. Understand Unicode hostnames and IDN conversion
A hostname written in an Indian script is an internationalized domain name. Browser-standard WHATWG URL processing converts Unicode characters in the hostname to an ASCII representation using Punycode when serializing the URL. In most cases, you can pass the real Unicode URL directly to Playwright; do not encode the entire URL as if every character were part of the hostname.
A URL has separate parts: scheme, hostname, path, query, and fragment. Keep each in its proper position. For example, Unicode in a path is not the same thing as Unicode in a hostname.
To inspect or validate a hostname with Node.js, use domainToASCII() and domainToUnicode() from the built-in node:url module. These helpers operate on the hostname, not a whole URL:
const { domainToASCII, domainToUnicode } = require('node:url');
const hostname = 'उदाहरण.भारत';
const asciiHostname = domainToASCII(hostname);
if (!asciiHostname) {
throw new Error('Invalid internationalized hostname');
}
console.log('ASCII hostname:', asciiHostname);
console.log('Unicode hostname:', domainToUnicode(asciiHostname));
An empty result from domainToASCII() indicates an invalid domain input. Node.js documents these helpers and its WHATWG URL and IDN behavior.
Check the domain’s exact spelling before capturing it. Visually similar characters or a typo can name a different host. Punycode conversion can help diagnose representation; it does not prove the domain is the intended one or that its server is reachable.
5. Diagnose navigation separately from the HTTP response
page.goto() can throw when the URL is invalid, the server is unreachable, an SSL error occurs, or navigation times out. A completed navigation that returns an HTTP 404 or 500 response is different: Playwright can return a response object normally. Check response?.status() and decide whether that page is useful to save.
For a quick diagnostic run, log the resolved URL as well as the status:
const response = await page.goto('https://उदाहरण.भारत', {
waitUntil: 'domcontentloaded',
});
console.log('Final URL:', page.url());
console.log('HTTP status:', response?.status());
A redirect can make the final URL differ from the one supplied. A screenshot of an error page may still be a valid image file, so use the status and final URL to interpret what it shows.
6. Common problems and fixes
| Symptom | Likely cause | What to check or change |
|---|---|---|
| Invalid URL error | The URL is missing a scheme, malformed, or contains a hostname that cannot be parsed. | Use the complete address, such as https://.... Verify the hostname spelling and, if useful, call domainToASCII() on the hostname alone. |
| Navigation times out | The server is slow or unreachable, or the chosen navigation milestone is not reached. | Confirm the address is reachable and select an appropriate waitUntil milestone. A timeout is a navigation failure, not evidence that IDN conversion failed. |
| SSL or certificate error | The target’s HTTPS connection has a certificate problem or cannot be established. | Check the exact host and its HTTPS configuration. Do not treat changing the Unicode hostname to Punycode as a certificate fix; both forms identify the same IDN when valid. |
| Image shows a 404 or 500 page | Navigation completed with an HTTP error response. | Inspect response?.status() and page.url(). Handle the status in your script rather than assuming goto() throws for HTTP errors. |
| Image is blank or content is missing | The page may not have rendered the desired content by the chosen navigation milestone. | Wait for a relevant element when you know its selector, or use an appropriate wait strategy. Confirm the element exists before taking a locator screenshot. |
| Screenshot is unexpectedly large or small | Full-page capture or device-pixel scale changed the output dimensions. | Review fullPage and scale. Use viewport capture or CSS scale when those dimensions suit your use case. |
7. Performance, reliability, and cost
- Capture only the required area. Viewport and element screenshots generally produce less image data than a very tall full-page capture.
- Pick the navigation milestone deliberately.
domcontentloadedcan avoid waiting for every resource, but a page’s important content may need a later or element-specific wait. Waiting longer can increase capture time. - Keep browser cleanup reliable. Close the browser in a
finallyblock so an exception does not leave the process running. - Check both failure classes. Handle navigation exceptions and inspect returned HTTP status codes. A saved file alone does not establish that the intended page loaded successfully.
- Plan storage and transfer. Full-page, high-resolution, or device-scale images can consume more memory, disk space, and bandwidth. Choose the smallest format and scale that meet the downstream need.
- Playwright has no per-screenshot API charge in this workflow. Account for the compute, browser runtime, storage, and network resources of the machine or service where you run it.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server. It accepts a URL in one GET request and returns an image or PDF. For an Indian-language hostname, pass the site’s complete URL as the url parameter. See the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://उदाहरण.भारत -o shot.webp
Cookie and consent banners are accepted before capture, and 60+ known consent platforms, newsletter popups, and chat widgets are removed; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up for 1,000 free screenshots a month, with no card.
FAQ
Do I need to convert an Indian-script domain to Punycode before calling Playwright?
Usually no. Pass the complete Unicode URL to Playwright. Convert only the hostname for diagnosis if you need to inspect its ASCII form.
Can Playwright save a screenshot when the page returns 404?
It can capture the page if navigation completes. Check the returned status so you can distinguish the error page from successful content.
Does fullPage: true capture content that loads only after scrolling?
It captures the page’s scrollable area, but content that depends on scrolling may require additional page interaction or waiting before capture.


