ScreenshotNeo

BlogHow-to

How to Screenshot a Webpage in Marathi with Python Playwright

Capture a webpage with Python Playwright, set the Marathi locale, save a full-page image, and troubleshoot Devanagari rendering.

By the ScreenshotNeo team4 October 20266 min read

Use Playwright’s Python API to open the page and call page.screenshot(). To emulate Marathi locale behavior, create the browser context with locale="mr-IN". To capture the entire scrollable page, pass full_page=True.

Locale and font rendering are separate concerns: Playwright’s locale setting affects navigator.language, the Accept-Language request header, and locale-sensitive formatting. It does not install a Devanagari font or guarantee that Marathi glyphs will render correctly. Playwright’s locale documentation describes the locale behavior.

1. Install Playwright and a browser

Install the Python package, then install the Chromium browser binary Playwright uses:

python -m pip install playwright
python -m playwright install chromium

These commands follow the official Playwright Python installation guide. Run the browser-install command in each environment where the script will run, such as a development machine or deployment container.

2. Capture a Marathi page with Python

This synchronous script sets the Marathi locale, waits for the page to load, and saves a full-page PNG. Replace the example URL with the page you want to capture.

from pathlib import Path
from playwright.sync_api import sync_playwright

URL = "https://example.com"
OUTPUT = Path("marathi-page.png")

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    context = browser.new_context(
        locale="mr-IN",
        viewport={"width": 1440, "height": 1000},
        device_scale_factor=1,
    )
    page = context.new_page()

    response = page.goto(URL, wait_until="networkidle", timeout=60_000)
    if response is not None and response.status >= 400:
        raise RuntimeError(f"Page returned HTTP {response.status}: {URL}")

    page.screenshot(path=str(OUTPUT), full_page=True)

    context.close()
    browser.close()

print(f"Saved {OUTPUT.resolve()}")

Playwright’s browser launches headless by default; the example makes that setting explicit. The screenshot guide documents viewport and full-page screenshots, and the Page screenshot API lists screenshot options.

Choose the capture area

  • Current viewport: omit full_page or set it to False. This captures only what is currently visible.
  • Full scrollable page: use full_page=True. This produces a taller image of the page.
  • One element: use a locator screenshot, for example page.locator(".article").screenshot(path="article.png"). Replace the selector with one that matches the element you need.

Wait for the right state

The example uses wait_until="networkidle", but some sites keep network requests open or continuously poll. If navigation times out, try wait_until="domcontentloaded" and then wait for a page-specific element:

page.goto(URL, wait_until="domcontentloaded", timeout=60_000)
page.locator("main").wait_for(state="visible", timeout=15_000)
page.screenshot(path="marathi-page.png", full_page=True)

For a page that paints content shortly after navigation, a brief explicit delay can help, but prefer waiting for a meaningful selector when one is available.

3. Use the asynchronous Python API

Use the async API when the surrounding program already uses asyncio or needs to coordinate multiple asynchronous tasks:

import asyncio
from playwright.async_api import async_playwright

async def main():
    async with async_playwright() as p:
        browser = await p.chromium.launch(headless=True)
        context = await browser.new_context(
            locale="mr-IN",
            viewport={"width": 1440, "height": 1000},
            device_scale_factor=1,
        )
        page = await context.new_page()
        response = await page.goto(
            "https://example.com",
            wait_until="domcontentloaded",
            timeout=60_000,
        )
        if response is not None and response.status >= 400:
            raise RuntimeError(f"Page returned HTTP {response.status}")

        await page.locator("body").wait_for(state="visible")
        await page.screenshot(path="marathi-page.png", full_page=True)
        await context.close()
        await browser.close()

asyncio.run(main())

4. Check Marathi text and image output

If the page uses locale-sensitive behavior, verify what the browser received and what the document declares:

print("Browser language:", page.evaluate("() => navigator.language"))
print("Document language:", page.locator("html").get_attribute("lang"))

The browser language reflects the emulated context locale. The document’s lang attribute is set by the site and may be absent or different; locale emulation does not rewrite the page’s HTML.

If Marathi appears as boxes, missing glyphs, or incorrect shaping, inspect the browser and operating-system fonts available in the runtime and check whether the site’s own font files and stylesheets loaded. The locale setting alone cannot resolve missing-font or page-CSS problems. The exact fix depends on the operating system, container image, and page, so confirm those details in the environment where the screenshot runs.

Format, scale, and output

Playwright can save screenshots to a path or return screenshot bytes for further processing. The screenshot API also provides controls such as image type, quality for JPEG, background handling, animations, and scale. Check the Page screenshot options for the supported values and constraints. PNG is a sensible default when text clarity matters; JPEG can reduce file size but is lossy.

5. Troubleshooting

Symptom Likely cause What to try
Executable doesn't exist or browser launch fails The Playwright package is installed but its browser binary is not present in this environment. Run python -m playwright install chromium in the same environment that runs the script.
Navigation times out The site keeps network activity open, responds slowly, or blocks automated traffic. Use wait_until="domcontentloaded", increase the navigation timeout when appropriate, and wait for a specific visible element.
Screenshot is blank or incomplete The page has not rendered its main content yet, or content appears after a client-side action. Wait for a known content locator before capturing. Check the returned response status and inspect the page in a headed browser if needed.
Marathi text displays as squares or missing glyphs A suitable font may not be available or loaded in the browser runtime, or the page’s font CSS did not load. Inspect the runtime’s available fonts and the page’s network and CSS behavior. Setting locale="mr-IN" does not install fonts.
Dates or numbers remain in another locale The page may explicitly format values, use its own language preference, or ignore the browser locale. Check navigator.language, the page’s language controls, and the page’s own formatting logic.
Full-page capture is extremely tall The page contains long feeds, repeated content, or lazy-loaded sections. Capture a specific locator or use a viewport screenshot if the entire document is unnecessary. For lazy content, scroll or otherwise trigger the page’s loading behavior before capture.
Element screenshot fails or captures the wrong area The locator matches no element, more than one element, or an element that is not visible. Use a selector that identifies one visible element and wait for it before calling locator.screenshot().

6. Reliability, performance, and cost

  • Reuse the browser for batches: launching a browser for every URL adds startup work. For a batch, launch once and create a page or context per job, then close them when finished.
  • Set bounded timeouts: navigation and selector waits should have limits so a stalled site does not hold a worker indefinitely.
  • Control page state: use a stable viewport and device scale factor when you need repeatable dimensions. Site content can still change between captures.
  • Keep full-page captures in scope: very long pages require more image memory and produce larger files than viewport or element captures.
  • Account for runtime costs: self-hosted Playwright requires a Python runtime and browser installation wherever the script runs. Resource use depends on the page, capture dimensions, and concurrency; the supplied research contains no benchmark figures.

7. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. Its API can return a screenshot with one GET request; see the ScreenshotNeo API documentation. Set YOUR_API_KEY to your key and replace the target URL as needed.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));

ScreenshotNeo accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up for 1,000 free screenshots a month, with no card required.

FAQ

Does locale="mr-IN" translate a page into Marathi?

No. It emulates browser locale behavior. The website must provide Marathi content or respond to the locale itself.

Can Playwright save a screenshot without writing a file?

Yes. The screenshot API can return image bytes; see the Page screenshot documentation for the buffer option.

Should I use sync or async Python?

Either works. Use sync for a straightforward script and async when integrating with an asynchronous application.