ScreenshotNeo

BlogAI agents

How to Screenshot a Website in Hindi and Save It as a PDF with an AI Agent

Give an AI agent a clear capture prompt, then use Playwright to save a full-page image and PDF. Here are runnable options, fixes, and an API shortcut.

By the ScreenshotNeo team4 October 20268 min read

To have an AI agent screenshot a website in Hindi and save it as a PDF, ask it to open the URL, wait for the page to render, capture the full scrollable page, and export that page as a PDF. Ask for both files explicitly: a screenshot is an image, while a PDF is a separate browser export.

For example, give the agent this prompt: Open https://example.com and wait for the page to finish rendering. Save a full-page PNG screenshot as website.png. Export the same page as website.pdf. Confirm that both files were created. If you want the prompt itself to be in Hindi, translate those instructions while keeping the URL and filenames unchanged. The workflow and commands below are the same either way.

With Playwright CLI, the core commands are:

playwright-cli screenshot --full-page --filename=website.png
playwright-cli pdf --filename=website.pdf

Run these in the agent’s browser session after it has opened the target page. See the Playwright CLI screenshot and PDF documentation for command details.

1. Choose the capture method

Use the method that matches how your agent controls a browser. Playwright CLI is a direct command workflow; the Playwright Page API is useful when you already have a Node.js browser automation script; Playwright MCP is for an MCP-connected agent; and Chrome Headless is an alternative command-line route.

Method Use it when Key choice
Playwright CLI Your agent can run browser CLI commands Viewport or full page; image format and output filename
Playwright Page API You want to automate capture in JavaScript Screenshot image versus PDF; print or screen media
Playwright MCP Your agent uses Playwright’s MCP tools Viewport, element, or full-page screenshot
Chrome Headless CLI You want Chrome command-line capture Screenshot or PDF, with additional render time if needed

2. Capture an image and PDF with Playwright CLI

First have the agent navigate to the target URL in its Playwright browser session. Then capture the page. The screenshot command saves an image; the PDF command creates a separate PDF.

playwright-cli screenshot --full-page --filename=website.png
playwright-cli pdf --filename=website.pdf

For just the visible viewport, omit --full-page:

playwright-cli screenshot --filename=viewport.png

The documented screenshot formats include PNG, JPEG, and WebP; use a filename extension that matches the desired output. The CLI can also capture a target element. Consult the official command reference for the supported element-target syntax in your installed version.

3. Use the Playwright Page API in Node.js

This example launches Chromium, opens a URL, waits for network activity to settle, writes a full-page PNG, and writes a PDF. It is a complete script for a basic public page. The caller supplies the URL as the first argument.

import { chromium } from 'playwright';

const url = process.argv[2];
if (!url) throw new Error('Usage: node capture.mjs https://example.com');

const browser = await chromium.launch({ headless: true });
try {
  const page = await browser.newPage();
  await page.goto(url, { waitUntil: 'networkidle', timeout: 60000 });
  await page.screenshot({ path: 'website.png', fullPage: true });
  await page.pdf({ path: 'website.pdf', printBackground: true });
  console.log('Created website.png and website.pdf');
} finally {
  await browser.close();
}

Install Playwright in your project and install its browser as described in the Playwright getting started guide. Save the script as capture.mjs and run node capture.mjs https://example.com.

page.pdf() renders using print CSS media by default. A site’s print stylesheet can change colors, hide elements, or rearrange the page. If the PDF should follow screen CSS instead, set the media mode before exporting:

await page.emulateMedia({ media: 'screen' });
await page.pdf({ path: 'website.pdf', printBackground: true });

Print rendering may modify colors. Playwright documents -webkit-print-color-adjust as a way for page styles to control print color adjustment. The PDF and full-page screenshot may also differ because PDF pagination and print CSS are not the same as a tall image. See the Playwright Page API reference.

4. Use an MCP agent or Chrome Headless

Playwright MCP

If your AI agent is connected to Playwright MCP, ask it to navigate to the URL and capture the viewport, a chosen element, or the full scrollable page. The screenshot tool is for visual appearance. If the agent needs to understand page structure or interactions, use an accessibility snapshot as a separate step; a screenshot does not expose that structure. See the Playwright MCP screenshot documentation.

Chrome Headless CLI

Chrome Headless supports image screenshots and PDF output from the command line. For example:

chrome --headless --no-sandbox --screenshot=website.png --window-size=1440,900 https://example.com
chrome --headless --no-sandbox --print-to-pdf=website.pdf https://example.com

Use --timeout when a page needs additional rendering time before capture. The exact Chrome executable name and sandbox requirements depend on the environment. Refer to the Chrome Headless command-line reference.

5. Pick the right scope and output

  • Viewport screenshot: captures only the currently visible browser area. Use this for a quick view of the top of a page.
  • Full-page screenshot: captures the scrollable page as one image. Use --full-page in the CLI or fullPage: true in the Page API.
  • Element screenshot: targets a specific element when the whole page is unnecessary. This is available in the CLI and MCP screenshot workflows.
  • PDF: a separate document export, typically paginated and affected by print CSS. Request it separately even if an image has already been saved.

Choose PNG for a lossless image, or JPEG/WebP when those formats suit your storage or delivery needs. The CLI documentation lists PNG, JPEG, and WebP. For a PDF intended to resemble the screen, use screen media in the Page API; for a print-oriented document, keep print media and inspect the resulting pagination.

6. Wait for dynamic pages and verify the files

A page’s initial navigation may finish before client-side content, lazy-loaded images, or animations are ready. Wait for a meaningful page condition when possible, such as a known heading or content selector. If no reliable selector exists, use a deliberate delay, but there is no single wait duration that works for every site. Chrome Headless also provides a timeout option for additional rendering time.

  1. Navigate to the exact URL and wait for the page’s main content to appear.
  2. For long pages, allow lazy-loaded content to load before taking the full-page capture. Scrolling through the page can trigger content that appears only on scroll.
  3. Capture the image and export the PDF as distinct steps.
  4. Confirm that both output files exist and have nonzero size; open them to check that the page content and layout are present.

A successful command alone does not confirm that a dynamic page has finished rendering. The agent should inspect the outputs or report any navigation and capture errors.

7. Troubleshooting

Symptom Likely cause What to do
The image shows only the top of the page Viewport capture was used Use --full-page in the CLI or fullPage: true in the Page API.
The PDF looks different from the browser PDF uses print CSS by default Use page.emulateMedia({ media: 'screen' }) before page.pdf() if screen styling is required. Check print styles and pagination.
The screenshot is blank or missing page content Capture happened before rendering, or navigation failed Wait for a meaningful selector or additional rendering time, then check the page and navigation result before recapturing.
Images are missing lower down the page Lazy loading may require scrolling Scroll through the page and wait for images or content to load before capturing the full page.
The output format is unexpected The filename extension or requested command does not match the desired output Use the screenshot command for an image and the PDF command for a PDF. Set an appropriate image extension such as PNG, JPEG, or WebP.
The agent cannot understand buttons or page structure from the image A screenshot provides visual output, not a structured interaction view Use an accessibility snapshot alongside the screenshot when the workflow needs page structure or interactions.
CLI command is unavailable The CLI or browser session is not set up in the environment Install and configure the documented Playwright CLI/browser workflow, or use the Page API script or Chrome Headless if available.

8. Reliability, performance, and cost

Full-page images can become large for long pages, and PDF output can span many pages. Capture only the viewport or a relevant element when that meets the task. Use a selector-based readiness condition for repeatable automation; fixed delays can be too short on slow pages and unnecessarily long on fast ones.

For reliability, use explicit output paths, close the browser in a finally block, set a navigation timeout appropriate to your environment, and check that outputs were created. Pages requiring authentication or specific regional conditions may need matching browser context settings. Do not treat an HTTP response or a completed navigation as proof that the visible content is correct.

Browser-based capture consumes compute and time to launch a browser, load the site, render content, and write files. The cited Playwright and Chrome references document capture behavior, not universal benchmarks or fixed costs. Actual runtime and infrastructure cost depend on the page, browser environment, and concurrency.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request can return an image or PDF; the API supports full-page capture and PDF options. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o website.pdf

For a PDF response, request the PDF output option documented by the API for your integration. The standard example below saves an image response:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o website.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
r.raise_for_status()
open("website.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('website.webp', res);

ScreenshotNeo accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for AI agents and MCP clients.

The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots; every feature is available on every plan. Create a free ScreenshotNeo account and get 1,000 screenshots a month with no card.

FAQ

Does a screenshot automatically become a PDF?

No. A screenshot is an image file. Run a separate PDF export command or call the PDF capture workflow.

Can an AI agent follow instructions written in Hindi?

The prompt can be written in Hindi if the agent supports it. Keep the URL, requested output filenames, and requirement for a full-page capture explicit.

Should I use a screenshot or an accessibility snapshot?

Use a screenshot to inspect visual layout. Use an accessibility snapshot when the agent needs page structure or interaction details.

Will every website render identically in a PDF?

No. Print CSS, pagination, loading behavior, and page-specific scripts can change the output. Choose print or screen media deliberately and inspect the resulting file.

Sources