How to Capture Website Screenshots of UPI Payment Instructions with an AI Agent
Use an AI agent and Playwright to capture UPI instructions at the right scope, protect sensitive details, and review the image before sharing.
To capture a UPI payment instructions page with an AI agent, have the agent open the authorized page in a browser, confirm the page and instructions, choose a viewport, element, or full-page screenshot, and save the image. Review the result before sharing it. Mask or omit sensitive account details, and never include or disclose a UPI PIN. A screenshot records visible page content; it does not prove that instructions or a beneficiary are genuine.
This guide uses Playwright because its browser tooling documents viewport, element, and full-page screenshots. Other agents may expose different controls, filenames, or storage behavior, so check the documentation for the browser runtime you use. See the Playwright MCP screenshot guide and the Page screenshot API.
1. Choose a safe capture scope
First decide what the screenshot needs to show. Capture the smallest area that communicates the instructions clearly.
| Scope | Use it when | Watch for |
|---|---|---|
| Viewport | The relevant instructions fit on screen. | Content below the fold will be missing. |
| Element | A specific instruction card or component contains everything needed. | Make sure the selected element includes its labels and context. |
| Full page | The instructions span multiple scroll positions. | A very tall image can make text difficult to read. Inspect it at its intended display size. |
Use an accessibility snapshot or equivalent page-structure inspection to locate content and interaction targets. A screenshot helps with visual review; it is not the best tool for identifying controls or extracting text. Playwright documents screenshots as complementary to accessibility snapshots.
UPI QR payment flows can show merchant information on the payment page, according to NPCI’s UPI FAQ. Treat a screenshot as a record of what the page displayed, not authentication of the recipient or a reason to make a payment. NPCI also says not to share the UPI PIN. OpenAI’s computer-use documentation warns that screenshots can contain sensitive page or account data.
2. Capture with Playwright MCP
If your AI agent is connected to Playwright MCP, ask it to navigate to the authorized page, inspect the page structure, and capture the relevant scope. The MCP guide documents viewport screenshots, a target element, and the full scrollable page. Exact tool names and arguments depend on the MCP client and version; consult that runtime’s documentation.
- Open the intended page in an authorized browser session. Do not ask the agent to bypass access controls.
- Confirm the page identity and that the visible content is the intended UPI instructions.
- Inspect the page structure to identify the relevant region and any sensitive fields.
- Choose viewport, element, or full-page scope. Mask or exclude sensitive details where the tool supports it.
- Save the screenshot and inspect the resulting image before sharing or publishing it.
A useful agent instruction is: “Open the authorized UPI instructions page. Confirm its identity and locate the instructions using the page structure. Capture only the instruction area if it contains all necessary context; otherwise capture the full page. Exclude or mask sensitive account details, never capture a UPI PIN, save the image, and tell me its path and any content that was unclear. Do not submit a payment.”
The instruction is deliberately explicit about the capture boundary and the prohibited payment action. A screenshot does not establish that the page is trustworthy, and the agent should not infer payment credentials.
3. Capture with Playwright code
For a repeatable capture outside an agent-managed MCP session, Playwright’s Page API can save a screenshot. This runnable Node.js example opens a page, waits for a known instruction element, and captures that element. Replace the URL and selector with values appropriate to your authorized site; inspect the selector and output before sharing.
import { chromium } from 'playwright';
const url = process.env.PAGE_URL;
const selector = process.env.INSTRUCTIONS_SELECTOR;
if (!url || !selector) {
throw new Error('Set PAGE_URL and INSTRUCTIONS_SELECTOR');
}
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage({ viewport: { width: 1280, height: 900 } });
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 30_000 });
const instructions = page.locator(selector);
await instructions.waitFor({ state: 'visible', timeout: 15_000 });
await instructions.screenshot({ path: 'upi-instructions.png' });
} finally {
await browser.close();
}
Install Playwright and its browser using the official Playwright getting-started instructions. This example assumes the page can be opened without a login. For an authenticated page, use an authorized session and protect any saved storage state as a secret; do not embed credentials in source code or screenshots.
To capture the current viewport, replace the element screenshot line with await page.screenshot({ path: 'upi-instructions.png' });. For the full page use await page.screenshot({ path: 'upi-instructions.png', fullPage: true });. The API also supports configurable output path, format and quality, masking selected locators, and scale. Check the API docs for exact options available in the installed version; image format and quality options vary by format.
For a sensitive field or other known page element, the screenshot API provides masking options. For example, with a locator already identifying the sensitive region, capture the page with await page.screenshot({ path: 'upi-instructions.png', mask: [page.locator('[data-sensitive]')] });. Verify the selector matches the intended content and inspect the saved image: a missed selector does not protect the underlying information. If the region should not appear at all, prefer a narrower element capture or hide it with a deliberate, reviewed stylesheet before capture.
4. Save, review, and share responsibly
- Check that the captured page is the intended site and shows the relevant instructions.
- Check that text remains readable at the size where the image will be used. If not, capture a smaller region or use a larger viewport.
- Inspect for names, account identifiers, transaction details, and other personal data. Crop or mask anything unnecessary.
- Never capture or disclose a UPI PIN. Do not ask an agent to enter, infer, or reveal one.
- Use a controlled destination for the image. Screenshots may contain account data, so apply the same access controls you would use for the page itself.
- Describe the screenshot as a record of displayed content. It cannot establish that payment instructions are authentic or that a payment should be made.
These are safety and readability recommendations based on the documented capture scopes and the cited payment and computer-use guidance. No particular website or agent workflow is presumed to have been tested here.
5. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| The screenshot is blank or shows a loading state. | Navigation or client-side rendering had not reached the instructions. | Wait for a stable, page-specific element to become visible, then capture. Confirm the page identity before relying on the output. |
| The element selector times out. | The selector is incorrect, the page structure changed, or the content is not available in this session. | Inspect the accessibility snapshot or DOM, update the selector, and confirm the intended page is loaded and accessible. |
| The image omits instructions. | The content is below the fold or outside the selected element. | Choose full-page capture or a container that includes the complete instructions, then inspect legibility. |
| The full-page image is hard to read. | The page is tall and the image is being displayed too small. | Capture the relevant element or split the content into clearly labeled captures. Avoid shrinking the image until text is unreadable. |
| Sensitive information remains visible. | The mask selector did not match, or the sensitive data was outside the capture boundary. | Correct the selector or use a narrower capture. Review the actual output before distributing it; do not assume masking succeeded. |
| The agent reports a different tool or option name. | MCP clients and runtime versions can expose different interfaces. | Check the documentation for the selected runtime and installed version. The generic workflow does not guarantee identical tool names or storage behavior. |
| The page requires authentication. | The browser session is not authorized or has not been authenticated. | Use an authorized browser session. Do not ask the agent to bypass access controls or put passwords and tokens in the screenshot. |
6. Performance, reliability, and cost
Capture only what you need: an element screenshot usually avoids producing an unnecessarily tall image, while full-page capture is useful when instructions span the document. Waiting for a specific visible element is more purposeful than capturing immediately after navigation. These choices follow the documented capture scopes; they are not performance benchmarks.
For repeatability, use a stable page URL, a page-specific selector, an explicit timeout, and a review step. A successful screenshot call only means an image was produced; it does not verify that the page content is correct. Dynamic content, authentication state, and page changes can affect what appears, so confirm the output on each run.
Playwright is browser automation software; this workflow has no per-screenshot service price specified here. Runtime, browser hosting, and storage may have costs in your environment. Protect saved images and any authenticated browser state as sensitive data. No accuracy, speed, or cost benchmark is claimed.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. Its MCP tools include take_screenshot, get_page_info, and capture_pdf for AI agents using Claude, Cursor, or another MCP client. Cookie and consent banners are accepted like a visitor and removed before capture, along with known newsletter popups and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Do not use any screenshot service to capture or disclose a UPI PIN or other sensitive data.
Make one GET request to capture an authorized public instructions page. Replace the URL with the page you are permitted to capture. The API accepts options for full-page or element capture, output formats, viewport and device presets, custom CSS and JavaScript, waits, headers and cookies, and other controls; see the ScreenshotNeo API documentation for parameter details. Treat output as page content, not payment verification.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed; an MCP server lets AI agents take screenshots. The Free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month with no card.
FAQ
Does a screenshot prove that a UPI payment request is genuine?
No. It records page content. Confirm the recipient and payment details through an appropriate trusted channel before taking any payment action.
Can an AI agent read a PIN from the page?
Do not ask it to do so. Never capture or share a UPI PIN.
Should I use a screenshot or an accessibility snapshot to find the instructions?
Use page structure or an accessibility snapshot to locate content and controls; use the screenshot to review visual presentation.
Will these Playwright steps work in every agent?
No. They describe a Playwright-based workflow. Verify the tool names, arguments, and file handling for the agent and runtime you selected.


