ScreenshotNeo

BlogHow-to

How to Schedule Screenshots of an Indian Insurance Portal Without Exposing Customer Data

Schedule portal screenshots with a narrow capture, masking before saving, and controls for credentials, storage, logs, and retention.

By the ScreenshotNeo team4 October 202610 min read

To schedule screenshots of an Indian insurance portal without exposing customer data, first confirm that the portal owner permits the automation and that your organization approves the purpose and recipients. Then minimize the page area and fields, mask known sensitive elements during capture, store the image and related artifacts in approved restricted storage, and verify representative outputs before enabling unattended runs. If a purpose-built report or sanitized page can meet the need, use it instead of creating a raw screenshot.

A screenshot mask only covers the elements you identify. It is not proof that every sensitive value has been found or removed. Treat screenshots, logs, traces, filenames, alerts, and temporary capture data as potentially sensitive artifacts.

1. Confirm permission and scope

Before writing automation, record the portal owner, the organization responsible for the account, the capture purpose, the minimum information needed, who may receive the output, and the approved storage and retention path. Check the portal’s terms and get the required organizational approval. The available sources do not establish permission to automate any particular insurer portal or account.

Insurance policy details can be available electronically through online portals, so do not assume a portal view is harmless simply because it is accessible after sign-in. IRDAI describes electronic policy records and portal access on its insurance repositories page. IRDAI’s guidelines index lists Information and Cyber Security Guidelines dated 2 September 2022; consult the underlying instrument and subsequent updates for requirements applicable to your organization. The Government of India’s summary of the DPDP Rules, 2025 describes the Rules and Act as a framework for personal-data use. Check the statutory text, commencement, applicability, amendments, and your organization’s advice rather than treating a general summary as a legal conclusion.

2. Prefer source-level minimization

  1. Ask whether an approved report, export, or purpose-built page can provide the needed record without customer-specific fields.
  2. Use a synthetic or demo account for development if the portal owner provides one.
  3. If a browser image is necessary, identify the smallest relevant element or viewport. Avoid full-page capture unless the task requires it.
  4. List sensitive fields that could appear, including policyholder names, contact details, policy or claim identifiers, addresses, payment details, and values rendered dynamically.

Source-level minimization avoids creating an unnecessary raw image. Post-processing a raw image leaves an additional unredacted artifact to protect and delete.

3. Capture and mask with Playwright

Playwright’s Page screenshot API supports masking locators with an overlay and applying a stylesheet during capture. Locator APIs support role, text, label, placeholder, and test-id selectors; prefer stable selectors grounded in the portal’s accessible labels or an explicit test contract. This capability does not establish that a particular portal permits Playwright, that its layout is compatible, or that masking covers every sensitive value.

Install Playwright and its Chromium browser in an approved environment:

npm init -y
npm install playwright
npx playwright install chromium

Save the following as capture.mjs. It takes a selector from the environment, authenticates with a storage-state file created through your approved process, masks known fields at capture time, and writes only the resulting image. The sample assumes the page’s selectors and authentication state have been validated for the authorized portal. Replace the example URL and selectors; do not put credentials directly in source code.

import { chromium } from 'playwright';

const portalUrl = process.env.PORTAL_URL;
const sectionSelector = process.env.SECTION_SELECTOR;
const outputPath = process.env.OUTPUT_PATH ?? 'portal-section.png';
const storageStatePath = process.env.STORAGE_STATE_PATH;

if (!portalUrl || !sectionSelector || !storageStatePath) {
  throw new Error('Set PORTAL_URL, SECTION_SELECTOR, and STORAGE_STATE_PATH');
}

const browser = await chromium.launch({ headless: true });
try {
  const context = await browser.newContext({
    storageState: storageStatePath,
    viewport: { width: 1280, height: 900 },
  });
  const page = await context.newPage();
  await page.goto(portalUrl, { waitUntil: 'domcontentloaded', timeout: 45_000 });

  const section = page.locator(sectionSelector);
  await section.waitFor({ state: 'visible', timeout: 20_000 });

  // Replace these with selectors for every sensitive field that might render.
  const masks = [
    page.getByTestId('policyholder-name'),
    page.getByTestId('policy-number'),
    page.getByTestId('customer-contact'),
  ];
  for (const locator of masks) {
    if (await locator.count()) await locator.first().waitFor({ state: 'visible', timeout: 5_000 }).catch(() => {});
  }

  await section.screenshot({
    path: outputPath,
    type: 'png',
    mask: masks,
    maskColor: '#000000',
    style: '[data-customer-sensitive] { visibility: hidden !important; }',
    animations: 'disabled',
    timeout: 30_000,
  });
  await context.close();
} finally {
  await browser.close();
}

The data-testid values and CSS attribute above are examples, not selectors known to exist on an insurer portal. Inspect the authorized page structure and replace them with selectors that have been reviewed. The sample waits for the capture section, but it cannot determine whether the page is the right customer’s record or whether a selector still targets the intended field after a redesign. Validate those conditions explicitly before scheduling.

Masking details and edge cases

  • Mask every known sensitive field. Add masks for identifiers and values that may appear in headings, tables, summaries, tooltips, or repeated rows. A mask affects matching elements; it cannot find information you have not identified.
  • Use a high-contrast opaque mask. The black overlay in the sample makes the selected region unreadable. Confirm the resulting pixels in a reviewed test capture.
  • Handle optional and repeated fields deliberately. A field may be absent on one policy and present on another. Repeated matches may require masking every locator rather than only the first. Confirm the installed Playwright version’s locator and screenshot behavior before relying on a pattern.
  • Consider frames and shadow content. Page-level selectors may not cover content in nested frames or components that need a different locator strategy. Include those cases in review; do not assume a page-wide selector sees every rendered value.
  • Wait for data to settle. A page can render shell content before policy data loads. Wait for a reliable, authorized page-state indicator, then capture. Avoid using a fixed delay as the only readiness check.
  • Do not save an unmasked fallback. If a selector fails or the capture step errors, fail the job and alert without attaching the page or image. Keep diagnostic output free of page text and customer values.
  • Keep full-page capture exceptional. More page area can include hidden-in-plain-sight sidebars, footer details, or additional records. Use the element screenshot shown above when it meets the purpose.

4. Schedule the job safely

Run the script with your organization’s approved scheduler or managed workflow. The scheduler is environment-specific; the Playwright documentation does not prescribe one or establish that a particular deployment is suitable. Apply these controls:

  • Keep browser credentials and storage state in the approved secret-management system, with narrowly scoped access and rotation. Do not commit them or print them to logs.
  • Give the job only the permissions it needs to reach the portal and write to the approved destination.
  • Use an explicit output directory with restricted access. Apply the organization’s approved encryption, transfer, and retention controls.
  • Prevent overlapping runs if they can reuse or overwrite output paths. Use unique, non-identifying filenames; avoid customer names and policy numbers in paths.
  • Alert on failures with a run identifier and error category, not a screenshot, page dump, URL containing personal data, or authentication material.
  • Review browser traces, crash dumps, temporary files, scheduler logs, and notification payloads under the same access and retention rules as the screenshot.
  • Set a purpose-based retention period and verify deletion, including backups and derived artifacts, according to approved organizational policy.

Start with a manual run in a non-production or synthetic environment where possible. Inspect successful captures, empty results, session expiry, access-denied pages, and portal errors. Repeat review after portal changes and whenever a selector, authentication flow, or page layout changes.

5. Review before unattended capture

  1. Confirm the automation is authorized and the test account or record is appropriate.
  2. Verify that the captured region contains only the intended information.
  3. Check that each mask covers the full value at normal and high-resolution rendering.
  4. Inspect empty, error, login-expiry, and changed-layout states. Confirm none can be mistaken for a valid record.
  5. Check image files and by-products, including logs and alerts, for accidental customer data.
  6. Have the appropriate owner approve the sample and the output destination before turning on the schedule.

These checks are operational recommendations. A screenshot API or browser library cannot certify that the capture is legally compliant or that every sensitive field has been removed.

6. Troubleshooting

Symptom Likely cause Fix
Login page appears instead of the portal Expired or invalid storage state, changed sign-in flow, or session restrictions. Refresh authentication only through the approved process; test session expiry explicitly. Do not put passwords or one-time codes in the script or logs.
Screenshot is blank or missing policy data The page shell loaded before its data, the wrong section selector was used, or the session lacks access. Wait for a stable authorized page-state indicator; verify the selector and account scope with a reviewed manual run. Treat blank output as a failed capture.
Customer information remains visible A field was not included in the mask, the selector changed, content is in a frame, or another part of the page exposes the value. Stop unattended runs, restrict access to existing output, investigate with approved access, update selectors, and review new samples before resuming. Do not assume masking was complete.
Mask selector times out The field is conditional, late-loading, absent, or has changed markup. Define which fields are required and optional. Validate the expected page state and fail closed if a required mask target is absent.
Capture times out intermittently Network delay, portal load variation, or a wait condition that never becomes true. Wait for a specific readiness condition, set bounded timeouts, and retry only under an approved policy. Avoid unbounded retries that may increase load or create duplicate artifacts.
Unexpected or stale image is picked up Fixed output path, overlapping executions, or a downstream process reading before capture completes. Use per-run temporary paths with restrictive permissions, publish only after success, and serialize runs where needed.
Alerts or logs contain sensitive data Full URL, page content, exception details, trace, or attached screenshot was included in diagnostics. Reduce diagnostics to an opaque run ID and sanitized error category; remove attachments and review retention for existing artifacts.

7. Performance, reliability, and cost

Capture only the needed element or viewport to reduce browser work, image size, and the amount of data handled. Full-page images and high-resolution output require more rendering and storage; use them only when needed. Wait for a defined page condition rather than adding a long fixed sleep to every run.

Portal availability, login expiry, dynamic content, and layout changes can make scheduled captures fail or become misleading. Use bounded timeouts, clear failure states, controlled retry behavior, and human review after changes. A successful image write does not prove the right page was captured or that masking was complete.

For a self-hosted browser workflow, account for browser compute, approved secret and storage services, monitoring, and review time. Exact costs depend on your infrastructure and schedule; the research available here does not establish a benchmark or cost estimate. Make sure retries and retention do not silently create extra copies.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. One GET request can return a PNG, JPEG, WebP, or PDF. Its consent-banner, popup, and chat-widget removal can be turned off step by step. Those cleanup features do not replace authorization, data minimization, or validation of customer-specific fields; do not send sensitive portal data to a service unless your organization has approved that processing and configuration.

For an authorized page that is appropriate to send to the API, this is the one-call pattern. See the ScreenshotNeo API documentation for options and configuration. Do not place customer credentials or sensitive query parameters in a URL unless approved.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed, and response headers indicate the page verdict and billing status. An MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Every feature is on every plan. For an insurance portal, first confirm approved use, access controls, and data handling; general website cleanup is not a guarantee that customer data is masked.

Sign up for 1,000 free screenshots a month, with no card.

FAQ

Does a black mask prove the screenshot is anonymized?

No. It covers matched elements. Unidentified fields, selector changes, frames, or other page regions can still expose data. Review the actual output and its surrounding artifacts.

Can I use this workflow for any Indian insurer portal?

The general workflow is not portal approval. Confirm the portal’s terms, account authorization, organizational requirements, and applicable current obligations for the specific use.

Should I capture the whole page to preserve context?

Only if the purpose requires it. Prefer a narrowly scoped element or a source-generated report that omits unnecessary customer details.

Can I attach failed screenshots to an alert for debugging?

Do not attach them by default. Alerts can create additional copies and recipients. Use sanitized error categories and inspect sensitive artifacts only through an approved, restricted process.