ScreenshotNeo

BlogHow-to

How to Automate Screenshots of Indian School Websites with Playwright

Capture authorized Indian school websites with Playwright. Learn viewport, full-page and element screenshots, batch runs, troubleshooting and privacy basics.

By the ScreenshotNeo team4 October 20267 min read

Use Playwright to open each authorized school website in a fresh browser context, wait for the page state your capture needs, then save a viewport, full-page, or element screenshot. The script below uses JavaScript and Chromium, sets a consistent desktop viewport and Indian English locale, and processes a supplied URL list. Create the output directory first and replace the example URLs with pages you are authorized to capture.

Playwright’s page.screenshot() saves screenshots to files or returns image bytes; it supports full-page and element captures. See the Playwright screenshots guide and Page API.

1. Install Playwright and prepare the output folder

Use a current Node.js installation. In an empty project directory, install Playwright and its Chromium browser:

npm init -y
npm install playwright
npx playwright install chromium
mkdir -p screenshots

For the ES module imports used below, either add "type": "module" to package.json or save the script as capture.mjs. Run it with node capture.mjs.

2. Capture a supplied list of school websites

This script creates a fresh context for each site, writes both viewport and full-page PNG files, and closes resources even if navigation or capture fails. The example domains are placeholders.

import { chromium } from 'playwright';

const targets = [
  { name: 'school-a', url: 'https://example.edu.in/' },
  { name: 'school-b', url: 'https://example-school.in/' },
];

const browser = await chromium.launch();
try {
  for (const target of targets) {
    const context = await browser.newContext({
      viewport: { width: 1365, height: 900 },
      locale: 'en-IN',
      timezoneId: 'Asia/Kolkata',
      colorScheme: 'light',
    });
    try {
      const page = await context.newPage();
      const response = await page.goto(target.url, {
        waitUntil: 'load',
        timeout: 45_000,
      });
      if (!response || !response.ok()) {
        console.warn(`${target.name}: navigation status ${response?.status() ?? 'no response'}`);
      }
      await page.screenshot({
        path: `screenshots/${target.name}-viewport.png`,
      });
      await page.screenshot({
        path: `screenshots/${target.name}-full.png`,
        fullPage: true,
      });
    } catch (error) {
      console.error(`${target.name}: ${error.message}`);
    } finally {
      await context.close();
    }
  }
} finally {
  await browser.close();
}

load means the page load event fired; it does not guarantee that all images, animations, or dynamically rendered content are ready. Choose a wait condition that matches the material you need. For a site-specific landmark, wait for a stable selector instead of adding a long fixed delay:

await page.goto(target.url, { waitUntil: 'domcontentloaded', timeout: 45_000 });
await page.locator('main').waitFor({ state: 'visible', timeout: 15_000 });

If the page relies on client-side rendering, wait for the relevant heading, banner, or content container. Avoid assuming that a selector exists on every school site; handle site-specific readiness in the target data or with a per-site condition.

3. Choose viewport, full-page, or element capture

Capture Playwright option Use it for Trade-off
Viewport Default screenshot A consistent first-screen view or visual review at a chosen screen size Content below the viewport is omitted
Full page fullPage: true A long-page record or broad page review Very long pages create large images and can expose more personal information
Element locator(...).screenshot() A specific header, notice, menu, or content block The selector must identify a visible element

For one element, wait until it is visible and capture it directly:

const notice = page.locator('main .notice').first();
await notice.waitFor({ state: 'visible', timeout: 15_000 });
await notice.screenshot({ path: `screenshots/${target.name}-notice.png` });

Use selectors appropriate to the particular site. A selector copied from one school’s markup may not exist on another. You can also capture bytes for an image-processing step instead of writing a file immediately:

const pngBytes = await page.screenshot({ fullPage: false });
// Pass pngBytes to the image-processing or storage code used by your application.

Playwright can select image format by file extension or the type option (PNG, JPEG, or WebP where supported by the API); JPEG and WebP can use a quality value. Transparency is supported for PNG. Consult the screenshot API options for version-specific details.

4. Configure the rendered page consistently

Viewport dimensions determine the page’s responsive layout. Keep them fixed when comparing captures. For mobile emulation, use a documented device descriptor or set viewport and device scale factor in the context; device settings can affect layout and rendering. Playwright also lets you emulate locale, timezone, color scheme, and other browser conditions. These settings influence browser presentation but do not make a site offer a translation it does not provide. See Playwright emulation.

const context = await browser.newContext({
  viewport: { width: 390, height: 844 },
  deviceScaleFactor: 2,
  isMobile: true,
  hasTouch: true,
  locale: 'en-IN',
  timezoneId: 'Asia/Kolkata',
  colorScheme: 'dark',
});

Use a fresh browser context to isolate cookies and local storage between independent runs. This avoids one capture inheriting consent choices or session state from another. Contexts are lightweight browser sessions; see Browser contexts.

5. Make batch runs predictable

  1. Maintain an explicit URL inventory supplied by the site owner or otherwise authorized for capture.
  2. Give each target a stable, filesystem-safe name. Avoid deriving file paths directly from arbitrary URL text.
  3. Choose one viewport, locale, timezone, color scheme, browser, and capture mode for each comparison set.
  4. Use a fresh context per site when you want independent cookies and storage. Reuse a context only when shared state is deliberate.
  5. Record failures per target and continue when one site times out or returns an error.
  6. Limit concurrency if running multiple pages at once; high parallelism can overload your machine or trigger site rate limits.

The sample uses sequential captures, which are easier to debug and place less simultaneous load on your machine. If you add concurrency, keep a small bounded worker pool, respect site restrictions, and ensure every context closes in a finally block.

6. Privacy and authorization

Capture only public pages and visual material needed for the task. Do not enter student or staff portals unless the operator expressly authorized that scope. Review images before sharing: a page may reveal student names, faces, results, contact details, or other personal information.

India’s Digital Personal Data Protection Act, 2023, section 9 addresses processing children’s personal data, including verifiable parental consent and restrictions on detrimental processing, tracking, and behavioural monitoring, subject to statutory qualifications and exceptions. This summary is not a legal determination for a particular capture project. Review the Act text and obtain appropriate guidance for your circumstances.

7. Troubleshooting

Symptom Likely cause What to do
Browser executable missing Playwright package installed without its browser binary Run npx playwright install chromium in the project environment.
Navigation timeout Slow server, stalled resource, or a page that never reaches the selected event Check the URL and network access; increase the timeout only when justified, or use domcontentloaded and wait for the specific content needed.
Screenshot is blank or incomplete Capture happened before client-side content appeared, or an error/blocked page loaded Inspect the response and page, wait for a meaningful selector, and check whether the site requires access you do not have.
Element locator times out Selector is absent, changed, hidden, or matches the wrong element Inspect the page markup, use a stable selector, and confirm visibility before capturing.
Full-page capture misses lazy content Content loads only as the page scrolls Use a site-appropriate scroll-and-wait routine before capture, then verify the resulting image. Do not assume every page responds the same way.
Images differ between runs Different browser, OS, fonts, viewport, color scheme, device scale, or dynamic page content Keep the capture environment and settings consistent, and account for changing content. Playwright describes these sources of screenshot comparison variation in its visual comparisons guidance.
Permission denied writing output Output directory is missing or not writable Create the directory and check the process user’s write access.

8. Performance, reliability, and cost

Browser startup is shared across targets in the example, while each context is isolated and closed after a site. Sequential processing reduces peak resource use and makes failures easier to attribute. Full-page images consume more memory and storage than viewport captures; use element or viewport captures when those contain the information required. Reuse a fixed browser setup for comparable runs, and retain the URL, timestamp, browser version, viewport, and capture mode alongside images if you need an audit trail.

Playwright itself is an open-source browser automation library; operating cost comes from the machine and infrastructure that run it, plus storage and maintenance. This workflow has no per-screenshot API charge, but it requires managing browser installation, execution, failures, and output storage.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. One GET request returns an image or PDF; its API documentation covers request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.edu.in/ -o school.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.edu.in/"},
    timeout=90,
)
r.raise_for_status()
open("school.webp", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.edu.in/',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await (await import('node:fs/promises')).writeFile('school.webp', bytes);
  • Cookie banners are accepted and removed before capture; newsletter popups and chat widgets are removed too. Each of these steps can be turned off.
  • Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers report the page verdict and billing status.
  • An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents.
  • 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000 screenshots.

Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.

FAQ

Does setting en-IN translate a school website?

No. It emulates a browser locale. The site must provide and select its own localized content.

Should every capture use a fresh context?

Use one when runs should have independent cookies and storage. Reuse a context only when retaining state is part of the workflow.

Can I compare screenshots pixel for pixel across machines?

That can be unreliable because browser and host environments affect rendering. Keep the environment consistent and account for dynamic content.