ScreenshotNeo

BlogHow-to

How to Bulk Screenshot Coaching Institute Websites for a Competitor Audit in India

Build a repeatable screenshot audit of coaching institute websites with Playwright, consistent capture settings and a clear comparison log.

By the ScreenshotNeo team4 October 202610 min read

To bulk screenshot coaching institute websites for a competitor audit, make a bounded list of public pages, check each site’s current terms and automation guidance, then use Playwright to capture equivalent pages at a consistent viewport. Save each image with a stable filename and record its URL, capture time, browser, viewport, and state in a CSV. Compare like page types and separate what is visibly present from your interpretation.

This workflow is for a proportionate visual audit, not a full-site crawl. A screenshot records what rendered in one browser at one time; it does not prove what every visitor saw or what the site looks like later.

1. Define the sample before capturing

Start with the audit question: for example, how institutes explain course offerings, present faculty credentials, disclose fees, or invite prospective students to enquire. Include only pages that answer it. A typical sample might include the home page, one course page, faculty, fees or admissions, and contact pages for each institute.

Field What to record
Institute The organization or brand name as shown on the sampled page.
Page type Home, course, faculty, fees, admissions, contact, or another defined category.
Canonical URL The exact public URL you intend to capture, including relevant path and query parameters.
Sample reason The audit question this page helps answer.
Interaction Initial state, or a specific simple interaction such as opening a navigation menu.

Keep page types comparable between institutes. If a page is unavailable, record that it was not found in the sampled pages; do not infer that the institute has no such information.

2. Check site rules and keep the capture proportionate

Review the current terms and published automation guidance for each target before using a script. Site rules differ. For example, the cited BrainBuzz Academy terms prohibit automated tools, bots, or scripts to access its platform, while Playwright Masters publishes its own terms. These examples do not establish rules for other sites. Public visibility alone is not blanket permission.

  • Capture only the public pages needed for the audit and space requests.
  • Avoid login-protected areas and pages containing personal data.
  • Do not try to defeat a CAPTCHA, bot check, access restriction, or other challenge. Stop if blocked.
  • Seek permission where required. This guide is not a legal opinion about a particular audit or website in India.

3. Set up a reproducible Playwright capture

Playwright can automate browser navigation and save ordinary, full-page, buffer, and element screenshots. Its browser automation supports Chromium, Firefox, and WebKit. The example below uses Chromium in headless mode and captures each URL in a CSV at the same viewport.

Install the dependencies

mkdir coaching-audit
cd coaching-audit
npm init -y
npm install playwright
npx playwright install chromium

Create urls.csv with a header row. Quote CSV values when they contain commas, and use canonical public URLs:

institute,page_type,url
Example Institute,home,https://example.org/
Example Institute,course,https://example.org/courses/example

Runnable Node.js script

Save as capture.mjs. The script saves full-page PNGs and a metadata CSV. It uses one browser and a fresh page for each URL, waits for DOM content and then a short settling interval, and records failures rather than stopping the whole batch.

import { chromium } from 'playwright';
import { readFile, writeFile, appendFile, mkdir } from 'node:fs/promises';

const inputPath = process.argv[2] ?? 'urls.csv';
const outputDir = process.argv[3] ?? 'captures';
const viewport = { width: 1440, height: 1000 };
const timestamp = new Date().toISOString().replaceAll(':', '-');

function parseCsvLine(line) {
  const values = [];
  let value = '';
  let quoted = false;
  for (let i = 0; i < line.length; i++) {
    const ch = line[i];
    if (ch === '"' && quoted && line[i + 1] === '"') {
      value += '"'; i++;
    } else if (ch === '"') {
      quoted = !quoted;
    } else if (ch === ',' && !quoted) {
      values.push(value); value = '';
    } else {
      value += ch;
    }
  }
  values.push(value);
  return values;
}

function csvEscape(value) {
  return `"${String(value ?? '').replaceAll('"', '""')}"`;
}

function slug(value) {
  return String(value).toLowerCase().replace(/[^a-z0-9]+/g, '-').replace(/^-|-$/g, '') || 'page';
}

const lines = (await readFile(inputPath, 'utf8')).split(/\r?\n/).filter(line => line.trim());
if (lines.length < 2) throw new Error('CSV needs a header and at least one data row');
const headers = parseCsvLine(lines[0]).map(v => v.trim());
const rows = lines.slice(1).map(line => {
  const values = parseCsvLine(line);
  return Object.fromEntries(headers.map((header, index) => [header, values[index] ?? '']));
});
for (const required of ['institute', 'page_type', 'url']) {
  if (!headers.includes(required)) throw new Error(`Missing CSV column: ${required}`);
}
await mkdir(outputDir, { recursive: true });
const metadataPath = `${outputDir}/metadata-${timestamp}.csv`;
const metadataHeader = ['institute', 'page_type', 'url', 'captured_at_utc', 'viewport', 'browser', 'status', 'file', 'notes'];
await writeFile(metadataPath, `${metadataHeader.join(',')}\n`);

const browser = await chromium.launch({ headless: true });
try {
  for (const [index, row] of rows.entries()) {
    let status = 'ok';
    let file = '';
    let notes = '';
    const capturedAt = new Date().toISOString();
    try {
      const url = new URL(row.url);
      if (!['http:', 'https:'].includes(url.protocol)) throw new Error('URL must use http or https');
      const page = await browser.newPage({ viewport });
      try {
        const response = await page.goto(url.href, { waitUntil: 'domcontentloaded', timeout: 45000 });
        await page.waitForTimeout(1500);
        const code = response?.status();
        if (code && code >= 400) notes = `HTTP ${code}`;
        file = `${String(index + 1).padStart(3, '0')}-${slug(row.institute)}-${slug(row.page_type)}-${timestamp}.png`;
        await page.screenshot({ path: `${outputDir}/${file}`, fullPage: true, animations: 'disabled', timeout: 45000 });
      } finally {
        await page.close();
      }
    } catch (error) {
      status = 'error';
      notes = String(error?.message ?? error).replaceAll('\n', ' ');
    }
    const values = [row.institute, row.page_type, row.url, capturedAt, `${viewport.width}x${viewport.height}`, 'Chromium (Playwright)', status, file, notes];
    await appendFile(metadataPath, `${values.map(csvEscape).join(',')}\n`);
    console.log(`${index + 1}/${rows.length}: ${status} ${row.url}${notes ? ` — ${notes}` : ''}`);
    // A brief pause makes the batch less aggressive. Increase it for larger audits.
    await new Promise(resolve => setTimeout(resolve, 1000));
  }
} finally {
  await browser.close();
}
console.log(`Metadata: ${metadataPath}`);

Run it with node capture.mjs urls.csv captures. The script’s viewport, navigation timeout, settling delay, and spacing are starting settings, not universal values. Adjust them consistently for every institute in the comparison. The fixed 1.5 second settling interval is not proof that all content has loaded; inspect outputs and document any site-specific wait condition needed.

Use a new timestamped run directory or retain the timestamp in filenames when taking repeat snapshots. That keeps historical captures from silently overwriting one another.

4. Choose capture state and dimensions deliberately

Viewport or full page

A viewport screenshot is useful for comparing the first screen and above-the-fold message at a fixed width and height. A full-page capture reveals content farther down the page, including sections that may otherwise be missed. For audit evidence, it can help to save both when the distinction matters. Full-page output can be very tall, and sticky or lazy-loaded elements may behave differently in a long capture.

Element capture

When the question concerns one component, such as a pricing panel or enquiry form, capture that element rather than comparing entire pages. Playwright’s official screenshot guide documents element-level screenshots.

const element = page.locator('main');
await element.screenshot({ path: 'main-section.png' });

Replace main with a selector that exists on the target. Selectors differ between sites, so check that the locator matches the intended content before relying on the image. An absent or ambiguous selector should be recorded as a capture issue rather than silently treated as comparable evidence.

Interactions and waits

Capture the initial state consistently. If the audit needs a menu, tab, or other simple interaction, record the interaction and save it as a separate state with a distinct filename. Use a selector wait for a known page element when a site renders asynchronously; use a delay only when there is a clear reason. Do not try to bypass access challenges. Record the wait condition in your audit notes so reviewers understand what they are seeing.

await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 45000 });
await page.locator('main').waitFor({ state: 'visible', timeout: 15000 });
await page.screenshot({ path: 'page.png', fullPage: true });

5. Organize and compare the evidence

Use a stable naming pattern such as institute_page-device_YYYY-MM-DD.png. Keep a spreadsheet or CSV with institute, page type, URL, capture timestamp in UTC, viewport dimensions, browser, status, interaction or wait condition, filename, and notes. The metadata is part of the audit workflow; Playwright does not automatically create this audit log.

Compare equivalent page types at the same viewport. Useful axes include:

  • First-screen headline, offer, and primary call to action.
  • Course and exam coverage, as shown on the sampled page.
  • Faculty credentials and visible credibility evidence.
  • Delivery mode and location.
  • Fees, financing disclosures, or enquiry steps where visible.
  • Mobile presentation, captured at the same mobile dimensions for each site.
  • Visible trust and policy cues.

Keep two columns in your notes: visible fact and interpretation. For example, “the sampled page shows a phone number above the fold” is an observation; “the institute makes it easy to enquire” is an interpretation. Label missing details as “not found on sampled page.” Avoid ranking claims when the pages, states, or capture conditions are not comparable.

6. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. One GET request can return a screenshot in PNG, JPEG, or WebP, or a PDF. Its capture options include full-page screenshots, viewport and device settings, element selectors, custom waits, and bulk capture of up to 100 URLs per call. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.org/ \
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.org/"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.org/'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) =>
  writeFile('shot.webp', Buffer.from(await res.arrayBuffer()))
);

Cookie and consent banners are accepted like a visitor and removed along with 60+ known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks, blank pages, and failed loads are never billed, and response headers identify the page verdict and billing status. An MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. All features are available on every plan. Use the service only where the target site’s rules allow your capture.

Sign up for 1,000 free screenshots a month, with no card required.

7. Troubleshooting

Symptom Likely cause What to do
Browser executable missing Playwright package is installed but its browser was not installed. Run npx playwright install chromium.
Navigation timeout The page is slow, waits indefinitely, or has a network issue. Check the URL and access manually, choose an appropriate timeout, and use a documented wait condition. Do not treat a timeout as a valid screenshot.
Screenshot is blank or incomplete The page may need more rendering time, a visible selector, or an interaction; a challenge or failed load may also be present. Inspect the page and response, adjust the wait for permitted ordinary rendering, and stop on challenges or blocks.
Full-page image is unusually tall The page contains long content, repeated sections, or lazy content. Use a viewport capture for first-screen comparisons, or capture a relevant element. Keep the chosen method consistent across the sample.
Images or fonts differ between runs Remote assets, personalization, network state, or site content changed. Retain timestamps, browser and viewport metadata; repeat only when useful and annotate differences.
CSV row fails or fields shift Commas or quote marks in a field were not escaped correctly, or the header does not match. Quote CSV fields containing commas and preserve the required institute, page_type, and url headers.
Target refuses or challenges automation Site policy or bot protection blocks the request. Stop. Check current terms and seek permission where appropriate; do not attempt to evade the restriction.

8. Performance, reliability, and cost

A local Playwright batch consumes your machine’s browser and network resources. Reusing one browser and opening one page at a time keeps the example simple and avoids adding concurrent load to target sites. For a larger allowed sample, space requests further apart and consider splitting work into smaller batches. Parallelism can reduce elapsed time, but it raises resource use and request volume; only use it when appropriate for the site and audit.

Reliability comes from preserving the capture context, handling each URL independently, keeping failed rows in the log, and inspecting the images. A successful file write does not establish that the right content rendered. Browser versions, responsive layout, geolocation, cookies, personalization, and time-dependent content can affect the result. Fix the viewport and browser configuration within a comparison run and record changes between runs.

The local script has no per-screenshot service charge, but uses local compute and bandwidth and takes setup and maintenance. ScreenshotNeo has a free plan of 1,000 shots per month without a card; paid plans are Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000. Yearly billing gives two months free. Only clean shots are billed; cache hits cost nothing. Select a workflow based on the number of permitted captures, need for repeat runs, and how much browser maintenance you want to handle.

FAQ

Can I infer an institute’s full strategy from one screenshot?

No. A screenshot is a dated record of a page in a particular browser state. Sample equivalent pages and treat conclusions as bounded by that evidence.

Should I capture every page on every institute website?

No. Define the question and sample only the public pages needed to answer it. A bounded sample is easier to compare and document.

Can I compare desktop and mobile layouts?

Yes. Capture each site at the same desktop dimensions and again at the same mobile dimensions, and label the two sets distinctly.

Does a public page mean automated capture is allowed?

No. Check the current terms and automation guidance for each target. Rules differ, and this workflow does not determine legality for a particular site or audit in India.

Primary references