How to Bulk Screenshot Indian SaaS Pricing Pages for a Market Research Report
Build a repeatable, auditable workflow for capturing Indian SaaS pricing pages, reviewing failures, and preserving evidence for a market research report.
For a comparable market research report, capture a canonical list of public pricing-page URLs with one documented browser profile, save each result under a stable filename, and keep a manifest that records the URL, UTC timestamp, region, viewport, capture mode, and outcome. Playwright is a practical do-it-yourself option: loop over the URL list, navigate, and save a screenshot. Review every image and log failures; a saved file alone does not prove that the page showed the intended Indian customer experience.
Choose viewport or full-page capture based on the evidence you need, and record consent, login, or bot-block states rather than silently altering research evidence. Playwright documents page screenshots and full-page capture in its screenshot guide and Page API.
1. Prepare a canonical URL inventory
Start with one row per company or product pricing page. Resolve redirects and confirm the destination is the pricing page you intend to compare. Give each row a stable ID so the report can trace a screenshot back to its source.
| Field | Purpose |
|---|---|
id |
Stable filename prefix, such as acme-cloud. |
company |
Company or product name used in the report. |
url |
Canonical public pricing-page URL. |
region |
Capture location or proxy region, if selectable; record unknown if not available. |
viewport |
Dimensions or named device profile. |
mode |
viewport or full-page. |
captured_at_utc |
Exact capture timestamp. |
image_path |
Saved image location, including the stable ID. |
status |
For example, ok, navigation_error, timeout, or review. |
notes |
Redirect, consent overlay, login wall, bot check, incomplete content, or manual intervention. |
For filenames, use a pattern such as company-product_YYYY-MM-DD_region_viewport_full.png. Avoid putting secrets or personal data into filenames or the manifest.
2. Choose the capture approach
Self-managed Playwright
Use Playwright when you want to control navigation, waits, viewport, output paths, and retry logic, and are prepared to install and maintain a browser runtime. Its screenshot API supports viewport screenshots and full-page screenshots. This gives you control of the workflow, but it does not guarantee that a site will load, allow automation, or serve the same content as it does to a visitor in India.
Hosted capture service
A hosted option can reduce browser-runtime maintenance if its batch controls, location options, result reporting, and storage fit your project. ScreenshotNeo provides a screenshot API and MCP server. The research sources also describe hosted capture offerings: the Apify bulk capture Actor describes per-URL results and errors, viewport presets, and proxy support; Capture.page describes hosted screenshots and browser sessions; Add Screenshots describes screenshot workflows and API access. These are provider descriptions, not independent comparisons. Before selecting a service, verify current pricing, region availability, URL limits, retention, export, retry behavior, and compatibility with the pages in your inventory.
| Decision | Questions to answer |
|---|---|
| Control | Can you set navigation, interactions, waits, viewport, and output format as required? |
| Batch results | Does each URL receive a clear success or error result? What are the concurrency and retry limits? |
| India representation | Can you choose a location suitable for your research question, and can you record it? |
| Repeatability | Can you save a profile and rerun it consistently or on a schedule? |
| Evidence handling | Where are images stored, how long are they retained, and can you export them? |
| Operations | Would you rather maintain a browser runtime or use a hosted interface? |
3. Standardize the capture profile
Decide these settings before the batch and keep them fixed across the inventory unless a documented exception is necessary:
- Viewport: choose explicit desktop dimensions, a mobile profile, or both. Do not compare pages captured at different widths without labeling that difference.
- Capture mode: viewport captures preserve the initial visible view; full-page captures can show content farther down the page but may be very tall and can interact differently with lazy-loaded sections. Record which mode each image uses.
- Format: PNG preserves sharp text and interface detail; JPEG is often smaller but lossy. Use one format consistently where the workflow supports it.
- Wait condition: use a navigation or page-ready condition, then a bounded pause only when needed. “Network idle” can be unsuitable for pages that keep background requests open. A fixed delay can still miss slow or lazy content.
- Lazy content: if below-the-fold content matters, consider a documented scroll-and-wait procedure before a full-page capture. Inspect the result for placeholders or missing sections.
- Location: record the actual selected capture region or proxy location. Do not assume an arbitrary server location represents an Indian visitor.
- Consent and overlays: retain the state relevant to the research question. If you accept a banner or remove an overlay, record that intervention because it changes the evidence.
There is no universally correct profile: choose settings that match the report’s intended audience and disclose them. A region control does not establish that every Indian buyer sees the same currency, tax, plan availability, or page variant.
4. Capture a URL list with Playwright
The following Node.js script reads a CSV with id,company,url columns, captures one viewport image per row, and writes a JSON Lines manifest. It records failures per URL and continues the batch. It uses Playwright’s Chromium browser. Install Node.js, then install Playwright and its browser:
npm init -y
npm install playwright
npx playwright install chromium
Save the following as capture-pricing.mjs. The small CSV reader supports quoted fields and escaped quotes; keep the input to these three columns, and avoid embedded line breaks inside values.
import { chromium } from 'playwright';
import { readFile, mkdir, appendFile } from 'node:fs/promises';
const inputPath = process.argv[2] ?? 'pricing-urls.csv';
const outputDir = process.argv[3] ?? 'captures';
const manifestPath = `${outputDir}/manifest.jsonl`;
const width = 1440;
const height = 1000;
const mode = 'full-page'; // Set to false below for viewport-only images.
const timeoutMs = 45000;
function parseCsvLine(line) {
const fields = [];
let value = '';
let quoted = false;
for (let i = 0; i < line.length; i++) {
const char = line[i];
if (quoted && char === '"' && line[i + 1] === '"') {
value += '"'; i++;
} else if (char === '"') {
quoted = !quoted;
} else if (char === ',' && !quoted) {
fields.push(value); value = '';
} else {
value += char;
}
}
fields.push(value);
return fields;
}
const lines = (await readFile(inputPath, 'utf8')).split(/\r?\n/).filter(Boolean);
if (lines.length < 2) throw new Error('CSV needs a header and at least one URL row');
const headers = parseCsvLine(lines[0]).map(x => x.trim());
const rows = lines.slice(1).map(line => {
const cells = parseCsvLine(line);
return Object.fromEntries(headers.map((key, i) => [key, (cells[i] ?? '').trim()]));
});
if (!headers.includes('id') || !headers.includes('url')) {
throw new Error('CSV headers must include id and url');
}
await mkdir(outputDir, { recursive: true });
const browser = await chromium.launch({ headless: true });
const context = await browser.newContext({ viewport: { width, height } });
try {
for (const row of rows) {
const startedAt = new Date().toISOString();
const safeId = (row.id || 'missing-id').replace(/[^a-zA-Z0-9_-]/g, '_');
const imagePath = `${outputDir}/${safeId}_${startedAt.slice(0, 10)}_unknown_${width}x${height}_full.png`;
const record = {
id: row.id ?? '', company: row.company ?? '', url: row.url ?? '',
captured_at_utc: startedAt, region: 'unknown',
viewport: `${width}x${height}`, mode, image_path: imagePath,
status: 'error', notes: ''
};
let page;
try {
const target = new URL(row.url);
if (!['http:', 'https:'].includes(target.protocol)) throw new Error('URL must use http or https');
page = await context.newPage();
const response = await page.goto(target.href, { waitUntil: 'domcontentloaded', timeout: timeoutMs });
await page.waitForTimeout(1500); // Bounded settling pause; adjust and document if required.
record.final_url = page.url();
record.http_status = response?.status() ?? null;
await page.screenshot({ path: imagePath, fullPage: mode === 'full-page', type: 'png' });
record.status = response && response.status() >= 400 ? 'review' : 'ok';
if (record.status === 'review') record.notes = `HTTP ${response.status()}`;
} catch (error) {
record.error = String(error?.message ?? error);
record.status = record.error.toLowerCase().includes('timeout') ? 'timeout' : 'navigation_error';
} finally {
await page?.close();
}
await appendFile(manifestPath, `${JSON.stringify(record)}\n`);
console.log(`${record.status}: ${record.url}`);
}
} finally {
await context.close();
await browser.close();
}
Create pricing-urls.csv:
id,company,url
sample-one,Sample One,https://example.com/pricing
sample-two,Sample Two,https://example.org/plans
Run it:
node capture-pricing.mjs pricing-urls.csv captures
The script labels the region unknown because a local Playwright browser does not automatically represent India. Run it from a location appropriate to the study or use a documented region/proxy setup if available. The sample creates a fresh page per URL, captures in sequence, and appends one manifest record per attempt. To use viewport-only images, set mode to viewport; the screenshot call checks for full-page. For a scheduled study, use a fresh output directory or include a run ID in paths so later captures do not overwrite earlier evidence.
5. Or skip the browser setup
ScreenshotNeo accepts a URL in one GET request and returns an image or PDF. This example saves a WebP capture of a pricing page; replace the URL with an entry from your inventory. See the ScreenshotNeo API documentation for request parameters.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
For a batch, loop over the canonical URL list, save each response with its stable ID, and write a manifest row with the response status and capture settings. ScreenshotNeo accepts parameters used by other screenshot APIs, supports bulk capture of up to 100 URLs per call, and provides a usage API. Select and record the location and capture parameters that match your study; do not infer a capture region from an image alone.
ScreenshotNeo removes known cookie/consent banners, newsletter popups, and chat widgets before capture, and each step can be turned off. That behavior is useful for clean page imagery but can alter research evidence: if consent UI or popups are part of the subject, turn the relevant cleanup off and document the setting. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed; response headers identify page verdict and billing state. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for AI agents. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Every feature is on every plan. Create a free account at ScreenshotNeo sign-up.
6. Review results and preserve evidence
After the run, check that every inventory row has a manifest result and expected image file. Inspect captures for:
- blank or partially rendered pages and loading placeholders;
- redirects to a home page, region selector, sign-in, or unrelated destination;
- bot checks or CAPTCHA screens;
- consent banners, popups, or chat overlays that obscure plan details;
- cut-off plan columns, missing below-the-fold sections, or lazy content that did not load;
- currency, billing period, or locale that differs from the report’s intended audience.
Record failed attempts and keep their original result. Retry only with a documented change, such as a longer bounded wait or a different capture location, and retain both attempts. A screenshot documents what the capture returned at that time and under those settings; it does not prove that every buyer saw the same page or could purchase at the displayed price.
When a screenshot supports a pricing claim, transcribe the visible plan name, currency, billing period, and observation date into the research dataset. Do not infer a price or tax treatment from a cropped, unreadable, or ambiguous image.
7. Performance, reliability, and cost
Batch throughput
The sample processes URLs sequentially, which is simple to audit and avoids creating many browser pages at once. Larger inventories take longer. If you add concurrency, cap it, preserve a separate result for every URL, and avoid overloading target sites. Browser startup, page load, and full-page image size all affect runtime; the sources do not establish a universal throughput figure.
Reliability
Use bounded navigation timeouts and per-URL error handling so one failure does not discard the whole run. A retry should be selective and logged, not an unrecorded replacement. Browser automation may encounter authentication, consent state, bot defenses, or regional delivery differences. A successful HTTP response or image file is not equivalent to a verified pricing page.
Cost and storage
Self-managed Playwright avoids per-shot service charges but uses compute, storage, and maintenance time. Hosted services may charge by usage or plan; compare current plan limits and retention before committing, because this research does not provide an independent price comparison across providers. Store only the captures and metadata needed for the report, and decide how long to retain them. For ScreenshotNeo, the stated plan prices are Free: 1,000 shots/month; Starter: $5 for 3,000; Growth: $15 for 15,000; Pro: $39 for 60,000; Scale: $99 for 250,000; Business: $249 for 1,000,000. Yearly billing gives two months free. Only clean shots are billed, so failed loads and cache hits do not consume billed shots.
8. Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| Browser executable missing | Playwright package installed without its browser binary. | Run npx playwright install chromium in the project environment. |
| Navigation timeout | Slow page, stalled request, or a wait condition that never completes. | Keep a bounded timeout, try a documented navigation condition such as domcontentloaded, and selectively retry. Inspect the resulting state instead of silently treating it as success. |
| HTTP 403, CAPTCHA, or bot page | The site blocks or challenges the automated request. | Record the block and capture context. Do not claim the screenshot reflects the pricing page; use an authorized manual review route if needed. |
| Wrong page in image | Redirect, locale routing, or region-specific delivery. | Record the final URL and selected location, then verify manually from the intended market context. |
| Plans missing below the fold | Viewport capture, lazy loading, or content that needs scrolling. | Use full-page mode or a documented scroll-and-wait routine, then inspect the output. |
| Consent banner obscures price | The banner is part of the page state captured. | Record it. If you dismiss it, note the interaction and preserve the original state when it matters to the research question. |
| Manifest row exists but image is absent | Navigation failed before screenshot output or file writing failed. | Check the row’s error and image path, correct filesystem permissions or path issues, then retry as a new logged attempt. |
| CSV row parsed incorrectly | Malformed quoting, embedded newline, or missing required field. | Use valid CSV quoting, keep one record per line, and ensure id and url headers exist. |
| Screenshot API returns an error or unexpected file | Invalid key, URL, unsupported option, or a page verdict such as a block or blank page. | Check the response and service documentation, record the verdict/billing headers where available, and review the image before including it. |
9. Short FAQ
Should a market report use full-page screenshots?
Use them when below-the-fold plans and details matter, but keep the mode consistent and inspect for lazy-loaded content. A viewport capture is often easier to compare as a first-screen view.
Does a screenshot prove that a price is available to every Indian customer?
No. It records one page response under one capture time, location, and browser state. Verify commercial terms separately when the report makes a purchase-availability claim.
Can I remove cookie banners from report evidence?
Only when that matches the research question. Record the state change; a removed banner cannot serve as evidence of the page as initially presented.
How should I compare mobile and desktop pricing?
Capture them as separate, labeled profiles with the same viewport settings across all products in each profile.


