ScreenshotNeo

BlogHow-to

How to Capture Screenshots of Multiple Indian Travel Booking Pages in Bulk

Capture travel booking pages in a repeatable batch with Playwright, consistent framing, unique filenames, and practical checks for reliable comparisons.

By the ScreenshotNeo team4 October 202611 min read

To capture screenshots of multiple Indian travel booking pages in bulk, keep a list of the exact page URLs and use a browser automation script to visit each one, wait for the content you need, and save a uniquely named image. Use the same viewport for comparable snapshots, or choose full-page captures when you need the entire scrollable page. The example below uses Playwright with Node.js and writes one PNG per URL.

This captures the state rendered in the browser at that moment. Results can vary with location, date, login state, selected itinerary, and dynamic loading. Before automating any named booking site, check its current terms and access rules; policies for particular sites are not established here.

1. Choose what each screenshot should show

Decide whether you need a viewport image or a full-page image before you start the batch:

  • Viewport: saves only the currently visible browser area. This is usually the better choice when comparing the same part of several results pages with consistent framing.
  • Full page: captures the page’s scrollable content in one image. This is useful for documentation, but pages with long results or dynamic sections can produce very tall images.
  • Element: captures a selected page element, such as a result card or price summary, when the relevant page has a stable selector. Verify the selector on each page; different sites generally use different markup.

Playwright documents viewport, full-page, and element screenshot capture in its screenshot guide and Page API. Choose one capture mode for the whole batch unless you intentionally record the mode in each filename or in a separate manifest.

2. Prepare the URL list and output folder

Make a text file named urls.txt, with one complete booking-page URL per line. Use pages you are allowed to access and that do not require a workflow your script is not designed to perform. For reproducibility, record the route, search criteria, and capture date separately if they matter to your comparison.

https://example.com/booking-page-one
https://example.com/booking-page-two

The example uses placeholder URLs. Replace them with the specific pages you need to capture. Do not assume that a search-results URL will preserve the same itinerary, location, availability, or prices on a later run.

3. Install Playwright

Install Node.js, create a working directory, and install Playwright. Its browser installation command downloads the browser binaries used by the script. See the official Playwright getting started guide for current prerequisites and setup details.

mkdir travel-page-captures
cd travel-page-captures
npm init -y
npm install playwright
npx playwright install chromium

4. Capture each page with Node.js

Save the following as capture.mjs beside urls.txt. It uses a fixed viewport, visits URLs sequentially, waits for the page load event, allows a short configurable settling interval, and writes a unique numbered PNG for each URL. Sequential capture is a conservative starting point for pages that load substantial content.

import { chromium } from 'playwright';
import { mkdir, readFile } from 'node:fs/promises';

const input = await readFile('urls.txt', 'utf8');
const urls = input
  .split(/\r?\n/)
  .map((line) => line.trim())
  .filter((line) => line.length > 0 && !line.startsWith('#'));

if (urls.length === 0) {
  throw new Error('urls.txt contains no URLs');
}

for (const url of urls) {
  let parsed;
  try {
    parsed = new URL(url);
  } catch {
    throw new Error(`Invalid URL in urls.txt: ${url}`);
  }
  if (!['http:', 'https:'].includes(parsed.protocol)) {
    throw new Error(`Only http and https URLs are supported: ${url}`);
  }
}

const outputDir = 'screenshots';
const fullPage = process.env.FULL_PAGE === '1';
const settleMs = Number.parseInt(process.env.SETTLE_MS ?? '1500', 10);
if (!Number.isFinite(settleMs) || settleMs < 0) {
  throw new Error('SETTLE_MS must be a non-negative integer');
}

await mkdir(outputDir, { recursive: true });
const browser = await chromium.launch({ headless: true });

try {
  const context = await browser.newContext({
    viewport: { width: 1440, height: 1000 },
    deviceScaleFactor: 1,
  });

  for (let i = 0; i < urls.length; i += 1) {
    const url = urls[i];
    const page = await context.newPage();
    const number = String(i + 1).padStart(3, '0');
    const filename = `${outputDir}/${number}.png`;

    try {
      const response = await page.goto(url, {
        waitUntil: 'load',
        timeout: 60_000,
      });
      if (!response) {
        console.warn(`No main-document response for ${url}`);
      } else if (!response.ok()) {
        console.warn(`Main document returned HTTP ${response.status()} for ${url}`);
      }

      // Replace this with a page-specific locator wait when you know
      // which visible element signals that the needed content is ready.
      if (settleMs > 0) {
        await page.waitForTimeout(settleMs);
      }

      await page.screenshot({ path: filename, fullPage });
      console.log(`Saved ${filename} from ${url}`);
    } catch (error) {
      console.error(`Failed to capture ${url}:`, error);
    } finally {
      await page.close();
    }
  }
} finally {
  await browser.close();
}

Run the batch from the directory containing both files:

node capture.mjs

To capture the full scrollable page instead of the viewport:

FULL_PAGE=1 node capture.mjs

To change the settling interval, in milliseconds:

SETTLE_MS=3000 node capture.mjs

The script keeps going after an individual navigation or screenshot error and reports that URL. Check the output log before treating the folder as a complete batch. The numbered filenames are unique within one run; if you need to preserve multiple runs, use a new output folder for each run or add a date-and-time run identifier.

5. Wait for the content that matters

A fixed delay is simple but does not prove that search results or prices have finished rendering. If a page exposes a stable, visible element that marks the content you need, wait for that locator after navigation. For example, replace the settling block in the script with:

await page.locator('main').waitFor({ state: 'visible', timeout: 20_000 });

main is only a generic example. Inspect the page and choose a selector that is meaningful for that specific site; a shared selector may not work across several different booking platforms. Playwright’s Locator API describes locator waits.

Use waitUntil: 'load' when the page’s load event is a reasonable starting point. Some pages continue fetching data after that event. Network-idle waiting can be unsuitable for pages with persistent requests, such as analytics or live updates. Prefer a page-specific readiness signal when possible, and inspect captures for incomplete content.

6. Make a batch comparable and auditable

  • Keep the viewport constant: use the same width, height, and device scale factor for every page when comparing visible layouts.
  • Choose one capture mode: viewport and full-page images answer different questions. Record the mode if you mix them.
  • Use unique names: include a sequence number or stable page identifier, plus a run date when preserving multiple runs. Avoid filenames derived directly from untrusted URL text.
  • Keep a manifest: record each source URL, output filename, capture time, viewport, capture mode, and any relevant search state. The script’s console output is a minimal record; a CSV or JSON manifest is useful for repeat work.
  • Check overlays and loading: review images for cookie banners, chat widgets, popups, blank regions, and content that has not loaded. A screenshot records what was rendered, including overlays.
  • Be consistent about browser state: a fresh context gives pages a separate browser session. If a workflow depends on consent, cookies, or login, handle that state deliberately and only where permitted.

7. Optional: capture with Python

For a Python-based workflow, install the Playwright package and Chromium, then use this sequential script. It uses the same one-URL-per-line input format and saves numbered PNG files.

python -m venv .venv
. .venv/bin/activate
pip install playwright
playwright install chromium
from pathlib import Path
from urllib.parse import urlparse
from playwright.sync_api import sync_playwright

urls = [
    line.strip()
    for line in Path('urls.txt').read_text(encoding='utf-8').splitlines()
    if line.strip() and not line.lstrip().startswith('#')
]
if not urls:
    raise SystemExit('urls.txt contains no URLs')

for url in urls:
    if urlparse(url).scheme not in {'http', 'https'} or not urlparse(url).netloc:
        raise SystemExit(f'Invalid HTTP(S) URL: {url}')

out = Path('screenshots')
out.mkdir(parents=True, exist_ok=True)

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    context = browser.new_context(
        viewport={'width': 1440, 'height': 1000},
        device_scale_factor=1,
    )
    for index, url in enumerate(urls, start=1):
        page = context.new_page()
        filename = out / f'{index:03d}.png'
        try:
            response = page.goto(url, wait_until='load', timeout=60_000)
            if response is not None and not response.ok:
                print(f'HTTP {response.status} for {url}')
            page.wait_for_timeout(1500)
            page.screenshot(path=str(filename), full_page=False)
            print(f'Saved {filename} from {url}')
        except Exception as exc:
            print(f'Failed to capture {url}: {exc}')
        finally:
            page.close()
    browser.close()

Run it with python capture.py. Change full_page=False to True for full-page images, and replace the fixed wait with a site-appropriate locator wait where you can identify the needed content.

8. Optional: capture with cURL or a direct HTTP client

cURL and ordinary HTTP clients download a URL response; they do not render a modern page like a browser, execute its client-side code, or capture the rendered viewport. They are appropriate only when the URL itself serves the image or document you need. For rendered booking pages, use browser automation such as Playwright or a screenshot service.

9. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request can return a PNG, JPEG, WebP, or PDF. For a batch, repeat the call for each URL or use its bulk capture option for up to 100 URLs per call. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Replace the example target URL with the booking-page URL you are authorized to capture. For a batch, make one request per URL or use the bulk endpoint described in the docs. Cookie banners are accepted before capture and more than 60 known consent platforms, newsletter popups, and chat widgets can be removed; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up free for 1,000 screenshots a month, with no card required.

10. Performance, reliability, and cost

Run sequentially first

Sequential navigation is easier to debug and avoids opening many browser pages at once. If the URL list is large and the sites permit your access pattern, modest concurrency can reduce total elapsed time, but it increases CPU, memory, and network use. Add concurrency only after the sequential run is producing complete, correctly named files. Do not infer a safe request rate for a particular site; check its current rules.

Plan for timeouts and partial batches

Each page can load at a different speed or fail independently. A timeout does not mean the page is empty, and a saved file does not guarantee it contains the content you expected. Keep per-URL success and failure records, review the images, and rerun only the failed entries after checking why they failed. Avoid silently overwriting a good prior run.

Estimate local resource use

Browser automation uses more resources than downloading static files because it starts a browser and renders pages. Full-page screenshots can use more memory than viewport captures, especially for long pages. Close each page after capture, as in the example, and close the browser when the batch ends.

Account for changing page state

Prices, availability, personalized results, and page layout may change between captures. A fixed viewport improves visual consistency but cannot freeze the site’s data. For comparisons, capture the pages close together in time and record the query context and timestamp. This is an implementation practice, not a guarantee that different sites expose equivalent results.

Consider hosted capture costs

Local Playwright has no per-screenshot API charge, but it uses your machine or server and requires browser installation and maintenance. ScreenshotNeo offers 1,000 shots per month free with no card; listed paid plans are Starter at $5 for 3,000, Growth at $15 for 15,000, Pro at $39 for 60,000, Scale at $99 for 250,000, and Business at $249 for 1,000,000. Yearly billing gives two months free, and every feature is available on every plan. Check the current product documentation for request configuration and plan details.

11. Troubleshooting

Symptom Likely cause What to do
No screenshots are created The input file is empty, the script is running from another directory, or the output path is not writable. Run the command beside urls.txt, check the script’s error output, and confirm the process can create the screenshots directory.
Invalid URL error A line is incomplete, has whitespace or typo issues, or is missing an HTTP(S) scheme. Use the full URL beginning with https:// or http://; remove accidental spaces and verify the page address.
Browser executable is missing Playwright’s package is installed but its browser binary is not. Run npx playwright install chromium, or playwright install chromium for Python, as appropriate.
Navigation times out The page is slow, continues loading, or cannot be reached from the machine running the script. Check the URL and connectivity, inspect the page manually, and adjust the navigation timeout only if the longer wait is appropriate. Do not treat a timeout as a successful capture.
Screenshot is blank or missing results The content may render after the load event, depend on a user action, or be unavailable in that session. Wait for a relevant visible locator, inspect the page state and console, and verify whether the page requires a permitted login or search step.
Cookie or chat overlay covers content The screenshot records overlays currently visible in the browser. Check the site’s terms and available consent controls. For captures where overlay removal is useful, ScreenshotNeo can accept consent banners and remove supported overlays before capture.
Files overwrite or are hard to match to URLs Names are reused between runs or do not identify the source. Use a run-specific folder and maintain a manifest mapping each URL to its image filename.
Full-page capture is excessively tall The page contains a long or expanding results list. Use viewport capture for a consistent comparison, or capture a specific element when an appropriate stable selector exists.
Some URLs fail while others succeed Network or page failures can be isolated to individual URLs. Keep the per-URL log, inspect failed pages individually, and rerun only after identifying the cause. Do not label the batch complete until every intended URL is accounted for.

12. FAQ

Can I capture pages that require a login?

Only if you are authorized to access and automate those pages. The sample script uses a fresh browser context and does not sign in. Any authenticated workflow needs deliberate session handling and careful protection of credentials.

Will the screenshots show the same prices every time?

No. A screenshot records the rendered state at capture time. Availability, prices, personalization, and other page data can change.

Should I use one browser page for every URL?

The sample creates and closes a page for each URL inside one browser context. This keeps the workflow simple while reusing the browser process. Use separate contexts if you need isolated browser state between pages.

Can I create PDFs instead of image files?

Playwright’s browser APIs and ScreenshotNeo support different capture workflows; ScreenshotNeo’s API can return PDFs with configurable PDF options. Consult the respective official documentation for the exact settings supported by your chosen method.

References