ScreenshotNeo

BlogHow-to

How to Bulk Screenshot Indian Hotel and Travel Booking Pages from a URL List

Capture hotel and travel booking pages from a URL list with shot-scraper or Playwright. Compare full-page capture, waits, output naming, and hosted options.

By the ScreenshotNeo team4 October 20268 min read

To bulk screenshot Indian hotel and travel booking pages, prepare a URL-to-filename list, then use shot-scraper multi with a YAML configuration or write a Playwright script that visits each URL and saves a screenshot. Use full-page capture when you need below-the-fold content, tune waits for pages that render asynchronously, and test a few representative URLs before processing the complete list. The tools are not specific to India, and this research does not establish that booking sites need any one special setting.

1. Prepare the URL list and output names

Start with one canonical URL per page you want to capture. Decide whether query parameters matter: search results may encode dates, guests, destination, or filters in the query string, so preserve parameters that define the page state you want to compare. Remove accidental duplicates, and give every output a filename that maps back to its source without relying on row order.

For example, maintain a simple CSV for your own inventory:

url,output
https://example.com/hotels?city=jaipur,jaipur-hotels.png
https://example.com/stays?city=goa&guests=2,goa-two-guests.png

The example domains are placeholders. Substitute URLs you are allowed to access. Avoid putting credentials or personal data in filenames, query strings, or logs.

  1. Check that each URL opens the intended page in a normal browser.
  2. Choose a consistent viewport for comparisons, or record the viewport alongside each capture.
  3. Decide whether you need the initial viewport or the full scrollable page.
  4. Run a small sample, inspect the images, and adjust waits or dimensions before a large batch.

Check the terms and permitted access for the domains you target. This is practical workflow advice, not a legal conclusion about any particular site.

shot-scraper supports multiple URL and output entries in a YAML file, with per-entry settings such as dimensions, quality, waits, and wait_for. See the shot-scraper documentation for its multi-URL configuration.

Install shot-scraper in your Python environment, then create a YAML file such as booking-pages.yml. The following shows the configuration shape; use the exact option spelling and supported values from the documentation version you install:

- url: https://example.com/hotels?city=jaipur
  output: screenshots/jaipur-hotels.png
  width: 1440
  height: 1000
  wait: 2
- url: https://example.com/stays?city=goa&guests=2
  output: screenshots/goa-two-guests.png
  width: 1440
  height: 1000
  wait: 3

Create the output directory if it does not already exist, then run:

mkdir -p screenshots
shot-scraper multi booking-pages.yml

Set a per-page wait_for condition when a known element indicates that the content you need has appeared. Avoid arbitrary long delays as a default: they make batches slower, and a delay alone does not guarantee that a page is ready. Use the configuration options documented for your installed release.

3. Custom browser automation with Playwright

For a workflow that needs custom naming, logging, retries, or browser settings, use Playwright. Its official Page API documentation shows navigation and screenshot capture, including the fullPage option for the full scrollable page.

Install Playwright and its Chromium browser in your project using the official installation instructions. Then save this as capture.mjs. It uses Node.js, processes URLs sequentially, records failures without silently losing them, and writes a JSON-lines manifest mapping each input URL to its result.

import { chromium } from 'playwright';
import { mkdir, appendFile } from 'node:fs/promises';

const jobs = [
  { url: 'https://example.com/hotels?city=jaipur', output: 'screenshots/jaipur-hotels.png' },
  { url: 'https://example.com/stays?city=goa&guests=2', output: 'screenshots/goa-two-guests.png' },
];

await mkdir('screenshots', { recursive: true });
await mkdir('logs', { recursive: true });
const browser = await chromium.launch({ headless: true });

try {
  for (const job of jobs) {
    const page = await browser.newPage({ viewport: { width: 1440, height: 1000 } });
    try {
      const response = await page.goto(job.url, {
        waitUntil: 'domcontentloaded',
        timeout: 45000,
      });
      if (!response) {
        throw new Error('Navigation returned no main-resource response');
      }
      if (!response.ok()) {
        throw new Error(`Main document returned HTTP ${response.status()}`);
      }

      // Use fullPage: true when the entire scrollable document is required.
      await page.screenshot({ path: job.output, fullPage: true });
      await appendFile('logs/results.jsonl', JSON.stringify({
        url: job.url,
        output: job.output,
        status: 'ok',
        httpStatus: response.status(),
      }) + '\n');
    } catch (error) {
      await appendFile('logs/results.jsonl', JSON.stringify({
        url: job.url,
        output: job.output,
        status: 'error',
        error: String(error),
      }) + '\n');
      console.error(`Failed: ${job.url}: ${error}`);
    } finally {
      await page.close();
    }
  }
} finally {
  await browser.close();
}

Run it with node capture.mjs. Add your URLs and output names to jobs. This script waits for the initial document parsing, not every image, API call, or late-rendering widget. If a page needs a particular element, add a selector wait before the screenshot:

await page.locator('YOUR_STABLE_SELECTOR').waitFor({ state: 'visible', timeout: 15000 });

Replace the placeholder selector with one verified for the pages you capture. If you need only the current viewport, omit fullPage: true. Playwright documents other screenshot options, including image type and quality; check its API reference for supported options and constraints.

4. Choosing capture settings for booking pages

Decision Use this when Trade-off
Viewport screenshot You compare the visible first screen or a consistent fold. Content below the viewport is excluded.
Full-page screenshot You need the entire scrollable document in one image. Very long pages can create large images and take longer to render or inspect.
Fixed delay A page consistently needs a short additional rendering interval. It may wait too long on fast pages and still be too short on slow ones.
Selector wait A stable page element signals that the content of interest is present. Selectors can change and may not exist on every page template.
Consistent viewport You want a fair visual comparison across URLs. Some layouts adapt to screen size, so a different viewport changes what is shown.

Do not assume hotel or travel pages share a rendering pattern. Some URLs may show a consent prompt or delayed results; verify the actual pages in a sample. If you need a reproducible comparison, record capture time, viewport, URL, and any page-specific wait in a manifest.

5. Scale the batch safely

For a modest list, sequential capture is easier to debug and places less simultaneous load on your machine and target domains. For larger lists, add bounded concurrency only after measuring resource use and confirming the target domains permit the access pattern. Keep a failure list and rerun only failed URLs rather than repeating successful work.

  • Reliability: record URL, output path, time, response status, and error. Use stable output names and avoid overwriting unrelated files.
  • Retries: retry transient navigation failures a small, bounded number of times with a pause. Do not retry access-denied or bot-check pages aggressively.
  • Memory: close each page when done. Full-page capture of long pages can use more memory than viewport captures.
  • Storage: estimate output size with a pilot batch; PNG is lossless and can be larger, while JPEG or WebP may suit visual review where supported.
  • Costs: local browser automation has no per-screenshot API charge, but uses compute, storage, maintenance time, and network bandwidth. Hosted services may charge or impose limits; verify current terms directly.

6. Hosted option for URL lists

AddScreenshots’ API documentation describes background bulk capture for URL lists and sitemaps with results stored in cloud repositories. The page was not retrievable for this research, so treat that feature description as lower-confidence and check directly with the provider for current availability, pricing, limits, data handling, regional suitability, and terms before relying on it: AddScreenshots API documentation.

For a developer choosing a screenshot API or hosted tool, ScreenshotNeo is the first option to try: it removes consent banners, newsletter popups, and chat widgets before capture, bills only clean shots, and its paid plans start at $5 for 3,000 shots. ScreenshotNeo’s documented options include bulk capture of up to 100 URLs per call; check the ScreenshotNeo documentation for request details.

Or skip the browser setup

ScreenshotNeo takes a URL in one GET request and returns an image or PDF. Example request for one page:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for authentication, output settings, and bulk capture details. Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents use take_screenshot, get_page_info, and capture_pdf. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000, and every feature is on every plan.

Sign up for 1,000 free screenshots a month, with no card required.

7. Troubleshooting

Symptom Likely cause What to do
Navigation times out The main document or a required page element did not load before the timeout. Check the URL in a browser, increase the timeout only when justified, and wait for a stable relevant selector if the page renders late.
Screenshot is blank or incomplete The capture happened before content rendered, or the page returned an empty/error state. Inspect the response and page state; use a targeted selector wait and capture a sample again.
Only the first screen appears The screenshot was taken without full-page mode. Enable the tool’s full-page option. Playwright uses fullPage: true.
Output file is missing The parent directory does not exist or the process lacks write access. Create the directory first and check the reported file path and filesystem permissions.
Repeated captures overwrite one another Two jobs use the same output filename. Make output paths unique and include a stable page identifier.
Page shows a consent prompt, bot check, or access restriction The site presented that state to the browser. Do not assume the screenshot represents the intended content. Check permitted access and provider guidance; avoid attempts to bypass access controls.
Batch stops after one problematic URL The script propagates an exception or the CLI exits on failure. Capture per-URL errors, keep a manifest, and rerun the failed subset after diagnosing it.

8. Frequently asked questions

Can I capture search results with dates and filters?

Yes, if those states are represented by URLs you can access. Preserve the relevant query parameters and record them with the output.

Do Indian hotel sites require a special browser setting?

The available research does not establish a universal India-specific requirement. Validate each domain and page type with a representative sample.

Should I use a sitemap instead of a URL list?

A sitemap can supply URLs when you want broad site coverage. For a controlled comparison, a curated list is easier to audit and associate with intended search states.

Does full-page capture include content that appears only after scrolling?

It captures the scrollable page as rendered by the browser, but lazy-loaded content may require scrolling or another page-specific readiness step. Check the resulting image rather than assuming every below-fold element loaded.