ScreenshotNeo

BlogHow-to

How to Screenshot Multiple Pages of an Indian University Website in Bulk

Capture an authorized list of university pages in one repeatable Playwright workflow, save full-page screenshots, and track every result in a manifest.

By the ScreenshotNeo team4 October 20269 min read

To screenshot multiple pages of an Indian university website in bulk, prepare an authorized list of URLs and use a browser automation tool such as Playwright to visit each page, capture it, and save the result under a predictable filename. Enable full-page capture when you need content below the fold. Keep a manifest that connects every URL to its screenshot and records failures, redirects, and capture settings.

This guide uses Node.js and Playwright for a runnable local workflow, then provides Python and cURL alternatives. The research behind this guide did not test any specific Indian university website, so first pilot a few representative pages. Login requirements, lazy-loaded content, consent banners, and site-specific behavior can affect results.

1. Build and check the URL list

Make a plain text file with one complete URL per line. Include only pages you are authorized to capture. University navigation and a sitemap, if the site provides one, can help you discover pages; a sitemap may be absent or incomplete, so check the resulting list against your intended scope.

https://www.example.edu.in/
https://www.example.edu.in/admissions
https://www.example.edu.in/academics/departments
https://www.example.edu.in/examinations/results

Replace these illustrative URLs with the actual pages you need. Before running the batch:

  • Remove duplicate URLs and accidental off-site links.
  • Decide whether the output should show the current viewport or the whole scrollable page.
  • Check whether pages require authentication. Do not put passwords or session cookies in a shared URL list.
  • Choose a manageable sample with different page types, such as a landing page, a long article, and a page with dynamic content.

2. Capture pages with Node.js and Playwright

Playwright’s page screenshot API supports full-page capture with fullPage: true. This script reads URLs from urls.txt, saves PNGs in screenshots/, and writes a JSON Lines manifest with the source URL, output filename, status, final URL, timestamp, and error when applicable.

Install

npm init -y
npm install playwright
npx playwright install chromium

Create capture.mjs

import { chromium } from 'playwright';
import { mkdir, readFile, appendFile } from 'node:fs/promises';

const urls = (await readFile('urls.txt', 'utf8'))
  .split(/\r?\n/)
  .map((line) => line.trim())
  .filter((line) => line && !line.startsWith('#'));

await mkdir('screenshots', { recursive: true });
await mkdir('logs', { recursive: true });

const browser = await chromium.launch({ headless: true });
const manifestPath = 'logs/manifest.jsonl';

try {
  for (let i = 0; i < urls.length; i += 1) {
    const url = urls[i];
    const filename = `${String(i + 1).padStart(4, '0')}.png`;
    const page = await browser.newPage({ viewport: { width: 1440, height: 1000 } });
    const record = {
      index: i + 1,
      requested_url: url,
      filename,
      captured_at: new Date().toISOString(),
      viewport: { width: 1440, height: 1000 },
      full_page: true
    };

    try {
      const response = await page.goto(url, {
        waitUntil: 'domcontentloaded',
        timeout: 45000
      });

      // Give client-side rendering a short opportunity to update the page.
      await page.waitForTimeout(1000);
      await page.screenshot({ path: `screenshots/${filename}`, fullPage: true });

      record.status = response ? response.status() : null;
      record.final_url = page.url();
      record.result = 'captured';
    } catch (error) {
      record.final_url = page.url();
      record.result = 'error';
      record.error = String(error);
    } finally {
      await appendFile(manifestPath, `${JSON.stringify(record)}\n`);
      await page.close();
    }
  }
} finally {
  await browser.close();
}

Run it

node capture.mjs

The script processes pages sequentially, which limits load on both your machine and the university site. A non-2xx HTTP response can still produce an image, so review the recorded status and the screenshots rather than treating file creation as proof that a page is correct. The manifest appends records; move or delete the old manifest before a new run if you want a fresh batch log.

3. Capture pages with Python and Playwright

Playwright’s Python API uses full_page=True. Install its package and Chromium browser, save the following as capture.py, and run it in the directory containing urls.txt.

python -m pip install playwright
python -m playwright install chromium
import asyncio
import json
from datetime import datetime, timezone
from pathlib import Path
from playwright.async_api import async_playwright

async def main():
    urls = [
        line.strip()
        for line in Path('urls.txt').read_text(encoding='utf-8').splitlines()
        if line.strip() and not line.lstrip().startswith('#')
    ]
    Path('screenshots').mkdir(exist_ok=True)
    Path('logs').mkdir(exist_ok=True)

    async with async_playwright() as p:
        browser = await p.chromium.launch(headless=True)
        try:
            for index, url in enumerate(urls, start=1):
                filename = f'{index:04d}.png'
                page = await browser.new_page(viewport={'width': 1440, 'height': 1000})
                record = {
                    'index': index,
                    'requested_url': url,
                    'filename': filename,
                    'captured_at': datetime.now(timezone.utc).isoformat(),
                    'viewport': {'width': 1440, 'height': 1000},
                    'full_page': True,
                }
                try:
                    response = await page.goto(url, wait_until='domcontentloaded', timeout=45000)
                    await page.wait_for_timeout(1000)
                    await page.screenshot(path=f'screenshots/{filename}', full_page=True)
                    record.update({
                        'status': response.status if response else None,
                        'final_url': page.url,
                        'result': 'captured',
                    })
                except Exception as error:
                    record.update({
                        'final_url': page.url,
                        'result': 'error',
                        'error': str(error),
                    })
                finally:
                    with Path('logs/manifest.jsonl').open('a', encoding='utf-8') as manifest:
                        manifest.write(json.dumps(record, ensure_ascii=False) + '\n')
                    await page.close()
        finally:
            await browser.close()

asyncio.run(main())
python capture.py

4. Use cURL for a single screenshot request

cURL is useful for checking a screenshot endpoint or saving a single response, but it does not itself automate a browser or loop through a URL inventory. For a batch, call it once per URL from a script and record each result. ScreenshotNeo’s API uses one GET request per screenshot; see the ScreenshotNeo API documentation for its parameters and response headers.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://www.example.edu.in/admissions \
  -o admissions.webp

Replace the example domain and filename. Keep API keys out of source control and shared scripts.

5. Choose capture settings and handle page behavior

Viewport or full page

A viewport screenshot captures what is visible at the configured browser dimensions. Full-page mode captures the page’s scrollable content in one image. Full-page images can be very tall and large, and some sites render sections only after scrolling; pilot those pages to see whether content appears as expected.

Wait for the right content

The example waits for domcontentloaded and then pauses briefly. This is a practical starting point, not a guarantee that every site’s content is ready. Pages that render client-side may need a longer wait or a selector-based wait for a specific heading or content region. Avoid assuming that waiting for every network request to finish is suitable: analytics, advertisements, or long-lived requests can prevent network-idle conditions.

Lazy-loaded sections

Some pages fetch images or content as the visitor scrolls. Full-page screenshot support does not prove that a particular site’s lazy-loaded elements have been activated. If sections are missing, test a workflow that scrolls through the page before capture, or wait for a known element. Review the resulting image rather than assuming a capture completed correctly.

Login and redirects

A page behind a login can redirect to a sign-in screen or show an access-denied page. The sample records the final URL and HTTP status to help identify this. If capture is authorized, use a properly managed authenticated browser context and protect its state; do not place credentials in the URL or manifest.

Filename and manifest design

Numbered filenames are stable within a fixed, ordered URL list. For lists that change, use a slug or a short hash of each normalized URL to avoid assigning a different page the same number across runs. Keep the original URL in the manifest so files remain traceable. Record the browser, viewport, capture mode, and time when screenshots will be compared later.

6. Run a reliable batch

  1. Confirm scope: verify the URL inventory and that you are authorized to capture it.
  2. Pilot: capture a few pages from different sections, including at least one long or dynamic page.
  3. Inspect: confirm the intended destination loaded, below-the-fold content is present where expected, and text is readable.
  4. Run sequentially first: increase concurrency only after checking resource use and the site’s acceptable request rate.
  5. Review the manifest: investigate errors, redirects, unusual status codes, and duplicate destinations.
  6. Retry selectively: retry failed URLs with a suitable timeout or wait condition instead of blindly recapturing every successful page.

For large batches, add retry limits and backoff to handle temporary network failures. Keep failed entries in the manifest so a partial run cannot be mistaken for a complete one. The sample scripts catch per-page errors and continue, but they do not automatically retry.

7. Performance, reliability, and cost

Local Playwright captures use your machine’s CPU, memory, browser installation, and network connection. Full-page screenshots of long pages consume more resources than viewport shots. Sequential processing is slower than parallel capture but is easier to monitor and places less simultaneous load on the target site. If you add concurrency, keep it modest and measure memory use and failure rates on a small batch first.

Reliability depends on the URL inventory, network conditions, page behavior, and whether the content is available without login. Save the manifest alongside the images, preserve error details, and make the run repeatable with the same URL ordering and settings. A screenshot is a point-in-time rendering; changes to page content, fonts, ads, or dynamic modules can make later captures differ.

Playwright is software, and the research does not establish a required physical product for this workflow. Local costs are the computer and time used to run and review the batch. Hosted screenshot APIs may charge according to their own plans and billing rules; check the provider’s documentation before scaling.

8. Troubleshooting

Symptom Likely cause What to do
Browser executable missing Playwright package is installed, but its Chromium browser was not installed. Run npx playwright install chromium or python -m playwright install chromium.
Navigation timeout Slow server, network problem, or a page that keeps loading resources. Check the URL in a browser, increase the navigation timeout where appropriate, and consider waiting for a specific content selector instead of all network activity.
Screenshot is blank or incomplete Client-side content had not rendered, a redirect occurred, or access was denied. Inspect the manifest’s final URL and status, wait for a meaningful selector, and confirm the page can be viewed in the same access context.
Images or sections are missing Content may load only after scrolling or after a delayed request. Inspect the page behavior, scroll through it before capture if necessary, and pilot that page type.
Too many or duplicate output files Input list contains repeated URLs, or old output remains from a previous run. Deduplicate the inventory and use a clean output directory or a run-specific directory.
Manifest shows a failure but the batch continues The sample intentionally records per-page errors and proceeds to later URLs. Review every manifest line and rerun only failed entries after correcting their cause.
Image is too tall or hard to review Full-page mode captures a long scrollable document as one image. Use viewport captures, or split the work into sections if the chosen tool and workflow support that capture strategy.

Or skip the browser setup

ScreenshotNeo takes a screenshot with one API request. This example captures an illustrative university admissions page as WebP; replace the target URL with a page you are authorized to capture. See the API documentation for available capture options and response headers.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://www.example.edu.in/admissions -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://www.example.edu.in/admissions"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://www.example.edu.in/admissions' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));

ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots, and the free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. You can use its API in a loop over your URL list and maintain the same kind of manifest shown above.

Sign up free for 1,000 screenshots a month, with no card required.

FAQ

Can I capture an entire university website automatically?

This workflow captures the URLs you provide. It does not discover every page or establish that the site has a complete sitemap; define and review the scope yourself.

Will full-page mode include content that appears only after scrolling?

Not necessarily. Lazy loading is site-specific. Test representative pages and add scrolling or content-specific waits if needed.

Can I tell whether a screenshot came from the right page?

Record the final URL and response status, then review the image. A saved file alone does not confirm that the intended content appeared.

Should I run every URL at once?

Start sequentially. If you later add parallelism, check resource use and the target site’s acceptable request rate, and keep failure logging and retry limits.