ScreenshotNeo

BlogHow-to

How to Schedule Screenshots of Indian University Admission Websites

Build a scheduled workflow that captures public Indian university admission pages, waits for useful content, and saves timestamped screenshots.

By the ScreenshotNeo team4 October 202610 min read

To schedule screenshots of Indian university admission websites, write a browser script that opens each public page, waits for its useful content, saves a viewport, full-page, or element screenshot, and run that script with a local operating-system scheduler or a hosted workflow such as GitHub Actions. Browser automation does the capture; the scheduler starts it. Those are separate jobs.

This guide uses Playwright with Node.js and GitHub Actions. It is designed for public admissions, notices, or announcements pages. Check each site’s published conditions and access instructions first. Do not target private application pages, expose credentials or session files, or try to bypass access controls. No single URL pattern, update cadence, or automated-access policy applies to every university.

1. Choose the pages and capture scope

List the exact public URLs you need. Admissions information may be spread across an admissions page, notices page, or announcements page, so verify the relevant pages individually. Keep their names with the URLs so each saved screenshot is identifiable.

Choose what the screenshot should show:

  • Viewport: the currently visible browser area; useful for a quick visual record.
  • Full page: the entire scrollable page, including below-the-fold content.
  • Element: a specific notice list or other region identified by a CSS selector.

Playwright documents screenshot-to-file, full-page, and element capture in its screenshots guide. Use a consistent viewport and browser version if you plan to compare images across runs; consistency is an implementation choice, not a guarantee that a site will render identically each time.

2. Create a Playwright screenshot script

Install Node.js, then create a project and install Playwright:

mkdir admission-screenshots
cd admission-screenshots
npm init -y
npm install playwright

Save the following as capture.mjs. It reads target pages from pages.json, waits for an optional page-specific selector, and saves timestamped PNG files. The default is full-page capture. Set CAPTURE_MODE=viewport or CAPTURE_MODE=element to change the scope; element mode uses each page’s selector.

import { chromium } from 'playwright';
import { mkdir, readFile } from 'node:fs/promises';

const pages = JSON.parse(await readFile(new URL('./pages.json', import.meta.url), 'utf8'));
const outputDir = process.env.OUTPUT_DIR ?? 'screenshots';
const mode = process.env.CAPTURE_MODE ?? 'full';
const timeoutMs = Number(process.env.PAGE_TIMEOUT_MS ?? 45000);
const width = Number(process.env.VIEWPORT_WIDTH ?? 1440);
const height = Number(process.env.VIEWPORT_HEIGHT ?? 1000);

if (!['full', 'viewport', 'element'].includes(mode)) {
  throw new Error('CAPTURE_MODE must be full, viewport, or element');
}

function timestamp() {
  return new Date().toISOString().replaceAll(':', '-').replace(/\.\d{3}Z$/, 'Z');
}

function safeName(value) {
  return value.toLowerCase().replace(/[^a-z0-9]+/g, '-').replace(/^-|-$/g, '') || 'page';
}

await mkdir(outputDir, { recursive: true });
const browser = await chromium.launch({ headless: true });
let failures = 0;

try {
  for (const item of pages) {
    const page = await browser.newPage({ viewport: { width, height }, deviceScaleFactor: 1 });
    try {
      const response = await page.goto(item.url, { waitUntil: 'domcontentloaded', timeout: timeoutMs });
      if (response && !response.ok()) {
        throw new Error(`Navigation returned HTTP ${response.status()}`);
      }

      if (item.readySelector) {
        await page.locator(item.readySelector).waitFor({ state: 'visible', timeout: timeoutMs });
      }

      const path = `${outputDir}/${safeName(item.name)}-${timestamp()}.png`;
      if (mode === 'element') {
        if (!item.selector) throw new Error('Element mode requires selector in pages.json');
        await page.locator(item.selector).screenshot({ path, animations: 'disabled', timeout: timeoutMs });
      } else {
        await page.screenshot({
          path,
          fullPage: mode === 'full',
          animations: 'disabled',
          timeout: timeoutMs,
        });
      }
      console.log(`Saved ${path}`);
    } catch (error) {
      failures += 1;
      console.error(`Failed ${item.name} (${item.url}): ${error.message}`);
    } finally {
      await page.close();
    }
  }
} finally {
  await browser.close();
}

if (failures > 0) process.exitCode = 1;

Create pages.json alongside it. Replace the example URL with the exact public page you want to capture. readySelector is optional; choose a selector that indicates the useful content is present. For element capture, provide selector too.

[
  {
    "name": "university-admissions",
    "url": "https://example.edu/admissions",
    "readySelector": "main"
  },
  {
    "name": "university-notices",
    "url": "https://example.edu/notices",
    "readySelector": ".notice-list",
    "selector": ".notice-list"
  }
]

The example checks HTTP response status when a response is available, but a successful status alone does not prove the page contains current or meaningful admission information. A selector wait helps with known page structures; verify selectors against each target and revisit them if a site changes its layout.

3. Wait for page content reliably

The script navigates with domcontentloaded, then optionally waits for a visible, page-specific selector. This separates document parsing from the signal that matters to your capture. If a notice area appears only after client-side rendering, use that area as readySelector.

Playwright supports navigation wait states including commit, domcontentloaded, load, and networkidle. Its API reference marks networkidle as discouraged for testing and recommends using assertions to assess readiness. A page-specific selector is often a better fit for this workflow. A fixed delay can be a fallback for a page you understand, but it can be too short on a slow run and waste time on a fast one. See the Playwright Page API.

For pages that update a notice list after initial load, select a stable element that appears with the content you need. If the page has no dependable selector, add a deliberate delay only after checking how the site behaves, and accept that a delay cannot confirm the content is complete.

4. Run and inspect a capture locally

Run the script once before scheduling it:

node capture.mjs

To capture just the first viewport or the configured notice elements:

CAPTURE_MODE=viewport node capture.mjs
CAPTURE_MODE=element node capture.mjs

On Windows PowerShell, set an environment variable for the current shell with $env:CAPTURE_MODE="viewport", then run node capture.mjs. Inspect the saved image, confirm the page and content are useful, and verify that filenames include the expected page name and UTC timestamp.

5. Schedule recurring runs

Option A: GitHub Actions

A hosted workflow can run while your personal computer is off. Add the script, pages.json, and a workflow file to a repository. Create .github/workflows/capture.yml:

name: Capture admission pages

on:
  schedule:
    - cron: '0 2 * * *'
  workflow_dispatch:

permissions:
  contents: write

jobs:
  capture:
    runs-on: ubuntu-latest
    steps:
      - name: Check out repository
        uses: actions/checkout@v4
      - name: Set up Node.js
        uses: actions/setup-node@v4
        with:
          node-version: 22
      - name: Install dependencies
        run: npm ci
      - name: Install Chromium
        run: npx playwright install --with-deps chromium
      - name: Capture pages
        run: node capture.mjs
      - name: Commit screenshots
        run: |
          git config user.name "github-actions[bot]"
          git config user.email "41898282+github-actions[bot]@users.noreply.github.com"
          git add screenshots
          if ! git diff --cached --quiet; then
            git commit -m "Save scheduled admission screenshots"
            git push
          fi

The schedule above runs daily at 02:00 UTC. GitHub documents scheduled workflows as POSIX cron schedules that run against the latest commit on the default branch. By default they run in UTC, and the documented shortest interval is once every five minutes. GitHub also supports specifying an IANA timezone. These are platform capabilities, not a recommendation to poll admission sites every five minutes. Choose a proportionate cadence after checking each site’s guidance and your actual need. See GitHub’s workflow schedule documentation.

This example commits screenshots to the repository, which makes them available with the code but grows repository history over time. For a private or long-lived archive, consider where the artifacts should live and how long to retain them. Do not commit personal data, credentials, cookies, or session files. The workflow needs permission to push commits; if repository policy disallows that, use the platform’s artifact storage or another destination you configure instead.

Option B: a local operating-system scheduler

A local scheduler is straightforward if the machine is available at the chosen time. It must be awake and able to reach the sites. Set the working directory to the project and run node capture.mjs using your operating system’s scheduler, such as cron, Task Scheduler, or launchd. Redirect logs to a file so failures are visible. Exact setup varies by operating system and how Node.js is installed; use the full path to the Node executable if the scheduler does not inherit your interactive shell’s PATH.

Local and hosted scheduling have different operational tradeoffs: local jobs depend on that machine being on and retaining the files locally; hosted jobs depend on workflow configuration, repository access, and where artifacts are stored. Choose based on where you need the archive and who needs access to it.

6. Make the archive useful

  • Use filenames that identify the institution or page and the capture time, such as university-admissions-2026-10-03T20-00-00Z.png.
  • Keep a record of the exact URL and selector used, so a changed page structure is easier to diagnose.
  • Choose a retention period and storage location that suit the archive. There is no universally best storage or retention policy established for this workflow.
  • Keep browser conditions consistent when visual comparison matters: viewport, device scale, and capture scope.
  • Target public pages and avoid collecting personal application information or placing secrets in a public repository.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. One GET request captures a URL as an image or PDF; use it when you want to avoid managing a browser installation. For automated recurring captures, call the API from your own scheduled job and save each response with a timestamp. The API call itself does not create the schedule.

See the ScreenshotNeo documentation for request details and options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

Replace the example target with a public page you are allowed to capture. Store the access key as a secret in your scheduler, not in source control. Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed; an MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000. Sign up free for 1,000 screenshots a month, no card required.

Troubleshooting

Symptom Likely cause What to check
Selector wait times out The selector is incorrect, the content did not load, or the site changed. Open the public page, inspect its current structure, and update readySelector. Confirm the element becomes visible without relying on a login.
Screenshot is blank or incomplete The useful content rendered after navigation, or the chosen scope excludes it. Wait for a meaningful page-specific selector; check whether full-page capture is needed. A fixed delay is only a fallback, not proof of readiness.
Navigation times out The site was slow or unavailable, or the timeout was too low for that run. Check the target in a browser and review the run logs. Adjust PAGE_TIMEOUT_MS for a known slow page rather than setting an unlimited wait.
HTTP error is reported The server returned a non-success status, or the response was unavailable. Check the URL and the site’s current public availability. Do not attempt to evade an access block.
Scheduled workflow does not run The workflow file is not on the default branch, its cron syntax is invalid, or the platform has not started the run yet. Check the workflow configuration and Actions run history. Scheduled runs use UTC by default unless a timezone is configured.
Workflow captures files but cannot commit them The job lacks write permission or repository policy prevents pushes. Review repository permissions and choose an approved artifact destination if committing output is unsuitable.
Images accumulate or repository grows Every run creates a new timestamped file and repository history retains prior commits. Set a retention policy and use storage suited to the archive’s access and retention needs.

Performance, reliability, and cost

Capturing multiple sites sequentially, as this script does, limits simultaneous browser work and makes per-page errors easier to identify, but total runtime grows with page count and load time. Browser installation and execution consume compute and storage. Keep the schedule no more frequent than the archive needs, and check each site’s published conditions and access instructions. This research does not establish a common update frequency or request policy for Indian universities.

Reliability depends on external pages, selectors, network access, browser availability, scheduler behavior, and artifact storage. Log failures and inspect the saved images periodically; a successful job can still capture stale or changed content. A timestamped file records what the browser rendered at that run, not whether the information is authoritative or current.

GitHub’s schedule interval limit is a platform constraint, not a target cadence. The cost of a local setup depends on the machine and storage you already use; a hosted workflow and archive have their own applicable limits and charges. No source here establishes a universal price or service-level reliability for the scheduling choices.

FAQ

Can I capture several universities in one run?

Yes. Add each public page as an entry in pages.json. Each page is handled independently, and a failure is logged while the script continues with the remaining entries.

Should I use a full-page screenshot?

Use it when below-the-fold content matters. For a quick record of the initial view, use viewport mode. For a notice list with a stable selector, element mode can focus the capture.

Does a scheduled screenshot tell me when an admission notice changed?

It preserves periodic visual records. Detecting changes requires an additional comparison step, such as comparing images or extracting and comparing page content; this script only captures and saves images.

What schedule should I use?

Choose an interval based on the page’s importance, your need for a record, and the site’s published guidance. The available evidence does not establish a standard update cadence for Indian university admission pages.

Can I capture a page that requires an applicant login?

This workflow is for public pages. Do not store applicant credentials or session data in a repository or use the process to bypass access controls.