ScreenshotNeo

BlogHow-to

Schedule Weekly Screenshots of Indian University Webpages with Java Selenium

Build a Java Selenium job that captures university webpages each week, saves organized screenshots, and handles scheduling, browser setup, and failures.

By the ScreenshotNeo team4 October 202611 min read

Use Selenium WebDriver with Java to open each university page, wait for the content you need, and save a PNG screenshot. Run the capture once a week either from an always-on Java process using ScheduledExecutorService or from an external scheduler that launches a one-shot Java program. A normal WebDriver screenshot captures the current viewport; capturing an entire long page needs a separate, validated approach.

This guide uses headless Chrome and Selenium Manager for browser-driver setup. Selenium documents Java WebDriver use and browser screenshots; Chrome and ChromeDriver major versions must match. Selenium Manager is the default driver and browser management route in Selenium bindings. Selenium WebDriver documentation · Chrome-specific guidance.

1. Choose what the screenshot should contain

Decide the capture area before you schedule the job. “Screenshot” can mean three different things:

Output What it captures Use it for
Viewport The visible browser area at the configured window size Consistent snapshots of the top of a page
Element A selected element, such as a news panel Tracking one section independent of the rest of the page
Full page The whole document, including content below the fold Archiving long pages; requires an explicit full-page strategy

The Java example below saves viewport images. Selenium supports element screenshots as well. Do not assume TakesScreenshot will reliably capture an arbitrarily long document: use a browser-specific full-page method or scroll-and-stitch strategy only after validating it against representative pages and their fixed headers, lazy-loaded images, and page length.

2. Set up a Java Selenium project

Create a Maven project and add Selenium Java as a dependency. Pin a Selenium version approved for your environment rather than relying on a floating version.

<dependency>
  <groupId>org.seleniumhq.selenium</groupId>
  <artifactId>selenium-java</artifactId>
  <version>4.27.0</version>
</dependency>

The version above is an example dependency pin. Update it deliberately to the version your project supports. Selenium Manager is used by current Selenium bindings to locate or manage the browser driver by default. If you manage ChromeDriver yourself, keep its major version aligned with the installed Chrome major version. The same Chrome guidance documents headless options and Java examples.

3. Write a one-shot screenshot job

This complete Java class reads the target URLs from command-line arguments, starts headless Chrome, waits for document readiness and a short settling interval, then stores one viewport PNG per URL in a dated directory. It logs each success or failure and continues to the next URL.

import java.io.IOException;
import java.nio.file.Files;
import java.nio.file.Path;
import java.time.Duration;
import java.time.LocalDate;
import java.time.LocalDateTime;
import java.time.format.DateTimeFormatter;
import java.util.List;

import org.openqa.selenium.OutputType;
import org.openqa.selenium.TakesScreenshot;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.chrome.ChromeDriver;
import org.openqa.selenium.chrome.ChromeOptions;
import org.openqa.selenium.support.ui.WebDriverWait;

public class WeeklyUniversityScreenshots {
    private static final Duration PAGE_TIMEOUT = Duration.ofSeconds(45);
    private static final Duration SCRIPT_TIMEOUT = Duration.ofSeconds(30);
    private static final Duration READY_TIMEOUT = Duration.ofSeconds(30);
    private static final Duration SETTLE_TIME = Duration.ofSeconds(2);
    private static final DateTimeFormatter RUN_ID =
            DateTimeFormatter.ofPattern("yyyyMMdd-HHmmss");

    public static void main(String[] args) throws InterruptedException {
        if (args.length == 0) {
            System.err.println("Usage: java WeeklyUniversityScreenshots <url> [<url> ...]");
            System.exit(2);
        }
        List<String> urls = List.of(args);
        Path outputDir = Path.of("screenshots", LocalDate.now().toString());
        try {
            Files.createDirectories(outputDir);
        } catch (IOException e) {
            throw new RuntimeException("Cannot create output directory " + outputDir, e);
        }

        ChromeOptions options = new ChromeOptions();
        options.addArguments("--headless=new", "--window-size=1440,1200");
        // In restricted Linux containers, the environment may require additional
        // Chrome flags or sandbox configuration. Apply only what that environment needs.

        for (int i = 0; i < urls.size(); i++) {
            String url = urls.get(i);
            String fileName = String.format("page-%02d-%s-%s.png",
                    i + 1, safeHost(url), RUN_ID.format(LocalDateTime.now()));
            Path destination = outputDir.resolve(fileName);
            WebDriver driver = null;
            try {
                driver = new ChromeDriver(options);
                driver.manage().timeouts().pageLoadTimeout(PAGE_TIMEOUT);
                driver.manage().timeouts().scriptTimeout(SCRIPT_TIMEOUT);
                driver.manage().window().setSize(new org.openqa.selenium.Dimension(1440, 1200));
                driver.get(url);

                new WebDriverWait(driver, READY_TIMEOUT).until(d ->
                        "complete".equals(((org.openqa.selenium.JavascriptExecutor) d)
                                .executeScript("return document.readyState")));
                Thread.sleep(SETTLE_TIME.toMillis());

                byte[] png = ((TakesScreenshot) driver).getScreenshotAs(OutputType.BYTES);
                Files.write(destination, png);
                System.out.printf("OK %s -> %s (%d bytes)%n",
                        url, destination, png.length);
            } catch (Exception e) {
                System.err.printf("FAILED %s: %s: %s%n",
                        url, e.getClass().getSimpleName(), e.getMessage());
            } finally {
                if (driver != null) {
                    try { driver.quit(); }
                    catch (Exception e) {
                        System.err.println("Browser cleanup failed: " + e.getMessage());
                    }
                }
            }
        }
    }

    private static String safeHost(String url) {
        try {
            String host = java.net.URI.create(url).getHost();
            if (host == null) return "page";
            return host.replaceAll("[^A-Za-z0-9.-]", "_");
        } catch (IllegalArgumentException e) {
            return "page";
        }
    }
}

Compile and run it with your project’s build tool. Pass each approved URL as an argument, for example:

java -cp target/classes:target/dependency/* WeeklyUniversityScreenshots \
  https://www.example-university.edu/ \
  https://www.example-university.edu/admissions/

Use actual university URLs only after checking the institution’s access rules. The example’s fixed two-second settling wait is a starting point, not a guarantee that every page’s dynamic content has finished loading. Replace it with a wait for a page-specific element when the content you care about has a stable selector.

Capture one element instead

After navigating and waiting for the page, locate the element and call its screenshot method. This saves the element’s rendered area rather than the current viewport.

import org.openqa.selenium.By;
import org.openqa.selenium.WebElement;

WebElement panel = driver.findElement(By.cssSelector("main .news-list"));
Files.write(destination, panel.getScreenshotAs(OutputType.BYTES));

Use a selector that belongs to the page you are capturing, and fail clearly if it changes. Selenium’s screenshot documentation covers element capture in addition to browser screenshots: WebDriver interactions.

4. Schedule the job once a week

There are two practical ways to make the capture recurring. Choose based on whether the machine and process will be continuously available and whether missed runs after downtime must be recovered.

Approach How it runs Tradeoff
In-process scheduler A Java process stays alive and runs the capture task every seven days Simple Java-only setup, but no run while the process or host is down
External scheduler A system or managed scheduler starts the one-shot Java program weekly Process exits after each run; configure the scheduler’s timezone, logs, retries, and missed-run behavior

In-process weekly scheduling

ScheduledExecutorService supports periodic tasks using fixed-rate or fixed-delay scheduling. The following example schedules the capture seven days after startup and again every seven days. Keep the process running for the schedule to continue.

import java.time.LocalDateTime;
import java.util.concurrent.Executors;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.TimeUnit;

public class WeeklyRunner {
    public static void main(String[] args) {
        ScheduledExecutorService scheduler = Executors.newSingleThreadScheduledExecutor();
        Runnable task = () -> {
            System.out.println("Starting weekly capture at " + LocalDateTime.now());
            try {
                WeeklyUniversityScreenshots.main(args);
            } catch (Exception e) {
                System.err.println("Capture run failed: " + e.getMessage());
                e.printStackTrace(System.err);
            }
        };
        scheduler.scheduleWithFixedDelay(task, 7, 7, TimeUnit.DAYS);
        Runtime.getRuntime().addShutdownHook(new Thread(scheduler::shutdown));
    }
}

This waits seven days before its first capture. To run immediately at startup and then weekly, call the task once before scheduling it, or schedule with an initial delay of zero. Fixed delay counts seven days after a run finishes; fixed rate targets a seven-day cadence from the initial start and can start late if a run takes longer than the period. The Java API describes both behaviors in the ScheduledExecutorService reference. Neither method persists its schedule across process restarts or makes the host available.

External scheduler

For a one-shot job, build a runnable JAR or use your deployment’s Java command, then configure the host’s scheduler to launch it weekly. For example, a Unix-like cron entry for 06:00 every Monday is:

0 6 * * 1 /usr/bin/java -jar /opt/university-capture/capture.jar \
  https://www.example-university.edu/ >> /var/log/university-capture.log 2>&1

Cron uses the host’s local timezone unless configured otherwise. Validate timezone, working directory, Java path, environment variables, output permissions, and the scheduler’s handling of missed runs for your environment. The example launches one process per run, so it does not need a resident Java scheduler.

5. Configure stable captures

  • Viewport: Set a consistent width and height. Responsive pages can show different navigation or content at different widths.
  • Wait condition: document.readyState reaching complete is a broad load signal. For delayed data, wait for a specific selector or a known condition, with a bounded timeout.
  • Settling delay: A short delay can allow animations or asynchronous rendering to settle, but avoid relying on an unnecessarily long fixed sleep.
  • Lazy content: Images or sections may load only after scrolling. Scroll the relevant areas or use a suitable full-page capture strategy, then verify the resulting image.
  • Cookie and consent UI: A banner can cover page content or change between visits. Handle it only in a way permitted by the site and appropriate for the archive’s purpose; record the policy used so weekly images remain comparable.
  • Output names: Include a stable page identifier and capture date. Avoid using arbitrary page titles as filenames without sanitizing them.
  • Storage: Choose a persistent destination with sufficient retention space. A local directory on an ephemeral host may disappear when the job environment is replaced.

6. Respect site rules and plan for operations

Check the terms and access guidance for each institution before automating it, and keep request frequency reasonable. Selenium cautions that some sites do not permit automation and others may block it; there is no uniform rule for Indian university webpages. The Selenium project says: “Selenium will let you do this, but please make sure you are familiar with the website’s terms of service as some websites do not permit it and others will even block Selenium.” See Selenium’s guidance on discouraged practices.

For a maintainable weekly archive, log the run start and end, URL, output path, and exception details. Alert on failed runs or missing output, retain a small number of recent logs, and periodically confirm that representative images still contain the expected page. These are operational practices, not guarantees from Selenium. Keep screenshots access-controlled if pages or captured data are not intended for public distribution.

7. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. One GET request returns an image or PDF; its API accepts common screenshot parameter names to make switching easier. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://www.example-university.edu/ \
  -o university.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://www.example-university.edu/"},
    timeout=90,
)
r.raise_for_status()
with open("university.webp", "wb") as f:
    f.write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://www.example-university.edu/'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('university.webp', new Uint8Array(await res.arrayBuffer()));

In Node.js environments without Bun, write the response bytes with the built-in node:fs/promises writeFile function.

  • Cookie and consent banners, newsletter popups, and chat widgets are removed before capture; each cleanup step can be turned off.
  • Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing. Response headers report the page verdict and billing status.
  • An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents and MCP clients.
  • 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000 screenshots.

Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.

8. Troubleshooting

Symptom Likely cause Fix
ChromeDriver cannot start or reports a session creation error Chrome and ChromeDriver major versions differ, or Chrome is unavailable Check the installed browser version and use the matching driver; let Selenium Manager resolve the driver where suitable, or install a compatible browser-driver pair.
Chrome exits immediately in a server environment Headless dependencies, permissions, sandboxing, or container resource limits Read the ChromeDriver and browser logs, install required system libraries, and adjust container-specific launch configuration only as needed.
Timeout while navigating The site is slow, has stalled resources, or automation is blocked Use a bounded page-load timeout, log the failing URL, check access rules and connectivity, and avoid increasing the timeout indefinitely.
Screenshot is mostly blank or missing a section Content rendered after the readiness check, requires scrolling, or depends on client-side data Wait for the specific content selector or condition; scroll to trigger lazy loading; validate with a representative page.
Cookie dialog covers the page Consent state differs across runs or the banner is injected after page load Decide a policy-compliant handling approach and wait for the resulting page state before capture.
A target disappears from future captures Markup or CSS selectors changed Fail the page explicitly when the expected element is absent, log the selector and URL, then update the selector after reviewing the page.
Weekly job silently stops An exception escaped a periodic task; periodic executions can be suppressed after abnormal completion Catch and log failures around each run, monitor the last successful capture time, and restart or alert through the deployment’s process supervisor.
Output file is missing or overwritten Directory permissions, unstable storage, or a non-unique filename Create the output directory before capture, check write permissions, use a timestamp or run ID, and store artifacts on persistent storage.

9. Performance, reliability, and cost

Local Selenium costs include the compute and storage for a browser process per capture, plus the engineering effort to maintain Java, Chrome, and compatible browser automation dependencies. Opening one browser for every URL isolates failures but adds startup time; reusing one browser can reduce startup overhead, but requires careful cleanup and recovery if a session becomes unhealthy. Measure the workload on the pages and host you intend to use rather than assuming a capture duration.

Keep captures sequential unless you have a reason and enough resources to run several browsers at once. Parallel sessions increase memory and CPU use and can create a higher request rate toward university sites. Bound navigation and element waits, set a sensible browser viewport, and avoid capturing more pages or full-page content than the archive needs.

For reliability, use an external scheduler if the job should be independently launched and monitored, and decide explicitly what should happen after a missed weekly run. In-process scheduling requires the Java process and host to stay alive; it does not catch up automatically after a restart. Verify artifacts and retain logs so a failed run is visible. Rendering may vary with browser version, fonts, network conditions, consent state, and responsive layout, so do not treat screenshots from different environments as pixel-identical.

10. Frequently asked questions

Does the program capture all Indian university websites the same way?

No. Page structure, access rules, scripts, content timing, and blocking behavior vary by site. Configure and validate each approved target individually.

Will a weekly scheduler run if the machine is shut down?

An in-process Java scheduler cannot run while its process is stopped. External schedulers also depend on their host and configuration; check whether your chosen scheduler makes up missed runs.

Can I save JPEG instead of PNG?

The shown WebDriver screenshot flow returns PNG bytes. If you need JPEG or another format, convert the image with an image library and document the conversion settings used.

Can I compare screenshots automatically?

Yes, but establish consistent viewport, browser, fonts, timing, and page state first. Treat dynamic dates, rotating banners, and other expected changes separately from meaningful layout changes.