ScreenshotNeo

BlogHow-to

How to Schedule Website Screenshots with Selenium Grid

Use an external scheduler to run Selenium jobs against Grid, capture pages in remote browsers, and save screenshots as artifacts.

By the ScreenshotNeo team4 October 202610 min read

Direct answer: Selenium Grid does not schedule screenshots on a calendar. An external scheduler—such as cron or a CI schedule—starts a script at the cadence you choose. That script requests a remote browser session from Grid, opens a page, waits for site-specific readiness, captures a screenshot, saves the image where your job runner can retain it, and quits the session. Grid routes browser sessions to available Nodes; your scheduler controls when the capture job runs.

This guide builds that workflow with Selenium Python, explains Grid topologies and capacity planning, and shows where screenshots and failure handling belong. The same separation applies to other Selenium language bindings.

1. Choose and start a Grid topology

Use Standalone for local development or a small single-machine setup. Use Hub and Node when one entry point should route sessions to machines with different browser or platform coverage. Use Distributed when Grid components need separate deployments. The Grid URL is the Standalone address, Hub address, or Router address, respectively; the documented default is http://localhost:4444. See Selenium’s Grid getting started guide and Grid endpoints.

For a local Standalone instance, download the Selenium Server JAR as described in Selenium’s documentation, then start it with Java:

java -jar selenium-server-<version>.jar standalone

Replace <version> with the JAR version you downloaded. The process must remain running while your capture script executes. Check that Grid responds before scheduling work:

curl --fail http://localhost:4444/status

You can also open http://localhost:4444 for the Grid UI. For Hub and Node or Distributed deployments, follow the documented startup commands and networking requirements for that topology. In particular, Nodes and Grid components must be able to reach the right Grid services; simply changing the client URL does not configure a distributed deployment.

2. Install Selenium and configure the capture job

The example below uses Python 3 and Selenium’s Python binding. It asks Grid for a Chrome session, navigates to a page, waits for a visible heading, and writes a viewport screenshot to a directory supplied by the job environment. The explicit heading check is an example of a page-specific readiness condition; choose a selector that indicates the content you need has actually appeared.

python -m pip install selenium

Save this as capture.py:

import os
from pathlib import Path
from urllib.parse import urlparse

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait

GRID_URL = os.environ.get("SELENIUM_GRID_URL", "http://localhost:4444")
TARGET_URL = os.environ.get("TARGET_URL", "https://example.com")
ARTIFACT_DIR = Path(os.environ.get("ARTIFACT_DIR", "artifacts"))


def safe_filename(url: str) -> str:
    host = urlparse(url).netloc.replace(":", "_") or "page"
    return "screenshot-" + host.replace("/", "_") + ".png"


def main() -> None:
    ARTIFACT_DIR.mkdir(parents=True, exist_ok=True)
    options = Options()
    options.set_capability("browserName", "chrome")
    # Request a predictable viewport for repeatable comparisons.
    options.add_argument("--window-size=1440,1000")

    driver = None
    try:
        driver = webdriver.Remote(command_executor=GRID_URL, options=options)
        driver.set_page_load_timeout(60)
        driver.get(TARGET_URL)

        # Replace this with a selector that represents readiness for your page.
        WebDriverWait(driver, 30).until(
            EC.visibility_of_element_located((By.TAG_NAME, "h1"))
        )

        output_path = ARTIFACT_DIR / safe_filename(TARGET_URL)
        if not driver.save_screenshot(str(output_path)):
            raise RuntimeError("WebDriver did not save the screenshot")
        print(f"Saved {output_path} for {TARGET_URL}; session={driver.session_id}")
    finally:
        if driver is not None:
            driver.quit()


if __name__ == "__main__":
    main()

Run the script with environment values appropriate to your deployment:

SELENIUM_GRID_URL=http://localhost:4444 \
TARGET_URL=https://example.com \
ARTIFACT_DIR=artifacts \
python capture.py

The code creates the output directory and always attempts to close a created session, including when navigation or capture fails. save_screenshot captures the current browser viewport. Selenium also supports screenshots of individual elements; the documented examples show the element screenshot API in several bindings (Selenium WebDriver screenshot examples). Do not assume this call produces a full-document page image: full-page behavior is browser- and binding-dependent, so confirm the API and output for the exact setup if you need that scope.

3. Put the cadence in an external scheduler

For a simple Unix-like host, cron can invoke the script every day at 06:00 UTC. Use absolute paths and ensure the cron environment can reach Grid and write the artifact directory:

0 6 * * * cd /opt/site-capture && SELENIUM_GRID_URL=http://grid.internal:4444 TARGET_URL=https://example.com ARTIFACT_DIR=/var/lib/site-capture/artifacts /usr/bin/python3 capture.py

On a CI runner, configure the platform’s scheduled workflow to run the same command and retain the artifact directory using that platform’s artifact mechanism. The exact schedule syntax, timezone rules, retention duration, and artifact limits depend on the scheduler or CI product. Selenium Grid supplies remote browser sessions; it does not provide the calendar trigger or automatically archive the screenshot files.

For a reliable scheduled job, decide explicitly:

  • Cadence and overlap: prevent a second run from piling up if a previous run is still active, unless parallel captures are intended.
  • Artifact naming and retention: include a date or run identifier if you keep repeated snapshots, and set retention in the runner or storage system.
  • Failure reporting: return a nonzero exit status on capture failures and have the scheduler surface failed runs.
  • Targets: validate URLs and keep credentials out of source code and logs.

4. Select browsers, versions, platforms, and viewport

Grid assigns a requested session to a Node whose available slot matches the requested capabilities. A request for Chrome cannot be satisfied by a Firefox-only slot. Add browser version or platform requirements when the Grid has matching Nodes and you need those environments. Selenium’s Grid architecture documentation explains slots and matching; the getting started guide demonstrates browser and platform capabilities.

In Python, set standard capabilities through browser options, for example:

options = Options()
options.set_capability("browserName", "chrome")
options.set_capability("browserVersion", "<version-available-on-your-grid>")
options.set_capability("platformName", "<platform-available-on-your-grid>")

Use the actual browser versions and platform names advertised by your Nodes. Unsupported or mismatched requests may remain queued or fail; requesting a capability does not install that browser on a Node. Viewport size is a browser window setting, separate from Grid’s browser and platform matching. For longitudinal visual comparisons, keep browser version, viewport, device scale settings, target state, and readiness checks consistent.

5. Schedule a matrix or many URLs

There are two common ways to scale a screenshot schedule:

  • One job, sequential captures: loop over URLs, creating and closing a session for each capture or reusing a session when pages can safely share state. Sequential work is easier to size and diagnose but takes longer.
  • Multiple workers or scheduled jobs: submit independent captures concurrently. Grid queues new session requests and the Distributor assigns matching requests to free slots. If all suitable slots are busy, requests can wait; queue timeouts and job-level timeouts need to be considered. See Grid components and the New Session Queue.

Keep each output path unique when workers run concurrently. Bound worker concurrency to the available browser slots and host resources; increasing job workers beyond capacity usually increases waiting rather than throughput. If you need multiple operating systems or browser versions, provision Nodes with those environments and request matching capabilities.

6. Plan capacity, performance, and reliability

Grid capacity depends on the browser and operating-system matrix, concurrent sessions, available machines, and machine power. Selenium gives approximately 1 GB of RAM per browser session as a sizing reference, not a guarantee; measure with your actual pages and browser versions. Monitor session creation time, queue delay, capture duration, memory, CPU, and failure rate. Start with conservative concurrency, then adjust from observed behavior. See Selenium’s Grid sizing guidance.

Pages with large assets, third-party scripts, animations, or delayed content can dominate capture time. Use a meaningful readiness condition rather than an arbitrary long sleep, and set timeouts that fit the slowest acceptable page. A navigation-complete event does not guarantee that asynchronous content, fonts, or images have settled. If the output is used for comparisons, stabilize page data and timing as well as browser settings.

For reliability, record each run’s target URL, requested capabilities, session ID when created, outcome, elapsed time, and artifact location. Retry transient infrastructure failures selectively; repeated retries against a persistently broken page can waste capacity and hide the underlying cause. Always close sessions in cleanup so Nodes release slots. Store artifacts outside ephemeral runner storage if they must survive the job.

7. Protect the Grid and control costs

A Selenium Grid endpoint is powerful infrastructure. Restrict access with network controls and authentication appropriate to your deployment; do not expose an unprotected Grid to arbitrary public clients. Selenium warns that external access can expose Grid infrastructure, internal applications and files, and permit third parties to run custom binaries. Read the official Grid security guidance.

Self-hosted cost includes the machines and operational work needed to run the Grid, browser Nodes, scheduler, and artifact storage. Capacity headroom matters: scheduling many targets in a short window may require more concurrent slots and memory. The cited Selenium material does not provide a universal node count or capture rate, so measure with your page set and workload before committing to a capacity estimate. CI or hosted infrastructure costs depend on the provider and configuration.

8. Troubleshooting common failures

Symptom Likely cause What to check or change
Connection refused or timeout creating a session Wrong Grid URL, Grid is stopped, route/firewall issue, or the job cannot resolve the Grid host. Check /status from the same runner, confirm the Router/Hub address and port, and verify network access between runner and Grid.
Session request waits and then times out No free matching slot, requested browser/platform is unavailable, or all matching slots are busy. Inspect Grid UI or status, verify the requested capabilities against registered Nodes, reduce concurrency, or add appropriate Node capacity.
Browser name or capability rejected Capability does not match Grid slots or is malformed. Use standard capability names and request only browser versions/platforms actually registered. Check the Grid’s advertised Node details.
Screenshot exists but is blank or incomplete Page-specific content was not ready, navigation produced an error page, or the page needs a different readiness condition. Check the current URL and page title, wait for a meaningful element or state, and record browser logs or a diagnostic screenshot where appropriate.
Wait for selector times out The selector is wrong, content is absent, hidden, delayed, or blocked. Confirm the selector locally, use a condition that matches the actual page, and inspect whether authentication, consent, or network restrictions affect rendering.
Session slots remain occupied after a failed run The script exited without closing its WebDriver session. Keep driver.quit() in a finally block and ensure worker cancellation also runs cleanup.
Artifact missing after a successful capture Screenshot was saved on the runner, but the runner did not retain or upload it. Check the path and permissions, then configure the scheduler or CI system to collect that directory before the job ends.
Scheduled run never starts or runs at an unexpected hour Schedule syntax, timezone, branch/ref, or scheduler policy differs from expectation. Inspect scheduler logs and its documented timezone and trigger rules; confirm the workflow is enabled and the target branch is selected.

9. Or skip the browser setup

If your goal is a clean website screenshot rather than managing remote browser sessions, ScreenshotNeo is a website screenshot API and MCP server. One GET request returns an image or PDF. Its API accepts common screenshot API parameter names, and its docs include the ScreenshotNeo API reference.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. An MCP server gives AI agents screenshot, page information, and PDF capture tools. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

Create a free ScreenshotNeo account and start with 1,000 screenshots a month, no card required.

FAQ

Does Selenium Grid run cron jobs?

No. Grid manages remote browser sessions. Use a separate scheduler or CI system to start your capture script at the desired times.

Does Grid save screenshot files for me?

No. The WebDriver client receives and saves the screenshot. Configure your runner or storage system to retain and retrieve the resulting file.

Can I schedule screenshots across different browsers?

Yes. Run jobs with browser capabilities that match the Nodes available in your Grid, and schedule enough time and capacity for the required matrix.

Will the example capture the entire page?

It captures the current viewport. Confirm full-document support for your specific browser and binding before relying on it.