How to use Selenium to capture website thumbnails on a schedule
Build a Selenium screenshot script, wait for pages to render, schedule recurring captures with GitHub Actions, and retain the resulting thumbnails.
To capture website thumbnails on a schedule, write a Selenium script that opens the target URL in headless Chrome, waits for the page state your thumbnail needs, saves a screenshot to a deliberate path, and always closes the browser. Then run that script with a scheduler such as GitHub Actions. Treat the viewport, readiness condition, output naming, and artifact retention as explicit choices: they determine whether each scheduled image is consistent and retrievable.
This guide uses Python for the Selenium implementation and GitHub Actions for scheduling. The same pattern applies on a local machine or another CI runner, but installation and storage setup depend on that environment. Only capture pages you are permitted to access.
1. Define what the thumbnail should show
Before writing code, decide what counts as a successful thumbnail:
- Target URL: use the canonical page URL and include any required query parameters.
- Viewport: set a fixed width and height, such as 1280×720, so captures have consistent framing. A viewport screenshot is not automatically a marketing thumbnail with a custom crop.
- Capture scope: use a viewport screenshot for the visible browser area, a full-page screenshot when the entire page is required, or an element screenshot when a specific component is the subject.
- Ready state: identify a page-specific signal that means the important content has rendered, such as a known element becoming visible.
- History and retention: choose whether new captures replace the previous file or have timestamped names, and where images must live after the job exits.
For pages that depend on asynchronous requests, a successful navigation alone may happen before the content you care about appears. Prefer a condition tied to that content over an arbitrary pause.
2. Install Selenium and prepare headless Chrome
Use a supported Python environment and install Selenium:
python -m pip install selenium
Current Selenium setups can manage browser drivers through Selenium Manager when a compatible browser is installed. In a CI image, verify that the browser is available and that the runner can obtain or use the matching driver. Chrome Headless runs without visible UI. Modern Chrome uses unified Headless; the older Headless implementation is available only as a separate chrome-headless-shell binary starting with Chrome 132.0.6793.0. Avoid copying obsolete flags from old snippets without checking the browser in your environment. See Chrome Headless mode documentation and the Selenium browser documentation.
3. Write a repeatable Selenium capture script
This script makes the output directory, configures a fixed headless viewport, waits for a page-specific element, writes a timestamped PNG, and quits Chrome even if a step fails. Replace the example URL and readiness selector with values for the site you are authorized to capture.
from datetime import datetime, timezone
from pathlib import Path
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
URL = "https://example.com"
READY_SELECTOR = "main" # Replace with a meaningful selector for this page.
OUT_DIR = Path("screenshots")
OUT_DIR.mkdir(parents=True, exist_ok=True)
# UTC timestamps sort consistently across machines and daylight-saving changes.
stamp = datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%SZ")
output = OUT_DIR / f"example-{stamp}.png"
options = Options()
options.add_argument("--headless")
options.add_argument("--window-size=1280,720")
# Keep the driver lifecycle inside try/finally so recurring jobs do not leave
# browser processes behind after a navigation or screenshot error.
driver = webdriver.Chrome(options=options)
try:
driver.set_window_size(1280, 720)
driver.get(URL)
WebDriverWait(driver, 30).until(
EC.visibility_of_element_located(("css selector", READY_SELECTOR))
)
if not driver.save_screenshot(str(output)):
raise RuntimeError(f"Screenshot was not saved: {output}")
print(f"Saved {output}")
finally:
driver.quit()
The example uses a 30-second explicit wait as a ceiling; tune it to the page and runner. It does not guarantee that every image, animation, or third-party widget has finished. If the desired thumbnail depends on lazy-loaded images, scroll the relevant content into view or use a page-specific readiness condition before capturing. If a selector is not present on every page state, choose a more reliable signal rather than waiting for it indefinitely.
Selenium’s Python API includes save_screenshot for the current browsing context. Selenium also supports taking a screenshot of an element when the whole viewport is not needed. See Selenium screenshot documentation.
Choose a filename strategy
- Keep history: use a timestamp or run identifier in each name, as above. Ensure the destination can hold the growing collection and has a cleanup policy.
- Keep only the latest: save to a stable name such as
latest.png. This is simple for downstream consumers, but each run overwrites the previous image. - Separate targets: include a stable, filesystem-safe page identifier in the filename. Do not build paths directly from untrusted URL text.
4. Schedule the script with GitHub Actions
A scheduled workflow runs the script from the repository’s default branch. Add a file such as .github/workflows/capture.yml and adapt the Python version and dependency setup to your project:
name: Capture website thumbnail
on:
schedule:
- cron: "17 6 * * *"
workflow_dispatch:
permissions:
contents: read
jobs:
capture:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.x"
- name: Install Selenium
run: python -m pip install selenium
- name: Capture thumbnail
run: python capture.py
- name: Upload screenshot
uses: actions/upload-artifact@v4
with:
name: website-thumbnail-${{ github.run_id }}
path: screenshots/*.png
if-no-files-found: error
The sample uploads each run’s PNG as a workflow artifact so it can be retrieved from the run. Artifact retention is configurable and depends on the repository and workflow settings; do not assume an artifact is permanent storage. If the image must remain available to a website or another service, add an authorized publishing or object-storage step and define its access and cleanup rules.
GitHub Actions scheduled workflows use POSIX cron syntax, default to UTC, can specify an IANA timezone, and have a shortest interval of five minutes. They run against the latest commit on the default branch and may be delayed during periods of high workflow load, so cron is suitable for recurring capture but not an exact-time guarantee. See GitHub workflow syntax and the schedule event documentation.
Schedule on a local machine or another runner
The capture script is independent of GitHub Actions. A local operating-system scheduler or another hosted runner can invoke it on a recurring basis. Compare options by schedule timing, browser and driver maintenance, network access to the target, and whether output survives the job. Make sure the scheduler runs from the expected working directory or use absolute paths for the script and output. The exact scheduler syntax and artifact configuration are platform-specific.
5. Viewport, full-page, and element captures
The sample saves the current browser view. Use the capture shape that matches the intended image:
- Viewport: set a fixed window size before navigation and capture after the desired content is ready. This is generally the simplest thumbnail shape.
- Element: locate the target element and call its screenshot method, for example
element.screenshot("card.png"). Wait for the element to be visible first, and account for elements outside the viewport or obscured by overlays. - Full page: Selenium’s standard screenshot captures the current viewport. Full-page behavior varies by browser and implementation; if the page must be complete, use a browser-supported approach and verify the output dimensions and lazy-loaded content in your environment.
Browser viewport dimensions and the saved image’s pixel dimensions can differ because of device scale settings. If exact output dimensions matter, inspect the produced file and account for browser scaling rather than assuming the window-size argument alone defines every pixel.
6. Or skip the browser setup
If you need scheduled screenshots without managing Selenium, Chrome, and a driver, ScreenshotNeo is a website screenshot API and MCP server. Its one-call API accepts a URL and returns an image or PDF. This cURL example saves a WebP screenshot; replace the URL and API key with your values. See the ScreenshotNeo API documentation for options and setup.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.com \
-o shot.webp
ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, with no card required.
7. Keep scheduled captures reliable
- Make runs independent: create directories as needed, use explicit waits, and close the driver in a
finallyblock. - Make failures visible: fail the job when the screenshot is missing, and retain enough run logs to diagnose navigation and selector failures.
- Check the result: verify that the expected file exists and is non-empty. If consumers depend on exact dimensions, add a separate validation step appropriate to your project.
- Handle changing pages: content, consent dialogs, authentication, and bot protections can change. Use only authorized access and do not try to bypass access controls.
- Control growth: timestamped files accumulate. Define retention or cleanup for both local files and uploaded artifacts.
- Mind timing: a scheduled trigger starts a job; page load and rendering add time, and hosted schedules may be delayed. Avoid relying on an exact minute.
8. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Chrome or driver fails to start | Browser missing, driver mismatch, or runner environment lacks required browser support. | Confirm Chrome is installed and Selenium can access a compatible driver. Review the runner’s browser setup and Selenium Manager behavior. |
| Screenshot is blank or missing expected content | Capture ran before asynchronous rendering completed, or the wrong URL/page state was loaded. | Wait for a meaningful page-specific element or state. Check navigation and page logs, and confirm the selector matches the current page. |
| Explicit wait times out | The selector is wrong, hidden, delayed, or absent in this page variant. | Inspect the page structure and choose a condition that is present for the expected state. Increase the timeout only when the content legitimately needs longer. |
| Output file is not found by the workflow | Script ran from a different working directory, wrote elsewhere, or used a different extension/path. | Use a known output directory, print the saved path, and make the artifact path match it. |
| Every run replaces the prior image | Filename is constant. | Use a timestamp or run identifier for history, or keep the stable name intentionally if only the latest image is needed. |
| Scheduled workflow has not started at the exact cron minute | Hosted scheduling can be delayed under load. | Allow for delay and design downstream steps to consume the latest completed capture rather than assuming exact-time execution. |
| Capture works locally but not in CI | Different browser availability, network policy, permissions, working directory, or environment configuration. | Compare the runner environment with local prerequisites, ensure the target is reachable from the runner, and inspect job logs. |
| Artifacts disappear or cannot be found later | Workflow artifacts have configured retention, or files were only stored on an ephemeral runner. | Check artifact retention settings and upload the file. Use persistent storage when longer availability is required. |
9. Performance, reliability, and cost
Each capture starts or uses a browser, navigates to a page, waits for readiness, and writes an image. The page’s complexity, network, assets, and wait condition determine much of the run time. A fixed viewport, a specific readiness selector, and avoiding unnecessary full-page work help keep the job focused; there is no universal capture time for arbitrary sites.
A hosted CI runner avoids maintaining an always-on local machine, but browser setup and artifact persistence remain part of the workflow. A local scheduler gives you control over the host and files but depends on that machine being available and networked at run time. Neither approach makes schedule start times exact. For a small number of recurring pages, evaluate the cost in runner minutes, storage, and maintenance against the frequency and retention you need. Selenium itself does not charge per screenshot; your execution and storage environment may have its own costs and limits.
FAQ
Can Selenium save a screenshot as JPEG?
The example uses Selenium’s PNG screenshot output. If another format is required, convert the resulting image with an image library as a separate step and validate the output.
Can I run the capture every five minutes?
GitHub Actions documents five minutes as the shortest schedule interval, but scheduled starts can be delayed. Choose a scheduler designed for the timing guarantees your use case requires.
Will the script keep screenshots after the job ends?
Files on a runner are not automatically durable. Upload them as artifacts or publish them to storage configured for your retention needs.
Does Selenium capture pages that require a login?
It can interact with pages your account is authorized to access, but authentication handling must be designed for the site and kept secure. Do not commit credentials to the repository.


