ScreenshotNeo

BlogHow-to

How to Schedule Website Screenshots in Python with APScheduler

Schedule recurring website screenshots with APScheduler and Playwright. Choose interval or cron timing, save viewport or full-page captures, and plan for restarts and slow pages.

By the ScreenshotNeo team4 October 202610 min read

Use APScheduler to decide when a Python function runs, and Playwright to open the page and capture it. Choose an interval trigger for a fixed elapsed cadence, such as every 30 minutes, or a cron trigger for calendar times, such as weekdays at 09:00. The example below uses APScheduler 3.x and Playwright’s synchronous API; it creates a viewport screenshot by default and can be switched to full-page capture.

1. Install APScheduler and Playwright

These examples use the APScheduler 3.x interface. Keep the major version below 4 so the BackgroundScheduler and add_job() code below matches the installed API. Playwright’s Python library supports synchronous and asynchronous APIs and runs browsers headlessly by default. Install the browser binary separately from the Python package. See the APScheduler 3.x guide and Playwright installation guide.

python -m venv .venv

# macOS/Linux
. .venv/bin/activate

# Windows PowerShell
# .venv\Scripts\Activate.ps1

python -m pip install "APScheduler>=3.10,<4" playwright
python -m playwright install chromium

On Linux deployments, install the browser’s operating-system dependencies in the same image or host that runs the job. Playwright documents python -m playwright install --with-deps chromium for environments where those dependencies can be installed automatically; see the browser installation documentation. When upgrading Playwright, check whether its required browser binaries also need to be installed again.

2. Create a screenshot function

Save this as scheduled_screenshot.py. The callable is a module-level function, which is useful if you later move its APScheduler job into a persistent job store. It makes the output directory, navigates to a URL, waits for the page load event, and closes the browser even if navigation or screenshot capture raises an error.

from pathlib import Path
from datetime import datetime, timezone
import logging

from playwright.sync_api import sync_playwright

logging.basicConfig(
    level=logging.INFO,
    format="%(asctime)s %(levelname)s %(message)s",
)

OUTPUT_DIR = Path("captures")
TARGET_URL = "https://example.com"


def capture_website(url: str = TARGET_URL, full_page: bool = False) -> None:
    """Capture a website and save a timestamped PNG."""
    OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
    timestamp = datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%SZ")
    output_path = OUTPUT_DIR / f"example-{timestamp}.png"

    with sync_playwright() as playwright:
        browser = playwright.chromium.launch()
        try:
            page = browser.new_page(viewport={"width": 1440, "height": 900})
            response = page.goto(
                url,
                wait_until="load",
                timeout=45_000,
            )
            if response is not None and response.status >= 400:
                raise RuntimeError(
                    f"Navigation returned HTTP {response.status} for {url}"
                )

            page.screenshot(path=str(output_path), full_page=full_page)
            logging.info("Saved screenshot: %s", output_path)
        finally:
            browser.close()

page.goto() waits for the configured navigation condition, but that does not guarantee that every application-specific widget or late-loading image is ready. If the page has a known stable element, wait for it before capturing:

page.locator("main article").wait_for(state="visible", timeout=15_000)
page.screenshot(path=str(output_path), full_page=full_page)

Use a selector that exists on the target page. For a fixed delay, use page.wait_for_timeout(2_000) sparingly; a selector is usually more reliable than guessing how long the page needs. Playwright supports viewport, full-page, buffer, and element screenshots; see its screenshot guide.

3. Schedule it with APScheduler

Add one of the following trigger choices to the same file. The runner starts a background scheduler and then keeps the process alive. A background scheduler does not keep capturing after the surrounding Python process exits.

Fixed elapsed interval

Use interval for a cadence such as every 30 minutes. The interval is when the job becomes due; it is not a guarantee that a capture finishes within 30 minutes.

from apscheduler.schedulers.background import BackgroundScheduler


def run_interval_schedule() -> None:
    scheduler = BackgroundScheduler(timezone="UTC")
    scheduler.add_job(
        capture_website,
        trigger="interval",
        minutes=30,
        id="example-site-screenshot",
        replace_existing=True,
        max_instances=1,
        coalesce=True,
        misfire_grace_time=300,
    )
    scheduler.start()
    logging.info("Screenshot schedule started; next run: %s", scheduler.get_job(
        "example-site-screenshot"
    ).next_run_time)

    try:
        # Keep this process running. A service manager can supervise it in production.
        import threading
        threading.Event().wait()
    except (KeyboardInterrupt, SystemExit):
        scheduler.shutdown(wait=True)


if __name__ == "__main__":
    run_interval_schedule()

Run it with python scheduled_screenshot.py. To test the capture function once without waiting for the trigger, temporarily call capture_website() from the main block.

Calendar schedule with cron

Use cron for wall-clock rules such as 09:00 Monday through Friday. Set the timezone explicitly if the schedule must follow local time across daylight-saving changes. APScheduler’s cron fields combine to determine matching fire times; its CronTrigger reference documents the supported fields.

from apscheduler.schedulers.background import BackgroundScheduler

scheduler = BackgroundScheduler(timezone="America/New_York")
scheduler.add_job(
    capture_website,
    trigger="cron",
    day_of_week="mon-fri",
    hour=9,
    minute=0,
    id="weekday-morning-capture",
    replace_existing=True,
    max_instances=1,
    coalesce=True,
    misfire_grace_time=900,
)
scheduler.start()

This snippet needs the same process-lifetime handling as the interval runner: wait in the foreground or let a service manager supervise a dedicated scheduler process. Cron means a local calendar time in the configured timezone; interval means elapsed time between runs. For interval behavior and its parameters, see the IntervalTrigger reference.

4. Choose capture scope and page readiness

Need Playwright approach Tradeoff
Visible browser area page.screenshot(path="capture.png") Bounded output; content below the viewport is not included.
Entire scrollable page page.screenshot(path="capture.png", full_page=True) Can produce tall, large images; pages with infinite scrolling may not have a meaningful end.
One component page.locator(".report").screenshot(path="report.png") Requires a selector that resolves to the intended element.
Process image bytes in memory image_bytes = page.screenshot() Lets the job upload or transform bytes without first writing a local file.

Use a deliberate readiness condition. wait_until="load" is suitable for many pages, but client-rendered sites can continue changing after the load event. Wait for a meaningful locator when available. A fixed delay can be a fallback, but it adds latency on every run and can still be too short under load.

For one target, a single job is simplest. For multiple sites, schedule one job per site when they need independent timing, retries, logging, or retention. A dispatcher job that iterates over a target list is simpler when they share the same policy. In either design, give each output a stable target identifier and a timestamp, and define how long old captures are retained.

5. Handle slow runs, missed schedules, and restarts

APScheduler 3.x permits one instance of a job by default. If the prior capture is still running when the next run becomes due, the later run may be treated as a misfire. Set max_instances, coalesce, and misfire_grace_time to match the desired behavior. The 3.x guide explains concurrent instances and missed executions and coalescing.

  • One capture at a time: keep max_instances=1 to avoid overlapping browser processes for a target.
  • Skip a backlog: coalesce=True combines multiple missed run times into one run instead of replaying every missed screenshot.
  • Allow a late run: choose a misfire_grace_time appropriate to the schedule. A job delayed longer than the grace period can be skipped.
  • Allow overlap only deliberately: increasing max_instances can increase memory and CPU use and may cause duplicate or out-of-order outputs.

The default in-memory job store loses its schedules if the process exits or crashes. A persistent store can preserve job definitions, but it does not keep the Python process alive; use a service manager or container supervisor to restart the scheduler process. For APScheduler 3.x, one common persistent option is a SQLite-backed SQLAlchemy job store:

from apscheduler.jobstores.sqlalchemy import SQLAlchemyJobStore
from apscheduler.schedulers.background import BackgroundScheduler

jobstores = {
    "default": SQLAlchemyJobStore(url="sqlite:///scheduler.sqlite")
}
scheduler = BackgroundScheduler(
    jobstores=jobstores,
    timezone="UTC",
)
scheduler.add_job(
    capture_website,
    trigger="interval",
    minutes=30,
    id="example-site-screenshot",
    replace_existing=True,
    max_instances=1,
    coalesce=True,
    misfire_grace_time=300,
)
scheduler.start()

Install the SQLAlchemy extra before using that store: python -m pip install "APScheduler>=3.10,<4[sqlalchemy]" may not be accepted by every shell/package parser because the extra syntax belongs directly after the package name. Prefer python -m pip install "APScheduler[sqlalchemy]>=3.10,<4". Give startup-created persistent jobs explicit IDs and use replace_existing=True so each restart does not add a duplicate. Persistent job stores serialize job details; use trusted storage and keep scheduled callables importable. See the APScheduler 3.x job store guidance.

Current APScheduler documentation describes a newer task, schedule, and data-store API. Do not copy current-generation examples into a 3.x installation or mix their startup methods with BackgroundScheduler. See the current guide and migration notes when choosing that API generation.

Or skip the browser setup

For recurring captures, call the ScreenshotNeo API from your scheduled Python job and save the response bytes. One GET request returns an image or PDF; the example saves a WebP screenshot of a target page. Add the request to your APScheduler callable, then schedule that callable with the interval or cron pattern above. See the ScreenshotNeo API documentation.

import requests

def capture_with_screenshotneo():
    r = requests.get(
        "https://api.screenshotneo.com/v1/shot",
        params={
            "access_key": "YOUR_API_KEY",
            "url": "https://example.com",
        },
        timeout=90,
    )
    r.raise_for_status()
    with open("shot.webp", "wb") as image_file:
        image_file.write(r.content)

ScreenshotNeo accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; those steps can be turned off. Bot checks, blank pages, and failed loads are not billed, and response headers report the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 screenshots per month with no card, and paid plans start at $5 for 3,000. Learn about ScreenshotNeo, then sign up for 1,000 free screenshots a month with no card.

6. Troubleshooting

Symptom Likely cause Fix
Browser executable is missing The Python package is installed, but the matching browser binary is not. Run python -m playwright install chromium in the same environment used by the scheduler.
Browser fails to launch on Linux Required operating-system libraries are absent, or the runtime image differs from the install environment. Install Chromium dependencies in the deployment image; consult Playwright’s browser and system dependency instructions.
Screenshot is blank or missing page content Navigation completed before the client-rendered content appeared, or the site returned an error page. Check the navigation response status, wait for a stable selector, and log navigation exceptions and output paths.
Job never runs again The Python process exited, or only a one-shot script invocation was used. Keep the scheduler process running under a supervisor. A persistent store preserves schedule data but does not run a stopped process.
Duplicate captures after restart A persistent store received a new job on each startup under a new ID. Use a stable explicit job ID and replace_existing=True.
Some runs are skipped A prior instance is still running, a run exceeded its misfire grace period, or coalescing collapsed missed runs. Review job duration and scheduler logs, then tune max_instances, misfire_grace_time, and coalesce to your policy.
Wrong local run time around a clock change The scheduler timezone was left implicit or set to UTC while the desired schedule is local wall-clock time. Set the scheduler timezone explicitly and verify the intended calendar rule.
Persistent job cannot be restored The callable is nested, unavailable at its saved import path, or its arguments cannot be serialized. Use a module-level importable function and simple serializable arguments; deploy the same module where the scheduler reads the job store.

7. Performance, reliability, and cost

Each local Playwright run launches a browser and loads the target page, so resource use and duration depend on the page and runtime environment. Reusing a browser can reduce repeated startup work for high-frequency captures, but then browser, context, and page cleanup becomes part of the long-running process design. Start with one browser per job for straightforward isolation. Limit concurrent jobs if CPU, memory, target-site rate limits, or output storage become constraints.

Full-page images can be much larger and take longer to capture than viewport images. Set the viewport intentionally, prefer a selector wait over an arbitrary long delay, and choose an output format and retention policy that match the downstream use. Log the scheduled target, start and finish times, duration, output file, and exceptions. Add retry behavior only for transient failures, with a bounded attempt count and delay so an unavailable site does not create an unending queue.

APScheduler and Playwright are software dependencies; local capture cost is the compute, storage, and operations required to keep the browser environment running. A persistent database and an always-on supervised process add deployment work. With ScreenshotNeo, the API handles the browser setup, and the stated plans are free for 1,000 shots per month, then $5 for 3,000, $15 for 15,000, $39 for 60,000, $99 for 250,000, and $249 for 1,000,000; yearly billing gives two months free. Every feature is on every plan.

FAQ

How do I take a screenshot of a website automatically every day?

Use an APScheduler cron trigger with the desired hour, minute, and timezone, and have the scheduled function navigate with Playwright and save the image.

How can I capture a full-page screenshot with Playwright?

Pass full_page=True to page.screenshot(). This captures the full scrollable page rather than just the current viewport.

How do I keep an APScheduler job after restarting my app?

Use a persistent job store and a stable job ID with replace_existing=True. Also run the scheduler under a process supervisor so it starts again after a failure.

Should I use interval or cron?

Use interval for elapsed cadence, such as every 30 minutes. Use cron for calendar rules, such as weekday mornings at a chosen local time.