ScreenshotNeo

BlogHow-to

How to Schedule Website Screenshots with Python on an Indian Linux Hosting Plan

Set up Playwright and cron to capture a website on a schedule, with practical checks for Indian shared hosting, VPS permissions, time zones, and failures.

By the ScreenshotNeo team4 October 202610 min read

Use Python with Playwright to open the site and save a screenshot, then schedule the script with cron or your hosting control panel. The main constraint on an Indian Linux hosting plan is whether it allows the browser binary and its required Linux libraries to run: a Python package installation by itself may not be enough. Check that first, then use a virtual environment, absolute paths, a writable output directory, and a log.

This setup works on shared hosting only when the provider permits the required browser process and scheduled task. If it does not, ask the host about headless Chromium support or use a Linux VPS where you control system packages. Requirements and resource limits vary by provider and plan.

1. Check whether your hosting plan can run Playwright

Before building the schedule, confirm these points with your host or by running a small launch test:

  • Python 3 is available and you can create a virtual environment.
  • Your account can install Python packages and download Playwright’s Chromium binary.
  • The browser’s Linux system dependencies are present or your plan lets you install them. Installing with pip cannot install operating-system libraries that need root access.
  • Headless browser processes are allowed, and the plan’s CPU, memory, process, and execution-time limits accommodate your capture frequency.
  • You have a scheduled-task feature or permission to use user crontab, and a writable directory with enough room for retained screenshots and logs.
  • You know what timezone the scheduler uses; do not assume a server in India runs cron in IST.

Playwright’s official documentation covers its Python library, browser installation, and Linux dependencies. Its browser binaries are version-specific, so reinstall the browser after upgrading Playwright. For headless Chromium-only use, its browser guide documents an option to install only the headless shell. Playwright Python library, browser installation and dependencies.

As one provider-specific example, Hostinger describes scheduled-task availability on some Web and Cloud plans and direct crontab customization on VPS; it describes root access for its self-managed Linux VPS. These are Hostinger’s stated plan details, not a general rule for Indian hosting. Check the features and restrictions of your own plan. Hostinger scheduled tasks, Hostinger Linux VPS.

2. Create a project and install Playwright

Run these commands in a shell as the account that will own the scheduled job. The example uses the home directory and Chromium:

mkdir -p "$HOME/site-shot/output"
cd "$HOME/site-shot"
python3 -m venv .venv
.venv/bin/python -m pip install --upgrade pip
.venv/bin/python -m pip install playwright
.venv/bin/python -m playwright install chromium

A virtual environment keeps this project’s Python packages isolated. Calling its interpreter by its full path later means cron does not need to activate the environment. Python’s documentation notes that virtual environments are tied to their installation and should not be copied or moved; recreate one if you move the project. Python venv documentation.

If browser installation reports missing shared libraries, the host must provide those operating-system dependencies. On a server where you have sudo rights, Playwright documents .venv/bin/python -m playwright install --with-deps chromium; on restricted shared hosting, this may fail because dependency installation requires system package access. Ask the provider whether Chromium dependencies are already installed. Do not try to solve a missing system library by repeatedly reinstalling the Python package.

After installing, a quick launch check is:

cd "$HOME/site-shot"
.venv/bin/python -c 'from playwright.sync_api import sync_playwright; p=sync_playwright().start(); b=p.chromium.launch(); print("Chromium launched"); b.close(); p.stop()'

3. Write a screenshot script

Save this as capture.py inside $HOME/site-shot. It creates a UTC timestamped full-page PNG and always closes the browser if navigation or capture raises an error.

from datetime import datetime, timezone
from pathlib import Path
from playwright.sync_api import sync_playwright

URL = "https://example.com"
PROJECT_DIR = Path(__file__).resolve().parent
OUTPUT_DIR = PROJECT_DIR / "output"
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)

stamp = datetime.now(timezone.utc).strftime("%Y%m%dT%H%M%SZ")
output_file = OUTPUT_DIR / f"example-{stamp}.png"

with sync_playwright() as p:
    browser = p.chromium.launch()
    try:
        page = browser.new_page(viewport={"width": 1440, "height": 1000})
        response = page.goto(URL, wait_until="domcontentloaded", timeout=60000)
        if response is not None and response.status >= 400:
            raise RuntimeError(f"Page returned HTTP {response.status}: {URL}")
        page.screenshot(path=str(output_file), full_page=True, animations="disabled")
        print(f"Saved {output_file}")
    finally:
        browser.close()

Playwright runs headless by default. The example waits for domcontentloaded rather than networkidle because analytics, chat, and other persistent requests can keep a page active. If a target renders its main content later, wait for a site-specific element before capturing, such as page.locator("main").wait_for(state="visible", timeout=15000). The screenshot API also supports viewport-only images, element screenshots, image format and quality, and screenshot bytes; see the screenshot API reference.

Useful capture variations

  • Viewport only: remove full_page=True. This captures the current viewport, not the entire scrollable document.
  • One element: use page.locator("main").screenshot(path=str(output_file)). The locator must match an element; an element inside a scrollable container captures only its currently visible content.
  • JPEG or WebP: use a matching filename extension such as .jpg or .webp; set quality=80 for lossy formats. PNG does not use the quality option.
  • Retina-sized output: set the browser context’s device_scale_factor, or use screenshot scale options as appropriate. Device scale can increase image dimensions and file size.
  • Wait for an application state: prefer a meaningful selector or explicit short delay for known client-side rendering over a blanket long wait.

4. Run it manually before scheduling

Use the same account and command that cron will use:

cd "$HOME/site-shot"
.venv/bin/python /home/USER/site-shot/capture.py

Replace /home/USER with the actual absolute home path. Check that a new file appears in output, opens as an image, and contains the expected page state. This catches path, permission, network, and browser launch problems before they become silent scheduled failures.

5. Schedule the capture with cron or a hosting panel

Use cron

Open the current account’s crontab:

crontab -e

For a daily run at 08:15, add:

15 8 * * * /home/USER/site-shot/.venv/bin/python /home/USER/site-shot/capture.py >> /home/USER/site-shot/cron.log 2>&1

Replace USER and confirm the absolute paths. This schedule means 08:15 in the cron table’s effective timezone, which may differ from India Standard Time. On implementations that support it, put CRON_TZ=Asia/Kolkata on its own line before the job to request IST. Verify support with the host: cron implementations differ. The Linux crontab manual documents CRON_TZ and notes that log timestamps use the daemon’s local timezone. Linux crontab manual.

Common schedule examples, with minute then hour fields:

Schedule Cron expression
Daily at 08:15 15 8 * * *
Every 6 hours at minute 0 0 */6 * * *
Weekdays at 09:00 0 9 * * 1-5
Every 30 minutes */30 * * * *

Frequent captures can create many files and consume account resources. Use a schedule appropriate to the purpose and add retention cleanup if needed. Cron jobs run as the owner of the crontab, in a limited environment; use absolute paths and log both standard output and errors. Cron environment and schedule details.

Use a hosting control panel

If the provider has a scheduled-task dashboard, enter the command in the format it specifies, using the full path to the virtual environment’s Python and script. Set the interval and timezone offered by the panel, and direct output to a log if the panel allows it. Some shared plans expose only a constrained list of intervals or impose execution limits; these are provider-specific. Do not also install a duplicate crontab entry unless you want the job to run twice.

6. Make recurring captures reliable

  • Keep paths stable: cron’s working directory is not your project directory. The script uses __file__ for output paths, and the cron command uses absolute paths.
  • Make failures visible: retain the log and inspect its latest output after the first scheduled run. Configure log rotation or truncate/archive old logs so they do not grow without limit.
  • Keep browser and library versions aligned: after changing Playwright versions, run its browser install command again.
  • Avoid overlapping runs: if one capture can take longer than the interval, use a lock mechanism or lengthen the interval. Overlapping Chromium processes can exceed hosting limits.
  • Control disk use: use a retention policy for dated images and monitor available storage. Back up screenshots only if they are valuable; output directories are not necessarily included in provider backups.
  • Handle page variability: use explicit readiness conditions for pages that hydrate after initial HTML. A full-page screenshot can be very tall and consume more memory and storage than a viewport capture.
  • Keep credentials out of source: if a target needs authentication, store secrets in a protected environment or secret mechanism supplied by the host, and explicitly pass them to the script. Cron may not inherit the same environment as an interactive shell.

7. Cost, performance, and choosing a plan

For local Playwright captures, the direct costs are the hosting plan, storage, and the resources consumed while Chromium runs. This research provides no benchmark that establishes a minimum RAM or CPU size: workload depends on target page complexity, frequency, image dimensions, concurrency, and retention. Start by confirming browser support on the current account. If it cannot launch Chromium or provide the needed libraries, compare plans on package and browser permissions, scheduler frequency and timezone, CPU and memory limits, process restrictions, disk and backup policy, and root access.

A VPS with root control is a plausible next step when a shared plan cannot provide the browser stack, but its capacity still needs to match the workload. With a self-managed VPS, you also take responsibility for operating system updates, Python and Playwright updates, browser dependencies, and disk monitoring.

8. Troubleshooting

Symptom Likely cause Fix
Executable doesn't exist or Playwright cannot find Chromium Browser binary was not installed for this environment, or Playwright was upgraded afterward. Run .venv/bin/python -m playwright install chromium as the same account that runs the job.
Browser launch error mentioning a shared library A required Linux system dependency is missing. Ask the host to install/provide the library, or use a plan with system dependency control. A pip reinstall alone cannot add OS packages.
Works in SSH but not cron Cron has a limited environment, different working directory, or different interpreter. Use absolute paths, the venv interpreter, explicit environment variables, and log stderr; test the exact cron command manually.
Runs at the wrong local time The server or cron daemon uses a different timezone, or CRON_TZ is unsupported. Check the provider’s scheduler timezone and implementation. Configure the supported timezone setting and verify the next run.
Screenshot is blank or content is missing The page has not finished rendering, blocks automation, or requires a login/cookie state. Wait for a meaningful element, inspect the log and response, and configure the required browser context state legitimately. Some bot checks prevent a usable capture.
Navigation times out Slow site, unstable network, or waiting condition never occurs. Check the URL from the host, use a longer but bounded timeout, and choose a suitable readiness signal instead of requiring network idle.
Permission denied when writing output Output path is outside the account’s writable area or permissions prevent writing. Use a directory owned by the cron account, such as the project folder under its home directory.
Disk fills over time Timestamped screenshots or logs accumulate without retention. Delete or archive old captures according to a retention policy; rotate or truncate logs and monitor quota.
Two captures run at once A previous browser run exceeded the schedule interval. Lengthen the interval or add a lock so a second run exits while the first is active.

9. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. One GET request returns an image or PDF, so a scheduled Python task can call the API instead of installing Chromium on the host. See the ScreenshotNeo API documentation for parameters and response behavior.

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
r.raise_for_status()
with open("shot.webp", "wb") as f:
    f.write(r.content)

Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. ScreenshotNeo also identifies page verdict and billing status in response headers. Learn more at ScreenshotNeo.

Create a free ScreenshotNeo account for 1,000 screenshots a month, with no card required.

10. FAQ

Can I schedule a screenshot from shared hosting without SSH?

Possibly, if the control panel offers scheduled tasks and your provider supplies a way to install or invoke Python packages and Chromium. Ask the provider about headless browser support and limits; panel access alone does not guarantee browser compatibility.

Do I need to keep my SSH session open?

No. Once saved in the account’s crontab or hosting scheduler, the task is launched by that scheduler. Keep its logs accessible so you can see whether it ran successfully.

Why does the screenshot show a different page than my browser?

The script starts a fresh browser context without your personal cookies, extensions, or saved login session. If the page depends on a consent choice or authentication, configure the required context state or use an authorized test account.

Will this capture pages behind a bot check?

Not reliably. If the site presents a CAPTCHA or bot check, the browser may not reach the intended page. Do not attempt to bypass access controls; capture only sites you are permitted to access.