Schedule Indian News Website Screenshots in Python Using Cron
Use cron to start a Python Playwright job on a schedule, then capture an Indian news page with deliberate readiness, timezone, and logging settings.
To schedule recurring screenshots of an Indian news website, put navigation and capture in a Python script, then have cron run that script at the times you choose. Cron starts the process; Playwright opens the page, waits for the content you need, and saves a viewport, full-page, or element screenshot. For a schedule based on Indian Standard Time (IST), configure the cron timezone where supported and separately set the browser timezone if the page’s rendered dates should also use IST.
The example below is an illustrative pattern, not a tested recipe for any particular news site. Replace the sample URL, output directory, and readiness condition with ones appropriate to the site you are authorized to capture.
1. Install Python and Playwright
Create a project and virtual environment under the account that will own the scheduled job. Use the same account to install dependencies, run the script, and write the output files.
mkdir -p /absolute/path/news-capture
cd /absolute/path/news-capture
python3 -m venv .venv
/absolute/path/news-capture/.venv/bin/python -m pip install playwright
/absolute/path/news-capture/.venv/bin/python -m playwright install chromium
Use the absolute path to this virtual environment’s Python executable in cron. Browser installation requirements vary by operating system; install the Chromium dependencies required by your host if launching the browser reports missing shared libraries. Playwright’s Python library documentation describes installation and browser setup: Playwright Python library.
2. Write the Python capture script
This synchronous example opens a public news homepage, waits for document navigation, waits for a site-specific headline locator, and saves a full-page PNG. A made-up selector such as article h2 may not match the page you choose: inspect that site and set a locator that identifies the content that matters to your capture.
from datetime import datetime
from pathlib import Path
import logging
import sys
from playwright.sync_api import TimeoutError as PlaywrightTimeoutError
from playwright.sync_api import sync_playwright
URL = "https://example.com/news/" # Replace with the selected news page
OUTPUT_DIR = Path("/absolute/path/news-capture/output")
READINESS_SELECTOR = "article h2" # Replace with a selector on the target site
logging.basicConfig(
level=logging.INFO,
format="%(asctime)s %(levelname)s %(message)s",
stream=sys.stdout,
)
def main() -> int:
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
stamp = datetime.now().strftime("%Y%m%d-%H%M%S")
output = OUTPUT_DIR / f"news-{stamp}.png"
with sync_playwright() as playwright:
browser = playwright.chromium.launch(headless=True)
try:
context = browser.new_context(
viewport={"width": 1365, "height": 900},
locale="en-IN",
timezone_id="Asia/Kolkata",
color_scheme="light",
)
page = context.new_page()
response = page.goto(
URL,
wait_until="domcontentloaded",
timeout=60_000,
)
if response is not None and response.status >= 400:
raise RuntimeError(f"Navigation returned HTTP {response.status}")
# This condition is site-specific. Change it to match the content
# that must be present before the capture is useful.
page.locator(READINESS_SELECTOR).first.wait_for(
state="visible",
timeout=30_000,
)
page.screenshot(path=str(output), full_page=True)
logging.info("Saved screenshot: %s", output)
context.close()
return 0
finally:
browser.close()
if __name__ == "__main__":
try:
raise SystemExit(main())
except PlaywrightTimeoutError as exc:
logging.exception("Timed out waiting for navigation or page content: %s", exc)
raise SystemExit(1)
except Exception:
logging.exception("Screenshot capture failed")
raise SystemExit(1)
Playwright supports saving screenshots to a path, capturing the full scrollable page with full_page=True, and capturing a specific element with a locator. Those options create different artifacts, so choose the one that fits the archive or monitoring task. See the Playwright screenshot guide.
The script exits with status 1 on a timeout or other failure so the failure appears in redirected logs and can be detected by surrounding job monitoring. A successful screenshot gets a timestamped filename rather than overwriting the last capture. If you prefer a fixed path, change the filename, but decide whether overwriting older evidence is acceptable.
3. Choose when the page is ready
wait_until="domcontentloaded" means the initial document has been parsed; it does not guarantee that headlines rendered later by JavaScript, images, advertisements, or feeds have finished loading. The locator wait provides a concrete condition, but it is only useful when the selector identifies the desired content on that particular site.
- Use a locator wait when a known headline container or article element signals that the page is useful.
- Use a short fixed delay only when the site has a known delayed update and no stable selector. Delays make every run slower and still do not prove content is ready.
- Use network idle cautiously. Pages with analytics, polling, or persistent requests may not become idle. It is not a universal news-page readiness rule.
- Check content when correctness matters. For a monitoring archive, consider logging a headline or page title alongside the image so an empty but technically successful screenshot is easier to spot.
Navigation can return no response object for some navigation types, so the sample checks the HTTP status only when a response is available. An HTTP error response is raised as a failed job rather than saved as if it were a normal capture.
4. Select the screenshot scope and rendering settings
| Choice | Playwright setting | When it helps |
|---|---|---|
| Visible viewport | page.screenshot(path=...) |
Captures the chosen window dimensions and what is visible without scrolling. |
| Entire scrollable page | page.screenshot(path=..., full_page=True) |
Creates a tall image of the complete page. Long pages can produce large files. |
| One component | page.locator("...").screenshot(path=...) |
Captures a headline list or article region without the rest of the page. |
Set the viewport intentionally. A desktop width and height may produce a different layout than a phone-sized viewport. Playwright browser contexts also accept locale and timezone settings; the sample uses en-IN and Asia/Kolkata so browser-rendered dates and locale-sensitive formatting follow those settings where the site uses them. Browser emulation settings do not change cron’s schedule timezone. See Playwright emulation.
For an element shot, replace the page screenshot call with a locator screenshot after waiting for that element. For example, use page.locator("main").screenshot(path=str(output)) after verifying that main selects the intended region. For a mobile layout, choose a mobile viewport or a documented device preset; changing only the browser timezone does not emulate a phone.
5. Schedule the script with cron
First run the exact command manually as the account that will own the crontab:
/absolute/path/news-capture/.venv/bin/python /absolute/path/news-capture/capture.py
Then edit that account’s crontab with crontab -e. This Cronie example runs at the start of every hour according to Asia/Kolkata:
CRON_TZ=Asia/Kolkata
0 * * * * /absolute/path/news-capture/.venv/bin/python /absolute/path/news-capture/capture.py >> /absolute/path/news-capture/capture.log 2>&1
The schedule fields are minute, hour, day of month, month, and day of week. Examples:
| Schedule | Meaning |
|---|---|
0 * * * * |
At minute zero, every hour |
*/15 * * * * |
Every 15 minutes |
30 6 * * * |
Daily at 06:30 |
0 8 * * 1-5 |
At 08:00, Monday through Friday |
CRON_TZ is documented by Cronie, but cron implementations differ. Verify that the daemon on your host supports it and confirm the effective timezone before relying on an IST schedule. Cronie checks entries each minute; its documentation describes special daylight-saving-time behavior, including skipped nonexistent local times and repeated times that may run twice. The five-field syntax and day-of-month/day-of-week matching rules can also vary across schedulers. Consult the host’s Cronie crontab manual and your system’s own cron documentation.
Cron runs under the crontab owner with a limited environment and configured shell. Use absolute paths for Python, the script, logs, and screenshot output. The output directory must exist or be creatable and writable by that account. The example redirects both standard output and errors into a log file; arrange log rotation if it will grow over time.
6. Keep scheduled captures reliable
- Prevent overlapping jobs. If a capture can run longer than the interval, use a host-appropriate lock mechanism or choose a longer interval. Otherwise, simultaneous browser processes may consume memory and overwrite shared output.
- Use unique output names. Timestamped files preserve history; add a retention policy so an archive does not grow indefinitely.
- Make failures visible. Keep a nonzero exit status on failure, log timestamps and exception details, and inspect the cron log or redirected application log.
- Account for site changes. A selector can stop matching after a redesign. Treat a readiness timeout as a signal to revisit the selector, not automatically as a browser failure.
- Keep capture frequency reasonable. Browser startup and page rendering use CPU, memory, and network. Full-page screenshots take more time and storage than a viewport or element capture. Schedule only as often as the use case requires and respect the site’s access rules.
No cron schedule guarantees a successful page load: the host may be offline, the site may be unavailable, or the page may challenge automated browsers. Retain logs and distinguish a missing output file, a navigation error, a readiness timeout, and an image that exists but contains unexpected content.
7. Troubleshoot common failures
| Symptom | Likely cause | What to check |
|---|---|---|
| Works in a terminal, not in cron | Cron has a smaller environment or different working directory. | Use absolute executable and file paths; run under the crontab owner; redirect output and errors; avoid relying on shell profile variables. |
ModuleNotFoundError: playwright |
The job is using another Python installation. | Call the virtual environment’s Python directly and install Playwright into that environment. |
| Browser executable or shared library error | Chromium or host dependencies are missing for that environment. | Run the Playwright browser installation command for the same user, then install the required operating-system dependencies for the host. |
| Timeout waiting for selector | The selector is incorrect, the page changed, content is delayed, or the request was challenged. | Inspect the target page and selector; log the page URL/title; choose a content-specific readiness condition and a suitable timeout. |
| Screenshot is blank or incomplete | Capture happened before the desired content rendered, or the site returned a challenge/error page. | Wait for the actual content; check the navigation status and logs; inspect the saved image. A completed navigation alone is not proof of useful content. |
| Files are missing or permission denied | The cron owner cannot write to the output directory, or the configured path is wrong. | Create the directory and test writing to it as the scheduled user; use an absolute path. |
| Job runs at the wrong local time | The server timezone differs, or the cron implementation ignores CRON_TZ. |
Check the host cron implementation and timezone settings. Browser timezone_id affects page rendering, not when cron starts the job. |
| Two captures appear around a clock change | Local civil time may repeat during a daylight-saving transition on implementations with that behavior. | Check the daemon’s timezone rules; use UTC scheduling if a fixed elapsed interval matters more than local wall-clock time. |
8. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server by Yorker Media. One GET request can return a screenshot or PDF; its API accepts the same parameter names used by other screenshot APIs to make switching easier. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://www.thehindu.com/ -o shot.webp
Cookie banners are accepted and removed before capture, along with known newsletter popups and chat widgets; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents and MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up for ScreenshotNeo and get 1,000 screenshots a month free, with no card.
FAQ
Does cron take the screenshot?
No. Cron starts the command on a schedule. Python and Playwright navigate to the page, decide when it is ready, and write the screenshot.
Does setting the browser timezone schedule the job in IST?
No. Playwright’s browser timezone controls page rendering. Cron’s schedule timezone is configured separately and depends on the cron implementation.
Should I capture the whole page or just the visible area?
Capture the viewport for a fixed-frame record, the full page for a scrollable archive, or a locator for one content region. The output size and what it documents differ.
Can I run this against any Indian news site?
The workflow is general, but readiness selectors, layouts, access controls, and site terms differ. Choose an appropriate target and verify the resulting capture for that site.


