How to Take Recurring Screenshots of Indian News Homepages in the Morning
Use Playwright to capture news homepages on a daily schedule. Set the scheduler’s timezone, choose a repeatable frame, and save dated files you can inspect later.
To take recurring screenshots of Indian news homepages in the morning, use a browser automation script to visit each URL, wait for the content you care about, save a dated screenshot, and run the script from a scheduler configured for the intended timezone. The example below uses Playwright with Node.js and leaves the site list, local capture time, viewport, and destination under your control.
Choose a viewport screenshot if you need a consistent record of the first screen. Choose a full-page screenshot if content below the fold matters. Playwright supports both, as well as screenshots of individual elements. Playwright screenshot documentation.
1. Decide what the archive should show
The title does not specify which publications or what “morning” means, so make those explicit configuration choices before scheduling:
- Sites: list the exact homepage URLs. Do not assume that a section page or regional edition is equivalent to the homepage.
- Time: choose a local wall-clock time, such as 07:00, and the timezone that defines it. For Indian Standard Time, use the IANA timezone name
Asia/Kolkatawhere the scheduler supports it. - Frame: set a fixed viewport for first-screen captures, or use full-page mode when the entire scrollable page is needed. Full-page images can be very tall and larger than viewport captures.
- Evidence: record capture time, requested URL, final URL, browser settings, and whether the page reached the expected state in a small manifest beside the images.
- Retention: choose a directory or object-storage prefix and a retention period that fits your purpose. Keep originals if later visual comparison matters.
Keep viewport dimensions, browser, capture mode, and timezone consistent across runs if you plan to compare images. Page redesigns, consent dialogs, lazy-loaded modules, network problems, and bot checks can all change what an image contains. No particular Indian news site or selector is established by this general workflow, so inspect captures and adapt waits site by site.
2. Install Playwright and create the capture script
Install Node.js, create a project folder, then install Playwright and its Chromium browser:
mkdir morning-news-shots
cd morning-news-shots
npm init -y
npm install playwright
npx playwright install chromium
Create capture.mjs. Replace the sample URLs with the homepages you intend to archive. This script creates a date-based folder, captures each site, and writes a JSON manifest. It uses a DOM-ready navigation condition and then waits for fonts and a short settling interval; tune these choices after inspecting actual pages.
import { chromium } from 'playwright';
import { mkdir, writeFile } from 'node:fs/promises';
import path from 'node:path';
const sites = [
{ id: 'publication-one', url: 'https://example.com/' },
{ id: 'publication-two', url: 'https://example.org/' },
];
const outputRoot = process.env.OUTPUT_DIR ?? './archive';
const fullPage = process.env.FULL_PAGE === '1';
const width = Number(process.env.VIEWPORT_WIDTH ?? 1440);
const height = Number(process.env.VIEWPORT_HEIGHT ?? 1000);
const settleMs = Number(process.env.SETTLE_MS ?? 1500);
// Set TZ in the process or scheduler to the timezone whose date labels you want.
const now = new Date();
const date = new Intl.DateTimeFormat('en-CA', {
timeZone: process.env.CAPTURE_TZ ?? 'Asia/Kolkata',
year: 'numeric', month: '2-digit', day: '2-digit',
}).format(now);
const runDir = path.join(outputRoot, date);
await mkdir(runDir, { recursive: true });
const browser = await chromium.launch({ headless: true });
const results = [];
try {
for (const site of sites) {
const context = await browser.newContext({
viewport: { width, height },
deviceScaleFactor: 1,
locale: 'en-IN',
timezoneId: process.env.CAPTURE_TZ ?? 'Asia/Kolkata',
});
const page = await context.newPage();
const startedAt = new Date().toISOString();
let error = null;
let status = null;
try {
const response = await page.goto(site.url, {
waitUntil: 'domcontentloaded',
timeout: 45_000,
});
status = response?.status() ?? null;
// Replace or supplement this generic wait with a site-specific locator
// when a stable headline or content container is known.
await page.evaluate(() => document.fonts?.ready);
await page.waitForTimeout(settleMs);
const filename = `${site.id}.${fullPage ? 'full' : 'viewport'}.png`;
await page.screenshot({
path: path.join(runDir, filename),
fullPage,
animations: 'disabled',
});
results.push({
id: site.id,
requestedUrl: site.url,
finalUrl: page.url(),
startedAt,
completedAt: new Date().toISOString(),
httpStatus: status,
screenshot: filename,
fullPage,
viewport: { width, height },
error: null,
});
} catch (e) {
error = e instanceof Error ? e.message : String(e);
results.push({
id: site.id, requestedUrl: site.url, startedAt,
completedAt: new Date().toISOString(), httpStatus: status, error,
});
console.error(`Capture failed for ${site.id}: ${error}`);
} finally {
await context.close();
}
}
} finally {
await browser.close();
}
await writeFile(
path.join(runDir, 'manifest.json'),
JSON.stringify({ captureDate: date, timezone: process.env.CAPTURE_TZ ?? 'Asia/Kolkata', results }, null, 2),
);
if (results.some((result) => result.error)) process.exitCode = 1;
The sample uses a generic settling delay, not a guarantee that every site’s news modules have finished. For a stable page, replace it with a selector that reflects the content you actually need:
await page.locator('main article h2').first().waitFor({ state: 'visible', timeout: 15000 });
That selector is only an example. Inspect each site’s markup and choose a locator that exists there; if the page is redesigned, the wait may need updating. A missing selector should fail visibly instead of silently saving an early or blank capture.
3. Run it manually before adding a schedule
CAPTURE_TZ=Asia/Kolkata node capture.mjs
Confirm that the dated directory contains one image per successful site and a manifest. Open the images and check that they show the expected edition, viewport, content state, and consent overlay behavior. If you need below-the-fold content, run with FULL_PAGE=1. To adjust the frame or settling interval:
VIEWPORT_WIDTH=1365 VIEWPORT_HEIGHT=900 SETTLE_MS=2500 CAPTURE_TZ=Asia/Kolkata node capture.mjs
Use a viewport capture for consistent first-screen comparisons. A full-page capture can include lazy-loaded material, but some sites only load that material after scrolling; the generic script does not scroll the page to trigger every lazy module. If such content matters, add a deliberate scrolling routine and verify the resulting image.
4. Schedule the script at the intended local time
The capture script and the scheduler are separate: the script captures, while the scheduler launches it. Configure the scheduler’s timezone field, not just the machine’s display clock. Cloud Scheduler defaults to UTC unless a timezone is selected, and its documentation warns that daylight-saving transitions can cause a wall-clock run to be skipped or repeated. Cloud Scheduler cron and timezone documentation.
Linux with systemd
For a machine that remains powered on, a systemd timer can launch the script each day. Set the timezone deliberately in the service environment and use the desired local time in the timer. Example for 07:00 Asia/Kolkata; replace the working directory and user with real values:
# /etc/systemd/system/morning-news-shots.service
[Unit]
Description=Capture selected news homepages
[Service]
Type=oneshot
User=YOUR_LINUX_USER
WorkingDirectory=/opt/morning-news-shots
Environment=CAPTURE_TZ=Asia/Kolkata
Environment=OUTPUT_DIR=/var/lib/morning-news-shots
ExecStart=/usr/bin/node /opt/morning-news-shots/capture.mjs
# /etc/systemd/system/morning-news-shots.timer
[Unit]
Description=Daily morning homepage screenshots
[Timer]
OnCalendar=*-*-* 07:00:00 Asia/Kolkata
Persistent=true
[Install]
WantedBy=timers.target
Install and enable the timer, then inspect its status and logs:
sudo systemctl daemon-reload
sudo systemctl enable --now morning-news-shots.timer
systemctl list-timers morning-news-shots.timer
journalctl -u morning-news-shots.service
Systemd time expressions use the current timezone unless one is specified. Check the installed systemd documentation and version if the named timezone expression is not accepted. systemd.time documentation.
Google Cloud Scheduler
For a hosted job, use a daily schedule such as 0 7 * * * and set the Cloud Scheduler timezone to Asia/Kolkata. The five cron fields are minute, hour, day of month, month, and day of week. Ensure the target actually runs the capture script and writes images to durable storage accessible after the job ends; an ephemeral job filesystem may disappear.
Kubernetes CronJob
Kubernetes supports a timeZone field for CronJobs in supported versions. Use that field for a local wall-clock schedule rather than embedding TZ or CRON_TZ in the cron expression. CronJob scheduling is approximate, so design the job to tolerate delayed or missed starts and avoid overwriting earlier images. See the Kubernetes CronJob documentation.
5. Make the archive repeatable and recoverable
- Use unique names. Include the capture date and a site identifier. If you can have more than one run a day, include time and timezone too.
- Keep a manifest. Store the requested and final URL, timestamp, viewport, full-page setting, response status where available, and error. This distinguishes a successful image from an incomplete run.
- Separate sites. Capture sequentially as in the sample when a small, predictable workload matters. Parallel pages can shorten a run but increase concurrent requests and resource use; do not overwhelm publishers.
- Set a timeout and report failures. A timeout prevents a single stalled navigation from holding a run indefinitely. Preserve per-site errors and make the overall process exit unsuccessfully if any capture failed so the scheduler can surface it.
- Plan for missed runs. A laptop that is asleep cannot capture. Use an always-on host or managed scheduler for unattended runs, and decide whether a missed date should be backfilled. A rerun should not silently overwrite the original; use a run timestamp or an explicit overwrite policy.
- Persist outputs. On hosted compute, upload the image and manifest to durable storage as part of the job. Set retention deliberately and restrict access if the archive is not public.
- Review samples. Inspect newly captured images regularly. A successful browser navigation does not prove that the expected headlines loaded or that a bot check was absent.
6. Capture options and tradeoffs
| Choice | Use it when | Tradeoff |
|---|---|---|
| Viewport screenshot | You compare what a reader sees above the fold. | Below-the-fold stories are omitted. |
| Full-page screenshot | You need a long-page record. | Images can be large; lazy modules may require scrolling before capture. |
| Fixed delay | A page has no reliable content selector. | May be too short on a slow run and waste time on a fast run. |
| Wait for selector | A stable content element indicates readiness. | Site redesigns can invalidate the selector; handle timeout explicitly. |
| Local scheduler | A machine is always on and local files are convenient. | Sleep, reboot, network loss, or disk limits can interrupt runs. |
| Managed scheduler | You need unattended execution independent of a workstation. | Configure timezone and durable output storage; account for the platform’s run behavior. |
Playwright’s screenshot API also allows image format and quality options, clipping, and element screenshots. The code uses PNG for lossless, easy inspection. If storage is a concern, choose another format supported by the API and compare legibility before changing the archive format. See the screenshot API guide.
7. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| No browser executable or launch error | Playwright package is installed but Chromium is not. | Run npx playwright install chromium in the deployment environment; install system dependencies if the host requires them. |
| Blank or partial image | The page had not rendered its main content, or the wait condition was too generic. | Wait for a site-specific visible content locator, inspect the manifest, and capture a debug screenshot during manual runs. |
| Navigation timeout | Slow network, stalled resource, or a page that never reaches the selected navigation condition. | Check connectivity and final URL; choose a navigation condition appropriate to the page, retain a finite timeout, and retry only with a bounded policy. |
| Bot check or CAPTCHA in the image | The publisher presented an interstitial or automated-traffic challenge. | Record that outcome and review the site’s access rules. Do not assume a screenshot was a normal homepage capture. |
| Selector wait times out | The selector is wrong, absent in that edition, or changed in a redesign. | Inspect the live DOM manually and update the selector per site; keep the failure visible rather than falling back silently. |
| Images or lower stories are missing | Lazy loading waits for scrolling or intersection with the viewport. | Use full-page mode and, if needed, scroll through the page before capture; verify that the page has finished loading the content of interest. |
| Runs occur at the wrong hour | The scheduler evaluates cron in UTC or another timezone. | Set the scheduler’s timezone explicitly and verify the next scheduled execution in its own console or timer listing. |
| Two images overwrite each other | Filenames only contain the date while more than one run occurs that day. | Add hour, minute, and timezone to the filename or create a unique run directory. |
| Job reports success but files are missing | The output path is ephemeral, unwritable, or different under the scheduler’s user. | Use an absolute writable path, check process permissions, and upload files to persistent storage before the job exits. |
8. Performance, reliability, and cost
For a handful of sites, sequential browser contexts keep resource use and request concurrency straightforward. Reusing a browser process avoids repeatedly starting Chromium, while a separate context per site isolates cookies and page state. For a much larger site list, measure the run duration on your own host before adding parallelism. No benchmark or guaranteed completion time is established here.
Reliability depends on the scheduler, host availability, network, publisher changes, and your readiness checks. Keep run logs, preserve failed-site details, and make missed-run handling explicit. A screenshot is a visual snapshot, not a guarantee that every headline or dynamic module was captured correctly.
With a DIY setup, cost depends on the compute and storage you already use or provision, plus your own maintenance. Estimate storage from the number of sites, image size, run frequency, and retention period; inspect actual output sizes rather than relying on a generic estimate. No site-specific legal permission or universal retention requirement is established by this workflow, so decide access and retention for your own purpose.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. One GET request takes a URL and returns a PNG, JPEG, WebP, or PDF. Add your API key and target homepage to this cURL request; the ScreenshotNeo documentation lists the request options for capture configuration.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Change the URL to the news homepage you want. For a Python script, the same request is:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
And with Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, with response headers indicating the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots, and every feature is on every plan.
Sign up for 1,000 free screenshots a month, with no card required.
FAQ
Should I capture every day, including weekends?
That depends on the archive’s purpose. Set the scheduler for every day or selected weekdays, then record the chosen cadence in the manifest or operating notes.
Can I compare screenshots pixel by pixel?
You can compare saved images, but changing headlines, advertisements, timestamps, and page layout naturally create differences. Keep the browser and viewport settings fixed and interpret diffs as visual evidence, not as a measure of news importance.
Does this identify the top stories automatically?
No. It stores visual pages. Extracting headlines or ranking stories is a separate task and requires site-specific parsing or another data source.


