How to Schedule Website Screenshots of Indian News Homepages in Multiple Languages
Capture Indian news editions on a reliable schedule with Playwright and GitHub Actions. Handle locales, timestamps, failures, and visual consistency.
Use a browser automation script to visit each news publisher’s actual language-edition URL, wait for the page content you need, and save a timestamped screenshot. Schedule that script with a recurring workflow such as GitHub Actions. Keep the URL, intended language, capture time, viewport, browser version, and result together so screenshots can be compared later.
A browser locale setting can affect language headers and formatting, but it does not guarantee that a publisher will serve a particular edition. Select real edition URLs explicitly. The example below uses Node.js, Playwright, and GitHub Actions; replace the example URLs with the publisher editions you have selected.
1. Choose the editions and capture policy
Start with a manifest rather than trying to infer language editions from one homepage. For each target, record its publisher, edition URL, intended language, and any region. Use the publisher’s own links or documentation to determine its edition URLs; this guide does not verify particular publishers’ current URLs or automation policies.
| Decision | Recommendation | Why it matters |
|---|---|---|
| Edition URL | Use an explicit URL for every language or regional edition. | A browser locale does not prove that the site will route to a specific edition. |
| Browser locale | Set a locale appropriate to the run, such as en-IN or a locale supported by your target audience. |
Playwright locale emulation affects navigator.language, the Accept-Language request header, and date and number formatting. A publisher can still choose its own content. |
| Viewport | Keep width and height constant across runs. | Responsive layout changes can look like editorial changes if viewport dimensions drift. |
| Capture type | Choose viewport screenshots for a fixed screen view, or full-page screenshots for a long homepage snapshot. | Full-page capture can produce very tall images and may expose lazy-loading or sticky-element behavior. |
| Readiness | Wait for a meaningful page condition, then capture. | Fixed short delays are unreliable when network or page load times vary. |
| Retention | Define where images and run records are stored, who can access them, and how long they are retained. | News pages can contain copyrighted material and personal data in comments or account-specific content. |
Before automating a target or redistributing its screenshots, check that publisher’s current access and reuse terms. This research did not establish publisher-specific automation permissions, robots policies, licensing, or republication rights.
2. Create a Playwright capture script
The following script reads a JSON manifest, starts Chromium, visits each edition, waits for the document and a configurable readiness selector, then writes a PNG and a JSON run record. It records a failed capture explicitly instead of silently treating a missing image as an unchanged page.
Create targets.json:
[
{
"publisher": "Example Publisher",
"edition": "English India",
"language": "en-IN",
"url": "https://example.com/",
"readySelector": "main"
},
{
"publisher": "Example Publisher",
"edition": "Hindi India",
"language": "hi-IN",
"url": "https://example.com/hi/",
"readySelector": "main"
}
]
These are placeholders, not verified news-site URLs. Replace them with the real URLs and selectors for your chosen publishers. If a site does not have a useful main element, use a selector that indicates the content you want, or remove the selector and rely on the navigation condition.
Create package.json:
{
"name": "scheduled-news-screenshots",
"private": true,
"type": "module",
"scripts": {
"capture": "node capture.mjs"
},
"dependencies": {
"playwright": "^1.0.0"
}
}
For reproducible scheduled runs, install dependencies once, commit the lockfile, and use npm ci in automation. The broad version range above is for initial setup; the lockfile pins the resolved package version.
Create capture.mjs:
import { chromium } from 'playwright';
import { mkdir, readFile, writeFile } from 'node:fs/promises';
import path from 'node:path';
const targets = JSON.parse(await readFile('targets.json', 'utf8'));
const outputDir = process.env.OUTPUT_DIR ?? 'artifacts';
const width = Number(process.env.VIEWPORT_WIDTH ?? 1440);
const height = Number(process.env.VIEWPORT_HEIGHT ?? 1000);
const timeoutMs = Number(process.env.PAGE_TIMEOUT_MS ?? 45000);
await mkdir(outputDir, { recursive: true });
const browser = await chromium.launch({ headless: true });
const results = [];
function safeName(value) {
return value.toLowerCase().replace(/[^a-z0-9]+/g, '-').replace(/^-|-$/g, '');
}
try {
for (const target of targets) {
const capturedAt = new Date().toISOString();
const stamp = capturedAt.replace(/[:.]/g, '-');
const base = `${safeName(target.publisher)}-${safeName(target.edition)}-${stamp}-${width}x${height}`;
const screenshotPath = path.join(outputDir, `${base}.png`);
const recordPath = path.join(outputDir, `${base}.json`);
const context = await browser.newContext({
viewport: { width, height },
locale: target.language,
colorScheme: 'light'
});
const page = await context.newPage();
page.setDefaultNavigationTimeout(timeoutMs);
page.setDefaultTimeout(timeoutMs);
const record = {
publisher: target.publisher,
edition: target.edition,
language: target.language,
url: target.url,
capturedAt,
viewport: { width, height },
browserVersion: browser.version(),
status: 'failed'
};
try {
const response = await page.goto(target.url, { waitUntil: 'domcontentloaded' });
record.httpStatus = response?.status() ?? null;
if (target.readySelector) {
await page.locator(target.readySelector).first().waitFor({ state: 'visible' });
}
// Optional settling time for client-rendered content. Prefer a real selector above.
const settleMs = Number(process.env.SETTLE_MS ?? 1000);
if (settleMs > 0) await page.waitForTimeout(settleMs);
await page.screenshot({ path: screenshotPath, fullPage: process.env.FULL_PAGE === '1' });
record.status = 'success';
record.screenshot = path.basename(screenshotPath);
} catch (error) {
record.error = String(error?.message ?? error);
} finally {
await writeFile(recordPath, JSON.stringify(record, null, 2));
results.push(record);
await context.close();
console.log(`${record.status}: ${target.publisher} / ${target.edition}`);
}
}
} finally {
await browser.close();
}
await writeFile(path.join(outputDir, 'run-summary.json'), JSON.stringify(results, null, 2));
if (results.some((result) => result.status !== 'success')) process.exitCode = 1;
The example uses a fresh browser context for each target. Playwright browser contexts are isolated and non-persistent by default, so this avoids carrying cookies or local storage from one edition to another. If your workflow intentionally requires a logged-in or consented state, design and protect that state explicitly rather than sharing it accidentally.
Install and run locally
npm install
npx playwright install chromium
npm run capture
To capture full pages or tune settings:
FULL_PAGE=1 VIEWPORT_WIDTH=1365 VIEWPORT_HEIGHT=900 SETTLE_MS=1500 npm run capture
The script accepts these environment variables:
| Variable | Default | Effect |
|---|---|---|
OUTPUT_DIR |
artifacts |
Directory for images and JSON records. |
VIEWPORT_WIDTH |
1440 |
Browser viewport width in CSS pixels. |
VIEWPORT_HEIGHT |
1000 |
Browser viewport height in CSS pixels. |
PAGE_TIMEOUT_MS |
45000 |
Navigation and selector timeout. |
SETTLE_MS |
1000 |
Optional delay after readiness selector; set to 0 to disable. |
FULL_PAGE |
unset | Set to 1 to capture the full scrollable page. |
3. Schedule it with GitHub Actions
Create .github/workflows/news-screenshots.yml:
name: Scheduled news screenshots
on:
schedule:
# Example: daily at 06:30 UTC. Change this to your desired capture time.
- cron: '30 6 * * *'
workflow_dispatch:
permissions:
contents: read
jobs:
capture:
runs-on: ubuntu-latest
timeout-minutes: 20
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 22
cache: npm
- run: npm ci
- run: npx playwright install --with-deps chromium
- name: Capture editions
run: npm run capture
env:
FULL_PAGE: '1'
VIEWPORT_WIDTH: '1440'
VIEWPORT_HEIGHT: '1000'
- name: Upload screenshots and run records
if: always()
uses: actions/upload-artifact@v4
with:
name: news-screenshots-${{ github.run_id }}
path: artifacts/
if-no-files-found: warn
retention-days: 14
GitHub scheduled workflows use UTC by default. GitHub Actions also supports an IANA timezone in the schedule configuration, and its documented shortest interval is five minutes. Choose the timezone deliberately and verify the workflow’s configuration when daylight-saving changes matter. Scheduled runs can be delayed under load, so treat a cron time as a trigger target rather than a precise capture timestamp; the script’s ISO timestamp records when each page was actually attempted.
The example uploads run artifacts with a 14-day retention setting. Set retention to fit your comparison window and storage needs. For a long-term archive, copy artifacts to storage you control and define access, retention, and deletion policies. GitHub’s scheduled-workflow syntax and timezone details are in the official workflow syntax documentation.
4. Capture multilingual pages consistently
A language edition and a browser locale are separate inputs. Keep the publisher’s actual edition URL in the manifest, and set Playwright’s locale for formatting and request behavior. Playwright documents that locale affects navigator.language, the Accept-Language header, and date and number formatting. It does not establish that any particular publisher uses those values to select content. See the Playwright locale and timezone documentation.
For a fair comparison across language editions:
- Use the same browser build, operating system image, viewport, color scheme, and capture mode.
- Record the language and URL for each image. Avoid relying on a filename that only says “India” when the URL actually identifies a regional edition.
- Use one consistent readiness selector per page where possible. Different selectors can mean screenshots represent different stages of loading.
- Capture at a consistent time relative to your schedule, while preserving the actual timestamp.
- Use a stable font environment. Indic script rendering depends on fonts available to the browser host; this workflow does not verify font coverage for any particular script or runner image.
- Expect ads, live headlines, rotating modules, personalization, and network timing to vary. A pixel difference does not by itself establish a site redesign.
Playwright supports viewport, element, and full-page screenshots. Its screenshot guide covers the capture API, and the Page API reference documents screenshot options.
Viewport, full-page, and element captures
- Viewport: the default in the script. Useful for comparing what appears above the fold. Set
fullPagetofalse. - Full page: set
FULL_PAGE=1. This captures the full page height, but exceptionally long pages can create large images and take longer. Lazy-loaded content may require scrolling or another site-specific readiness strategy before capture. - Element: replace the page screenshot call with
await page.locator('main').screenshot({ path: screenshotPath })to capture one element. Choose a selector that exists and is visible on every target.
For pages that load content as you scroll, add an intentional scroll-and-wait routine for the content you need. Do not assume a full-page screenshot triggers every site’s lazy-loading behavior in the same way. Keep such logic specific to the target and record it alongside the capture configuration.
5. cURL, Python, and ScreenshotNeo options
Playwright is the browser-based method when you need to control the browser environment, locale, or page readiness. cURL and Python’s requests library make direct HTTP requests; they do not run browser JavaScript or create a rendered browser screenshot by themselves. A hosted screenshot API can perform the browser capture for you. ScreenshotNeo accepts a URL in one GET request and returns an image or PDF; its options and API details are in the ScreenshotNeo documentation.
For example, these direct-request commands retrieve a ScreenshotNeo screenshot. Supply your API key and use the returned content as the image file. Configure the desired output format and capture settings according to the API documentation.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
with open("shot.webp", "wb") as f:
f.write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
The Node.js example uses Bun’s file writer. For standard Node.js, save the response body with the filesystem API:
import { writeFile } from 'node:fs/promises';
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
Or skip the browser setup
Use ScreenshotNeo when you want a screenshot API call instead of maintaining your own browser installation. One GET request captures the requested URL. The product accepts cookie and consent banners like a visitor, then removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status. An MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000 shots. Every feature is available on every plan.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
See the API documentation for request options, and ScreenshotNeo for the product. Sign up for 1,000 free screenshots a month with no card.
6. Reliability, performance, and cost
Reliability
- Preserve failures: keep per-target records even when navigation, readiness, or screenshot capture fails. The example exits unsuccessfully if any target fails, so the workflow is visibly marked failed.
- Distinguish HTTP status from readiness: a response can arrive with an error status, and a page can return successfully while its content is incomplete. Review
httpStatus,status, anderrortogether. - Use bounded timeouts: avoid waiting forever for pages that never reach a condition. Pick a timeout that fits the number of targets and workflow runtime limit.
- Keep configuration with outputs: the JSON record stores the URL, edition, language, viewport, timestamp, browser version, HTTP status, and outcome. Also preserve the manifest and workflow version used for a run if you need stronger auditability.
- Control environment drift: browser version, operating system, fonts, hardware, settings, and headless mode can change rendering. Stable environments make it easier to tell rendering noise from a page change. Playwright’s visual comparisons guidance explains why environment consistency matters.
- Plan for partial success: one inaccessible edition should not erase successful captures from the others. The script writes each record as it goes and always produces a summary if the process reaches completion.
Performance and storage
Runtime grows with the number of target pages, their load times, and screenshot height. Begin sequentially for a modest manifest: it limits browser load and makes failures easier to diagnose. If you later add concurrency, cap it and account for the publishers’ terms and the resource limits of your runner. Full-page PNGs can be substantially larger than viewport captures. Choose image format and archive retention based on whether you need pixel comparisons, human review, or long-term records.
GitHub-hosted runner availability and included minutes depend on the account and current GitHub plan; this research did not establish your account’s allowance or cost. Check the current GitHub pricing and usage details for your account before increasing capture frequency or target count. Hosted screenshot vendors are another option: ScreenshotAPI.net advertises recurring hourly, daily, weekly, or custom captures, stored captures, and email alerts. Those are vendor claims; its current pricing, retention details, locale controls, and terms were not verified in this research. Compare provider locale and environment controls, retention, alerts, exports, terms, and cost before choosing.
7. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| The wrong language edition appears. | The URL points to another edition, a redirect overrides the requested edition, or the publisher ignores the browser locale. | Verify the final URL and use the publisher’s explicit edition URL. Treat locale emulation as a browser setting, not an edition selector. |
locator.waitFor times out. |
The readiness selector does not exist, is hidden, or appears after the timeout. | Inspect the page structure, choose a stable visible content selector, or omit the selector and use a suitable navigation condition. Increase the timeout only when the target genuinely needs longer. |
| The image shows a loading shell or missing headlines. | The chosen readiness condition fires before client-rendered content is ready. | Wait for a content-specific selector. Add a short settling delay only as a fallback; arbitrary sleeps alone can still be too short or waste time. |
| Navigation times out on a page that continues loading. | Long-lived requests, ads, or streaming content prevent a network-idle condition. | The example waits for domcontentloaded and a selector instead of requiring all network activity to stop. Select a meaningful page condition and keep a bounded timeout. |
| Full-page capture is incomplete or unusually tall. | Lazy-loaded sections have not appeared, the page has dynamic height, or sticky elements behave differently during capture. | Test the target manually, scroll to load required sections, wait for them, and consider capturing a specific element or the viewport instead. |
| Indic text appears as boxes or the layout differs. | The runner may not have the needed fonts, or its font and browser environment differs from prior runs. | Use a consistent environment with suitable script fonts installed, then compare captures from the same browser and host configuration. This workflow does not validate any specific runner’s font coverage. |
| The workflow does not run at the expected local time. | The schedule is UTC by default, timezone configuration may differ, or scheduled execution was delayed. | Check the workflow timezone and cron expression. Store the actual capture timestamp and account for daylight-saving changes and queue delays. |
| The action cannot find Chromium. | The browser binary was not installed in the runner environment. | Run npx playwright install --with-deps chromium after installing dependencies, as in the workflow example. |
| The run succeeds but no artifact is retained. | The artifact upload step did not execute, found no files, or retention expired. | Inspect the upload step and its logs, retain the if: always() condition, and choose an explicit retention period or external archive. |
| Images differ despite no obvious editorial change. | Browser, operating system, font, viewport, headless mode, dynamic ads, timestamps, or live modules changed. | Compare run metadata, stabilize the browser environment, and interpret pixel differences as evidence for review rather than proof of a redesign. |
8. Frequently asked questions
Does changing Playwright’s locale select an Indian language edition?
No. It changes browser locale behavior, including language headers and formatting. Use the publisher’s actual edition URL to select the content.
Can I run captures more often than once a day?
Yes. GitHub Actions documents a shortest scheduled interval of five minutes. Account for workflow delays, runtime, runner usage, and the publisher’s terms.
Should I use full-page screenshots for every homepage?
Only if the entire page is useful to your task. Full-page images take more time and storage, and lazy-loaded or dynamic content may need target-specific handling.
Can screenshots be republished publicly?
This workflow guide does not determine reuse rights. Check the relevant publisher’s current terms and permissions before publishing or distributing captures.
What is the simplest way to keep multilingual captures comparable?
Use explicit edition URLs, a fixed browser environment and viewport, the same readiness rule, and a timestamped record for every result.


