How to Schedule Website Screenshots with a Rotating Proxy
Schedule repeatable website screenshots with Playwright and a rotating proxy. Configure proxy sessions, save timestamped captures, and troubleshoot failed runs.
Direct answer: use a scheduler to start a Playwright script, configure the browser with your proxy provider’s endpoint and credentials, wait for a repeatable page condition, then save the screenshot with a UTC timestamp. The scheduler triggers the work; the rotating proxy routes browser traffic and controls how its exit session changes. Use this only for monitoring you are authorized to perform, and respect the target site’s terms and rate limits.
This guide uses Node.js and Playwright for the browser capture, with a cron example for scheduling. The provider-specific proxy hostname, port, credentials, and session controls must come from your proxy provider. Playwright supports HTTP(S) and SOCKS v5 proxies, optional credentials, and proxy configuration at browser-launch or context scope. Playwright proxy documentation
1. Decide what each scheduled screenshot should record
Write down these choices before automating. They affect whether two captures can be meaningfully compared and how much traffic each run creates.
- Target: one exact URL, including query parameters if they define the page state.
- Cadence and timezone: choose the least frequent interval that answers the monitoring question; make the scheduler’s timezone explicit.
- Capture scope: viewport for a fixed-screen view, full page for content below the fold, or a selector for one component. Playwright supports all three. Playwright screenshot documentation
- Viewport and rendering: set a stable viewport and retain the same browser version, operating system, fonts, and rendering environment for visual comparisons. Playwright notes that host OS, browser version, settings, hardware, power state, and headless mode can change rendered pixels. Playwright visual comparison guidance
- Proxy session behavior: decide whether each scheduled run should use a fresh exit or a pinned session. If one page load has multiple dependent requests, check whether your provider can pin a session so those requests do not rotate between exits.
- Retention and access: choose where captures live, how long to keep them, and who can read them. Screenshots and logs can contain personal, account, or customer information.
2. Install Playwright and configure proxy secrets
Use a supported Node.js installation, create a project, and install Playwright. The Chromium install command downloads its browser runtime.
mkdir scheduled-shots
cd scheduled-shots
npm init -y
npm install playwright
npx playwright install chromium
Set the target URL and proxy details as environment variables in your shell or scheduler’s secret store. Do not commit proxy credentials to source control. The proxy URL should use the scheme and endpoint your provider documents, such as an HTTP proxy URL or a SOCKS endpoint. The username and password below are passed separately, so URL-encode unusual characters if instead you place credentials in a proxy URL.
export TARGET_URL='https://example.com/'
export PROXY_SERVER='http://proxy-provider-host:3128'
export PROXY_USERNAME='your-proxy-user'
export PROXY_PASSWORD='your-proxy-password'
export OUTPUT_DIR='./captures'
export CAPTURE_MODE='full'
For SOCKS, use the scheme and format supported by your provider and Playwright, for example socks5://host:port. Some providers encode a rotating or sticky-session instruction in the username or in a provider-specific endpoint; follow that provider’s current documentation. Do not guess the syntax.
3. Create a screenshot script
Save the following as capture.mjs. It launches one browser for a run, passes proxy authentication to Playwright, waits for the document to load and for a stable selector, and saves a timestamped image. Set READY_SELECTOR if the page has a reliable element that indicates the content is ready. Otherwise the script uses the document load event. It records the final URL and navigation status in a JSON sidecar file to make redirects and partial failures easier to inspect.
import { chromium } from 'playwright';
import { mkdir, writeFile } from 'node:fs/promises';
import path from 'node:path';
const targetUrl = process.env.TARGET_URL;
const proxyServer = process.env.PROXY_SERVER;
const proxyUsername = process.env.PROXY_USERNAME;
const proxyPassword = process.env.PROXY_PASSWORD;
const outputDir = process.env.OUTPUT_DIR ?? './captures';
const captureMode = process.env.CAPTURE_MODE ?? 'full';
const readySelector = process.env.READY_SELECTOR;
const timeoutMs = Number(process.env.NAVIGATION_TIMEOUT_MS ?? 60000);
if (!targetUrl || !proxyServer) {
throw new Error('Set TARGET_URL and PROXY_SERVER before running.');
}
if (!['full', 'viewport'].includes(captureMode)) {
throw new Error('CAPTURE_MODE must be "full" or "viewport".');
}
const timestamp = new Date().toISOString().replaceAll(':', '-');
const safeHost = new URL(targetUrl).hostname.replaceAll(/[^a-zA-Z0-9.-]/g, '_');
const basename = `${safeHost}-${timestamp}`;
const imagePath = path.join(outputDir, `${basename}.png`);
const metadataPath = path.join(outputDir, `${basename}.json`);
await mkdir(outputDir, { recursive: true });
const proxy = { server: proxyServer };
if (proxyUsername) proxy.username = proxyUsername;
if (proxyPassword) proxy.password = proxyPassword;
const browser = await chromium.launch({ headless: true, proxy });
let context;
try {
context = await browser.newContext({
viewport: { width: 1440, height: 1000 },
deviceScaleFactor: 1,
});
const page = await context.newPage();
page.setDefaultNavigationTimeout(timeoutMs);
const startedAt = new Date().toISOString();
let response = null;
let navigationError = null;
try {
response = await page.goto(targetUrl, { waitUntil: 'load', timeout: timeoutMs });
if (readySelector) {
await page.locator(readySelector).waitFor({ state: 'visible', timeout: timeoutMs });
}
} catch (error) {
navigationError = String(error);
}
// Preserve a failed run as metadata, but do not label an error page as a valid capture.
if (navigationError || !response || response.status() >= 400) {
const result = {
targetUrl,
finalUrl: page.url(),
startedAt,
finishedAt: new Date().toISOString(),
httpStatus: response?.status() ?? null,
outcome: 'failed',
error: navigationError ?? `HTTP ${response.status()}`,
};
await writeFile(metadataPath, JSON.stringify(result, null, 2));
console.error(JSON.stringify(result));
process.exitCode = 1;
} else {
await page.screenshot({
path: imagePath,
fullPage: captureMode === 'full',
type: 'png',
animations: 'disabled',
});
const result = {
targetUrl,
finalUrl: page.url(),
startedAt,
finishedAt: new Date().toISOString(),
httpStatus: response.status(),
outcome: 'captured',
imagePath,
captureMode,
viewport: { width: 1440, height: 1000 },
};
await writeFile(metadataPath, JSON.stringify(result, null, 2));
console.log(JSON.stringify(result));
}
} finally {
await context?.close();
await browser.close();
}
Run it once manually before creating the schedule:
node capture.mjs
Check the image and metadata file. A successful navigation status alone does not prove that the intended content rendered; inspect the screenshot and use a page-specific ready selector when possible. The example does not retry automatically: an immediate retry using a new rotating exit can hide a persistent target or proxy problem and create extra traffic.
4. Schedule the script with cron
For a Linux host using cron, install dependencies in a stable project directory and use absolute paths. This example runs daily at 09:00 UTC. Cron’s timezone handling depends on the host and its configuration; confirm the machine timezone or set the scheduler timezone explicitly in your environment.
0 9 * * * cd /opt/scheduled-shots && /usr/bin/node /opt/scheduled-shots/capture.mjs >> /var/log/scheduled-shots.log 2>&1
Configure the environment variables in a protected service environment or wrapper script readable only by the job account. Avoid placing secrets directly in a crontab that may be visible to other users. For a CI scheduler or managed timer, configure its schedule and timezone in that runner’s current documentation, inject the same secrets through its secret store, and make sure the output directory is persisted or uploaded; an ephemeral runner may discard local files when the job ends.
Prevent overlapping runs if a previous capture can last longer than the schedule interval. A lock or scheduler-level concurrency limit avoids multiple browser sessions competing for resources or sending a burst of requests to the target.
5. Configure capture, proxy, and wait behavior
Proxy scope and authentication
The sample applies the proxy at browser launch, which routes every context in that browser through it. Playwright also permits per-context proxy configuration when a single browser process needs isolated contexts with different proxy settings. Its proxy configuration supports HTTP(S) and SOCKSv5 proxies, optional username/password, and a bypass list for hosts that should avoid the proxy. Check your provider’s protocol support and routing policy before using these options. Proxy configuration options
Rotation itself is normally controlled by the provider, for example by an endpoint or session setting. A browser launch does not schedule proxy rotation. For multipage or authenticated flows, ask whether a sticky session is available and how its lifetime works. Keep the same session for the requests that need continuity; use fresh sessions only when the authorized monitoring design calls for them.
Choose the right readiness signal
waitUntil: 'load': a reasonable general starting point, but pages can continue updating after the load event.- Wait for a selector: set
READY_SELECTORto an element that appears when the monitored content is ready. This is usually more meaningful than sleeping a fixed number of seconds. - Short delay: if the content is known to update shortly after load and there is no stable selector, add a documented, bounded delay. A fixed delay can make runs slower and still miss unusually late content.
- Network idle: can suit pages that settle their requests, but analytics, long polling, or persistent connections may prevent the page from becoming idle. Prefer a specific page condition when available.
For lazy-loaded content, a full-page screenshot may not include assets that only load after scrolling. If those assets matter, scroll through the page in a controlled way, wait for the content, then capture; account for the additional requests. For pages that require login, use a dedicated authorized account and protect its storage state and screenshots as credentials or sensitive data.
Capture scope and image options
| Need | Playwright approach | Tradeoff |
|---|---|---|
| Fixed visible area | page.screenshot({ path, fullPage: false }) |
Predictable dimensions; content below the fold is omitted. |
| Entire scrollable document | page.screenshot({ path, fullPage: true }) |
Captures below the fold; can create very tall and larger files. |
| One component | page.locator('selector').screenshot({ path }) |
Smaller, focused comparison; selector changes can break the capture. |
| PNG | type: 'png' |
Lossless output, often larger. |
| JPEG | type: 'jpeg', quality: 80 |
Smaller lossy output; the quality option applies to JPEG. |
Playwright’s screenshot API also accepts options such as a clipping rectangle, JPEG quality, transparent background, and disabling or allowing animations. Check the screenshot documentation and API reference for the exact options supported by the installed version. Keep the selected format, viewport, browser build, and device scale factor constant if you compare images over time.
6. Store, name, and compare each run
The script uses an ISO UTC timestamp in both filenames and JSON metadata. This makes runs sortable and distinguishes scheduled time from actual completion time. For durable monitoring, write captures to storage that survives host replacement, define a retention period, and avoid overwriting the previous image. Keep run metadata alongside the image, including:
- target URL and final URL after redirects;
- scheduled time, start time, and finish time in UTC;
- HTTP status and outcome, including navigation, proxy, or timeout errors;
- browser version, viewport, capture mode, and any selected ready condition;
- proxy session identifier only if useful and safe to retain; never log its password or full credential-bearing URL.
For pixel comparisons, compare captures from the same rendering environment. Dynamic ads, clocks, personalized content, font loading, and animation can cause differences unrelated to a real layout change. Disable animations where appropriate, use an authorized stable account, and compare a specific element when full-page differences are too noisy. Playwright’s visual comparison guidance explains why environment differences affect pixels. Visual comparison consistency notes
7. Reliability, performance, and cost
- Keep cadence proportional to the question. Each browser visit loads a page and its dependent resources. A full-page capture or scrolling to trigger lazy assets may add time and traffic. Do not increase frequency or rotate identities to defeat access controls; handle a challenge, denial, or rate limit as a run outcome and seek permission or an approved access method.
- Bound runtime. Use navigation and selector timeouts, and ensure the scheduler has a job timeout larger than the expected capture duration but finite. Record failures rather than saving an error page under a successful filename.
- Retry carefully. Retry only transient infrastructure failures, with a small bounded retry count and backoff. Do not retry 403/429 responses, CAPTCHA or bot checks, or other access-control outcomes in a way that evades the site’s rules.
- Control concurrency. Browser processes consume memory and CPU. Start with sequential captures, then increase concurrency only when the target’s rules, proxy quota, host capacity, and provider limits allow it. Avoid overlapping schedules.
- Budget for the full pipeline. Self-hosted costs can include the runner or server, proxy traffic or sessions, storage, and image-diff processing. Hosted browser products have service-specific limits and pricing; verify current terms before choosing one. Cloudflare’s Browser Run documentation distinguishes quick screenshot actions from scripted browser automation and links to its current limits and pricing. Cloudflare Browser Run documentation
- Secure outputs. Restrict storage access, redact sensitive URLs or headers in logs, and define deletion rules. Treat proxy credentials and authenticated browser state as secrets.
8. Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| Proxy connection refused or timed out | Wrong endpoint or port, unavailable proxy, firewall rule, or unsupported scheme. | Check the provider’s endpoint and protocol, test network access from the scheduled host, and verify the provider account permits that connection. |
| Proxy authentication error | Missing or incorrect username/password, expired credentials, or provider-specific session syntax in the wrong field. | Update secrets from the provider dashboard, check encoding requirements, and use its documented session format. Do not print credentials while debugging. |
| Browser starts but page never loads | Proxy routing or DNS issue, target outage, TLS problem, or navigation timeout too short. | Inspect the recorded error and final URL; validate the proxy from the same host; increase timeout only when the page legitimately needs more time. |
| HTTP 403, 429, CAPTCHA, or bot check | The target refused or limited automated access. | Stop repeated attempts, reduce request volume, review the site’s rules, and request permission or use an approved monitoring interface. Do not switch exits to evade the refusal. |
| Screenshot is an error page despite a 200 response | The application rendered an error state client-side, redirected, or showed incomplete content. | Check the final URL and image, then wait for a page-specific selector or verify expected text before treating the run as successful. |
| Selector timeout | The selector changed, content is conditional, or the readiness assumption is wrong. | Inspect the page structure, update the selector, or use a more stable readiness condition; do not simply remove all waits. |
| Blank, partially styled, or inconsistent image | Capture happened before content or fonts settled, resources failed, or rendering environment changed. | Wait for a reliable element and needed assets; check proxy/resource errors; keep browser, OS, viewport, and device scale stable. |
| Images below the fold are missing | Lazy-loaded assets were never requested before capture. | Scroll through relevant sections and wait for their images to load before full-page capture. |
| Files disappear after a CI job | The runner’s filesystem is temporary. | Upload artifacts or write to durable object storage as part of the job, and confirm retention and access controls. |
| Two runs overlap or create a traffic burst | Previous browser work exceeded the interval or several targets start at once. | Set a concurrency limit or lock, stagger jobs, and lengthen the interval within the monitoring requirement. |
9. Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. One GET request captures an image or PDF; the API handles browser setup for the capture. See the ScreenshotNeo API documentation for parameters and response details.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
- Bot checks, blank pages, failed loads, timeouts, and cache hits cost nothing; response headers report the page verdict and whether it was billed.
- An MCP server provides
take_screenshot,get_page_info, andcapture_pdftools for Claude, Cursor, and other MCP clients. - The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Every feature is on every plan.
Sign up for ScreenshotNeo’s free 1,000 screenshots a month, with no card required.
FAQ
Does a rotating proxy run the schedule?
No. The scheduler starts your script. The proxy provider controls network routing and any rotation or session behavior.
Should the proxy rotate on every request?
That depends on the provider and the authorized task. A page may make many related requests; use provider-documented session pinning if a consistent session is needed during one load.
Can I use a full-page capture for visual monitoring?
Yes. It includes the scrollable page, but can be much taller than a viewport capture and may need scrolling first to trigger lazy-loaded assets.
Does this make proxy-based access appropriate for every website?
No. A proxy does not grant permission. Follow the website’s rules and stop if it blocks or rate-limits the job.


