Puppeteer or Selenium for Scheduled Website Screenshots
Compare Puppeteer and Selenium for recurring website captures, with runnable scripts, scheduling guidance, failure handling, and a clear tool-selection guide.
Short answer: For a small scheduled screenshot job written in Node.js, start with Puppeteer. Its core workflow is compact: launch a browser, navigate to a page, wait for the page state that matters, save a screenshot, and close the browser. Choose Selenium when your project needs its broader language bindings, an existing WebDriver workflow, or remote and distributed browser execution through Selenium Grid.
Neither library runs a recurring schedule by itself. A system scheduler or CI workflow starts your script; Puppeteer or Selenium handles the browser session and capture. There is no general performance winner established by the documentation: rendering depends on the browser, page, environment, and capture requirements.
1. How to choose
| Need | Good starting point | Reason |
|---|---|---|
| A modest recurring job in JavaScript with a known browser | Puppeteer | The standard package manages a compatible Chrome download, and the screenshot flow is direct. |
| An existing Selenium project or a language beyond JavaScript | Selenium | Selenium has bindings for more languages and uses the W3C WebDriver workflow. |
| Remote browsers, multiple browser versions, or parallel distributed runs | Selenium Grid or an appropriately managed remote browser | Grid routes WebDriver scripts to remote browser instances and supports distributed execution. |
| One screenshot from a single machine | Either, locally | Grid is not required just to capture a page. |
Use the same browser build, fonts, viewport, device scale factor, locale, timezone, and authentication state that matter to your output. Different environments are not guaranteed to produce pixel-identical screenshots.
For reference, Puppeteer’s guide says to use Page.screenshot() for capture (Puppeteer screenshots). Selenium describes WebDriver as controlling a browser locally or remotely (Selenium WebDriver).
2. A scheduled Puppeteer screenshot in Node.js
This runnable example captures a full page after navigation, writes a timestamped PNG, and closes the browser even if navigation or capture fails. Save it as capture.mjs.
import puppeteer from 'puppeteer';
const targetUrl = process.env.TARGET_URL ?? 'https://example.com';
const outputPath = `./screenshots/page-${new Date().toISOString().replaceAll(':', '-')}.png`;
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
await page.setViewport({ width: 1440, height: 1000, deviceScaleFactor: 1 });
await page.goto(targetUrl, {
waitUntil: 'networkidle2',
timeout: 45_000,
});
await page.screenshot({ path: outputPath, fullPage: true });
console.log(`Saved ${outputPath}`);
} finally {
await browser.close();
}
Install and run it with:
npm init -y
npm install puppeteer
mkdir -p screenshots
TARGET_URL=https://example.com node capture.mjs
The Puppeteer package downloads a compatible Chrome for Testing browser by default. If your package manager blocks install scripts and skips the browser download, install it explicitly with npx puppeteer browsers install. Use puppeteer-core instead when you manage the browser yourself or connect to a remote browser; it does not download Chrome. See the Puppeteer installation guide.
Choose a readiness condition deliberately
networkidle2 is a convenient example, not a universal definition of “ready.” Analytics, polling, streaming, and other long-running requests can prevent idleness; an idle network also does not prove that the visual component you need has rendered. If the page has a meaningful selector, wait for it instead:
await page.goto(targetUrl, { waitUntil: 'domcontentloaded', timeout: 45_000 });
await page.locator('[data-report-ready="true"]').wait();
await page.screenshot({ path: outputPath, fullPage: true });
Replace the selector with one that indicates the content you need. For pages without a stable selector, a deliberate short delay after navigation may help, but it is less precise than waiting for application state.
Capture one element
For a chart, report panel, or other region, wait for it and use the element handle’s screenshot method. Puppeteer scrolls the element into view if needed.
const chart = await page.waitForSelector('#revenue-chart', { timeout: 15_000 });
if (!chart) throw new Error('Chart did not appear');
await chart.screenshot({ path: outputPath });
See the Puppeteer screenshot guide for page and element capture details.
3. A scheduled Selenium screenshot in Python
Selenium works well when the job is already part of a WebDriver codebase or Python is the better fit for its surrounding automation. This example uses Selenium’s Chrome WebDriver, waits for a meaningful element, saves a timestamped PNG, and quits in a finally block. Save it as capture.py.
import os
from datetime import datetime, timezone
from pathlib import Path
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
url = os.environ.get('TARGET_URL', 'https://example.com')
out_dir = Path('screenshots')
out_dir.mkdir(parents=True, exist_ok=True)
stamp = datetime.now(timezone.utc).strftime('%Y%m%dT%H%M%SZ')
output_path = out_dir / f'page-{stamp}.png'
driver = webdriver.Chrome()
try:
driver.set_window_size(1440, 1000)
driver.set_page_load_timeout(45)
driver.get(url)
WebDriverWait(driver, 20).until(
EC.visibility_of_element_located((By.CSS_SELECTOR, 'main'))
)
if not driver.save_screenshot(str(output_path)):
raise RuntimeError('WebDriver did not save the screenshot')
print(f'Saved {output_path}')
finally:
driver.quit()
Install and run:
python -m pip install selenium
TARGET_URL=https://example.com python capture.py
Change main to a selector that represents the page state you actually need. WebDriver’s screenshot APIs capture the current browsing context, and Selenium also documents element screenshots. See Selenium screenshot documentation.
Use a remote WebDriver when the browser belongs elsewhere
When you have Selenium Grid, point a remote driver at its address and provide browser options. The following is a connection pattern; GRID_URL must be the address of your own Grid.
import os
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
options = Options()
options.add_argument('--headless')
driver = webdriver.Remote(
command_executor=os.environ['GRID_URL'],
options=options,
)
try:
driver.get(os.environ.get('TARGET_URL', 'https://example.com'))
driver.save_screenshot('page.png')
finally:
driver.quit()
Remote WebDriver runs the browser on a remote computer, while Grid routes scripts to browser instances. Grid brings infrastructure and security responsibilities; Selenium’s documentation warns that it must be protected from external access. Do not add Grid to a one-machine job without a need for remote execution, parallelism, or a browser matrix. See Remote WebDriver and Selenium Grid.
4. Schedule the script outside the browser library
The recurring workflow is: scheduler fires → script starts → browser session opens → page navigates and reaches the chosen ready state → image is saved → session closes → job reports success or failure. Use cron, a CI schedule, or your existing job scheduler to trigger the process. Keep the scheduler configuration separate from capture code so you can run the same script manually when diagnosing a failure.
For a Unix-like host, a cron entry can invoke the script at 06:30 daily:
30 6 * * * cd /srv/site-captures && /usr/bin/env TARGET_URL=https://example.com /usr/bin/node capture.mjs >> /var/log/site-capture.log 2>&1
Use absolute paths in scheduled environments because their working directory and PATH may differ from an interactive shell. Set timezone expectations explicitly: cron commonly follows the host’s configured timezone, while timestamped artifacts in the examples use UTC.
Operational checklist
- Pin the browser automation package and deliberately plan browser updates; verify that the browser binary and operating system dependencies exist on the runner.
- Choose and keep a fixed viewport, device scale factor, and output format when comparing captures over time.
- Write outputs to durable storage if the runner’s local disk is temporary. Use a stable naming policy, such as UTC timestamp plus site or job identifier.
- Prevent overlapping runs from writing the same filename. Use unique timestamps, per-run directories, or a scheduler-level concurrency limit.
- Keep credentials out of source code and logs. Restrict access to authenticated screenshots and any session data.
- Record the target URL, run time, elapsed time, outcome, and failure reason. Retain a failure screenshot or browser log when it helps diagnosis.
- Always close the browser or driver in cleanup code, including on navigation, wait, or file errors.
- Define whether a failed run should be retried, how many times, and with what delay. Avoid unbounded retries that can create a backlog or overload a site.
5. Scheduling, readiness, and reliability edge cases
Pages that never become network-idle
Long-lived connections and periodic requests can keep the network active. Use a meaningful selector or known application-ready state rather than increasing the timeout indefinitely. If a required image or chart loads after the main content, wait for that resource or its containing component before capture.
Pages that look ready but are visually incomplete
A successful navigation does not guarantee that fonts, images, client-rendered components, or lazy-loaded content are complete. Wait for the specific element or state your screenshot needs. For a full-page capture, check whether the page loads content only as it scrolls; if so, the page may need an explicit scroll-and-wait strategy tailored to that site.
Authentication and changing content
For a protected page, establish the required login or session state before capturing and treat screenshots as potentially sensitive artifacts. Dynamic timestamps, rotating banners, personalized content, and A/B tests can make successive screenshots differ even when the automation works correctly.
Overlapping jobs and partial artifacts
Two runs with the same output path can overwrite one another. Give every run a unique path or serialize the job. If downstream systems consume files while a capture is being written, save to a temporary filename and move it into place after a successful capture.
Browser updates and reproducibility
Pin package versions and decide how browser updates reach the runner. A package update can change browser downloads or behavior, and missing OS libraries can prevent launch. Keep the browser and automation library compatible, and review updates in the environment that produces the production captures.
6. Performance, reliability, and cost considerations
There is no documentation-backed universal speed comparison between Puppeteer and Selenium for scheduled screenshots. Measure your own representative pages in the intended runner. Record navigation time, time spent waiting for readiness, capture duration, memory use, and failure rate; compare equivalent browser builds and settings.
For a small job, local execution keeps the architecture simple but makes you responsible for browser installation, updates, dependencies, and artifact storage. Remote Grid can distribute browser work and cover browser/version combinations, but adds servers, network paths, security, and maintenance. Choose it when the workload or coverage requirement warrants that operating cost.
Neither tool charges per screenshot as a library feature; your costs come from the machine or remote browser infrastructure, storage, scheduling platform, and engineering time. Keep concurrency bounded, avoid capturing more often than the output needs, and define artifact retention so recurring jobs do not accumulate unnecessary files.
7. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Puppeteer launches but reports that Chrome is missing | Install scripts were blocked or the browser download was skipped. | Run npx puppeteer browsers install, or configure the managed browser path intentionally. |
| Browser starts locally but not in CI | The runner lacks a browser binary, required OS dependencies, or the expected permissions. | Provision browser dependencies in the runner and verify the headless launch there, not only on a developer machine. |
| Navigation times out on an active site | A network-idle condition is waiting on long-running traffic, or the page is slow. | Wait for domcontentloaded plus a meaningful selector, and set a bounded timeout appropriate to the job. |
| Screenshot is blank or missing a component | The capture happened before client rendering or lazy content completed. | Wait for the actual component or application state; for lazy content, use a site-specific scroll and wait strategy. |
| Screenshot dimensions or crop are unexpected | Viewport, full-page behavior, device scale factor, or element bounds differ from expectations. | Set viewport and scale explicitly; use page capture for the document or element capture for a component, and inspect the resulting dimensions. |
| Scheduled job cannot find files or environment variables | The scheduler uses a different working directory or environment than an interactive shell. | Use absolute paths and configure variables in the scheduler’s environment. |
| Images overwrite each other or disappear | Runs collide on a fixed filename, or the runner uses temporary storage. | Use unique names, control overlap, and persist artifacts to storage that survives the job. |
| Remote Selenium connection fails | The Grid address is wrong, unreachable, or the requested browser capability is unavailable. | Check the configured Grid URL, network access, and available browser instances; keep the Grid protected. |
| Browser processes accumulate | Cleanup is skipped after an exception. | Close Puppeteer in finally and call Selenium’s quit() in finally. |
8. Or skip the browser setup
If you need screenshots without provisioning and maintaining a browser runner, ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request can return an image or PDF. Its cookie and consent handling accepts banners like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status.
Example cURL request:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo API documentation for request options and response details. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Every feature is available on every plan. Sign up for 1,000 free screenshots a month, with no card required.
9. Frequently asked questions
Can Puppeteer run on a schedule by itself?
No. Run its script from a scheduler such as cron or a CI schedule; Puppeteer performs the browser automation and capture.
Can Selenium save a screenshot automatically?
Yes. WebDriver can save a screenshot of the current browser context, and Selenium also documents screenshots of individual elements.
Does Puppeteer only support Chrome?
The current Puppeteer FAQ says that, from version 23 onward, the project supports Chrome and Firefox. Browser protocol defaults vary; check the official FAQ for current details.
Do I need Selenium Grid for one recurring capture?
No. A local WebDriver session can capture a page. Grid is for routing browser work to remote instances, including distributed or cross-browser workloads.
Which tool produces more accurate screenshots?
Neither has a universal accuracy advantage. Define the browser, fonts, viewport, scale, locale, timezone, readiness state, and authentication needed for the result, then run that configuration consistently.
