How to run a daily Puppeteer screenshot job for competitor product prices
Build a daily Puppeteer job that waits for product prices, saves dated screenshots, and reports failed runs.
Use Puppeteer to open each product page at a fixed viewport, wait for a selector that contains the displayed price, and save a timestamped screenshot. Run the script daily with a scheduler such as GitHub Actions, and make failed captures visible instead of silently recording incomplete pages. The examples below use Node.js.
1. Check the target and prepare configuration
Before automating a target, check its terms and any applicable permission requirements. A robots.txt file is a request for automated clients to follow, not access authorization: the IETF’s RFC 9309 states that robots rules “are not a form of access authorization.” Do not try to bypass access controls. Use a restrained schedule and avoid unnecessary concurrency.
Keep URLs, readiness selectors, and identifiers in one configuration list. A price selector must match the target page’s actual markup; there is no universal product-price selector. Inspect pages you are permitted to access and update selectors when their markup changes.
2. Install Puppeteer
mkdir daily-price-capture
cd daily-price-capture
npm init -y
npm install puppeteer
Save the script below as capture-prices.mjs. The installed Puppeteer package provides the browser it uses. If your deployment environment has separate browser-installation requirements, follow the instructions for that Puppeteer version.
3. Create the capture script
import puppeteer from 'puppeteer';
import { mkdir } from 'node:fs/promises';
import path from 'node:path';
const targets = [
{
id: 'product-a',
url: 'https://example.com/products/product-a',
priceSelector: '[data-testid="price"]',
},
{
id: 'product-b',
url: 'https://example.com/products/product-b',
priceSelector: '.product-price',
},
];
const outputDir = process.env.OUTPUT_DIR ?? 'screenshots';
const timeoutMs = Number(process.env.SELECTOR_TIMEOUT_MS ?? 15000);
const viewport = { width: 1280, height: 900, deviceScaleFactor: 1 };
const timestamp = new Date().toISOString().replaceAll(':', '-');
await mkdir(outputDir, { recursive: true });
const browser = await puppeteer.launch({ headless: true });
const failures = [];
try {
for (const target of targets) {
const page = await browser.newPage();
try {
// Set the viewport before navigation for consistent layout.
await page.setViewport(viewport);
const response = await page.goto(target.url, {
waitUntil: 'domcontentloaded',
timeout: 45000,
});
if (response && response.status() >= 400) {
throw new Error(`Navigation returned HTTP ${response.status()}`);
}
await page.waitForSelector(target.priceSelector, {
visible: true,
timeout: timeoutMs,
});
const priceText = await page.$eval(
target.priceSelector,
element => element.textContent?.trim() ?? ''
);
if (!priceText) throw new Error('Price selector matched, but its text was empty');
const filename = `${target.id}-${timestamp}.png`;
const screenshotPath = path.join(outputDir, filename);
await page.screenshot({ path: screenshotPath, fullPage: true });
console.log(JSON.stringify({ status: 'ok', id: target.id, priceText, screenshotPath }));
} catch (error) {
const message = error instanceof Error ? error.message : String(error);
failures.push({ id: target.id, url: target.url, error: message });
console.error(JSON.stringify({ status: 'failed', id: target.id, error: message }));
} finally {
await page.close();
}
}
} finally {
await browser.close();
}
if (failures.length) {
console.error(JSON.stringify({ failedTargets: failures }));
process.exitCode = 1;
}
Replace the example URLs and selectors with targets you are permitted to capture. This script saves a full-page PNG per successful target, checks that the matching price text is nonempty, logs each result, and exits with a failure status if any target fails. Keep the timestamp in UTC so filenames sort consistently across runs. If you need to retain the extracted price as structured data, write the logged result to your chosen database or file store as a separate step.
4. Choose readiness and screenshot settings
Wait for the condition that means the price is ready
domcontentloaded waits for initial HTML parsing, not for every client-rendered price to appear. The explicit waitForSelector is the readiness check for this job. Puppeteer documents selector waiting and page screenshots in its screenshot guide and Page API.
For a site where the price changes after the selector first appears, wait for a more specific selector or add a site-specific condition that verifies the displayed value. A fixed sleep is less dependable: it can waste time on fast pages and still be too short on slow ones. networkidle can also be unsuitable for pages with persistent network activity. Pick the condition that fits the page and retain a timeout so a stalled page becomes a reported failure.
Keep captures comparable
- Set the same viewport and device scale factor for every run. Puppeteer notes that some sites do not expect their layout to be resized after navigation, so set the viewport first.
- Use
fullPage: trueto include the whole document. For a compact record, usefullPage: falseto capture the viewport only. - Use a stable target ID in each filename and a UTC timestamp for each run. Do not use the product’s price as its only identifier; prices change.
- If only the price region matters, wait for its selector and capture the element with
await (await page.$(target.priceSelector)).screenshot({ path: screenshotPath }). Check that the element exists before callingscreenshot().
5. Schedule the job with GitHub Actions
Create .github/workflows/daily-capture.yml:
name: Daily product screenshots
on:
schedule:
- cron: '17 7 * * *'
workflow_dispatch:
jobs:
capture:
runs-on: ubuntu-latest
timeout-minutes: 20
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 22
cache: npm
- run: npm ci
- run: node capture-prices.mjs
- name: Upload screenshots
if: always()
uses: actions/upload-artifact@v4
with:
name: product-screenshots-${{ github.run_id }}
path: screenshots/
if-no-files-found: ignore
The cron expression requests a run daily at 07:17 UTC. GitHub Actions scheduled workflows use POSIX cron, run from the default branch, and can be delayed during high load; GitHub warns that queued runs can also be dropped. Scheduling away from the start of an hour reduces exposure to the busy top of the hour, but does not guarantee an exact start time. See GitHub’s schedule event documentation. Use workflow_dispatch to run the workflow manually.
The artifact step is an example for keeping run outputs available through GitHub Actions. Choose retention and a longer-term storage destination to match your needs; the appropriate retention period depends on your deployment and comparison workflow. Pin action versions according to your repository’s maintenance and security practices.
6. Make missed or failed runs visible
- Check the workflow’s run status and logs. The script exits unsuccessfully if any target fails, so one missing price is visible as a failed job.
- Use the scheduler’s notification or alerting setup to notify the responsible person when a run fails. The destination depends on your deployment.
- Monitor for a missing daily record as well as an explicit failure. A scheduled event may be delayed or dropped, so a successful-run alert alone does not prove that every day has an artifact.
- Keep the per-target error in the logs. A selector timeout, HTTP error, or empty price text points to different causes.
7. Reliability, performance, and cost
Sequential captures, as shown, limit simultaneous load on target sites and make failures easier to attribute. Runtime grows with the number of targets and each page’s load and readiness time. Set navigation and selector timeouts based on the sites you monitor, and keep the overall workflow timeout long enough for the complete list. If the job becomes too slow, first remove unnecessary targets and use a readiness condition that reflects the page; add concurrency only when appropriate for the target site and your execution limits.
Browser automation consumes runner time and produces files that need storage and retention management. The dossier establishes no universal safe request rate, target-site policy, or scheduler cost for this workload, so check the applicable service limits and costs for your own account. Avoid capturing more often than the daily record requires.
8. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Selector wait times out | The selector changed, the price is rendered later, or the page did not reach the expected state. | Inspect the permitted page, update the selector, and use a page-specific readiness condition. Keep the timeout and let the run fail visibly. |
| Screenshot has no price although navigation succeeded | Initial HTML parsing finished before the price rendered, or the selector matches a hidden element. | Wait for the correct visible price element and verify its text before saving. |
| Navigation returns an HTTP error | The server returned an error response for the requested URL. | Check the URL and response status, then investigate the target’s availability and access requirements. Do not attempt to bypass access controls. |
| Layout differs between runs | Viewport, scale factor, page state, or timing differs. | Set a stable viewport before navigation, use consistent screenshot options, and wait for the page-specific price condition. |
| Workflow starts late or a day is missing | Scheduled events can be delayed or dropped during high load. | Choose a cron minute away from :00, monitor for missing records, and use manual dispatch when needed. A cron schedule is not an exact-time guarantee. |
| Browser remains running after an error | Cleanup was skipped in a script without a finally path. | Keep browser closure in finally, as in the example, and close each page after its target completes. |
| Artifact is missing | No capture succeeded, the output path differs, or the upload step could not find files. | Check the logged screenshot paths and workflow artifact step; retain if: always() so failure logs and any successful captures remain available. |
Or skip the browser setup
ScreenshotNeo takes a screenshot with one GET request. See the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(({ writeFile }) => writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
Replace the example URL with the product page you are permitted to capture. ScreenshotNeo accepts cookie banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each cleanup step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, with page verdict and billing information in response headers. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000. ScreenshotNeo is a website screenshot API and MCP server from Yorker Media.
Create a free ScreenshotNeo account for 1,000 screenshots a month with no card.
FAQ
Should the daily job save the price as well as the screenshot?
For reliable numeric comparisons, save the extracted price text alongside the image. A screenshot preserves visual context, while structured records are easier to compare over time.
What time zone does the example schedule use?
The example cron schedule is interpreted by GitHub Actions in UTC. The filename timestamp is also UTC.
Can the capture run more than once a day?
GitHub Actions supports scheduled workflows as frequently as once every five minutes, but this price-monitoring example is configured for one daily run. More frequent collection should be justified by the use case and target-site requirements.


