ScreenshotNeo

BlogComparisons

Apify Screenshot Actor vs DIY Puppeteer for Recurring Website Captures

Compare Apify’s hosted screenshot Actor with DIY Puppeteer for recurring captures: scheduling, control, costs, retention, runnable code, and failure handling.

By the ScreenshotNeo team4 October 202611 min read

Short answer: Choose Apify’s hosted Website Screenshot and PDF Actor when you need recurring captures of public pages and its built-in batching, scheduling, and managed runs fit your workflow. Choose DIY Puppeteer when you need browser interactions or processing the Actor does not expose, and your team can operate the scheduler, browser runtime, storage, retries, and maintenance.

This is a feature and operations comparison, not a speed, reliability, or cost benchmark. The Apify Store listing for reestri/web-screenshot identifies a community maintainer; it should not be described as an Apify-maintained first-party Actor. The older repository-backed Website Screenshot Generator is described as an example Actor, a separate project from the current Store listing. [Current Store listing] [Example Actor repository]

1. The decision at a glance

Choose When it fits What your team owns
Apify hosted Actor Public anonymous pages; standard screenshot options; batches and scheduled runs are useful; listed per-capture charges and storage behavior fit. Input configuration, schedule, checking per-URL results, and deciding how long to retain downloaded outputs.
DIY Puppeteer Custom site-specific interactions, processing, or deployment and storage choices not covered by the Actor. Browser setup, triggering, concurrency, retries, logs, output pipeline, retention, and ongoing maintenance.

Do not infer that either choice can capture every site. The Actor documents anonymous public-site capture; it does not log in, solve CAPTCHAs, rotate proxies, or defeat automated-browser blocks. Puppeteer gives you browser control, but that does not grant permission or guarantee access.

2. What the Apify screenshot Actor provides

The current Store Actor accepts a list of public URLs and can return per-URL records with file links and error details. Its published controls include viewport, full-page, and element screenshots; image and PDF formats; device presets; dark mode; wait options; and optional banner hiding. A run can include up to 1,000 URLs. Concurrency defaults to two and can be configured up to five. Full-page height defaults to a 15,000-pixel maximum and can be limited or reported as truncated. [Actor input and output documentation]

Scheduling is configured through Apify schedules, which can start Actors or tasks at specified times. [Apify schedules] The managed workflow reduces the amount of recurring-run plumbing you need to build, but you still need to inspect individual records and handle output retention.

Pricing and storage to check before scheduling

The Store listing currently states $0.002 per stored viewport or element screenshot, $0.004 per stored full-page image, $0.004 per PDF, and a $0.002 default run-start fee for a run configured with 2 GB memory. These are listed prices at the time covered by the research; account-level compute, storage, or other usage may also affect the total. The listing says unsuccessful captures are not charged, but URLs can be skipped after the configured maximum run cost is reached. Check the live listing and your account pricing before committing to volume. [Actor pricing] [Apify platform pricing]

Files in the unnamed run store are subject to plan retention. If these screenshots are an archive, download them or move/copy them into a named store with an explicit retention policy. A successful Actor run can still contain failed URL records, so the run status alone is not proof that every page was captured.

3. DIY Puppeteer: recurring capture workflow

Puppeteer provides browser automation and screenshot APIs; it does not itself provide the recurring scheduler and output operations around them. A dependable recurring job needs a trigger, bounded concurrency, per-URL result tracking, retry policy, durable storage, and monitoring. These are implementation responsibilities inferred from the Puppeteer API and Apify’s managed workflow features. [Puppeteer screenshot guide] [Puppeteer configuration guide]

Install

npm init -y
npm install puppeteer

The puppeteer package is the straightforward choice when you want Puppeteer to manage its compatible browser installation. If using puppeteer-core, provide the browser executable and configuration yourself; it ignores Puppeteer’s standard configuration files and environment variables. Consult the configuration guide for browser setup and proxy environment options. [Puppeteer configuration]

Runnable Node.js capture script

Save as capture.mjs. It captures a list with limited concurrency, records failures per URL, and writes PNG files into a local directory. Use a durable object store or database instead of local disk when the job runs on disposable workers.

import puppeteer from 'puppeteer';
import { mkdir, writeFile } from 'node:fs/promises';
import { createHash } from 'node:crypto';

const urls = process.argv.slice(2);
if (urls.length === 0) {
  console.error('Usage: node capture.mjs https://example.com [https://example.org ...]');
  process.exit(2);
}

const outDir = 'captures';
await mkdir(outDir, { recursive: true });
const concurrency = Math.min(Number(process.env.CONCURRENCY || 2), 5);
const browser = await puppeteer.launch({ headless: true });
let next = 0;
const results = [];

async function worker() {
  while (true) {
    const index = next++;
    if (index >= urls.length) return;
    const url = urls[index];
    const page = await browser.newPage();
    try {
      await page.setViewport({ width: 1440, height: 1000, deviceScaleFactor: 1 });
      await page.goto(url, { waitUntil: 'networkidle2', timeout: 45000 });
      const bytes = await page.screenshot({ type: 'png', fullPage: true });
      const id = createHash('sha256').update(url).digest('hex').slice(0, 16);
      const path = `${outDir}/${id}.png`;
      await writeFile(path, bytes);
      results[index] = { url, ok: true, path };
    } catch (error) {
      results[index] = { url, ok: false, error: String(error) };
    } finally {
      await page.close();
    }
  }
}

try {
  await Promise.all(Array.from({ length: concurrency }, () => worker()));
} finally {
  await browser.close();
}
console.log(JSON.stringify(results, null, 2));
if (results.some((result) => !result.ok)) process.exitCode = 1;

Run it manually with node capture.mjs https://example.com https://example.org. Connect the command to your host, container platform, or operating-system scheduler at the desired interval. Add a lock or idempotency key if overlapping runs could overwrite outputs or overload target sites. Puppeteer documents Page.screenshot() as the screenshot method. [Puppeteer screenshots]

Adapting the script safely

  • Viewport capture: remove fullPage: true.
  • Element capture: locate a stable selector with page.locator('main').screenshot(); handle a missing selector as a per-URL failure.
  • Wait strategy: networkidle2 is convenient for many pages but may never settle on pages with persistent network activity. Use domcontentloaded or load and then wait for a meaningful selector, or use a bounded delay where appropriate.
  • Full-page limits: extremely long pages can consume substantial memory or hit browser/image limits. Capture sections or a viewport if a single tall image is not essential.
  • Output names: the sample hashes the URL to avoid unsafe filenames. For a real archive, include a timestamp or content version and keep a manifest mapping keys back to URLs.
  • Retries: retry transient navigation failures with a small bounded count and backoff. Do not endlessly retry access denials, CAPTCHA pages, or persistent application errors.

4. cURL, Python, and Node.js alternatives

These examples call Puppeteer’s browser automation through a small local HTTP worker. They are useful when a separate scheduler or workflow runner triggers captures over HTTP. The worker must be deployed and protected by your own authentication; do not expose an unauthenticated browser endpoint publicly.

Minimal local worker in Node.js

npm install express puppeteer
// worker.mjs
import express from 'express';
import puppeteer from 'puppeteer';

const app = express();
app.use(express.json({ limit: '16kb' }));
const browser = await puppeteer.launch({ headless: true });

app.post('/capture', async (req, res) => {
  const { url } = req.body ?? {};
  let parsed;
  try { parsed = new URL(url); } catch { return res.status(400).json({ error: 'Invalid URL' }); }
  if (!['http:', 'https:'].includes(parsed.protocol)) return res.status(400).json({ error: 'Only HTTP(S) URLs are supported' });
  const page = await browser.newPage();
  try {
    await page.setViewport({ width: 1440, height: 1000 });
    await page.goto(parsed.href, { waitUntil: 'networkidle2', timeout: 45000 });
    const image = await page.screenshot({ type: 'png', fullPage: true });
    res.type('png').send(image);
  } catch (error) {
    res.status(502).json({ error: String(error) });
  } finally {
    await page.close();
  }
});

app.listen(3000, '127.0.0.1');

cURL trigger

curl --fail-with-body -X POST http://127.0.0.1:3000/capture \
  -H 'Content-Type: application/json' \
  -d '{"url":"https://example.com"}' \
  -o example.png

Python trigger

import requests

response = requests.post(
    "http://127.0.0.1:3000/capture",
    json={"url": "https://example.com"},
    timeout=60,
)
response.raise_for_status()
with open("example.png", "wb") as image:
    image.write(response.content)

Node.js trigger

const response = await fetch('http://127.0.0.1:3000/capture', {
  method: 'POST',
  headers: { 'content-type': 'application/json' },
  body: JSON.stringify({ url: 'https://example.com' }),
  signal: AbortSignal.timeout(60_000),
});
if (!response.ok) throw new Error(await response.text());
const image = Buffer.from(await response.arrayBuffer());
await import('node:fs/promises').then(({ writeFile }) => writeFile('example.png', image));

For a one-off command-line screenshot, Puppeteer also has a CLI documented by the project; the recurring schedule and archive remain your responsibility. [Screenshot guide]

5. Recurring-run design: reliability, performance, and cost

Scheduling and missed runs

Apify schedules can start Actors or tasks at configured times. In a DIY system, choose a scheduler that matches your host and deployment, then record intended and actual run times. Decide whether a missed interval should run late, be skipped, or trigger a backfill. Prevent duplicate overlap with a lease, lock, or run identifier where duplicates are harmful.

Per-page failure handling

Store one outcome per URL: status, timestamp, final URL if redirected, error category, and output location. A run can partly succeed. Retry only transient errors, cap attempts, and preserve the first failure for diagnosis. Alert on a rising failure count or missing scheduled run, rather than treating any one failed page as proof the whole system is down.

Concurrency and memory

More browser pages consume more memory and can increase load on the pages being captured. The Actor’s published maximum concurrency is five; that is a product limit, not a performance guarantee. In DIY Puppeteer, begin with low concurrency, observe memory and completion behavior in your own environment, and increase gradually. Apify documentation gives 1,024 MB as the minimum memory for Actors using Puppeteer or Playwright for real browser rendering. [Actor memory guidance]

Cost comparison

For Apify, estimate the listed per-stored-output charges and run-start fee, then include any account-level platform compute, storage, and retention implications. For DIY, there is no Apify Actor event fee, but browser compute, hosting, storage, engineering time, and maintenance still have costs. Which is cheaper depends on capture count, cadence, page length, memory, retention, and the value of operating the pipeline; the available evidence does not establish a universal cheaper option.

6. Options and edge cases that change the choice

Requirement Apify Actor DIY Puppeteer implication
Public pages, ordinary screenshots Ready input, batch records, formats, device and wait controls. Simple to implement, but you supply recurring operations.
Custom page interaction or post-processing Limited to the Actor’s published input behavior. Code can implement custom interactions, subject to target behavior and access rules.
Authenticated or internal pages Not supported by the Actor’s anonymous-only scope; private-network targets are refused. Possible only where authorized and securely configured; access is not guaranteed.
Cookie banner handling Optional CSS-based hiding; it does not accept or reject consent. Implement site-specific consent behavior only where appropriate; hiding a banner is different from recording a consent choice.
Very long full-page output Height limit applies; output may be truncated. WebP has a documented height constraint for very long images. Browser and image limits still apply; split content or choose another format/viewport.
Long-term archive Unnamed run storage has plan retention limits; move or download files. You choose storage, naming, retention, backups, and access controls.

7. Troubleshooting

Symptom Likely cause What to do
Actor run is successful but some URLs have no image Run completion does not imply every URL succeeded; failures are recorded per URL. Inspect each record’s status and error, retry eligible URLs, and keep partial results.
Some URLs are skipped The configured maximum run cost was reached. Review the cost limit and split batches or adjust the limit after estimating output charges.
Full-page image is incomplete Configured maximum height was reached or the output was truncated. Inspect truncation metadata; increase the allowed height if supported, or capture sections/viewports.
Cookie notice remains visible The Actor’s optional hiding only targets supported CSS behavior; it does not accept consent. Check whether the option is enabled and whether the page’s notice is hideable under the Actor’s documented behavior.
Puppeteer cannot launch a browser Browser binary missing, incompatible, or unavailable in the runtime. Install the browser required by your Puppeteer package, verify executable configuration, and allocate enough memory.
puppeteer-core ignores config It does not use Puppeteer’s standard configuration files or environment variables. Set launch options explicitly and follow the configuration guide.
Navigation times out on a page that appears loaded Persistent requests prevent network-idle conditions, or the site is slow. Use a different waitUntil condition, then wait for a page-specific selector with a timeout.
Screenshot is blank or blocked The site may require interaction, deny automated access, or fail to load. Inspect the page and error details. Do not assume either tool can bypass access controls or CAPTCHA challenges.
Scheduled captures disappear later Files were stored in an unnamed run store with limited retention. Download or copy desired files to storage with a retention policy suited to the archive.

8. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. A single GET request returns an image or PDF; its parameter names also work with those used by other screenshot APIs. See the API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
  • Cookie banners are accepted like a visitor, and 60+ known consent platforms, newsletter popups, and chat widgets are removed before capture; each step can be turned off.
  • Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing; response headers report the page verdict and billing status.
  • An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
  • 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000. Every feature is on every plan.

Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.

9. FAQ

Is the Apify screenshot Actor maintained by Apify?

The current Store listing names a community maintainer. Do not confuse it with the separate repository-backed example Actor.

Can I use the Actor for a daily list of URLs?

Yes, it supports batches and Apify schedules. Keep the batch within its published limits and inspect each URL record after every run.

Does DIY Puppeteer guarantee more reliable captures?

No comparative reliability test is available. Reliability depends on the implementation, runtime, sites, and operational monitoring.

Which is better for authenticated pages?

The documented Actor workflow is anonymous-only. DIY may support authorized authentication flows, but security and site access constraints remain your responsibility.

Should I choose based on the listed per-image price alone?

No. Compare the complete workload: run starts, platform usage, browser hosting, engineering, concurrency, and how long you must retain the outputs.

Sources