Lessons from Running Headless Browsers in Production
Reliable browser automation depends on matching browser binaries to library versions, runtime dependencies, and the exact headless mode you deploy.
Reliable headless browser runs start with a reproducible environment: pin the automation library, install the browser binaries that version expects, include the operating system dependencies, and use the same browser channel in CI and production. Then make failures diagnosable with browser logs and traces. There is no universal memory, throughput, reliability, or cost figure for browser workers; measure those against your workload and runtime.
This guide focuses on operating Playwright and Puppeteer in CI or deployed services. It covers versioning, browser modes, containers, caching, test reliability, troubleshooting, and when a screenshot API may be a better fit than running browser infrastructure yourself.
1. Treat the library and browser as one deployable unit
Playwright releases expect particular browser binaries. A package update can therefore require a browser reinstall. Keep the automation package version and its browser installation in the same reproducible build or deployment process. Avoid building an image with one Playwright version and later changing the package independently.
For a Chromium-based Playwright setup on Linux, install the package and its expected browser and system dependencies as part of the build:
npm ci
npx playwright install --with-deps chromium
Pin dependencies in your lockfile and rebuild the image when you intentionally update Playwright. Apply the same principle to Puppeteer: inspect how the deployed package obtains its browser and where the browser cache lives, then make that behavior part of the image or startup contract.
For version-specific installation details and supported browser channels, see the Playwright browser documentation.
2. Choose the headless implementation deliberately
“Headless Chromium” can refer to different implementations. Playwright documents both its headless shell and a newer Chromium headless channel. The newer channel runs the real Chrome browser and is described as more suitable for higher-accuracy end-to-end and extension testing. The shell may be a better fit when its behavior matches the workload and resource constraints.
Choose based on what the job must reproduce, then validate that exact mode. Run CI and deployed workers with the same browser channel so differences in rendering or browser capabilities do not appear only after release.
import { chromium } from 'playwright';
const browser = await chromium.launch({
// Use this when you need the newer Chrome headless implementation.
channel: 'chromium',
headless: true
});
try {
const page = await browser.newPage();
await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
console.log(await page.title());
} finally {
await browser.close();
}
If you omit the channel, Playwright uses its default browser configuration. Do not assume the default and an explicitly selected channel behave identically; verify the configuration you intend to ship. Playwright relays Chrome documentation’s description of new headless as the “real Chrome browser,” and says it is more suitable for high-accuracy end-to-end or extension testing. See the browser documentation.
3. Build the runtime your cloud actually provides
A browser package alone does not guarantee that a browser can launch. The operating system needs compatible libraries and other runtime dependencies, and the process needs access to the browser binary and cache directory. Check these details for the actual container base image and cloud runtime you deploy.
Puppeteer’s cloud troubleshooting guide specifically notes that Google Cloud Run’s default Node.js runtime does not include the system packages required by Headless Chrome; its guidance is to use a custom Dockerfile and install the dependencies. Other cloud environments can have different package and browser-cache behavior, so validate their own runtime rather than carrying over assumptions.
A deployment checklist:
- Install the browser version expected by the automation package.
- Install required OS packages in the image or runtime setup.
- Confirm the browser cache path is writable and persists for the lifetime you expect.
- Launch the browser during image validation or a deployment smoke check.
- Run the same browser channel and relevant launch options in CI and production.
- Keep image and automation-library updates deliberate and reproducible.
See Puppeteer’s cloud troubleshooting guidance for Cloud Run and Google environment cache considerations.
4. Measure browser caching before adopting it
Restoring a browser cache may take about as long as downloading the browser, and Linux system dependencies cannot be cached as browser binaries. Playwright therefore does not recommend caching browsers by default in CI.
Measure both restore and install time in your own CI environment. If caching wins for your workflow, key the cache to the Playwright version so a package update cannot silently reuse incompatible binaries. Recheck the result when the runner, image, or dependency installation changes.
| Choice | Operational tradeoff | Good next step |
|---|---|---|
| Install during each CI run | Predictable version pairing; repeats download work. | Measure total install time and keep package and browser installs together. |
| Restore a browser cache | May reduce downloads, but cache restoration can cost similar time; OS dependencies still need installation. | Compare restore and install times; key by Playwright version. |
| Bake browser into a container image | Provides a controlled runtime image, which must be rebuilt when browser or library versions change. | Build both package and expected browser into the same image pipeline. |
Playwright’s CI guidance describes its caching recommendation and version-keying advice.
5. Make automation checks resilient and observable
Use locators and web-first assertions for user-visible state. Playwright’s migration guidance discourages relying on ElementHandle and recommends locators, which resolve against current page state, together with assertions that wait for the expected condition.
import { test, expect } from '@playwright/test';
test('account page shows the signed-in state', async ({ page }) => {
await page.goto('https://example.com/account');
await expect(page.getByRole('heading', { name: 'Your account' })).toBeVisible();
});
For failures that happen only in CI or intermittently, collect traces and browser launch logs as artifacts. Playwright supports traces through its test runner, and documents DEBUG=pw:browser for browser launch diagnostics.
DEBUG=pw:browser npx playwright test
Configure the test runner to retain a trace when a test fails, then attach the trace and relevant logs to the failed CI job. This gives you evidence about the browser actions and environment for post-mortem investigation instead of relying on a failure message alone. See Playwright CI and its Puppeteer migration guidance.
6. Decide who owns the browser operations
Teams asking whether to use a VPS or Kubernetes cluster or a hosted browser service are raising a real operational choice, but one deployment model does not win for every workload. Compare the requirements that matter to your job:
| Decision area | Questions to answer |
|---|---|
| Browser fidelity | Do you need Chromium, Firefox, or WebKit? Does the job require the headless shell or newer Chromium headless? |
| Version lifecycle | Who pins the library and browser, installs updates, and rebuilds the runtime? |
| Runtime dependencies | Which OS packages and cache paths does the target image or cloud runtime provide? |
| CI startup | Does installing browsers or restoring a cache improve total job time in this runner? |
| Failure investigation | Can the system retain traces, launch logs, and enough context to reproduce a failed run? |
| Ownership | Does the team want to operate browser workers, or use a service for the screenshot task? |
A public discussion has used the wording “How are you guys running Playwright/Puppeteer in production?” as an example of this decision. It is an anecdotal question, not evidence that a particular provider or deployment pattern is better. Make the choice from your workload and operational requirements.
7. Troubleshooting common production failures
| Symptom | Likely cause | What to check or change |
|---|---|---|
| Browser executable is missing after a package update | The installed browser does not match the automation package, or installation was omitted from the new build. | Reinstall the browser expected by the pinned package in the same build or deployment process. |
| Browser fails to launch in a container | Required system libraries are missing, or the browser cannot access its executable or cache directory. | Inspect runtime packages and paths. For Cloud Run, use a custom image with the required dependencies. |
| Tests pass locally but render differently in CI | Different browser channel, binary version, or runtime environment. | Pin and align package, browser installation, channel, and relevant image between environments. |
| CI cache does not make jobs faster | Cache restore time is comparable to downloading; system dependencies still install separately. | Measure both paths. Remove browser caching if it does not improve the job, or key a useful cache by Playwright version. |
| Intermittent launch failure has little diagnostic detail | Browser logs and execution artifacts were not retained. | Run with DEBUG=pw:browser and save test traces or other failure artifacts. |
| A check flakes while page content changes | The check is tied to an outdated element handle or timing assumption rather than current user-visible state. | Use a locator with a web-first assertion for the expected state. |
8. Performance, reliability, and cost: measure the workload
The cited setup guidance does not establish universal memory requirements, throughput, success rates, or cost per browser. Browser choice, page behavior, concurrency, image size, operating system, and runtime all affect a particular deployment. Avoid sizing or budgeting from a generic figure presented without the same workload and environment.
For a useful baseline, measure your own representative jobs in the target runtime: startup and navigation time, job completion and failure reasons, resource use at the planned concurrency, and the cost of the machines or service used. Compare the same pages, browser channel, and output requirements. Retain failures and rerun representative cases when changing the browser version, image, or concurrency.
Or skip the browser setup
If your job is to produce a website screenshot rather than operate a general-purpose browser worker, ScreenshotNeo provides a screenshot API and MCP server for developers. A single GET request returns a PNG, JPEG, WebP, or PDF. The API uses familiar screenshot parameters, and the ScreenshotNeo documentation covers the available options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
Cookie banners are accepted and removed before capture, along with known newsletter popups and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and billing status. AI agents can use its MCP server tools, including take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up free for 1,000 screenshots a month, with no card required.
FAQ
Should I use Playwright or Puppeteer in production?
The available setup guidance does not establish a universal winner. Choose based on browser and test requirements, deployment environment, and the operational model your team can support.
Does headless mode guarantee the same result as a headed browser?
No. Headless implementations and channels can differ. Choose the channel that matches the required fidelity and test it in the runtime you will deploy.
Is browser caching always faster in CI?
No. Cache restoration can take about as long as downloading, and it does not replace installing Linux system dependencies. Measure both paths in your runner.
Can an API replace a browser worker for every automation task?
No. A screenshot API fits screenshot and PDF capture workflows; tasks that need general browser interaction or custom execution may still need an automation library and browser runtime.


