Why Pyppeteer Gets Stuck in Docker and How to Fix It
Find whether Pyppeteer is stuck downloading Chromium, launching it, or waiting on a page—and fix the container issue safely.
Pyppeteer can appear to hang in Docker for several different reasons. The process may still be downloading Chromium, the executable may be missing a shared library, Chromium may be blocked by its sandbox configuration, the container may have too little shared memory, or your script may already have launched the browser and be waiting on navigation or a selector.
Start by identifying the exact operation that stops. Add logs immediately before and after launch(), page creation, navigation, and selector waits. Then check the browser download and executable, system libraries, sandbox and container identity, shared memory, and child-process cleanup in that order.
1. Identify the symptom before changing Docker flags
Run the container with visible application and browser output. Pyppeteer exposes debugging and output options through its launcher API; dumpio=True routes browser process output to the parent process.
import asyncio
import logging
from pyppeteer import launch
logging.basicConfig(level=logging.DEBUG)
async def main():
print("before launch", flush=True)
browser = await launch(
headless=True,
dumpio=True,
# Add flags only after you have evidence they are needed.
args=[]
)
print("after launch", flush=True)
page = await browser.newPage()
print("after newPage", flush=True)
await page.goto("https://example.com", {"waitUntil": "networkidle2", "timeout": 60000})
print("after goto", flush=True)
await page.waitForSelector("body", {"timeout": 30000})
print("after selector", flush=True)
await browser.close()
asyncio.get_event_loop().run_until_complete(main())
Interpret the last message:
- No output after
before launch: investigate Chromium installation, executable permissions, missing libraries, sandbox setup, and container resources. after launchappears but navigation does not finish: the browser started; investigate DNS, outbound access, the target site, navigation wait conditions, and your timeout.- Navigation completes but a selector wait does not: the page may render different markup in the container, require authentication, or never create that selector.
Pyppeteer’s launcher reference documents dumpio, logging controls, and launch arguments. Keep these boundary logs while diagnosing instead of adding several unrelated flags at once.
2. Check Pyppeteer’s Chromium download and executable path
Pyppeteer downloads Chromium on first use unless you install it ahead of time with pyppeteer-install. In a container, that first-use download often happens in a build stage or runtime filesystem that is later discarded, or under a different user and cache directory.
Install Chromium during the image build
FROM python:3.11-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt \
&& pyppeteer-install
COPY app.py .
CMD ["python", "app.py"]
Build and run the same image that will execute the job. Confirm the binary and cache are present inside the runtime container:
docker build -t pyppeteer-job .
docker run --rm -it pyppeteer-job sh
python -c "import pyppeteer; print(pyppeteer.__file__)"
find / -type f -name 'chrome*' 2>/dev/null | head
If your build uses one user and runtime uses another, make the Chromium cache readable by the runtime user or configure a cache location that both stages preserve. Also verify that the container has write access if Pyppeteer is expected to download at runtime.
Use an explicit executable only when you control compatibility
browser = await launch(
headless=True,
executablePath="/usr/bin/chromium",
dumpio=True,
)
Pyppeteer exposes executablePath, but its documentation recommends the bundled Chromium in practice and does not guarantee compatibility with other browser versions. If you select a system browser, record the browser version, the Pyppeteer version, and the image digest so a rebuild is reproducible.
3. Check shared libraries in the actual image
A browser file can exist and still fail immediately because a shared library is missing. The Puppeteer troubleshooting guide discusses this failure mode for its bundled Chrome for Testing in Docker; use it as adjacent Chromium guidance, not as a Pyppeteer-specific dependency list.
docker run --rm -it pyppeteer-job sh
which chromium || true
ldd /usr/bin/chromium | grep 'not found' || true
cat /etc/os-release
Inspect the executable that your process actually uses. Do not copy a package list from an unrelated base image without checking the missing-library output. Choose a maintained browser-capable base image or install the exact libraries reported by the binary, then rebuild and repeat the launch-boundary test.
4. Check sandbox settings and container identity
Chromium’s sandbox is a security boundary. Whether it can run depends on the user, kernel, capabilities, and container configuration. A root process and a non-root process do not have the same requirements.
The current Puppeteer Docker guidance describes a sandboxed image that requires SYS_ADMIN. Playwright’s Docker documentation shows a separate configuration in which a root-run image disables Chromium sandboxing. These are adjacent project examples, not universal Pyppeteer settings.
Prefer a sandboxed configuration where possible
Run the browser as a non-root user and provide the capability and image setup required by your environment. Verify the result from browser logs. If your threat model permits no sandbox, the commonly considered flag is:
browser = await launch(
headless=True,
args=["--no-sandbox", "--disable-setuid-sandbox"],
dumpio=True,
)
--no-sandbox changes the security posture. Treat it as a deliberate containment decision for trusted pages and an appropriately isolated workload, not as a harmless universal fix. If adding it makes launch succeed, document why it is acceptable and plan a sandboxed configuration for untrusted content.
5. Check shared memory and memory limits
Chromium uses shared memory for several browser operations. The Playwright Docker guide recommends --ipc=host because Chromium can run out of shared memory and crash without enough capacity. Apply that as adjacent Chromium-in-container guidance and first measure your own container limits.
docker run --rm --shm-size=1g pyppeteer-job
# Or, where your isolation policy allows it:
docker run --rm --ipc=host pyppeteer-job
Also inspect the container’s memory limit and whether the kernel killed the process:
docker inspect <container> --format '{{json .HostConfig.ShmSize}}'
docker stats --no-stream <container>
docker events --since 10m | grep -i -E 'oom|kill' || true
Increase shared memory or memory limits when logs show renderer crashes, out-of-memory events, or failures that disappear under a larger limit. Do not use a larger setting to mask a missing-library or sandbox error.
6. Check child-process cleanup and PID 1 behavior
Long-running workers can accumulate orphaned Chromium processes when the container has no init process. Official Puppeteer and Playwright Docker guidance recommends an init process; Playwright explicitly connects this with handling zombie processes at PID 1.
docker run --rm --init pyppeteer-job
This is most relevant when repeated jobs become slower, process counts grow, or old browser processes remain after an exception. It is less conclusive for one isolated launch delay. Always close the browser in a finally block:
browser = None
try:
browser = await launch(headless=True, dumpio=True)
page = await browser.newPage()
await page.goto("https://example.com", {"timeout": 60000})
finally:
if browser is not None:
await browser.close()
7. Separate browser launch from navigation waits
There is no single timeout flag that proves a Pyppeteer Docker launch problem. Log each asynchronous boundary and set an explicit timeout for each operation.
await page.goto(
"https://example.com",
{"waitUntil": "domcontentloaded", "timeout": 60000}
)
await page.waitForSelector("main", {"timeout": 30000})
Use domcontentloaded when waiting for every network request is unnecessary. Use networkidle2 only when the page’s background traffic is known to settle. A page can keep connections open indefinitely, and a selector can be absent because the site serves different content to the container’s user agent or requires credentials.
8. A repeatable Docker diagnostic setup
- Build an image that runs
pyppeteer-installduring the build. - Run the same image and user in which the production job executes.
- Enable
dumpio=Trueand Python logging. - Log before and after
launch(),newPage(),goto(), and every selector wait. - Confirm the executable path and run
lddagainst that exact binary. - Check user identity, sandbox errors, capabilities, memory, and
/dev/shm. - Add
--initfor repeated jobs and close every browser in cleanup code. - Only then test a narrowly justified launch argument, recording its security and reproducibility impact.
9. Troubleshooting common errors
| Observed symptom | Likely cause | Fix |
|---|---|---|
| Stalls before any browser output | First-use Chromium download, unwritable cache, or blocked network | Run pyppeteer-install at build time; verify the cache and runtime user. |
| Executable not found | Build-stage browser was not copied into the runtime image or the configured path is wrong | Inspect the runtime container and remove or correct executablePath. |
| Shared library error | Image lacks a dependency required by the selected browser | Run ldd on the actual executable and install the reported libraries or use a suitable base image. |
| “Running as root without –no-sandbox” or sandbox failure | Container identity and sandbox configuration are incompatible | Prefer a non-root sandboxed setup; use no-sandbox only after evaluating isolation and trust. |
| Browser starts, then renderer crashes | Insufficient memory or shared memory | Inspect OOM events and /dev/shm; adjust limits based on evidence. |
| Repeated jobs leave Chromium processes | Missing init process or incomplete cleanup | Run with --init, close browsers in finally, and inspect process counts. |
goto() never returns |
Navigation wait condition, DNS, outbound access, or target-site behavior | Log before and after navigation, set a timeout, test connectivity, and choose an appropriate waitUntil. |
| Selector wait times out | Markup differs, content is delayed, authentication is missing, or the selector is wrong | Capture HTML or a diagnostic screenshot, verify the selector, and wait for the application state you actually need. |
10. Performance, reliability, and cost considerations
- Build once, reuse the image: downloading Chromium during image creation removes a variable network step from every job.
- Reuse a browser carefully: one browser with controlled pages can avoid repeated startup cost, but close pages and restart the browser when memory or process counts grow.
- Pin the environment: keep the Python, Pyppeteer, browser, and base-image versions together. An
executablePathoverride increases your responsibility for compatibility. - Use bounded waits: explicit navigation and selector timeouts prevent one page from occupying a worker forever.
- Size resources from observations: monitor memory and shared memory during representative pages; larger limits do not repair a bad executable or sandbox configuration.
- Make retries selective: retry transient navigation or network failures, but do not repeatedly retry a deterministic missing-library or permission error.
11. Or skip the browser setup
If your goal is a clean website image rather than maintaining Chromium in every container, ScreenshotNeo provides a GET-based screenshot API and an MCP server. See the API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Cookie and consent banners are accepted and removed before capture, along with more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing status. An MCP server lets Claude, Cursor, and other MCP clients take screenshots. The Free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000 shots. Sign up free.
12. FAQ
Does installing Chromium on the host fix a Pyppeteer container hang?
Only if the container can access that executable and its compatible libraries. Containers normally need the browser and dependencies inside the image or an explicitly mounted, supported path.
Should I always add --no-sandbox?
No. First determine whether the process is running as root and whether a sandboxed configuration is possible. Disabling the sandbox changes isolation and should match your threat model.
Why does the same script work locally?
The local machine may have a cached bundled browser, more shared memory, different libraries, a different user, or unrestricted outbound networking. Compare those conditions with the runtime container.
Can a longer timeout solve the problem?
Only when the operation is slow but progressing. Boundary logs tell you whether the timeout belongs on browser launch, navigation, or a selector wait.


