ScreenshotNeo

BlogEngineering

Why Your AI-Built Scraper Works Locally and Breaks in the Cloud

AI-built scrapers often fail in the cloud because browsers, libraries, networking, and startup settings differ from a laptop. Here’s how to diagnose and fix them.

By the ScreenshotNeo team29 September 202611 min read

Why Your AI-Built Scraper Works Locally and Breaks in the Cloud

An AI-built scraper can work on your laptop and fail in AWS Lambda, Cloud Run, or Docker because the deployed runtime is a different machine with a different browser, Linux libraries, user permissions, sandbox, shared-memory limit, network path, certificates, environment variables, startup process, and timeout budget. The fix is to make that runtime explicit and reproducible: pin the browser and automation versions, build and run the same production image locally, verify its dependencies and network access, and collect diagnostics before adding retries.

The browser is part of the application. A local browser installation does not prove that Chromium and its required libraries exist in the deployed image. Google Cloud’s guidance for browser automation on Cloud Run says to install Chromium in the container and grant the required permissions. Playwright also uses browser builds matched to its release, so mismatched local and container versions can fail at launch. Google Cloud: Browser and OS automation in Cloud Run; Playwright browser management.

This guide uses Playwright with Node.js for the runnable example, but the diagnostic sequence applies to other browser-driven scrapers too.

1. First identify which layer is failing

“The scraper failed” is not a diagnosis. Separate failures into five layers before changing code:

A cloud scraper depends on the browser runtime, network path, and page readiness as separate layers.
A cloud scraper depends on the browser runtime, network path, and page readiness as separate layers.
Layer Typical evidence First question
Browser launch Executable missing, shared library error, sandbox denial, Chromium crash Does this image contain the expected browser and OS libraries?
Navigation and network DNS failure, TLS error, timeout, connection refused, unexpected status Can the runtime reach this destination with its proxy and certificate settings?
Page readiness Selector absent, empty content, content still loading Did the page load, and is the expected element actually ready?
Site response CAPTCHA, bot check, consent wall, denial, alternate content Did the target serve the same page to this cloud request?
Process and platform Invocation terminated, startup failure, memory or time limit exceeded Did the platform start the runtime and give it enough time and resources?

Save the exact exception and browser stderr. Also record the page URL, response status, console errors, and failed requests. These tell you whether to investigate Chromium, the network, the target page, or serverless startup.

2. Reproduce the production runtime locally

Build the same container image that you deploy, and run that image with the same user and security settings. Pin Playwright and its browser together in the dependency lockfile. Playwright documents that its browser builds are version-specific; after updating Playwright, rebuild the image so it receives the matching browser. Playwright: Docker; Playwright: Browsers.

A minimal Node.js example can be run from a project where Playwright is installed and the browser dependencies are included in the runtime:

const { chromium } = require('playwright');

(async () => {
  const browser = await chromium.launch({ headless: true });
  const page = await browser.newPage();
  page.setDefaultNavigationTimeout(30_000);

  try {
    const response = await page.goto('https://example.com', {
      waitUntil: 'domcontentloaded',
    });
    console.log({
      status: response && response.status(),
      title: await page.title(),
      url: page.url(),
    });
    await page.locator('h1').waitFor({ state: 'visible', timeout: 10_000 });
    console.log(await page.locator('h1').innerText());
  } finally {
    await browser.close();
  }
})();

For diagnostics, log the Playwright version, browser executable path, OS release, current user, installed shared libraries and fonts, and environment-variable names. Do not print secret values. Verify that the image contains the browser before invocation; downloading browsers during a serverless invocation adds a network dependency and makes cold starts less predictable.

When a container fails only under production-like security settings, investigate those settings rather than weakening them blindly. For untrusted crawling, Playwright recommends running as a non-root user and using its recommended seccomp profile. Its Docker guidance also recommends --init for process handling and --ipc=host with Chromium because insufficient shared memory can cause crashes. Test the actual deployment image digest, not just a developer image with different layers. Playwright Docker guidance.

3. Check container and serverless constraints

Browser automation starts child processes and uses substantial memory. A container that works interactively can behave differently when its process runs as PID 1, has a small shared-memory allocation, or is terminated at the platform’s invocation deadline.

  • Use Docker’s --init option where applicable so child processes are handled properly.
  • For diagnosis, test Chromium with --ipc=host where your environment supports it. If this changes stability, review the container’s shared-memory allocation and runtime constraints.
  • Keep the browser launch and page work inside a bounded invocation. Close pages and browsers in finally blocks so exceptions do not leave processes behind.
  • Measure cold-start time separately from page navigation. Do not spend the entire invocation budget downloading a browser or waiting for an unnecessarily strict network-idle condition.
  • For AWS Lambda, inspect the wrapper script and runtime startup logs before debugging page selectors. AWS warns that invocations can fail if a wrapper script does not successfully start the runtime process. AWS Lambda custom runtimes.

Cloud Run requires Chromium to be installed in the container and the necessary permissions granted. Treat the container image as the deliverable: it should include the browser, its libraries, fonts required by the pages you capture, and your application’s pinned dependencies. Google Cloud: Browser and OS automation in Cloud Run.

4. Fix network assumptions, especially localhost

localhost always means “this network namespace.” On a laptop, a scraper may reach a local development server at http://localhost:3000. Inside a container, that address usually refers to the container itself. In a remote cloud runtime, it refers to that remote instance. The local service is not reachable unless you deliberately connect the networks.

The same localhost address and TLS route can mean different things from a laptop and a remote container.
The same localhost address and TLS route can mean different things from a laptop and a remote container.
  1. Run DNS resolution and a TLS request from inside the deployed container, using the same hostname the scraper will use.
  2. Replace laptop-only localhost URLs with the service hostname reachable from the runtime. In Docker, configure the intended host mapping or service network; publish ports when a process must accept connections from outside its container.
  3. Inspect outbound firewall rules, DNS, proxy variables, and required ports. A page that loads in a browser on your laptop may be blocked from a cloud network.
  4. For HTTPS errors, check the container’s CA certificates and any corporate TLS interception. If the environment requires a custom CA, configure it explicitly; do not turn off certificate verification as a general fix.

Playwright’s Docker guidance covers host mappings, published ports, and container networking. Use the address appropriate to the process that makes the request, not an address copied from a local development setup. Playwright: Docker.

5. Compare configuration and wait behavior

A scraper may behave differently because the cloud injects different variables, uses another configuration file, or overrides values on the command line. Playwright Test documents this precedence for its options: configuration file, environment variables, then command-line arguments. Check all three sources for the effective values. Playwright Test configuration.

Timeouts that passed on a fast laptop can fail under cloud latency or cold starts. Set navigation and action timeouts deliberately, and give the platform invocation deadline enough room for browser startup plus the page operation. Waiting for networkidle can be a poor fit for pages with analytics, polling, or long-lived connections; prefer a specific readiness condition, such as a visible result selector, when that is what the scraper needs.

Playwright exposes navigation, action, download, proxy, and custom certificate authority controls. Use the relevant documented option for the failure you observe. A custom CA or proxy change is appropriate only when the runtime’s network path requires it. Playwright Browser API; Playwright Page API.

6. Follow a repeatable diagnostic sequence

  1. Classify the failure. Save the exception, stderr, page console, response status, and request failures. Distinguish launch, navigation, readiness, blocked content, and process termination.
  2. Prove the image contents. Log the Playwright version, browser path, OS, libraries, fonts, user, and environment-variable names. Confirm browser installation at build time.
  3. Run the production image locally. Use the intended user, --init, resource limits, and security profile. Test shared-memory behavior where supported.
  4. Test networking in the runtime. Resolve the host, make a TLS request, inspect proxy settings, and replace invalid localhost assumptions.
  5. Compare effective configuration. Check config files, environment variables, command-line overrides, certificate settings, and each timeout.
  6. Check startup and invocation limits. In Lambda, verify wrapper exit status and runtime startup. Confirm browser launch and page work fit within memory and time limits.
  7. Then tune readiness and retries. Wait for the data you need, and retry only bounded transient errors after the runtime is known to be correct.

7. Common errors and fixes

Error or symptom Likely cause Fix
“Executable doesn’t exist” or browser launch cannot find Chromium The image lacks the Playwright-matched browser, or the app expects a different path. Install the matching browser during image build, pin the Playwright version, and print the executable path in the deployed runtime.
Missing shared object or library error A required Linux library is absent from the image. Use the documented browser installation/dependencies for the chosen image and inspect installed libraries in the same image.
Chromium crashes or exits under load Shared memory, process handling, memory limits, or user permissions differ from local development. Try --init and, where supported for diagnosis, --ipc=host; review memory and security settings with the intended user.
Sandbox or permission denied The browser is running as an incompatible user or under container restrictions. Use the intended non-root setup and recommended seccomp configuration for untrusted pages; grant the documented Chromium permissions in Cloud Run.
Connection refused to localhost The service is outside the scraper’s container or remote host. Use a reachable service hostname, configure a Docker host mapping/network, and expose the needed port.
TLS certificate verification failure The container lacks a trusted CA or the cloud path intercepts TLS. Install/configure the correct CA and inspect proxy settings. Keep verification enabled unless you have a narrow, controlled reason to change it.
Navigation timeout but the page eventually loads Cloud latency, slow cold start, strict timeout, or an unsuitable wait condition. Separate startup from navigation, set a bounded realistic timeout, and wait for the required selector instead of global network silence.
Selector timeout or empty extraction The target served different content, a bot check, or the data is rendered later. Inspect screenshot, title, URL, response, console, and relevant requests before changing selectors or retrying.
Lambda invocation fails before browser logs appear Wrapper script did not start the runtime, or startup exceeded the platform’s allowance. Check wrapper exit status, executable permissions, and runtime startup logs first.

8. Performance, reliability, and cost

Browser startup and page loading are separate costs. Reuse a browser process across multiple pages within a warm worker when the platform’s process model allows it, while creating and closing page contexts per job so cookies and state do not leak between jobs. On platforms that freeze or terminate processes between invocations, verify that reuse is safe rather than assuming it.

Use a readiness condition tied to the data you need; cap navigation, action, and overall job time; and bound retries. Retrying a deterministic missing executable or certificate error wastes invocation time and can multiply cloud charges. Record durations for browser launch, navigation, readiness, and extraction separately so a slow page is distinguishable from a slow cold start.

Cloud cost depends on the runtime, memory allocation, invocation duration, and any external browser service; the dossier provides no comparable price or benchmark, so estimate with your own workload and provider pricing. Keep the exact image digest and dependency lockfile with each deployment to make a failure reproducible. Do not infer a general browser crash rate or remediation time from an isolated incident.

9. Do it yourself: deployment checklist

  • Pin the automation package and browser version together.
  • Install Chromium and required OS libraries and fonts during the image build.
  • Run the deployed image locally under the intended user and security settings.
  • Use --init; diagnose shared memory with --ipc=host where available.
  • Replace laptop-only localhost URLs with addresses reachable from the runtime.
  • Verify DNS, outbound access, proxy configuration, and certificate trust inside the runtime.
  • Compare configuration, environment variables, command-line overrides, and timeout values.
  • For Lambda, confirm the wrapper starts the runtime before investigating page logic.
  • Log launch, navigation, console, response, and request failure details without logging secrets.
  • Add bounded retries only after correcting the underlying environment or transient condition.

10. Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. One GET request captures a URL as PNG, JPEG, WebP, or PDF. Its capture flow accepts cookie and consent banners like a visitor, then removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers say which verdict applied and whether it was billed.

For a static visual capture, this avoids packaging Chromium and its Linux dependencies in your scraper. It does not replace browser automation when you need to extract data or interact with a page. See the API documentation for options and response details.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com \
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.com',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

ScreenshotNeo also offers an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. It has 1,000 screenshots a month free with no card; paid plans start at $5 for 3,000. Every feature is on every plan. Sign up for 1,000 free screenshots a month, no card required.

FAQ

Why did generated scraper code omit the cloud setup?

Generated code commonly reflects the environment where it was written. Browser installation, Linux dependencies, permissions, network topology, startup wrappers, and cloud time limits live outside the page automation logic, so verify each explicitly.

Should I add more retries when deployment starts failing?

Only after identifying a transient failure. Retries cannot install a missing browser, repair an invalid localhost address, or start a broken runtime wrapper.

Can I use the same container image for local and cloud runs?

That is the goal: build one pinned image and run the same image digest locally and in deployment, with the cloud’s actual user, network, and resource constraints represented during diagnosis.

When is a screenshot API enough?

Use one when the needed result is a rendered image or PDF. Keep browser automation when you need custom interactions, page data extraction, or control that a screenshot endpoint does not provide.