ScreenshotNeo

BlogEngineering

How Caching Works in Stagehand and Where It Breaks

Stagehand has server-side result caching and agent action replay. Learn how to tell them apart, configure each by version, and investigate misses or skipped custom tools.

By the ScreenshotNeo team29 September 202610 min read

How Caching Works in Stagehand and Where It Breaks

Stagehand caching refers to more than one mechanism. The server-side cache for act(), extract(), and observe() stores inference results on Browserbase and is distinct from agent action replay, which can replay browser steps from a local cache directory. The configuration depends on your Stagehand version: the v3 reference uses serverCache, while Browserbase’s August 21, 2026 v4 changelog describes cache with a hit-count threshold. Start by identifying the version, cache type, and browser environment before debugging a miss.

1. Identify which Stagehand cache you mean

There are at least two behaviors developers may call “the cache.” Treating them as interchangeable is a common source of confusion.

Stagehand server-side inference caching and agent action replay reuse different things and need separate troubleshooting.
Stagehand server-side inference caching and agent action replay reuse different things and need separate troubleshooting.
Mechanism What it reuses Where it applies Configuration clue
Server-side inference cache, v3 Results from act(), extract(), and observe() Browserbase environment; the v3 reference says it has no effect locally serverCache
Server-side inference cache, v4 Results from act(), observe(), and extract() Browserbase servers cache: { threshold: n }, or cache: false
Agent action replay cache A sequence of agent actions recorded for later replay Agent workflows; v3 reports reference cacheDir Agent caching and a local cache directory

The v3 API reference says server caching is enabled by default, applies only when env: "BROWSERBASE", and can be overridden per act(), extract(), or observe() call. Its changelog says repeated calls with the same inputs can return without consuming LLM tokens. [Stagehand v3 API reference](https://github.com/browserbase/stagehand/blob/main/packages/docs/v3/references/stagehand.mdx) · [Stagehand changelog](https://github.com/browserbase/stagehand/blob/main/CHANGELOG.md)

The v4 changelog describes a configurable threshold and cache-status metadata. Those names and semantics should not be copied into v3 code without checking the version-specific API. [Browserbase: Configurable caching in Stagehand](https://www.browserbase.com/changelog/caching-configurable)

2. Configure server caching in Stagehand v3

In v3, serverCache is an instance option and can be overridden on individual operations. You need a Browserbase run to benefit; setting it in a local run does not turn on an equivalent local inference cache.

import "dotenv/config";
import { Stagehand } from "@browserbasehq/stagehand";

const stagehand = new Stagehand({
  env: "BROWSERBASE",
  apiKey: process.env.BROWSERBASE_API_KEY,
  model: "openai/gpt-4o",
  serverCache: true, // v3 default is true
});

try {
  await stagehand.init();
  const page = stagehand.context.pages()[0];
  await page.goto("https://example.com");

  const first = await stagehand.extract("Get the page title");
  const second = await stagehand.extract("Get the page title");
  console.log(first, second);

  // Turn off server caching for just this operation.
  const fresh = await stagehand.extract("Get the page title", {
    serverCache: false,
  });
  console.log(fresh);
} finally {
  await stagehand.close();
}

Install the package and provide the credentials through environment variables before running this example. The v3 reference lists BROWSERBASE_API_KEY and model-provider keys such as OPENAI_API_KEY. Keep keys out of source control. Use the actual model configuration and method signature appropriate to your installed v3 release.

To disable server caching for the whole instance, use serverCache: false in the constructor. To disable it for a single operation, use the per-call serverCache: false override. The documented per-call control applies to act(), extract(), and observe(). This is useful when the page state or action result must be freshly inferred; it does not make a local run use Browserbase’s cache.

3. Configure server caching in Stagehand v4

Browserbase’s v4 changelog gives the threshold as the number of identical results the server must observe before serving a cached result. For example, an instance threshold of two means the server observes two identical results before serving that result from cache. A per-step threshold of one makes the second matching call a hit, as in the published example. A call can override the instance setting or set cache: false for that call.

A v4 cache threshold controls how many matching results are observed before the server serves a cached result.
A v4 cache threshold controls how many matching results are observed before the server serves a cached result.
const stagehand = await Stagehand.create({
  browser,
  cache: { threshold: 2 },
});

// This step uses its own threshold. The changelog's example
// says the second identical call is served from cache.
await stagehand.act("click the Sign in button", {
  cache: { threshold: 1 },
});

// Bypass the cache for one operation.
const result = await stagehand.extract("Read the current account status", {
  cache: false,
});

console.log(result.metadata?.cache);

This is the shape shown in Browserbase’s v4 changelog; use it only with a Stagehand version whose API supports it. The v4 result metadata is described as containing a cache status (HIT, MISS, or DISABLED), a miss reason, and tokens saved by a hit. Inspect the returned metadata in your installed version rather than assuming every method returns an identical object shape.

The changelog also says model configuration is outside the cache key, so changing models does not invalidate the cache according to that description. It does not document every cache-key input, expiration period, or invalidation rule. Don’t infer that a changed page, prompt, browser state, or other parameter will always hit—or always miss—unless the version’s documentation confirms it.

4. Distinguish inference caching from agent replay

Agent replay is about repeating a workflow, not simply reusing an inference result. Stagehand’s v3 reference describes cacheDir as a directory for caching action observations. A replay may save work when the workflow is repeated, but the replay’s correctness depends on all essential steps being represented and on the page still being in a compatible state.

A GitHub report about agent caching says custom tool actions were not recorded or replayed in the reported case, so essential work such as entering runtime credentials could be skipped during replay. That issue is marked open. Treat it as a specific report, not proof that all releases behave this way. If a workflow depends on custom tools, verify those steps execute on every run and test cache hits against the exact Stagehand release you deploy. [Stagehand issue #1558](https://github.com/browserbase/stagehand/issues/1558)

Keep the two diagnostics separate: server-cache status tells you about inference-result reuse; an agent replay problem can occur even if the server cache is behaving as documented.

5. Troubleshoot a cache that always misses

  1. Check the installed Stagehand version. Read the dependency lockfile or package metadata. Confirm whether the code and docs you are following are for v3 or v4. Do not mix serverCache with v4’s cache threshold examples.
  2. Check the browser environment. For v3 server caching, confirm env: "BROWSERBASE". The v3 reference explicitly says local environments are unaffected.
  3. Check configuration scope. Look for an instance-level setting and a per-call override. A call-level disabled setting can explain why one operation does not use the cache even when the instance enables it.
  4. Read the returned cache metadata. In v4, distinguish HIT, MISS, and DISABLED, and inspect the miss reason. A disabled call is different from a miss after an eligible lookup.
  5. Compare genuinely repeated operations. Repeat the same workflow in the same version and environment, then inspect the result metadata. The full cache key is not established by the sources here, so do not assume that visually similar calls are cache-identical.
  6. Test custom tools separately. If the browser is replaying agent actions but a tool’s side effect is absent, inspect the agent replay path and verify whether the custom tool ran. Do not diagnose that symptom as an act() server-cache miss.

A historical issue reported Stagehand 3.1.0 calls to act(), extract(), and observe() returning repeated misses on Browserbase despite serverCache: true. The issue is closed, but the retrieved issue page does not establish the fix or the first corrected release. If your symptoms match, treat the report as a lead: log your exact version, environment, operation, configuration, and cache metadata, then check current version-specific documentation and issue history. Do not assume the old report describes every current release. [Stagehand issue #1767](https://github.com/browserbase/stagehand/issues/1767)

6. Errors and edge cases

Symptom Likely explanation What to do
No cache benefit on a local run v3 server caching only applies to Browserbase. Use Browserbase for the documented server cache, or evaluate local action replay separately.
Every call reports MISS Version or environment mismatch, changed operation inputs, a threshold not yet reached, or a reported release-specific issue. Verify version and environment first; in v4 inspect miss reason and threshold.
DISABLED status Cache bypassed by the instance or a per-call option. Inspect both configuration scopes. Remove the override only when reuse is desired.
Agent replay skips credential entry or another custom action A reported agent-cache limitation with custom tools may apply. Check the issue and release status, and ensure critical side effects are performed or validated outside an unsafe replay assumption.
Method appears to ignore the setting Using the wrong version’s option name, or applying a per-call option to an unsupported method. Use the reference for the installed major version and the specific operation.

Do not place non-repeatable side effects behind an assumption that a cached result will execute them again. A cached inference result and a replayed action sequence are different things. For workflows that write data, submit forms, or use external tools, make idempotency and state validation explicit in the workflow itself.

7. Performance, reliability, and cost

A server-cache hit can avoid an inference call and token use; Browserbase’s v4 description says a hit skips inference entirely and reports tokens saved. That is a mechanism, not a published latency benchmark or guaranteed percentage saving. Actual benefit depends on how often eligible calls repeat and on the configured threshold. A higher threshold requires more identical results before serving the cache, so it may be appropriate when establishing a repeatable result matters more than getting an early hit. The changelog’s threshold examples are configuration examples, not a performance study.

For reliability, assume cached results can be inappropriate when your workflow expects current page state. Use a per-call bypass when fresh inference is required, and validate critical state after actions. For agent replay, test both a cold run and a repeated run, including custom tools and external side effects. Keep logs of version, environment, cache configuration, result status, and relevant operation inputs so a reported miss can be reproduced.

For cost accounting, separate inference-token savings from browser execution, hosting, and other service charges. The cited sources explain avoided inference and token behavior, but do not establish an all-in cost model or guarantee that all other work is free. Measure the metrics your application actually exposes and review the plan terms for the services you use.

8. Or skip the browser setup

If your job is to capture a page as an image or PDF rather than run an interactive Stagehand workflow, ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. A single GET request returns PNG, JPEG, WebP, or PDF. Its response distinguishes clean shots from bot checks, blank pages, failed loads, and cache hits; only clean shots are billed. That is a different job from Stagehand inference caching, but it can remove the need to configure and maintain a browser for screenshot capture. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://stripe.com \
  -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://stripe.com',
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
const bytes = new Uint8Array(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', bytes));

Cookie banners and consent prompts, newsletter popups, and chat widgets are removed before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are never billed. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for AI agents. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

FAQ

Does Stagehand caching work in local environments?

The v3 server-cache reference says no. It applies to env: "BROWSERBASE". Agent action replay is a separate mechanism.

How do I disable Stagehand server caching for one call?

In v3, pass serverCache: false to a documented supported method. In the v4 changelog’s API, pass cache: false. Match the option to your installed version.

Why is Stagehand cacheStatus always MISS?

Verify version, Browserbase environment, call-level overrides, and—on v4—the threshold and miss reason. A closed report documented recurring misses in v3.1.0, but does not establish the fix or affected release.

Why are custom tool calls skipped when an agent cache replays?

An open issue reports that behavior for a specific case. Verify your release and confirm that essential tool side effects execute on replay before relying on the cached workflow.

Does switching models clear the v4 cache?

Browserbase says model configuration is excluded from the cache key. The available description does not specify every key component or invalidation rule.

Sources and limits

Configuration and behavior above follow the [Stagehand v3 API reference](https://github.com/browserbase/stagehand/blob/main/packages/docs/v3/references/stagehand.mdx), [Stagehand changelog](https://github.com/browserbase/stagehand/blob/main/CHANGELOG.md), and Browserbase’s [August 21, 2026 v4 caching changelog](https://www.browserbase.com/changelog/caching-configurable). The reported limitations are attributed to [issue #1767](https://github.com/browserbase/stagehand/issues/1767) and [issue #1558](https://github.com/browserbase/stagehand/issues/1558). These issue reports do not prove a universal current defect. The cited material does not establish every cache-key input, expiration policy, invalidation rule, or general performance result.