Jev Computer Use: How to Add a Decision Gate to an Agent
Jev Computer Use adds a typed decision gate to an agent’s action loop. Learn what it does, how to wire it in, and where human review and verification still matter.

Jev Computer Use is a decision-layer pattern for computer-use agents: an agent proposes an action, Jev evaluates that proposal against the current state and a fixed question, and the host runtime decides whether to execute, pause for a person, or stop. Jev does not click, type, or operate the computer itself. Your agent still needs an executor and a separate way to verify what happened.
The useful mental model is observe → decide → validate → act → verify. Jev supplies a typed judgment inside that loop. It does not replace the model that observes the interface, the code that performs an action, or the policy that authorizes it. TypeSafe AI describes a decision time of roughly 70–500 ms; treat that as vendor-reported guidance, not a latency guarantee for your network or application. TypeSafe AI’s Jev computer-use overview describes the pattern and its intended gates.
1. What Jev does—and does not do
A computer-use agent commonly has a perception or planning component that reads a screen or structured interface and proposes an action such as clicking a button. Jev can judge whether that proposal is safe, choose among a bounded set of allowed actions, or decide whether the loop should escalate. The surrounding program then applies its own rules and either calls a registered executor or stops.
Jev returns structured decisions rather than free-form instructions to execute. That makes it possible to define the answer space in advance, but a valid answer is not proof that an action is safe or correct. The application must still enforce permissions, validate the target, and verify the outcome.
| Component | Responsibility |
|---|---|
| Observer or planner | Reads the current UI or tool state and proposes a candidate action. |
| Jev | Answers a bounded decision question about that state. |
| Policy and validator | Checks confidence, candidate identity, freshness, scope, permissions, and confirmation requirements. |
| Executor | Performs an allowed click, typing action, file operation, or other registered operation. |
| Verifier | Independently checks whether the desired state change occurred. |
This division matters most for destructive actions. A model’s proposal and Jev’s approval are inputs to your policy, not authorization by themselves. Keep hard rules—such as allowed directories, account boundaries, and required human confirmation—in ordinary code.
2. The computer-use loop
- Observe. Collect the relevant state, such as the current page, selected window, available controls, and proposed action. Prefer structured state where available, and include only the context needed for the decision.
- Enumerate legal candidates. Give the decision layer a constrained set of choices. Each candidate should identify a specific operation and target rather than leaving the model to invent a command.
- Ask Jev. Submit the state and a fixed typed question: for example, whether the proposed action is safe, which candidate is appropriate, or whether the loop should stop.
- Validate before execution. Reject stale observations, unknown candidate IDs, targets outside the permitted scope, and actions disallowed by your policy. Apply a confidence threshold that you have tuned against your own logs. Route low-confidence or destructive actions to a human.
- Execute through a registered tool. Dispatch only the validated candidate to a known executor. Jev selects; it does not generate arbitrary shell scripts or operate the interface.
- Verify independently. Read the new state and check for the expected result. A successful tool return or click receipt is not proof that the task completed.
- Continue, stop, or escalate. Repeat only when the verified state supports another step. Stop on completion, uncertainty, policy failure, or a repeated loop.

3. A documented Jev API decision gate
The Jev computer-use guidance describes sending the proposed step and screen state to the decision API, then reading the typed answer and confidence before execution. The request shape below uses a yes/no decision (“noul”) to gate a proposed action. This is the API call; your runtime must still implement the validation, executor, and verification steps around it.
curl https://jevtypesafeai.com/api/v1/decide \
-H "Authorization: Bearer $JEV_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"state": {
"screen_summary": "Settings page for the current account",
"proposed_action": "Click Save after changing notification preference",
"target": "Save button",
"reversible": true
},
"questions": {
"allow": {
"type": "noul",
"instructions": "Is this proposed action within the user's requested task, directed at the described target, and safe to execute without additional confirmation?"
}
}
}'
Keep the API key on the server or in a protected runtime environment. Do not embed it in browser-delivered code, logs, screenshots, or task context. The endpoint and request format are documented in the Jev API reference and the computer-use guide.
In production, parse the response according to the API’s current schema and fail closed on request errors, missing answers, unexpected values, or invalid confidence data. This example intentionally leaves response-field names out of the execution logic: confirm the current response schema in the API reference and write a parser for it rather than guessing fields.
Choose the right question type
- Yes/no (noul): Use for a bounded gate such as “should this action proceed?” A yes/no result does not replace fixed policy checks.
- Choice: Use when the runtime has already enumerated options, such as selecting one of several permitted UI actions. Keep candidate identifiers stable and validate the chosen identifier against the submitted set.
- Score: Use when the application needs an ordered assessment, such as prioritizing candidates. Define how score ranges map to behavior in your own code; do not infer universal thresholds.
The API supports sending focused typed questions with state. Ask only questions that help decide the next transition. Combining several related judgments in one request can avoid making separate decision calls, but do not combine unrelated decisions if their policies, context, or review requirements differ.
4. Safety controls that belong in your runtime
A dependable gate is fail-closed: if the request times out, returns an unusable result, or the observation may have changed, do not execute the pending action. Re-observe and ask again, or escalate. Avoid “best effort” fallbacks that silently run the proposed action when Jev is unavailable.

- Freshness: Attach an observation identifier or timestamp to each candidate. Invalidate the decision if the window, page, selection, or target changed after the observation.
- Candidate identity: Match the returned selection to the exact candidates submitted. Never treat a natural-language paraphrase as an executable command.
- Scope: Enforce allowed applications, sites, files, accounts, and operation types in code.
- Destructive effects: Require human confirmation for deletion, sending, purchasing, permission changes, or other consequential actions according to your policy—even if the model is confident.
- Least privilege: Give executors only the permissions needed for the task. A decision gate does not make a broadly privileged tool safe.
- Audit trail: Record the observation reference, candidates, decision, policy result, executor receipt, and verification outcome. Redact secrets and sensitive content.
- Independent verification: Check the resulting state, not just whether the executor returned success.
Confidence thresholds are application-specific. Tune them using representative logs that include confusing, stale, ambiguous, and adversarial cases. The source material does not establish a universal accuracy guarantee or a single threshold appropriate for all agents.
5. Platforms and implementations
The official API pattern can sit behind any host agent that can make an API request, but platform coverage for the actual interaction depends on your executor and observation layer. The API does not provide a universal desktop automation runtime.
| Approach | What the research supports | Important limit |
|---|---|---|
| Official Jev API pattern | Typed decision gates for a host agent; execution and policy remain yours. | You integrate observation, execution, verification, and confidence policy. |
| CUA-JEV reference framework | Windows UI Automation, browser DOM, Excel COM, CLI, MCP, and file APIs. | Its published runs are bounded case studies, not repeated benchmarks. Generalization to arbitrary tasks and macOS/Linux desktop support are not established. |
| jev-use community implementation | macOS Accessibility-tree observation and local read/act/check loop, with voice or typed commands described by the project. | Requires macOS permissions and a TypeSafe key; app accessibility coverage varies. The repository is a community implementation, not the official API runtime. |
The jev-use repository describes its macOS approach as reading the Accessibility tree rather than sending screenshots. Its README also describes data sent to the Jev endpoint, including command, app/window names, labelled targets, and recent actions, and says secure text fields are excluded. Those are implementation-specific statements; review the repository and verify current data handling before deployment.
Choose an implementation by checking its observation channel, supported operating systems, action breadth, escalation behavior, verification, privacy, and evaluation history. A successful short recording demonstrates that a path worked once; it does not establish a general success rate. The research dossier describes four bounded Windows case studies of 18–21 actions each and explicitly cautions that these are not repeated speed, cost, or success-rate benchmarks.
6. Latency, reliability, and cost
Adding a decision request before each action adds a network round trip. TypeSafe AI reports roughly 70–500 ms decision time; observed end-to-end delay also depends on network conditions and the host’s observation and execution steps. The community jev-use repository reports roughly 0.3–1.5 seconds per complete loop, including Accessibility reading, selection, execution, and a follow-up check. That is a project-reported figure, not an independent benchmark.
To keep a loop responsive, keep each state concise, ask focused questions, avoid duplicate checks when a single request can answer closely related questions, and set explicit timeouts. Do not remove safety checks merely to reduce latency. For repeated or high-impact tasks, measure end-to-end time and escalation frequency on your own workload.
Reliability depends on the complete loop. UI state can change between observation and action; controls can be inaccessible; the executor can act on the wrong target; and the desired change may not persist. Re-observation, freshness checks, strict candidate validation, bounded retries, and independent verification address different failure modes. Treat timeouts and malformed or missing decisions as a stop condition.
The dossier supplies no independently established universal cost or performance benchmark. Estimate your own cost from the provider’s current terms, number of decisions, amount of state sent, and fallback rate. Log calls and escalations so you can compare the value of a gate against the extra latency and cost in your workflow.
7. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| The action executes despite a failed decision call. | Error handling falls through to the executor. | Make API errors, timeouts, and invalid responses block execution. Provide a deliberate human-review path. |
| The agent selects a control that moved. | The decision used stale UI state or an unstable target label. | Bind candidates to the observation; re-read the interface immediately before execution and reject stale targets. |
| A gate approves an action outside the user’s request. | The question is vague, or scope is enforced only by the model. | Include the relevant user intent and candidate details, then enforce task scope and permissions in ordinary code. |
| The agent repeats the same action. | The loop checks tool completion rather than the resulting state, or it has no progress limit. | Verify the expected state change; detect repeated observations/actions and stop or escalate. |
| Accessibility automation cannot find a control. | The application does not expose a useful Accessibility tree, or required macOS permission is missing. | Check OS permissions and app accessibility support; use a supported observation/execution channel where necessary. |
| Confidence is high on a risky action. | Confidence is being treated as authorization or a universal correctness score. | Keep destructive-action policy separate and require confirmation where appropriate. Calibrate thresholds against your own cases. |
| Latency is unexpectedly high. | The total includes observation, network call, execution, and verification, not just the decision. | Measure each stage separately, reduce unnecessary context and calls, and preserve safety gates. |
8. Or skip the browser setup
For tasks that need a website screenshot rather than a computer-use action loop, ScreenshotNeo is a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF; the MCP tools let AI agents take screenshots, get page information, and capture PDFs. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits cost nothing; response headers report the page verdict and billing status. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000.
Sign up free for 1,000 screenshots a month, with no card required.
9. Frequently asked questions
Does Jev control the computer itself?
No. Jev returns a decision. The host agent and its registered executor perform the action, and the host should verify the result.
Should I gate every action?
That depends on the risk, latency, and cost of your workflow. The documented computer-use pattern allows checking each proposed step, but your policy should distinguish routine reversible steps from actions that require a person.
Can Jev replace a computer-use model?
No. The pattern assumes another component observes the interface and proposes actions. Jev contributes a bounded decision at selected points in that loop.
Does a high confidence value guarantee the action is correct?
No. Confidence is useful for routing and escalation when calibrated for your task. It is not a correctness guarantee or a substitute for fixed permissions and verification.
Is there a universal success rate for Jev computer use?
The cited project evidence does not establish one. The published Windows examples are bounded case studies, and the macOS timing is repository-reported. Evaluate the full loop on representative tasks before relying on it.


