How to Build an AI Code Generation Feature in a Browser Playground
Build a secure browser IDE where users prompt AI, review file patches, run code in an isolated sandbox, and see a live preview.
An AI code-generation feature needs more than a text box connected to a model. The reliable design is a browser client backed by a trusted application server, a typed file-operation contract, an isolated execution sandbox, and a preview loop that lets users inspect, approve, run, and repair changes.
Architecture at a glance
Use these components:
- Browser client: prompt and chat panel, project tree, code editor, diff viewer, terminal or log panel, and preview iframe or preview URL.
- Trusted application server: authenticates users, stores project metadata, calls the Responses API, validates model output, streams events to the browser, and owns billing, audit logs, rate limits, and approval state.
- Structured generation contract: require a typed list of file operations such as
create,replace,delete, andrename, plus a short explanation. - Execution plane: start an isolated sandbox per project or job. Mount only project files, expose only required capabilities, and keep secrets outside the workspace.
- Preview loop: run the development server in the sandbox, expose its port, and return a preview URL to the browser.
The model should never receive your application API key. Keep that key on the server and broker access through your own authenticated endpoint. See the Responses API documentation and OpenAI’s security guidance.
Define a safe file-operation contract
Whole-file text is easy to prompt but difficult to review and easy to overwrite accidentally. Ask for patches represented as data instead:
{
"summary": "Add a responsive pricing card",
"operations": [
{
"op": "create",
"path": "src/components/PricingCard.tsx",
"content": "..."
},
{
"op": "replace",
"path": "src/App.tsx",
"content": "..."
}
],
"checks": ["npm run lint", "npm test"]
}
Validate every response before showing it:
- Parse strict JSON or a typed equivalent; reject prose outside the contract.
- Allow only the four expected operations.
- Normalize paths and reject absolute paths, empty paths, null bytes, and any
..traversal. - Resolve the path beneath the project root and verify it remains inside that root.
- Reject files above your size limit and cap the number of operations per request.
- Reject edits to protected files such as deployment credentials, server configuration, or policy files unless your product explicitly supports them.
- Check that a
replaceoperation includes the expected file version or hash so stale edits do not silently overwrite newer work.
Render a human-readable diff. Require approval before deletes, renames, large replacements, publishing, purchases, account changes, or transmission of sensitive data.
Call the model from your application server
Send the user’s request together with selected files, diagnostics, project constraints, and the contract definition. Stream progress to the browser, but buffer the complete response on the server until validation succeeds. The Responses API supports server-sent streaming events.
Node.js server example
import OpenAI from "openai";
import express from "express";
const app = express();
app.use(express.json());
const openai = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });
const contract = {
type: "object",
additionalProperties: false,
properties: {
summary: { type: "string" },
operations: {
type: "array",
items: {
type: "object",
additionalProperties: false,
properties: {
op: { type: "string", enum: ["create", "replace", "delete", "rename"] },
path: { type: "string" },
content: { type: ["string", "null"] },
to: { type: ["string", "null"] }
},
required: ["op", "path", "content", "to"]
}
},
checks: { type: "array", items: { type: "string" } }
},
required: ["summary", "operations", "checks"]
};
app.post("/api/generate", async (req, res) => {
const { prompt, files, diagnostics = [], constraints = {} } = req.body;
if (typeof prompt !== "string" || prompt.length > 8000) {
return res.status(400).json({ error: "Invalid prompt" });
}
const response = await openai.responses.create({
model: process.env.OPENAI_MODEL,
input: [
{ role: "system", content: "Return only the requested file-operation object. Never use paths outside the project root." },
{ role: "user", content: JSON.stringify({ prompt, files, diagnostics, constraints }) }
],
text: { format: { type: "json_schema", name: "file_patch", schema: contract, strict: true } },
stream: false
});
const patch = JSON.parse(response.output_text);
validatePatch(patch, files);
res.json(patch);
});
function validatePatch(patch, files) {
if (!Array.isArray(patch.operations)) throw new Error("operations must be an array");
for (const operation of patch.operations) {
if (operation.path.startsWith("/") || operation.path.split("/").includes("..")) {
throw new Error("unsafe path");
}
if (operation.content && Buffer.byteLength(operation.content, "utf8") > 1_000_000) {
throw new Error("file too large");
}
}
}
app.listen(3000);
In production, send streaming events from the server to the browser with your framework’s SSE or WebSocket support. Do not apply partial text as files. Keep partial output in memory or temporary storage until the complete object parses and passes validation.
Python server example
import json
import os
from pathlib import PurePosixPath
from flask import Flask, jsonify, request
from openai import OpenAI
app = Flask(__name__)
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
@app.post("/api/generate")
def generate():
body = request.get_json(force=True)
prompt = body.get("prompt", "")
if not isinstance(prompt, str) or len(prompt) > 8000:
return jsonify(error="Invalid prompt"), 400
payload = {
"prompt": prompt,
"files": body.get("files", []),
"diagnostics": body.get("diagnostics", []),
"constraints": body.get("constraints", {})
}
response = client.responses.create(
model=os.environ["OPENAI_MODEL"],
input=[
{"role": "system", "content": "Return only a validated file-operation object."},
{"role": "user", "content": json.dumps(payload)}
]
)
patch = json.loads(response.output_text)
validate_patch(patch)
return jsonify(patch)
def validate_patch(patch):
if not isinstance(patch.get("operations"), list):
raise ValueError("operations must be a list")
for operation in patch["operations"]:
path = PurePosixPath(operation["path"])
if path.is_absolute() or ".." in path.parts:
raise ValueError("unsafe path")
content = operation.get("content") or ""
if len(content.encode("utf-8")) > 1_000_000:
raise ValueError("file too large")
if __name__ == "__main__":
app.run(port=3000)
cURL request
curl https://your-app.example/api/generate \
-H 'content-type: application/json' \
-H 'authorization: Bearer USER_SESSION_TOKEN' \
--data-binary @request.json
Build the browser interaction loop
- The user enters a request such as “Add a dark-mode toggle and keep the existing tests passing.”
- The client sends the prompt, selected files, current file versions, diagnostics, and constraints to your server.
- The server streams progress events such as
queued,generating,validating, andready_for_review. - The client displays the explanation and a side-by-side diff. Keep generated checks separate from file changes.
- After approval, the server applies operations atomically. If any operation fails, roll back the transaction and show the error.
- Start or reuse the project’s sandbox, run the requested checks under limits, and stream logs.
- Return a preview URL. Refresh the iframe only after the server reports that the development server is listening.
- Send compiler and runtime diagnostics in a follow-up model turn tied to the same project and sandbox session.
Associate conversation state and execution state explicitly. Continuing a model response does not automatically restore browser-session variables, processes, or runtime state.
Run generated code in an isolated sandbox
Use a sandbox when code must run, install packages, access multiple files, create artifacts, or expose a preview. OpenAI’s sandbox guidance describes this as the right boundary for work that depends on a workspace rather than prompt context.
For each project or job:
- Create an isolated filesystem and mount only that project’s files.
- Run as an unprivileged user with CPU, memory, process, disk, wall-clock, and output limits.
- Restrict outbound network access to an allowlist, or disable it for projects that do not need it.
- Keep application keys, cloud credentials, signing keys, and user secrets in a vault or trusted proxy.
- Treat generated code, package install scripts, repository files, terminal output, and preview content as untrusted input.
- Expire idle sessions and clean up processes, temporary files, and network rules.
Choose an ephemeral workspace for isolation and lower idle cost, or a persistent workspace for faster iterative repair with retained dependencies. Persistent sessions require stronger cleanup and tenancy controls.
Expose a live preview
Start the project’s development command inside the sandbox, wait until the expected port is listening, and expose that port through a controlled preview URL. Return only the URL and metadata needed by the browser.
const job = await sandbox.exec({
command: "npm run dev -- --host 0.0.0.0 --port 4173",
timeoutMs: 30_000,
env: { NODE_ENV: "development" }
});
await sandbox.waitForPort(4173, { timeoutMs: 30_000 });
const preview = await sandbox.exposePort(4173);
return { previewUrl: preview.url };
Use a distinct preview origin, add frame restrictions appropriate to your product, and never let preview JavaScript call your control-plane endpoints with ambient credentials. Snapshot only required artifacts such as build output, logs, or a lockfile.
Security checklist
- Keep the model and application API keys on the trusted server.
- Authenticate every project and sandbox operation and enforce per-user quotas.
- Separate control-plane services (auth, billing, approvals, tracing, audit, recovery) from sandbox compute.
- Require confirmation for destructive edits and other consequential actions.
- Use version checks to prevent stale patches from overwriting newer edits.
- Redact secrets from prompts, generated files, logs, diagnostics, and preview responses.
- Cap prompt size, selected-file size, patch size, command duration, log volume, and artifact size.
- Log who requested, approved, executed, and published each change.
Performance, reliability, and cost
Keep interaction responsive
- Stream status and model deltas while buffering the patch for validation.
- Send only relevant files, symbols, and diagnostics; summarize unchanged files.
- Reuse a warm persistent sandbox for repair loops when the security model permits it.
- Install dependencies from a cache, but key the cache by lockfile and runtime version.
- Debounce prompts in the editor and cancel superseded generations.
- Run lint, type checks, and tests in parallel when they do not mutate shared state.
Make failures recoverable
- Use idempotent job IDs and retry network failures with bounded backoff.
- Persist the approved patch and sandbox session identifier before execution.
- Take a project snapshot before applying a destructive patch.
- Return structured error codes instead of only terminal text.
- On timeout, preserve logs and offer “retry,” “open diff,” and “reset sandbox” actions.
Budget explicitly
Track model tokens, sandbox wall time, CPU and memory, dependency downloads, storage, and preview traffic per project. Managed Code Interpreter pricing reported in May 2025 was $0.03 per container; treat that figure as historical and verify current pricing before budgeting. Model and tool availability change, so date any price shown in your product documentation.
Common errors and fixes
| Error | Likely cause | Fix |
|---|---|---|
| Model returns prose instead of a patch | Loose output instructions or missing schema validation | Use strict structured output, parse it server-side, and reject anything that fails. |
| Patch writes outside the project | Absolute path or traversal sequence | Normalize paths, reject .., and verify the resolved path stays under the project root. |
| Changes disappear after another user edit | Stale file version | Include hashes or versions in the request and require a rebase or regenerated patch. |
| Preview iframe is blank | Server is not listening, wrong host binding, or blocked frame origin | Bind to 0.0.0.0, wait for the port, verify the exposed URL, and inspect frame headers. |
| Packages can read secrets | Secrets were mounted into the sandbox | Remove them from the environment and use a narrow trusted proxy for required operations. |
| Repair turn cannot find the previous process | New sandbox session or lost runtime mapping | Persist the session ID and explicitly restore runtime state before the follow-up turn. |
| Logs overwhelm the browser | Unbounded process output | Cap bytes, stream chunks, redact secrets, and store the full log server-side. |
Or skip the browser setup
If your feature only needs a clean screenshot of a generated preview or published URL, ScreenshotNeo provides a single GET request that returns PNG, JPEG, WebP, or PDF. Cookie and consent banners are accepted and removed before capture, along with more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and billing result. Its MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf.
See the ScreenshotNeo API documentation for all options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
You can also request full-page or element captures, dark mode, device presets, custom viewport and retina scale, PDF output, custom CSS and JavaScript, click and wait actions, blocked resources, headers, cookies, user agent, timezone, geolocation, resizing, caching, signed links, async webhooks, bulk capture, and usage data. Every plan includes every feature. The free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
FAQ
Should generated code run in the browser?
Only for trusted, strictly limited frontend code. Use a server sandbox when projects can install packages, execute commands, access files, or handle private data.
Should I ask the model for complete files?
Use structured operations for reviewable changes. Whole-file output can be acceptable for tiny generated projects, but it increases overwrite and merge risk as the project grows.
How do I support iterative fixes?
Keep the approved project and sandbox session associated with the conversation. Feed compiler and runtime diagnostics into a follow-up turn and explicitly restore runtime state.
What must be confirmed by the user?
Ask for confirmation before destructive file operations, publishing, purchases, account changes, or transmitting sensitive information.
When should a workspace be persistent?
Use persistence when repeated repairs benefit from retained dependencies and process state. Use ephemeral sessions when isolation and predictable cleanup matter more.


