How to Generate Images with AI at Render Time
Learn how to generate AI images when a page needs them, stream previews, handle failures, control costs, and ship a reliable runtime workflow.
Short answer: generate the image from your application server when the feature needs it, show a pending state immediately, optionally stream partial images, then store and display the final image with an accessible description. Use the OpenAI Image API for a single prompt-to-image request. Use the Responses API with its image-generation tool when image creation is part of a conversation or a multi-step exchange.
1. The runtime image-generation flow
Runtime generation means the browser does not contain a finished asset at deploy time. A user action or page-specific condition causes your application infrastructure to request an image from an image-generation API.
- Collect a prompt and any page context on the server.
- Choose the Image API for one-shot generation, or the Responses API tool for conversational and multi-step work.
- Return a job or pending response to the browser instead of blocking the entire page.
- Decode the base64 image data, or forward partial images when streaming is enabled.
- Store the final asset when it will be reused, and return a stable URL or protected download route.
- Render the image with useful alternative text and a fallback for failures.
Complex prompts may take up to two minutes to process, so a loading state is part of the design, not an edge case.
2. Choose the right OpenAI API shape
| Use case | Recommended API | Reason |
|---|---|---|
| One prompt creates one image | Image API | Direct request and image result with minimal orchestration. |
| Chat, iterative editing, or several turns | Responses API with image-generation tool | The tool can create or edit images in the conversation context. |
| User needs visual feedback during a long generation | Either API with streaming | Partial images can arrive before the final result. |
Keep API keys on your server. The browser should call your own endpoint; never embed an OpenAI key in client-side JavaScript.
3. A complete server endpoint (Node.js)
The following Express-style handler illustrates the request boundary. Adapt the SDK call to the current OpenAI SDK version and model documented in the official image guide.
import express from "express";
import OpenAI from "openai";
const app = express();
app.use(express.json({ limit: "32kb" }));
const openai = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });
app.post("/api/runtime-image", async (req, res) => {
const prompt = String(req.body?.prompt || "").trim();
if (!prompt || prompt.length > 4000) {
return res.status(400).json({ error: "prompt is required and must be 4,000 characters or fewer" });
}
try {
const result = await openai.images.generate({
model: "gpt-image-1",
prompt,
size: "1024x1024",
quality: "high",
output_format: "webp"
});
const item = result.data?.[0];
if (!item?.b64_json) throw new Error("The API returned no image data");
res.json({
mime_type: "image/webp",
data_url: `data:image/webp;base64,${item.b64_json}`
});
} catch (error) {
console.error("image_generation_failed", {
name: error?.name,
status: error?.status,
request_id: error?.request_id,
message: error?.message
});
res.status(error?.status === 429 ? 429 : 502).json({
error: "Image generation is temporarily unavailable"
});
}
});
app.listen(3000, () => console.log("Listening on http://localhost:3000"));
For production, write the decoded bytes to object storage and return a short-lived or application-authorized URL. A data URL is convenient for a demo but increases response size and keeps the image in memory.
4. Call the Image API directly
cURL
curl https://api.openai.com/v1/images/generations \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-image-1",
"prompt": "A technical illustration of a request flowing through a server into a finished image",
"size": "1024x1024",
"quality": "high",
"output_format": "webp"
}'
Python
import base64
import os
from openai import OpenAI
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
result = client.images.generate(
model="gpt-image-1",
prompt="A technical illustration of a request flowing through a server into a finished image",
size="1024x1024",
quality="high",
output_format="webp",
)
image_bytes = base64.b64decode(result.data[0].b64_json)
with open("runtime-image.webp", "wb") as file:
file.write(image_bytes)
Node.js
import OpenAI from "openai";
import { writeFile } from "node:fs/promises";
const client = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });
const result = await client.images.generate({
model: "gpt-image-1",
prompt: "A technical illustration of a request flowing through a server into a finished image",
size: "1024x1024",
quality: "high",
output_format: "webp"
});
await writeFile("runtime-image.webp", Buffer.from(result.data[0].b64_json, "base64"));
Image results are base64 encoded. PNG is the default; JPEG and WebP are available. JPEG can be faster than PNG, while WebP often reduces delivery size. Choose a format that fits transparency, image content, and downstream browser support.
5. Connect the page to your endpoint
<button id="generate">Generate</button>
<p id="status" role="status"></p>
<img id="result" alt="" hidden>
<script>
const button = document.querySelector("#generate");
const status = document.querySelector("#status");
const image = document.querySelector("#result");
button.addEventListener("click", async () => {
button.disabled = true;
image.hidden = true;
status.textContent = "Generating image…";
try {
const response = await fetch("/api/runtime-image", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
prompt: "A calm editorial illustration of a developer designing a runtime image workflow"
})
});
const body = await response.json();
if (!response.ok) throw new Error(body.error || "Request failed");
image.src = body.data_url;
image.alt = "Generated editorial illustration of a runtime image workflow";
image.hidden = false;
status.textContent = "Image ready";
} catch (error) {
status.textContent = "The image could not be generated. Try again.";
} finally {
button.disabled = false;
}
});
</script>
Do not make a critical page shell wait for generation unless that delay is acceptable. Render the rest of the page first and reserve a predictable area for the image to prevent layout shifts.
6. Stream partial images while generation continues
Both API shapes support streaming. You can request zero to three partial images. Fewer partials may arrive when the final image completes quickly, so the client must treat partials as optional previews and always handle the final result.
The streaming reference describes server-sent events, including image_generation.partial_image events containing base64 data, a partial index, output format, quality, and size. A typical flow is:
- Open an authenticated server-side stream.
- Forward only the fields your browser needs over your own SSE endpoint.
- Replace the preview when a newer partial arrives.
- Swap in the final image and close the stream.
- Show a retry control if the stream ends without a final image.
Keep partial images in memory or a temporary cache. Persist the final asset only after completion. The Image API reference and Responses API reference contain the current event and parameter names.
7. Output controls and design decisions
| Control | What it changes | Practical guidance |
|---|---|---|
| Size | Aspect ratio, pixel count, token use, and delivery weight. | Use the smallest size that fits the component; current GPT Image 2.5 guidance lists 1024×1024, 1536×1024, and 1024×1536 as recommended dimensions. |
| Quality | Visual detail, latency, and token consumption. | Use lower quality for previews and higher quality for final exports when the product permits. |
| Format | Encoding size, transparency support, and decode behavior. | Use PNG for lossless or transparent output; JPEG or WebP when delivery size and speed matter. |
| Background | Opaque versus transparent result where supported. | Request transparency only when the consuming component needs it. |
| Prompt context | Composition, style, subject, and constraints. | Build prompts from validated fields rather than concatenating untrusted HTML or arbitrary system instructions. |
Custom dimensions have constraints on edge multiples, aspect ratio, and total pixels. Check the current guide before relying on dimensions outside the recommended sizes.
8. Caching, storage, and duplicate requests
- Create a normalized cache key from the model, prompt, input images, size, quality, format, and background settings.
- Deduplicate concurrent identical requests so a double click does not create two paid generations.
- Store final bytes with a content type and immutable name, then serve through a CDN or your own authorized route.
- Keep prompts and request IDs with the asset when you need reproducibility or support diagnostics.
- Set an expiration policy for disposable user-generated images and a deletion path for user data.
Caching is an application choice; the APIs do not prescribe your object store, CDN, or framework.
9. Latency, reliability, and cost
Latency and eventual cost generally rise with image token usage. Larger dimensions and higher quality usually consume more tokens. The current guide lists GPT Image 2.5 token rates of $8 per million image input tokens, $2 per million cached image input tokens, $30 per million image output tokens, $5 per million text input tokens, and $1.25 per million cached text input tokens. These are token rates, not a fixed price per image; model, prompt, quality, and output size change actual usage. Cached image-input pricing applies to the Responses image-generation tool, not direct Images API requests.
- Show progress immediately and communicate that a complex request can take up to two minutes.
- Use a request timeout longer than your normal generation window, but cap it so abandoned browser tabs do not consume resources forever.
- Retry transient rate-limit and server failures with exponential backoff and jitter.
- Do not blindly retry quota errors, invalid parameters, or user-correctable prompts.
- Log the provider request ID, status code, model, size, quality, and your cache key without logging secrets.
- Measure p50 and p95 generation time, completion rate, retry rate, bytes delivered, and cost per completed asset.
10. Moderation and user-facing failures
Image prompts and generated images are filtered under the service content policy. The image-generation moderation option defaults to auto; low is less restrictive where supported. A blocked request can identify whether input or output moderation caused the block and may include coarse categories.
Keep the user message short and neutral, such as “This image could not be generated.” Put moderation details in developer logs, support workflows, or analytics. If your product needs its own moderation signal for text or image inputs, the separate Moderation API can classify them; it does not replace image-generation policy filtering.
11. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| 401 or 403 response | Missing, invalid, or server-side key configuration. | Load the key from the server environment, verify the project, and never expose it to the browser. |
| 429 response | Rate limit or exhausted quota. | Back off and retry transient limits; check quota before retrying repeatedly. |
| Request times out | Complex prompt, large output, or an overly short client timeout. | Use a pending job or stream, increase the server timeout within reason, and let the user retry. |
| No image in a successful response | Code assumes a result item always exists. | Validate the response shape and log the request ID before returning a generic error. |
| Only a preview appears | The client treats a partial event as final. | Track partial and final events separately and replace the preview with the final image. |
| Images are too expensive | Repeated prompts, large sizes, or high quality on every request. | Deduplicate, cache final assets, lower preview settings, and inspect usage fields. |
| Image is blocked | Input or output moderation policy. | Show a generic failure, record moderation details for support, and let the user revise the request. |
| Broken image URL after refresh | Data URL was held only in page memory. | Persist final bytes and return a stable, authorized asset URL. |
12. Or skip the browser setup
If your runtime feature needs screenshots of a generated page, dashboard, or preview, ScreenshotNeo provides a website screenshot API and MCP server. It accepts a URL and returns PNG, JPEG, WebP, or PDF. Cookie and consent banners, newsletter popups, and chat widgets are removed before capture; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, with verdict and billing information in response headers. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.
See the ScreenshotNeo API documentation for all options.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Plans include 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
13. FAQ
Should generation happen during the initial page load?
Only when the page cannot be useful without the image and the delay is acceptable. Otherwise load the interface first and generate in the background.
Can I guarantee that three partial images arrive?
No. You can request up to three, but the final generation may complete before all requested partials are emitted.
Which format should I store?
Use PNG for lossless or transparent assets. Use JPEG or WebP when smaller delivery payloads matter and transparency is unnecessary.
How should I retry a blocked prompt?
Do not repeat it unchanged. Return a neutral message and ask for a revised prompt; reserve automatic retries for transient rate-limit or server failures.
Is the listed token rate the price of one image?
No. It is a per-million-token rate. Actual cost depends on model, prompt, quality, dimensions, and token usage returned by the API.


