How to Batch-Generate Images with an API
Build an asynchronous image-generation batch, track its job, and map each result back to its prompt. Includes runnable OpenAI examples and practical guidance for retries and limits.

To batch-generate images with an API, create one generation request for each desired image, submit those requests together to an asynchronous batch endpoint, save the returned job ID, and retrieve the output when the provider reports completion. Each prompt is still its own request; a batch is a way to submit and process many requests together, not a synchronous call that turns one prompt into a set. For large, non-urgent runs, batching can reduce cost and simplify submission. For an interactive preview or a result needed now, use ordinary synchronous requests.
This guide uses OpenAI’s Batch API for a runnable example because its documented batch endpoints include image generation and image edits. It also explains the provider-neutral workflow, Google Gemini’s batch option, result reconciliation, selective retries, operational limits, and cost decisions. Check the linked provider documentation before implementation: supported models, limits, prices, and API behavior can change.
1. Decide whether a batch fits your job
Batch processing trades immediate responses for asynchronous completion. OpenAI advertises a 50% discount versus synchronous APIs and a 24-hour completion window. Google’s Gemini Batch API describes a 50% cost reduction and a 24-hour target. These are provider-published terms, not independent speed measurements or a promise that a particular job will finish at an exact time. A provider may complete sooner, but design your workflow to tolerate waiting.
| Workload | Usually a good fit | Why |
|---|---|---|
| Generate a catalog of many images overnight | Batch | Individual results are not needed immediately, and reduced batch pricing may matter. |
| Let a user change a prompt and see a preview | Synchronous request | The application needs to show a result in its interaction loop. |
| Produce a large set with occasional failures | Batch, with per-request reconciliation | You can collect successful outputs and retry only failed inputs. |
| Make a decision based on one generated image | Synchronous request | Waiting for a batch adds operational work without a clear benefit. |
Estimate the full cost from the model and output settings you will actually use. A stated batch discount does not tell you the total bill by itself. Confirm the current model price, whether your chosen image endpoint and features support batches, account-specific quotas, and any other charges before submitting a large job.
2. Design inputs that can be traced
Keep a durable copy of the exact input data and assign each desired image a stable, unique caller-controlled key. Do not rely on output order to match input order. Output records may arrive in a format or order that makes an explicit key essential, and partial failures mean some prompts may have no image result.

For every request, validate these items before submission:
- The endpoint and model are supported by the provider’s current batch guide.
- The request body matches that endpoint’s normal schema, including required prompt and image settings.
- Every custom key is unique and remains associated with the original prompt.
- Prompt length, parameter values, reference-image inputs, and file sizes meet current limits.
- The input file is valid JSONL if the provider requires JSONL: one complete JSON object per line, with no surrounding array.
- Your application has a place to persist the source file, job ID, submitted time, state, output location, and error records.
OpenAI’s current Batch guide documents JSONL with one request per line and lists /v1/images/generations and /v1/images/edits among supported endpoints. Google’s Gemini Batch API accepts inline requests for smaller payloads or a JSON Lines input file for larger collections; each request follows its GenerateContent request structure. The request envelope and result format differ by provider, so do not copy one provider’s JSONL shape into another provider’s API.
3. Build and submit an OpenAI image batch
The following Python example creates a JSONL file for two image-generation requests and submits it using the OpenAI SDK. It demonstrates the lifecycle through batch creation; the next section shows how to retrieve a completed result. Use an installed, current version of the SDK and follow the current OpenAI Batch API guide for SDK method details and any changes. The image request body must use parameters supported by the selected image model and endpoint; confirm those details in the image generation guide.
import json
from pathlib import Path
from openai import OpenAI
client = OpenAI() # Reads OPENAI_API_KEY from the environment.
# Keep this mapping in durable storage in a real application.
prompts = [
{"key": "product-blue-mug-v1", "prompt": "A blue ceramic mug on a pale stone table, soft morning light"},
{"key": "product-red-mug-v1", "prompt": "A red ceramic mug on a pale stone table, soft morning light"},
]
input_path = Path("image_batch.jsonl")
with input_path.open("w", encoding="utf-8") as f:
for item in prompts:
line = {
"custom_id": item["key"],
"method": "POST",
"url": "/v1/images/generations",
"body": {
"model": "gpt-image-1",
"prompt": item["prompt"],
},
}
f.write(json.dumps(line) + "\n")
uploaded = client.files.create(file=input_path.open("rb"), purpose="batch")
batch = client.batches.create(
input_file_id=uploaded.id,
endpoint="/v1/images/generations",
completion_window="24h",
)
print("batch_id:", batch.id)
print("status:", batch.status)
The sample uses one request per prompt and a caller-supplied custom_id. It intentionally keeps output-format choices minimal: add only parameters supported by the chosen model. Store the returned batch.id before your process exits. Submission acceptance means the provider accepted the job; it does not mean the images are ready.
Use the batch API as a state machine
- Prepare: validate every line and persist the exact input file and key-to-prompt mapping.
- Upload and submit: record the upload ID, batch ID, timestamp, endpoint, and model.
- Track: query the provider’s documented batch status mechanism or use a supported completion notification. Handle only documented states.
- Retrieve: when complete, fetch the output and error records from the locations or file IDs the provider returns.
- Reconcile: match each result’s request key to your input record and store the image payload or its durable storage location.
- Retry selectively: create a new request only for inputs that failed for a retryable reason.
Do not build a tight polling loop. Use a reasonable interval, back off between checks, and stop when a documented terminal state is reached. Persist state so a worker restart does not lose the job identifier or submit the same batch again by accident.
4. Retrieve images and map results to prompts
The batch output is not necessarily a folder of files named after your prompts. Treat it as records to inspect. For each record, check whether it represents a successful response, parse the response body according to the image endpoint’s documented schema, decode or download the image data, and associate it with the input’s custom_id. Also read the error output. A completed batch can contain request-level failures; job completion alone does not prove every image was generated.
A simple reconciliation design uses a database table keyed by your own prompt key, with fields for batch ID, provider request key, state, output location, error category, and retry count. Write successful results idempotently: if a worker sees the same completed record again, it should update or confirm the existing record rather than create duplicate work. Keep the raw provider record for diagnosing parsing or schema changes.
For a small batch, a polling script can inspect the job until it reaches a terminal state, then retrieve the output file using the SDK’s current file-content method. The exact method and response shape are SDK-version-specific, so use the current Batch API documentation for those calls. Avoid hard-coding assumptions such as “the first output line belongs to the first prompt” or “all output records contain an image.”
5. Limits, configuration, and provider choice
OpenAI
The OpenAI Batch guide currently documents up to 50,000 requests and a 200 MB input file per batch, along with model-specific queued-token constraints. Those are published limits, not a guarantee that every account can submit a batch at the maximum. Check the guide and your account’s limits immediately before implementation. Image payloads, endpoint support, and available models can have additional requirements. The image-generation guide covers the supported generation and editing paths.
Google Gemini
Gemini’s batch API has its own request limits, input modes, operation tracking, and response behavior. Its documentation describes inline requests for smaller payloads and JSONL input for larger collections, and currently mentions webhook notifications for completed events. Check the current Gemini Batch API guide for the exact setup, and the rate limits guide for account and model constraints. Do not assume that OpenAI’s request cap, 24-hour window, JSONL envelope, or output structure applies to Gemini.
| Decision axis | What to verify |
|---|---|
| Urgency | Can the application wait for asynchronous completion, or does it need the image in the request-response path? |
| Cost | Check current model pricing and batch eligibility. Both providers describe a 50% batch reduction relative to standard or synchronous pricing; verify the applicable terms for your model and account. |
| Scale | Request and file caps, queued-token limits, concurrent jobs, rate limits, and account tier. |
| Features | Exact model, image generation versus editing, reference inputs, output encoding, and supported parameters. |
| Operations | Job-state retrieval, webhook support, output and error formats, retention, and selective retry strategy. |
| Data handling | Confirm current data-residency and retention terms for your account and region. The provider guides cited here do not establish a comparable cross-provider policy. |
6. Reliability, performance, and cost controls
Batching improves submission efficiency and may lower per-request API cost, but it does not make each generation instantaneous. Plan around the provider’s stated completion window, not an assumed throughput rate. The cited documentation does not establish independent latency benchmarks or a guaranteed completion time for a particular workload.

- Make submissions recoverable: persist the input artifact and job ID before handing work to another process.
- Prevent accidental duplicates: track a stable application batch key and check whether it has already been submitted before retrying a submission whose response was lost.
- Separate job and request status: a finished job may have failed records. Count successful, failed, and unaccounted-for inputs before declaring the run complete.
- Retry narrowly: retry transient or otherwise retryable failures only, after checking the provider’s error guidance. Do not resubmit successful generations automatically; that can duplicate work and charges.
- Keep output storage in view: generated images and retained input/output files need a storage and cleanup policy suited to your application. Check provider retention terms rather than assuming a permanent download link.
- Control run size: split work when a batch nears provider limits, when you need smaller recovery units, or when the workflow benefits from staged review.
Batch pricing can reduce generation charges, but your total operating cost may also include storage, downstream image processing, orchestration, and the engineering time needed to track asynchronous jobs. Estimate from expected successful outputs and account for retries. Verify current prices before committing: the 50% figures are provider descriptions and can depend on the applicable model and terms.
7. Troubleshooting common failures
| Symptom | Likely cause | What to do |
|---|---|---|
| Batch creation rejects the input | Malformed JSONL, wrong endpoint envelope, unsupported endpoint, or invalid body field. | Validate one JSON object per line, inspect the first reported line or request, and compare the request shape with the provider’s current batch guide. |
| Upload succeeds but submission fails | Wrong purpose or file reference, batch limit exceeded, endpoint mismatch, or account constraint. | Record the provider error and request ID; check file size, endpoint compatibility, current limits, and account quota before retrying. |
| Job stays queued or in progress | Asynchronous work has not reached a terminal state; capacity and timing vary. | Use documented status checks or notifications, back off between polls, and allow the published processing window. Do not resubmit just because the result is not immediate. |
| Some prompts have no image | Partial request-level errors, output parsing assumptions, or missing reconciliation keys. | Read both output and error records; join by custom key and retry only failures that are safe to retry. |
| Authentication or permission error | Missing, invalid, or insufficiently authorized API credentials. | Check the environment variable and project access; do not print secrets in logs. Follow the provider’s authentication guidance. |
| Quota or rate-limit error | Account or model limits, including queued work, were exceeded. | Check account limits and the provider’s rate-limit documentation, reduce or stage submissions, and retry according to the provider’s guidance. |
| Server error or connection loss during submission | Transient service/network issue; the client may not know whether the job was accepted. | Log the request ID and inspect existing jobs or submission records before retrying. Use an application idempotency strategy to avoid duplicate batches. |
| Image data cannot be decoded | Code assumed a particular response encoding or treated an error record as a successful image. | Validate the response status and schema first, then decode using the endpoint’s documented output format. |
For OpenAI image-generation failures, handle HTTP status codes or SDK exception types, log request IDs, and consult the provider’s error guidance for authentication, quota, rate-limit, and server failures. Keep prompts and credentials out of diagnostic logs unless your data policy explicitly allows them.
8. Or skip the browser setup
For a website screenshot in an image-generation workflow, use ScreenshotNeo, a website screenshot API and MCP server from Yorker Media. A single GET request returns a PNG, JPEG, WebP, or PDF. It is for capturing web pages; it does not generate new artwork from text prompts.
For example, capture a reference page as WebP:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for request options. Cookie banners are accepted like a visitor and 60+ known consent platforms, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and whether the request was billed. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. Sign up for the free plan.
Frequently asked questions
Does one batch prompt create multiple images?
Not by definition. A batch groups individual API requests for asynchronous processing. To request several outputs, check the selected model’s supported parameters or create multiple request records.
Can I use batch generation for an interactive image editor?
Usually, a synchronous request is a better fit when the user is waiting for a preview or making rapid prompt changes. Batch is designed for work that can wait.
Should I retry the whole batch if one image fails?
No. Reconcile per-request results and retry only appropriate failed inputs. Repeating successful requests can create duplicate work and charges.
Are the 50% discounts guaranteed for every image model?
They are provider-published batch pricing claims. Check current model pricing, eligibility, and account terms before calculating a budget.
Can I compare OpenAI and Gemini limits directly?
Compare the current documentation for the models and accounts you will use. Their limits, input formats, state handling, and quotas are provider-specific.


