How to Use a Python Image Generation SDK
Generate images with Python, decode base64 responses, save files safely, edit references, handle options, and troubleshoot production integrations.
Direct answer: install the official OpenAI Python SDK, set OPENAI_API_KEY, call client.images.generate(), base64-decode result.data[0].b64_json, and write the bytes in binary mode. Use client.images.edit() when you have a reference image or mask. Model names and supported options change, so verify the current image guide and API reference before shipping.
This guide covers setup, complete runnable examples, output controls, edits, streaming, file safety, reliability, costs, and troubleshooting.
1. Set up the Python SDK
- Create an API key in the OpenAI dashboard.
- Export it as an environment variable. Keep it out of source files, notebooks committed to Git, and client-side code.
- Install the official package using the command in the OpenAI Python quickstart. Check the current guide instead of copying an old package version.
export OPENAI_API_KEY='your_api_key_here'
python -m pip install openai
The SDK reads OPENAI_API_KEY when you construct OpenAI(). In CI, provide it through your secret store.
2. Generate an image and save it
The response contains base64 JSON image data. Decode it and write the bytes with wb; this preserves binary data and PNG alpha.
import base64
from pathlib import Path
from openai import OpenAI
client = OpenAI()
result = client.images.generate(
model='gpt-image-2',
prompt='A small red fox reading a book in a sunlit library',
)
image_bytes = base64.b64decode(result.data[0].b64_json)
out = Path('fox.png')
out.write_bytes(image_bytes)
print(f'saved {out} ({len(image_bytes)} bytes)')
Run it with python generate.py. Match the filename extension to the requested output format.
Write better prompts
Describe the subject, action, setting, composition, lighting, style, and constraints. Put required text in quotes and inspect the result because generated lettering may need iteration. Store the exact prompt and settings with each asset.
3. Output settings
The image API exposes model-dependent controls such as size, quality, output_format, and background. Verify allowed values in the image generation guide and image API reference.
result = client.images.generate(
model='gpt-image-2',
prompt='An isometric illustration of a Python data pipeline',
size='1024x1024',
quality='high',
output_format='png',
background='transparent'
)
| Setting | Guidance |
|---|---|
| Size | Use the smallest dimensions that meet the display need. |
| Format | PNG supports lossless detail and alpha; JPEG is broadly compatible; WebP is efficient when supported. |
| Quality | Use lower quality for drafts and higher quality for final assets when supported. |
| Background | Request transparency for compositing and confirm that the model and format preserve alpha. |
4. Save files safely in an application
Validate prompts, create the destination directory, decode with validation, and atomically rename a completed temporary file.
import base64
import os
import tempfile
from pathlib import Path
from openai import OpenAI
client = OpenAI()
def generate_to_file(prompt: str, destination: str) -> Path:
if not prompt.strip():
raise ValueError('prompt must not be empty')
target = Path(destination)
target.parent.mkdir(parents=True, exist_ok=True)
response = client.images.generate(model='gpt-image-2', prompt=prompt)
try:
encoded = response.data[0].b64_json
except (AttributeError, IndexError) as exc:
raise RuntimeError('image response did not contain b64_json data') from exc
data = base64.b64decode(encoded, validate=True)
fd, temporary = tempfile.mkstemp(dir=target.parent, suffix=target.suffix)
try:
with os.fdopen(fd, 'wb') as file:
file.write(data)
file.flush()
os.fsync(file.fileno())
Path(temporary).replace(target)
except Exception:
Path(temporary).unlink(missing_ok=True)
raise
return target
print(generate_to_file('A blue ceramic mug on a wooden desk', 'output/mug.png'))
5. Edit a reference image or use a mask
Use client.images.edit() when supplying an existing image. A mask guides a localized edit, but GPT Image masking is guidance rather than a pixel-perfect boundary. Leave a margin around the area to change and review the output.
import base64
from openai import OpenAI
client = OpenAI()
with open('original.png', 'rb') as source:
edited = client.images.edit(
model='gpt-image-2',
image=source,
prompt='Replace the background with a quiet mountain lake; keep the subject unchanged',
# mask=open('mask.png', 'rb'),
)
with open('edited.png', 'wb') as output:
output.write(base64.b64decode(edited.data[0].b64_json))
Use generate for prompt-only creation and edit for source images, variations, or masked changes. Keep originals so edits can be retried.
6. Streaming and batches
The API documents partial-image events and a completion event carrying base64 content. Streaming is useful for progressive previews, but a completed response is simpler for a file-save script. For many prompts, queue jobs with bounded concurrency and retry only transient failures.
7. cURL and Node.js
curl https://api.openai.com/v1/images/generations \
-H 'Authorization: Bearer $OPENAI_API_KEY' \
-H 'Content-Type: application/json' \
-d '{"model":"gpt-image-2","prompt":"A red fox reading in a library"}'
const OpenAI = require('openai');
const client = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });
const result = await client.images.generate({
model: 'gpt-image-2',
prompt: 'A red fox reading in a library'
});
// Decode result.data[0].b64_json and write the bytes to a file.
8. Reliability, performance, and cost
- Retries: retry rate limits and temporary server or network errors with capped exponential backoff. Do not retry malformed requests or policy rejections blindly.
- Timeouts: give the overall job more time than one HTTP attempt and persist job state before calling the API.
- Concurrency: use a queue and bounded workers to smooth bursts and respect account limits.
- Storage: write decoded bytes, not base64 text. Store dimensions, format, prompt, model, and request ID when audits matter.
- Cost: current cost depends on model, size, quality, and pricing. Check the live pricing page before budgeting; use smaller or lower-quality drafts where suitable.
- Privacy: review current data controls for image inputs and outputs. Model ZDR compatibility does not prove that your organization has ZDR enabled.
9. Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
| Authentication error | The key is unset, misspelled, or unavailable to the process. | Export it in the same runtime or configure the secret; never print it. |
ModuleNotFoundError |
The package is installed in another Python environment. | Run python -m pip install openai with the interpreter running the script. |
| Invalid model or option | The catalog changed or the option is unsupported. | Check the live guide and API reference, then remove or replace the option. |
No b64_json |
The response failed or has a different shape. | Inspect the complete error and guard data[0] as shown above. |
| Corrupt or empty file | Base64 was treated as text or a partial file was exposed. | Use b64decode, open with wb, validate length, and rename atomically. |
| 429 or intermittent 5xx | Rate limits, bursts, or transient service failure. | Reduce concurrency and use capped exponential backoff. |
| Mask boundary is inaccurate | The mask guides the edit rather than enforcing exact pixels. | Expand the mask, describe preserved areas, and iterate from the original. |
| Text or layout is wrong | Generative output is probabilistic. | Specify layout and text clearly, generate alternatives, and review final assets. |
10. Production checklist
- Keep keys in a secret manager or environment variable.
- Validate models and options against the current reference.
- Log prompts, settings, and request IDs without sensitive image data.
- Decode with validation and write atomically.
- Distinguish transient errors from invalid requests and policy responses.
- Set concurrency, timeout, storage, and spend limits.
- Handle reference images and masks under your data policy.
- Review assets when exact text, branding, or boundaries matter.
11. Or skip the browser setup
If you need to capture a generated image or web page for a report, test, or agent workflow, ScreenshotNeo provides a website screenshot API. Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the verdict and billing status. Its MCP server lets Claude, Cursor, and other MCP clients use take_screenshot, get_page_info, and capture_pdf.
See the ScreenshotNeo API docs for all options.
curl -G 'https://api.screenshotneo.com/v1/shot' -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get('https://api.screenshotneo.com/v1/shot', params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'}, timeout=90)
open('shot.webp', 'wb').write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Every feature is on every plan, including full-page and element capture, device and retina settings, custom CSS and JavaScript, waits, blocking rules, headers and cookies, geolocation, PDFs, caching, signed links, async webhooks, bulk capture, and a usage API. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
12. FAQ
Can I avoid writing base64 to disk?
Yes. Decode b64_json in memory and write the resulting bytes directly.
Should I use generate or edit?
Use generate for prompt-only images and edit when supplying a reference image or mask.
Do I need streaming to save a file?
No. A completed response is simpler. Streaming is useful for progressive previews.
How do I handle model changes?
Validate the model and option set at deployment, link to the live reference, and treat unsupported values as configuration errors.


