Image Generation SDKs for Node.js, Python, PHP, and Ruby
Compare image-generation SDKs for Node.js, Python, PHP, and Ruby, with provider maps, runnable examples, model limits, and integration guidance.

Short answer: there is no universal image-generation SDK that offers identical features in Node.js, Python, PHP, and Ruby. Runway documents image-generation SDKs for Node.js and Python. Cloudinary provides broader media SDK quick starts for all four languages. Amazon Bedrock exposes image models through general AWS SDKs, where your code sends a model-specific JSON body. OpenAI separates a direct Image API for one-off generations and edits from a Responses API workflow for conversational, multi-step image work.
Choose the provider from the workflow and model you need, then verify the exact language package, runtime, model schema, output format, and account or regional availability before coding. Package availability alone does not prove that every image operation is supported.
Provider and language coverage
| Provider or route | Node.js | Python | PHP | Ruby | What the SDK actually represents |
|---|---|---|---|---|---|
| Runway | Documented | Documented | Not listed in reviewed SDK documentation | Not listed in reviewed SDK documentation | Dedicated API clients for Runway image generation |
| Cloudinary | Quick start | Quick start | Quick start | Ruby/Rails quick start | Broad programmable-media SDKs, including image operations; not a universal generative-model client |
| Amazon Bedrock | AWS SDK; JavaScript Nova Canvas example | Image-generation examples | Stability Image Core example | AWS SDK exists; no Ruby image example identified in the reviewed material | General cloud invocation with model-specific request and response schemas |
| OpenAI image APIs | Check current official client support for your language | Image API for direct generation/editing; Responses API for multi-turn and multi-step flows | |||
Runway’s reviewed documentation maps POST /v1/text_to_image to client.textToImage.create in Node.js and client.text_to_image.create in Python. It states Node.js 18+ and Python 3.8+ compatibility. Cloudinary’s SDKs cover media management and transformations as well as image delivery. AWS lists SDKs for PHP, Python, and Ruby, but the image examples reviewed are model-specific rather than a single Bedrock image abstraction. The AWS documentation explicitly says, “The request body is model-specific.”
How to select an SDK
1. Match the workflow
- One prompt, one image: use a provider’s direct image endpoint when available.
- Edit an existing image: confirm image-input and mask support, plus the accepted MIME types.
- Multi-turn creative work: OpenAI documents the Responses API for image generation inside conversations and iterative editing.
- Media storage and transformations: Cloudinary’s SDKs are designed for programmable media workflows, not as a generic front end to every generative model.
- Many model vendors in one cloud account: Bedrock can be useful, but your application must construct each model’s native JSON payload.
2. Check language and package status
Use the official package for your runtime and inspect its release and compatibility notes. A vendor may document an SDK for a language while documenting the image method only for another. Treat “SDK exists” and “image generation is documented” as separate checks.

3. Compare output controls
Confirm dimensions, quality, format, compression, transparency or background behavior, and whether the response is bytes, base64, or a URL. OpenAI documents controls for dimensions, quality, format, compression, and background. Bedrock examples return base64 image data that the application decodes and stores.
4. Verify model capabilities
Before implementation, check input and output modalities, model identifier, account access, region, request limits, and streaming behavior. Bedrock’s model metadata and documentation are authoritative for these details; a generic AWS client does not make every model interchangeable.
Node.js: dedicated and general SDK patterns
Runway SDK pattern
Runway’s documented Node.js client includes TypeScript bindings and requires Node.js 18 or newer. Install the package named in the current Runway SDK documentation, configure the API key using the environment mechanism recommended there, and call the text-to-image method:
import RunwayML from '@runwayml/sdk';
const client = new RunwayML({
apiKey: process.env.RUNWAYML_API_SECRET
});
const result = await client.textToImage.create({
model: 'gen4_image',
promptText: 'A small glass greenhouse beside a misty alpine lake at sunrise',
ratio: '1024:1024'
});
console.log(result);
Method names, model IDs, and accepted ratio values can change, so copy the current options from Runway’s official SDK page before shipping.
Generic HTTP call from Node.js
When a provider has no supported package for your language, use its HTTPS API directly. Keep the key in an environment variable, set a timeout, and save the response according to the provider’s documented shape:
const controller = new AbortController();
const timeout = setTimeout(() => controller.abort(), 90_000);
try {
const response = await fetch('https://provider.example/v1/images', {
method: 'POST',
headers: {
'Authorization': `Bearer ${process.env.IMAGE_API_KEY}`,
'Content-Type': 'application/json'
},
body: JSON.stringify({
prompt: 'A small glass greenhouse beside a misty alpine lake at sunrise',
size: '1024x1024',
output_format: 'png'
}),
signal: controller.signal
});
if (!response.ok) throw new Error(`HTTP ${response.status}: ${await response.text()}`);
const data = await response.json();
console.log(data);
} finally {
clearTimeout(timeout);
}
Python: SDK calls and base64 output
Runway Python pattern
The reviewed Runway documentation maps the endpoint to client.text_to_image.create and states Python 3.8+ compatibility:
import os
from runwayml import RunwayML
client = RunwayML(api_key=os.environ["RUNWAYML_API_SECRET"])
result = client.text_to_image.create(
model="gen4_image",
prompt_text="A small glass greenhouse beside a misty alpine lake at sunrise",
ratio="1024:1024",
)
print(result)
Bedrock-style response handling
Bedrock image examples construct the model-native request, invoke the model, decode base64 data, and write a file. The exact model ID and JSON fields must come from the selected model’s documentation:
import base64
import json
import boto3
bedrock = boto3.client("bedrock-runtime", region_name="us-east-1")
body = {
"taskType": "TEXT_IMAGE",
"textToImageParams": {
"text": "A small glass greenhouse beside a misty alpine lake at sunrise"
},
"imageGenerationConfig": {
"width": 1024,
"height": 1024,
"numberOfImages": 1
}
}
response = bedrock.invoke_model(
modelId="MODEL_ID_FROM_CURRENT_DOCUMENTATION",
body=json.dumps(body),
contentType="application/json",
accept="application/json",
)
payload = json.loads(response["body"].read())
image_bytes = base64.b64decode(payload["images"][0])
with open("generated.png", "wb") as image_file:
image_file.write(image_bytes)
PHP: Bedrock and direct HTTP integration
A reviewed Bedrock PHP example uses the AWS SDK’s invokeModel operation, sends a JSON body for Stability AI Stable Image Core, and reads the first item in the returned images array. The model’s current schema is required:
<?php
require 'vendor/autoload.php';
use Aws\BedrockRuntime\BedrockRuntimeClient;
$client = new BedrockRuntimeClient([
'region' => getenv('AWS_REGION') ?: 'us-east-1',
'version' => 'latest',
]);
$body = json_encode([
'prompt' => 'A small glass greenhouse beside a misty alpine lake at sunrise',
'output_format' => 'png'
], JSON_THROW_ON_ERROR);
$result = $client->invokeModel([
'modelId' => 'MODEL_ID_FROM_CURRENT_DOCUMENTATION',
'body' => $body,
'contentType' => 'application/json',
'accept' => 'application/json',
]);
$payload = json_decode((string) $result['body'], true, 512, JSON_THROW_ON_ERROR);
file_put_contents('generated.png', base64_decode($payload['images'][0], true));
For a provider without a PHP SDK, use Guzzle or PHP’s cURL extension, preserve the documented headers, and validate the response before decoding it.
Ruby: AWS SDK or raw HTTP
AWS publishes a Ruby SDK, but the reviewed material does not establish a Ruby-specific image-generation example. You can still use the general Bedrock runtime client if the chosen model is available in your account. Construct the model-specific body exactly as documented:
require 'json'
require 'aws-sdk-bedrockruntime'
client = Aws::BedrockRuntime::Client.new(region: ENV.fetch('AWS_REGION', 'us-east-1'))
body = {
taskType: 'TEXT_IMAGE',
textToImageParams: { text: 'A small glass greenhouse beside a misty alpine lake at sunrise' },
imageGenerationConfig: { width: 1024, height: 1024, numberOfImages: 1 }
}.to_json
response = client.invoke_model(
model_id: ENV.fetch('BEDROCK_MODEL_ID'),
body: body,
content_type: 'application/json',
accept: 'application/json'
)
payload = JSON.parse(response.body.read)
File.binwrite('generated.png', [payload.fetch('images').first].pack('m0'))
OpenAI image workflows
OpenAI documents the Image API as the straightforward choice for generating or editing a single image from one prompt. Its Responses API is intended for image generation in conversations, image inputs in context, and iterative edits. Decide which workflow you need before choosing a client method, then verify current SDK support for your language.
from openai import OpenAI
client = OpenAI()
result = client.images.generate(
model="gpt-image-1",
prompt="A small glass greenhouse beside a misty alpine lake at sunrise",
size="1024x1024",
)
print(result)
Keep output handling explicit: store returned bytes or decode the documented base64 field, record the model and request ID, and avoid assuming that a URL is permanent.
cURL smoke test
Use a direct HTTP request to separate authentication or network problems from SDK problems. Replace the endpoint and JSON fields with those from the provider you selected:
curl -sS https://provider.example/v1/images \
-H "Authorization: Bearer $IMAGE_API_KEY" \
-H "Content-Type: application/json" \
-d '{"prompt":"A small glass greenhouse beside a misty alpine lake at sunrise","size":"1024x1024"}'
Output, retries, and production practices
- Validate before decoding: check HTTP status, content type, required JSON keys, and base64 validity.
- Use bounded retries: retry transient 408, 429, and 5xx responses with exponential backoff and jitter. Do not blindly retry malformed requests or permission errors.
- Make jobs idempotent: attach your own request ID and persist the prompt, model, dimensions, and provider response metadata.
- Control concurrency: respect provider quotas; a worker queue prevents bursts from turning into repeated 429 responses.
- Measure the real workload: compare current model pricing, quota behavior, output size, and latency for your prompts. The reviewed sources do not provide a comparable cross-provider benchmark.
- Protect secrets: use environment variables or a secret manager, never browser-exposed keys.
- Store artifacts deliberately: choose a retention period, content type, and deterministic naming scheme; scan or moderate generated content where your application requires it.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| 401 or 403 | Missing, wrong, expired, or unauthorized key | Check the environment variable, account permissions, selected region, and model access. |
| 400 with “invalid body” | Payload belongs to a different model | Read the selected model’s native schema; Bedrock request bodies are model-specific. |
| 404 model or endpoint | Wrong model ID, API version, or region | Copy the current identifier from official documentation and verify regional availability. |
| 429 | Rate or concurrency limit | Queue work, reduce parallel requests, honor retry headers, and use exponential backoff. |
| Response has no image | Code assumed a URL or wrong response field | Log a redacted response shape and follow the provider’s documented bytes, URL, or base64 format. |
| Image file is corrupt | Base64 was not decoded correctly or an error body was saved | Check status and content type before decoding; use strict base64 validation. |
| SDK import or runtime error | Unsupported runtime or package version | Use the documented Node.js/Python minimum, reinstall the official package, and check release notes. |
| Long or inconsistent latency | Queueing, large dimensions, or provider load | Set a timeout, measure each phase, cap dimensions, and move asynchronous work to a queue. |
Or skip the browser setup
If your application needs screenshots of generated-image galleries, documentation, dashboards, or landing pages, ScreenshotNeo provides a single website screenshot API request. Its clean-shot step accepts cookie and consent banners, then removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and responses identify the result with X-Page-Verdict and X-Billed headers.

See the ScreenshotNeo API documentation for all options, including full-page capture with lazy images, CSS-selector element capture, dark mode, device presets, retina scale, PDF output, custom CSS and JavaScript, clicks, waits, blocked resources, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, caching, signed links, asynchronous jobs, bulk capture, usage reporting, and the OpenAPI specification.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also provides an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
FAQ
Is there one SDK that works equally in all four languages?
No. Cloudinary covers all four with broad media SDKs, while Runway documents dedicated image SDKs for Node.js and Python. Bedrock uses general cloud SDKs plus model-specific payloads.
Should I choose Python or Node.js for image generation?
Choose the language already used by your service and the provider with the clearest official client for your chosen model. The language alone does not determine model capability or cost.
Does an AWS SDK guarantee that a Bedrock model will work?
No. Confirm model access, region, modalities, request schema, output format, and quotas for the exact model.
When should I use a raw HTTP request?
Use it for a smoke test, for a language without a supported client, or when the provider’s SDK lags behind a newly released endpoint. Keep the same authentication, timeout, validation, and retry discipline.
Are SDK versions and model IDs stable?
No. Recheck official documentation immediately before implementation because packages, runtime requirements, model identifiers, pricing, quotas, and regional availability change.


