ScreenshotNeo

BlogHow-to

How to Build a Blog from Images with AI and React

Build an image-to-blog workflow with React, a server-side vision API, human review, and accessible publishing.

By the ScreenshotNeo team4 October 202611 min read

Build the workflow as five parts: a React interface for image selection and review, a server endpoint that calls an AI vision service, storage for the image, an editable draft, and an explicit publish action. The AI can describe visible content and propose text; it cannot reliably infer missing context such as who is pictured, where a photo was taken, or what an event means. Keep the result as a draft until a person reviews it.

This guide uses React in the browser and a Node.js/Express server for the AI call. The server example is intentionally a small integration pattern: choose a currently supported vision model and confirm its input limits and pricing in the provider documentation before deploying. OpenAI’s guide distinguishes image analysis from image generation and editing; this workflow analyzes an uploaded image to draft text. OpenAI images and vision guide.

1. Choose the content model

Keep media, generated text, editorial changes, and publication state distinct. A practical post record might include:

{
  "id": "stable-record-id",
  "slug": "coastal-walk",
  "imageUrl": "https://media.example.com/coastal-walk.jpg",
  "imageAlt": "A narrow path beside a rocky coastline",
  "title": "A walk along the coast",
  "body": "The edited article text...",
  "generationStatus": "review",
  "createdAt": "2026-10-04T12:00:00.000Z",
  "publishedAt": null
}

This is an implementation design, not a required schema. Use a stable ID even if the slug changes. Store the original media reference separately from any transformed display image, and retain editorially approved title, body, alt text, and publication timestamps. Avoid treating model output as a source of truth for identity, location, dates, or other facts that the uploader has not supplied.

2. Build the React upload and review interface

The UI should show the selected image, generation progress, errors, and an editable draft. This component posts a file plus optional user context to the application server. It expects the server endpoint shown in the next section.

import { useState } from "react";

export default function ImagePostComposer() {
  const [file, setFile] = useState(null);
  const [preview, setPreview] = useState("");
  const [context, setContext] = useState("");
  const [draft, setDraft] = useState(null);
  const [status, setStatus] = useState("idle");
  const [error, setError] = useState("");

  function chooseFile(event) {
    const nextFile = event.target.files?.[0];
    setDraft(null);
    setError("");
    setFile(nextFile ?? null);
    setPreview(nextFile ? URL.createObjectURL(nextFile) : "");
  }

  async function generate(event) {
    event.preventDefault();
    if (!file) return;
    setStatus("generating");
    setError("");

    const form = new FormData();
    form.append("image", file);
    form.append("context", context);

    try {
      const response = await fetch("/api/draft-from-image", {
        method: "POST",
        body: form
      });
      const result = await response.json();
      if (!response.ok) throw new Error(result.error || `Request failed: ${response.status}`);
      setDraft(result);
      setStatus("review");
    } catch (err) {
      setError(err instanceof Error ? err.message : "Could not generate a draft.");
      setStatus("error");
    }
  }

  async function publish(event) {
    event.preventDefault();
    if (!draft) return;
    setStatus("publishing");
    setError("");
    try {
      const response = await fetch("/api/posts", {
        method: "POST",
        headers: { "Content-Type": "application/json" },
        body: JSON.stringify({ ...draft, imageAlt: draft.imageAlt })
      });
      const result = await response.json();
      if (!response.ok) throw new Error(result.error || `Publish failed: ${response.status}`);
      setStatus("published");
    } catch (err) {
      setError(err instanceof Error ? err.message : "Could not publish the post.");
      setStatus("review");
    }
  }

  return (
    <form onSubmit={draft ? publish : generate}>
      <label>Choose an image
        <input type="file" accept="image/*" onChange={chooseFile} required />
      </label>
      {preview && <img src={preview} alt="Preview of the selected image" />}
      <label>Context the image cannot provide
        <textarea value={context} onChange={e => setContext(e.target.value)}
          placeholder="Add names, location, event, or the point you want the post to make" />
      </label>
      {draft && (
        <section aria-label="Review generated draft">
          <label>Title
            <input value={draft.title} onChange={e => setDraft({ ...draft, title: e.target.value })} />
          </label>
          <label>Image alternative text
            <input value={draft.imageAlt} onChange={e => setDraft({ ...draft, imageAlt: e.target.value })} />
          </label>
          <label>Draft body
            <textarea rows={12} value={draft.body} onChange={e => setDraft({ ...draft, body: e.target.value })} />
          </label>
        </section>
      )}
      {error && <p role="alert">{error}</p>}
      <button disabled={!file || status === "generating" || status === "publishing"}>
        {status === "generating" ? "Generating…" : draft ? "Publish reviewed post" : "Create draft"}
      </button>
      {status === "published" && <p>Post published.</p>}
    </form>
  );
}

In a real component, revoke each object URL with URL.revokeObjectURL when the preview is replaced or the component unmounts. Add a separate regenerate action if users need one; publishing should remain a deliberate action after edits. The example preview alt text describes its interface role. For the published article, write alt text that conveys the image’s useful content, or use an empty alt value for purely decorative imagery.

3. Call the vision service from your server

Keep the private provider credential on the server. The browser uploads to your application; your server validates the request, supplies the image and context to the vision API, and returns editable draft fields. The following Express example accepts a multipart file in memory and illustrates the boundary. The provider request is shown as pseudocode because exact SDK methods, model names, limits, and response shapes can change; use the current official API reference when wiring it up.

import express from "express";
import multer from "multer";

const app = express();
const upload = multer({
  storage: multer.memoryStorage(),
  limits: { fileSize: 8 * 1024 * 1024, files: 1 }
});

app.post("/api/draft-from-image", upload.single("image"), async (req, res) => {
  if (!req.file) return res.status(400).json({ error: "Choose an image file." });
  if (!req.file.mimetype.startsWith("image/")) {
    return res.status(415).json({ error: "The uploaded file must be an image." });
  }

  try {
    const context = String(req.body.context || "").slice(0, 4000);
    // Convert req.file.buffer to an input format supported by your selected vision API.
    // Call the provider here using a server-only credential and a current model.
    // Prompt it to return a JSON object with title, imageAlt, and body.
    const generated = await createVisionDraft({
      imageBytes: req.file.buffer,
      mimeType: req.file.mimetype,
      context,
      instruction: "Describe only visible details. Do not guess identities, location, dates, or causes. Return a concise title, useful alt text, and a clearly labeled draft that can be edited."
    });

    return res.json({
      title: generated.title,
      imageAlt: generated.imageAlt,
      body: generated.body
    });
  } catch (error) {
    console.error("Draft generation failed", error);
    return res.status(502).json({ error: "The draft service could not process this image. Try again or write the draft manually." });
  }
});

app.listen(3000);

createVisionDraft is an adapter you implement using your chosen provider’s current SDK or HTTP API; it is deliberately not presented as a runnable vendor call. For a runnable integration, consult the current OpenAI image input guide for supported image inputs and endpoint choices, then implement this adapter with the selected API and model. Do not put an API key in a React bundle. Add authentication and authorization to both endpoints, validate file size and actual image content, and define retention and deletion behavior for uploaded media. These are application responsibilities; the cited guide does not establish a complete security design.

4. Store the image and draft

For a short-lived prototype, a local file may be enough. A deployed app generally needs durable media storage or a media delivery service. The application can upload the file, save the resulting stable URL with the post draft, and optionally create resized derivatives for delivery. Cloudinary’s tutorial is a directly relevant example of an upload and captioning flow using React and Express: Create a blog from an image using AI and React. Compare current storage, transformation, privacy, and pricing terms before choosing a provider.

  • Keep an original or a clearly designated canonical asset, and reference it from the post record.
  • Restrict accepted formats and size; reject malformed files even if their filename ends in an image extension.
  • Do not expose private storage credentials in browser code. Use a server upload or a narrowly scoped upload mechanism.
  • Make retries idempotent where possible so a repeated publish request does not create duplicate posts.
  • Keep generated copy as draft data until an editor approves it.

5. Review, then publish an accessible post

Before publishing, check that the title and copy match the intended audience, remove unsupported claims, and correct names, places, dates, and context using information from the author. Review the alternative text independently: it should describe the image’s relevant content, not repeat the surrounding article. Use alt="" when an image is decorative and adds no information. React documents image alternative text, dimensions, and lazy loading in its image element reference.

function BlogImage({ src, alt, width, height, eager = false }) {
  return (
    <img
      src={src}
      alt={alt}
      width={width}
      height={height}
      loading={eager ? "eager" : "lazy"}
      decoding="async"
    />
  );
}

Supply intrinsic width and height when known so the browser can reserve the image’s space and reduce layout shift. Lazy-load below-the-fold images; keep the lead image eager when it is needed immediately. For post content, store structured or sanitized content rather than injecting untrusted model-generated HTML directly into the page.

6. Plain React or Next.js?

Plain React is enough when your app already handles routing, storage, and image delivery. Next.js is an optional React framework with documented image optimization and metadata conventions. Choose based on the requirements of the application, not because AI drafting itself requires a framework.

Decision Plain React Next.js
Image rendering Use the browser image element or your chosen delivery layer. next/image supports responsive image handling and optimization; remote sources need explicitly permitted patterns.
Image dimensions Pass width and height when known. Provide dimensions or use fill with a layout that preserves aspect ratio.
Sharing previews Implement metadata with your routing/server setup. Route metadata and dynamic Open Graph images are documented options.

See the official Next.js image guide for remote patterns, sizing, and loading, and Next.js metadata and Open Graph image guide for App Router metadata conventions. For remote images, permit only the hosts and paths the application needs.

7. Add post-specific sharing metadata

Sharing metadata is optional. In a Next.js App Router project, define metadata for each post route and, if useful, generate a post-specific Open Graph image. Keep the title, description, and preview image aligned with the reviewed published content. A plain React application can implement equivalent metadata through its server or routing layer; the Next.js convention is framework-specific.

8. Performance, reliability, and cost

Performance

  • Send a suitable image size to the vision service. Resizing before analysis can reduce upload time and payload, but preserve detail that matters to the task.
  • Use object storage or a media delivery layer rather than sending every reader’s image through the drafting endpoint.
  • Keep draft generation asynchronous from the UI’s perspective: show progress, retain the selected file on recoverable errors, and allow a manual fallback.
  • Use responsive delivery and lazy loading for non-critical published images; reserve dimensions to reduce layout movement.

Reliability

  • Handle unsupported formats, oversized files, network interruptions, provider timeouts, and malformed model output as separate cases.
  • Validate the returned fields and show a recoverable error if the model response cannot be parsed.
  • Use request IDs or idempotency keys for retries where supported by your application design.
  • Keep an editorial path that works without AI, so a provider outage does not prevent a post from being written or published manually.
  • Do not assume generated alt text is accessible or accurate without review.

Cost

There is no universal per-post cost to quote here: it depends on the chosen model, image input, request volume, storage, and delivery options. Check the provider’s current pricing and limits before launch. Reduce unnecessary regeneration, enforce upload limits, and measure usage at the server boundary. Also account for media storage and image delivery separately from AI inference.

9. Troubleshooting

Symptom Likely cause Fix
Request returns 400 or says no image Multipart field name differs from the server’s expected image, or no file was selected. Append the file with the exact field name and check req.file.
Request returns 413 or upload is rejected File exceeds the application or proxy size limit. Set an intentional limit at each layer and show it in the interface; resize when appropriate.
Request returns 415 The content is not an accepted image format or its MIME type is unsupported. Validate the actual file, allow only formats supported by your storage and vision provider, and explain accepted types.
Provider rejects the image Unsupported encoding, dimensions, or payload format; the app may also have sent bytes in the wrong representation. Check the current provider input guide, normalize supported formats if needed, and pass the correct MIME type.
Browser reports a CORS error The browser is calling a different origin without matching CORS configuration. Prefer a same-origin application endpoint or configure the server’s allowed origins deliberately.
Draft is empty or malformed The response shape was assumed incorrectly or the model did not return the requested structure. Parse and validate server-side, handle invalid output, and allow retry or manual authoring.
Draft invents a name, place, or event The image does not supply that context, or the prompt did not constrain claims sufficiently. Remove the claim, provide verified context, and require human review before publication.
Published image is broken The saved URL is temporary, private, or inaccessible from the public site. Store a durable reference, verify public delivery permissions, and avoid expiring preview URLs in published records.
Page shifts while images load Image dimensions are absent or layout sizing is not reserved. Set width and height or use an aspect-ratio container with the chosen image component.

10. Or skip the browser setup

If your next step is capturing a web page as an image for a post, ScreenshotNeo is a website screenshot API and MCP server. Its one-call API returns an image or PDF; see the ScreenshotNeo API docs.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo removes cookie banners, popups, and chat widgets before capture; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000. Sign up for free.

FAQ

Can AI reliably write a finished post from pixels alone?

No. It can propose a description or draft, but authors must supply and verify context that is not visible. The OpenAI image guide advises: “Account for the limitations of the model when using answers.”

Does this workflow require Next.js?

No. React can provide the interface in many application setups. Next.js adds framework-specific image and metadata options if those solve a need in your project.

Should the generated description become the image alt text automatically?

Treat it as a suggestion. Edit it for accuracy and purpose, and use empty alt text for decorative images.