How to Automatically Generate Alt Text for Images
Generate alt text as a draft, then review it against the image’s purpose and page context. This guide covers custom pipelines, WordPress, Shopify, and common edge cases.
Short answer: use an image-captioning tool to draft descriptions, but do not publish every draft automatically. First decide what each image does on its page, then review the candidate description for accuracy, relevance, and whether it duplicates nearby text. Some images need empty alt text, and some need a description of their function rather than their appearance.
For a custom pipeline, send an image to an image-captioning API, save the result as a draft, and route it for review before updating your site. WordPress plugins and Shopify apps can automate parts of that workflow. The right choice depends on your publishing platform, privacy needs, and how you will review and correct the output.
1. Decide what the image needs to communicate
Alt text is a text alternative for an image in its page context. It is not simply a list of visible objects. The same photo might need different alt text in a product listing, a news article, and a decorative banner because each page uses it differently. W3C groups common cases into informative, decorative, functional, text-containing, complex, grouped, and image-map images. W3C Images Tutorial
| Image role | What to do | Example approach |
|---|---|---|
| Informative | Describe the information the page relies on. | For a photo illustrating a story about a flooded road, describe the road and flooding that matter to the story. |
| Decorative | Use an empty alt value: alt="". |
A flourish that adds no information should not create extra screen-reader output. |
| Functional | Describe the action or destination of the control. | A linked magnifying glass might use “Search” if that is its function, rather than “magnifying glass icon.” |
| Text in an image | Include the meaningful words in text, or provide an equivalent nearby. | A button image should expose its label as text or an appropriate alternative. |
| Complex image | Give a concise label and put the detailed equivalent in nearby text. | For a chart, summarize its subject in alt text and provide data or trends in a table or description. |
| Grouped images | Consider the group’s combined meaning and avoid repeating equivalent information. | A set of adjacent images may need one concise explanation in surrounding text rather than repetitive descriptions. |
WCAG 2.1 Success Criterion 1.1.1 requires text alternatives that serve an equivalent purpose for non-text content, with exceptions and special cases including decoration, controls, tests, sensory content, and CAPTCHA. The author has to decide what equivalent purpose means in context; a generated visual caption cannot make that decision reliably on its own. WCAG 2.1 and Section508.gov guidance
2. Choose an automation route
| Route | Fits best when | Check before adopting |
|---|---|---|
| Built-in authoring feature | Editors need occasional descriptions while working in documents or presentations. | Whether processing is local or cloud-based, whether it runs automatically or on request, and how reviewers edit the result. Microsoft describes local generation on Copilot+ PCs and user-requested cloud generation in some Microsoft 365 workflows. Microsoft FAQ |
| WordPress plugin | You need to process media-library backlogs or draft descriptions on upload. | Bulk and upload workflows, page context, external image processing, provider settings, and revision or undo controls. The WordPress AI Alt Text Generator listing describes bulk and upload-time generation and discloses sending images and metadata to an external service. |
| Shopify app | You manage product or blog imagery in a Shopify store. | Whether product information informs descriptions, how bulk updates work, review and rollback options, and current pricing. The Shopify AltText.ai listing describes automatic and bulk workflows; marketplace plan details can change. |
| Image-caption API | You control a custom publishing or asset pipeline. | Supported languages, confidence handling, data retention, failure behavior, service lifecycle, and migration path. The Microsoft Azure Vision Image Analysis 4.0 page says the service is deprecated and scheduled to retire on September 25, 2028; check its current migration guidance before starting a new integration. Microsoft Learn |
These are examples of available routes, not a complete market comparison. Check current product documentation and policies before sending images to an external provider.
3. Build a custom draft-and-review pipeline
A safe custom workflow keeps the generated string separate from the published alt attribute until a person or a clearly defined rule approves it. The example below uses Python and Microsoft’s Azure Vision Image Analysis API to request a caption. The service lifecycle is time-sensitive: Microsoft states that Image Analysis 4.0 is deprecated and scheduled to retire on September 25, 2028, so use this example to understand the workflow and check the linked migration guidance before adopting it for a new production system.
Python example: request a caption and save a review draft
Install the HTTP client with python -m pip install requests. Set VISION_ENDPOINT to the endpoint for your resource and VISION_KEY to its key. Do not commit keys or put them in browser-side code.
import json
import os
from pathlib import Path
import requests
endpoint = os.environ["VISION_ENDPOINT"].rstrip("/")
key = os.environ["VISION_KEY"]
image_path = Path("image.jpg")
url = f"{endpoint}/computervision/imageanalysis:analyze"
params = {"api-version": "2024-02-01", "features": "caption"}
headers = {
"Ocp-Apim-Subscription-Key": key,
"Content-Type": "application/octet-stream",
}
with image_path.open("rb") as image_file:
response = requests.post(
url,
params=params,
headers=headers,
data=image_file,
timeout=60,
)
response.raise_for_status()
result = response.json()
caption = result.get("captionResult", {}).get("text", "").strip()
confidence = result.get("captionResult", {}).get("confidence")
# Save a draft for review; do not silently publish it as alt text.
draft = {
"image": str(image_path),
"candidate_alt": caption,
"confidence": confidence,
"status": "needs_review",
}
Path("alt-draft.json").write_text(json.dumps(draft, indent=2), encoding="utf-8")
print(json.dumps(draft, indent=2))
The endpoint shape and API version above follow Microsoft’s Image Analysis guidance. Confirm the resource endpoint, supported API version, request limits, and migration path in the current API documentation. The code intentionally records confidence as review metadata; it does not treat confidence as proof that the description is correct.
Make the pipeline production-safe
- Inventory before generating. Record each image’s URL or asset ID, page, current alt value, and likely role. Preserve existing reviewed text.
- Exclude cases that should not be captioned. Mark decorative images, repeated logos, and images whose meaning is already conveyed nearby so they can be assigned empty alt or handled in context.
- Send only necessary data. Use a server-side worker, restrict access to credentials, and check what image data and metadata leave your environment.
- Store candidates separately. Keep generated text, provider metadata, status, and the existing published value in separate fields so you can review and undo a change.
- Review in page context. Show the image alongside its heading, adjacent copy, link or control, and intended role. Correct omissions, guesses, and redundant wording.
- Pilot a small batch. Measure how many candidates need correction and what kinds of mistakes recur. Expand only after the review workload and rollback path are clear.
- Monitor provider changes. Recheck API lifecycle, authentication, supported languages, limits, costs, and error formats before relying on a service.
4. Add alt text in WordPress or Shopify
WordPress
- Review the plugin’s current documentation and settings. Confirm whether generation happens on upload, in bulk, or both.
- Check the external-service disclosure and determine which provider receives images and metadata.
- Try a small batch with varied image types, including decorative images, screenshots with text, product photos, and charts.
- Review the draft against each image’s page role, correct it, and verify the published media field.
- Keep a record of changed asset IDs and previous values so incorrect bulk updates can be reversed.
The WordPress plugin listing is one example that describes bulk processing and generation on upload. Plugin features and service terms can change, so confirm the current listing and plugin settings before use.
Shopify
- Check whether the app uses product title, product description, or other catalog context, and whether that context can introduce incorrect assumptions.
- Review permissions, external processing, bulk behavior, current pricing, and how generated changes can be edited or rolled back.
- Run a pilot across different product categories and image roles, then review results on the actual product pages.
- Keep decorative or redundant imagery from receiving noisy descriptions, and make linked or functional images communicate their purpose.
The AltText.ai app listing describes automatic and bulk updates using product information. Treat marketplace features and pricing as changeable, and judge the result against the page context rather than the app’s accessibility claims alone.
5. Review candidates with a consistent checklist
- Purpose: Does this image need a text alternative, or should its alt value be empty?
- Relevance: Does the text communicate the information or function this page relies on?
- Context: Does it make sense beside the page heading, adjacent copy, and link or control?
- Accuracy: Did the model guess an identity, attribute, relationship, or action that is not supported by the image?
- Duplication: Does the text repeat a caption, link label, or nearby sentence without adding useful information?
- Text and data: If the image contains words or complex data, is there an equivalent text representation where needed?
- Language: Is the wording in the language expected by the page and assistive technology audience?
- Reviewability: Can an editor correct the candidate and restore the prior value if a batch is wrong?
Machine-generated descriptions can be incomplete, overly general, or imprecise; Microsoft advises reviewing and revising them. Section508.gov also cautions that a computer-generated visual description can miss the image’s relevant content. Do not infer correctness from fluent wording or a confidence score alone. Microsoft guidance; Section508.gov
6. Handle edge cases deliberately
- Decorative images: Set
alt=""when the image contributes no information or function. A visual-caption API will usually return objects even when those objects should not be announced. - Images that link: Describe what happens when activated or where the link goes. “A blue arrow” may describe appearance while missing the control’s purpose.
- Text embedded in an image: Include meaningful text in an equivalent text alternative or nearby content. OCR may help extract words, but verify them, especially for small, stylized, or low-contrast text.
- Charts, diagrams, and maps: Use concise alt text to identify the subject and provide detailed values, relationships, or directions in nearby text or a table. A caption alone cannot carry every data point.
- Product images: Use verified catalog details to disambiguate the item, but do not let metadata add attributes that the image does not show or that the page has not established.
- Repeated images: If the same image appears beside equivalent text, avoid announcing the same information twice. Review each occurrence in its own context.
- Uncertain or sensitive content: Route low-confidence, ambiguous, or sensitive images to a person. A system should be able to leave the field unchanged or mark it for review instead of guessing.
- Existing alt text: Do not overwrite human-reviewed values by default. Use explicit selection or a separate draft field.
7. Troubleshooting common problems
| Problem | Likely cause | Fix |
|---|---|---|
| Every image gets a visual caption, including decoration. | The workflow treats a blank alt field as missing text rather than a role decision. | Classify image purpose first; set decorative images to empty alt and exclude them from caption generation. |
| Descriptions are fluent but irrelevant to the page. | The generator sees pixels without enough page context, or the review focuses only on visible objects. | Review beside page content and revise to communicate the information or function the page relies on. |
| Captions contain incorrect details. | The model inferred objects, identities, or relationships incorrectly. | Reject or correct the draft; route ambiguous cases for review and never use confidence alone as approval. |
| Bulk processing overwrites good alt text. | The job updates all records, including previously reviewed values. | Skip non-empty fields by default, retain previous values, log updates, and test rollback on a small batch. |
| API request returns unauthorized. | The key or endpoint is wrong, the key is missing, or the credential was sent in the wrong header. | Check the provider’s current authentication instructions, resource endpoint, secret configuration, and key validity. Never expose a secret in client code. |
| API returns a validation or unsupported-feature error. | The API version, feature name, image format, or request body does not match current service requirements. | Compare the request with current API documentation and verify the endpoint and supported input formats. |
| Requests fail intermittently or take too long. | Network issues, provider throttling, large inputs, or transient service failures. | Use bounded timeouts, retry only transient failures with backoff, cap attempts, and move failed items to a reviewable retry queue. |
| The integration is built on a retiring API. | Service lifecycle information was missed during implementation. | Review the provider’s retirement and migration notice before launch and plan a replacement path. Microsoft’s cited Image Analysis 4.0 service is scheduled to retire September 25, 2028. |
| Generated alt text duplicates captions or nearby copy. | The generator is unaware of surrounding content. | Present the page context in the review interface and remove unnecessary repetition. |
8. Performance, reliability, and cost
Caption generation adds provider latency and an external dependency to the publishing flow. For large media libraries, process asynchronously in bounded batches rather than blocking an upload request. Track pending, successful, failed, and review-required items so a temporary outage does not silently lose work.
- Retries: Retry transient network and server failures with exponential backoff and a maximum attempt count. Do not repeatedly retry authentication, invalid input, or other permanent errors.
- Idempotency: Key jobs by asset ID, content hash, and model or configuration version so reruns do not create duplicate work or overwrite approved edits.
- Staleness: Regenerate only when the image or relevant context changes, and retain the prior approved value until a new candidate is reviewed.
- Throughput: Respect the provider’s current request limits. Use a queue and bounded concurrency rather than assuming unlimited parallel requests.
- Observability: Record status, latency, provider error category, and review outcome without logging image contents or secrets unnecessarily.
- Cost: Image-caption providers and platform apps may charge according to current plans or usage. Check live pricing, batch behavior, and whether retries or failed requests are billable before estimating a large backlog.
- Privacy: Images may contain personal, confidential, or regulated information. Review retention, processing location, and provider terms; use local processing where appropriate and available.
- Lifecycle: Confirm that the API is supported for the expected lifetime of your integration and budget time for migration.
9. Capture source images for a review workflow
If your editorial pipeline needs reference screenshots of pages that contain images, a website screenshot API can capture the page for human review. ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. It captures a URL as PNG, JPEG, WebP, or PDF. Its cookie-banner, popup, and chat-widget cleanup can help make a page capture easier to inspect, but a screenshot does not determine correct alt text; the description still needs review in page context.
For a custom image-description pipeline, use a captioning service and review the candidate as described above. For a screenshot of the source page, ScreenshotNeo’s one-request endpoint is documented at ScreenshotNeo API documentation. It is a page-capture option, not an alt-text generator.
10. Frequently asked questions
Can AI write alt text for all my images?
It can draft descriptions for many images, but it cannot reliably decide whether an image is decorative, functional, or meaningful in a particular page context. Use a review step and explicit handling for special cases.
Should decorative images have generated alt text?
No. When an image is purely decorative and adds no information or function, use an empty alt value so it does not add noise for assistive technology.
Can I bulk add alt text in WordPress?
WordPress plugin listings include options for bulk processing and generation on upload. Check the plugin’s current workflow and external-service disclosure, then pilot and review the results before applying them across a media library.
Can I automatically add alt text to Shopify product images?
Shopify app listings include tools that describe automatic and bulk updates, sometimes using product information. Verify current capabilities and pricing, and review each description on the product page.
Is a generated image caption automatically WCAG-conforming?
No. WCAG’s requirement is about an equivalent purpose for the non-text content. A fluent caption can still be irrelevant, incomplete, or wrong for that purpose.
Or skip the browser setup
If you need a screenshot of a page as part of an image review workflow, call ScreenshotNeo’s API with the page URL. This captures the page; it does not generate alt text.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot. Bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. See the API documentation and ScreenshotNeo.


