How to Add AI-Generated Subtitles to Videos
Generate accurate AI subtitles, fix transcription errors, sync timing, choose SRT or burned-in captions, and publish accessible videos.
Short answer: import your video into an automatic transcription tool, select the spoken language, generate captions, review every line, correct timing and important sounds, then export either a selectable subtitle file or a video with captions burned in. Automatic captions are a draft: YouTube advises reviewing and editing every machine-generated line.
1. Choose the right subtitle output
| Output | What viewers get | Best use | Trade-off |
|---|---|---|---|
| SRT or SBV track | Captions can be switched on or off | YouTube, most players, translations | Player must support the file |
| WebVTT | Web captions with limited positioning and styling | HTML5 video and web players | Support varies by player |
| TTML/DFXP | Timed text with richer styling and positioning | Broadcast and supported enterprise workflows | More complex authoring |
| Burned-in captions | Text is part of the picture and always visible | Social clips, silent autoplay, presentations | Cannot be turned off, searched, or restyled after export |
W3C defines captions as synchronized text for dialogue and relevant non-speech audio. Closed captions are selectable; open captions are permanently visible. For accessibility, keep an editable timed-text file even when you also export a burned-in version.
2. Prepare the video for transcription
- Use the clearest available audio. Reduce music and background noise before transcription if they mask speech.
- Identify the spoken language and dialect. Select that language explicitly in your caption tool.
- Separate overlapping speakers when possible. Two people speaking at once, long silence, accents, unsupported languages, and very long videos can reduce accuracy.
- Keep the original video and audio. You will need them to check a questionable word or speaker change.
- Decide whether names, profanity, and meaningful sounds such as [applause] or [door closes] should be represented.
3. Generate captions in common tools
YouTube Studio
- Open YouTube Studio and select the video.
- Open Subtitles and choose the video language.
- Open the automatic caption track after processing.
- Edit the transcript and timings line by line, or upload a prepared SRT, SBV, WebVTT, or TTML file.
- Preview the result on the video page before publishing.
YouTube notes that automatic-caption quality varies with accents, dialects, background noise, overlapping speech, silence, language, and video length. Its guidance says: “You should always review automatic captions and edit any parts that haven’t been properly transcribed.” See the YouTube caption help.
Adobe Premiere Pro
- Open Window → Text.
- Generate a transcript, choose the language, and enable speaker labels when useful.
- Transcribe the required audio.
- Create a caption track from the transcript.
- Correct words and timings in the Text panel, then style or translate the caption track.
- Export a caption file or include captions in the rendered video.
Premiere Pro also documents caption translation workflows. Adobe announced automatic translation for 27 languages in Premiere Pro 25.2; check the current version’s documentation before planning a multilingual release.
Descript
- Import the video or audio.
- Choose the transcription language and let Descript create a transcript.
- Edit the transcript as text; the media follows your edits.
- Generate and style captions.
- Export an SRT/WebVTT-style subtitle file or a captioned video.
Descript advertises subtitles in 30+ languages and claims 95% automatic-transcription accuracy. Treat that percentage as a vendor claim, not an independent benchmark; your review effort depends on the recording.
CapCut
- Import the clip.
- Choose Auto Caption (also called Recognise Subtitles).
- Select the spoken language and generate captions.
- Correct the transcript, split or merge lines, adjust timing, style the text, and export.
CapCut describes this feature as AI speech-to-text. Short-form clips can be edited quickly, but inspect every proper noun and number before posting.
4. Review the transcript systematically
Automatic transcription is a starting point. Review with the video paused and then watch the complete captioned result at normal speed.
- Words: correct names, places, technical terms, contractions, numbers, and homophones.
- Punctuation: add sentence endings and commas that clarify meaning without creating unnatural fragments.
- Speakers: split turns consistently and add speaker labels when the audience needs them.
- Sound: include meaningful non-speech audio such as [laughter], [music], or an alarm.
- Profanity: apply the destination platform’s policy consistently; do not let automatic masking change the meaning.
- Language: retain accents and dialect-specific words instead of silently standardizing a speaker’s voice.
- Numbers and symbols: verify dates, prices, units, URLs, code, and product names character by character.
5. Fix timing and readability
Captions must appear with the speech, remain long enough to read, and disappear when the thought ends. Section 508 guidance warns that speech above 180 words per minute (about three words per second) may be too fast for captions. Faster dialogue may require shorter lines, carefully chosen breaks, or slightly condensed wording that preserves meaning.
- Break at natural grammar boundaries rather than after a random character count.
- Avoid covering a speaker’s face, on-screen labels, or essential demonstrations.
- Keep a caption on screen long enough for a viewer to read it comfortably.
- Do not create one-frame flashes or leave stale text after the speaker has stopped.
- Check cuts, music changes, laughter, and speaker changes for clean transitions.
- Preview on a phone and a desktop. Small screens expose crowded lines and poor contrast.
For exact timing, use the editor’s waveform and frame-accurate controls. If a tool’s automatic segmentation is poor, split or merge cues manually instead of changing the spoken words.
6. Edit an SRT file directly
An SRT file contains a sequence number, a start and end timestamp, one or more text lines, and a blank line:
1
00:00:01,000 --> 00:00:04,200
Welcome to the product demo.
2
00:00:04,400 --> 00:00:07,800
Today we will configure the API.
Use commas for milliseconds in SRT timestamps. Keep numbering sequential, use UTF-8 encoding, and save the file with the .srt extension. WebVTT uses a WEBVTT header and periods in timestamps:
WEBVTT
00:00:01.000 --> 00:00:04.200
Welcome to the product demo.
7. Export burned-in captions with FFmpeg
Burning captions into a video makes them visible everywhere, but the result is no longer selectable. Install FFmpeg, then run:
ffmpeg -i input.mp4 -vf "subtitles=captions.srt:force_style='FontName=Arial,FontSize=22,PrimaryColour=&Hffffff&,OutlineColour=&H000000&,BorderStyle=3,Outline=2,Shadow=0,Alignment=2,MarginV= forty'" -c:a copy output-captioned.mp4
Replace the invalid placeholder in MarginV= forty with a number before running (for example, MarginV=40). A safer portable command is:
ffmpeg -i input.mp4 -vf "subtitles=captions.srt" -c:a copy output-captioned.mp4
Use a selectable caption track for accessibility and translation whenever the destination supports it. Export a burned-in copy when captions must remain visible in a feed or player that ignores subtitle files.
8. Translate and publish
- Correct the source-language captions first.
- Duplicate the reviewed track for each target language.
- Translate meaning, names, idioms, and sound cues; do not translate raw machine output without review.
- Re-time lines after translation because word order and length change.
- Publish selectable tracks with clear language labels. Upload the burned-in version only when permanent visibility is required.
- Keep the source video, editable transcript, final timed-text files, and exported videos together for later corrections.
9. Or skip the browser setup
If your subtitle tutorial, documentation, or QA process needs clean screenshots of web pages, ScreenshotNeo can capture them with one GET request. It is a website screenshot API and MCP server, not a subtitle generator.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo API documentation for the available options. Cookie banners, newsletter popups, and chat widgets are removed before the shot; bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP server lets Claude, Cursor, and other MCP clients take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
10. Troubleshooting
| Problem | Likely cause | Fix |
|---|---|---|
| Many wrong words | Noise, music, accent, or wrong language | Clean the audio, select the correct language, and re-transcribe a shorter segment. |
| Two speakers are merged | Overlapping speech or missing speaker detection | Separate turns manually and add speaker labels during review. |
| Captions drift over time | Variable frame rate, edited media, or a shifted track | Regenerate from the final edit, then check sync at the beginning, middle, and end. |
| Text is too fast to read | Long cues, rapid speech, or poor line breaks | Split cues at grammatical boundaries and keep reading speed near the 180-WPM guidance. |
| Boxes or garbled characters appear | Wrong encoding or unsupported font | Save as UTF-8, use a common font, and verify punctuation after import. |
| SRT will not upload | Malformed timestamps, numbering, or blank lines | Check HH:MM:SS,mmm syntax, sequential indexes, and blank lines between cues. |
| Captions cover important content | Fixed placement or unsafe margins | Move the track, increase margins, or reposition the on-screen subject before export. |
| Automatic captions are unavailable | Unsupported language, processing state, audio quality, or video length | Wait for processing, upload your own timed-text file, or transcribe locally in a supported tool. |
11. Performance, reliability, privacy, and cost
- Performance: shorter clips and clean, single-speaker audio generally process faster and require fewer corrections.
- Reliability: keep the original media and editable captions; re-exporting from a flattened video makes future fixes harder.
- Accuracy: do not publish a machine transcript without a human pass, especially for names, medical or legal terms, numbers, and accessibility-critical content.
- Privacy: check where each service stores uploads and whether your project permits cloud transcription. For sensitive footage, use an approved local or enterprise workflow.
- Cost: platform-native captioning may be included, while paid editors and translation services charge by subscription, usage, or export. Compare the cost of correction time, not only the generation price.
FAQ
Are AI subtitles accurate enough to publish automatically?
They are useful drafts, but accuracy varies with audio, speakers, language, and vocabulary. Review every line before publication.
Should I use SRT or burned-in captions?
Use SRT, WebVTT, or another selectable track when accessibility, search, translation, or user controls matter. Use burned-in captions when the destination may ignore caption files or text must always be visible.
Can I edit subtitles without re-exporting the video?
Yes. Edit and replace the subtitle track when it is selectable. Burned-in text requires a new video render.
What should captions include besides speech?
Include meaningful sounds, speaker changes, and audio events needed to understand the scene, while avoiding irrelevant background noise.
How do I make captions readable on phones?
Preview at phone size, use high contrast, keep lines short, avoid covering faces or key graphics, and leave each cue on screen long enough to read.


