How to Add Audio and Subtitles to Generated Videos
Add, sync, and mix audio in generated videos, then choose between selectable SRT or WebVTT tracks and captions burned into the picture.

To add audio and subtitles to a generated video, first check the exported video’s duration, frame rate, and audio streams. Prepare and align a voice-over, music, or effects track; create and review a timestamped transcript; then either attach an SRT or WebVTT subtitle track or render the text into the picture as burned-in captions. Keep a clean master without burned-in text if you may need to correct captions or publish other languages later.
This workflow works whether your video came from an AI video generator or another editing pipeline. The important steps are the same: confirm what is already in the file, make audio fit the picture, check captions against the final mix, and verify the result in the player where people will watch it.
1. Inspect the generated video
Before editing, find out how long the video is, its frame rate, and whether it already contains audio. This prevents accidental double audio and helps you spot a duration mismatch before adding tracks. FFmpeg’s ffprobe can inspect a local file:
ffprobe -v error -show_entries format=duration \
-show_entries stream=index,codec_type,codec_name,duration,r_frame_rate \
-of json generated.mp4
Look for a video stream, any existing audio stream, and the overall duration. If the file already has generated sound, decide whether to keep it, replace it, or mix it with narration or music. Save a copy of the original export before making changes.
2. Prepare and add audio
Use a voice-over for narration, music for a bed, or sound effects for specific events. Trim the audio to the intended section, remove unwanted silence where appropriate, and listen for clipping or distracting background noise. If dialogue and music coexist, lower the music enough that speech remains easy to understand. Normalize or adjust levels using your editor’s audio tools, then listen to the rendered result rather than relying only on waveform appearance.

Mux’s documented audio-track workflow accepts audio URLs pointing to M4A, WAV, or MP3 files. See its text and subtitle tracks documentation and audio tracks documentation for hosted-video workflows. For a local file, a simple FFmpeg command can combine video and audio. This example copies the video stream, encodes audio as AAC, and stops when the shorter input ends:
ffmpeg -i generated.mp4 -i voiceover.wav \
-map 0:v:0 -map 1:a:0 -c:v copy -c:a aac -shortest with-audio.mp4
If the video already has an audio track you want to preserve, map and mix streams deliberately instead of replacing it. For example, use an audio filter such as amix after confirming the inputs’ timing and volume. Test the opening, middle, and ending: a track can begin correctly yet drift, end early, or mask speech later.
Audio alignment checklist
- Confirm the audio starts at the intended moment, including any deliberate lead-in.
- Check sync at the beginning, middle, and end, especially for longer clips.
- Trim or loop music intentionally; avoid an abrupt cutoff unless it is part of the edit.
- Keep dialogue intelligible over music and effects.
- Listen to the final exported file on the target playback device or player.
3. Create and time the transcript
Automatic speech recognition can provide a first draft, but review it against the final audio. Correct names, punctuation, speaker changes, and timing. Include meaningful non-speech cues such as [music] or [applause] when producing captions for accessibility. Clear speech generally gives automatic captioning a better starting point; music, background noise, and long silence can reduce caption quality, as Mux notes in its caption-generation guidance.
Use captions when you intend to convey relevant sound information for deaf and hard-of-hearing viewers. Use subtitles when the primary purpose is translating spoken dialogue. Mux describes subtitles as on-screen text for translation and captions as including sound cues; platforms often use the terms loosely, so follow the destination’s labels. YouTube likewise describes subtitles and captions as a way to reach viewers who are deaf or hard of hearing and viewers who speak another language (YouTube Help).
For each cue, note a start time, an end time, and the text. Keep cues readable and synchronized; split long speech into manageable chunks rather than placing a full paragraph onscreen at once. Review timings against the mixed audio, since adding an intro, trimming the picture, or changing audio timing can make an otherwise good transcript inaccurate.
4. Save an SRT or WebVTT file
FFmpeg supports both SRT and WebVTT subtitle formats. Mux accepts either format for text tracks. An SRT file uses numbered cues and comma-separated milliseconds:
1
00:00:00,800 --> 00:00:03,200
A short opening line.
2
00:00:03,500 --> 00:00:06,100
A second line, timed to the narration.
A WebVTT file starts with WEBVTT and uses a period before milliseconds:
WEBVTT
00:00:00.800 --> 00:00:03.200
A short opening line.
00:00:03.500 --> 00:00:06.100
A second line, timed to the narration.
Use consistent language metadata, such as en or en-US, and label speakers or include sound cues where they matter. If you publish multiple languages, keep one track per language rather than combining translations into one crowded cue.
5. Choose selectable subtitles or burned-in captions
| Choice | Viewer control | Revision and delivery | Best fit |
|---|---|---|---|
| Selectable SRT/WebVTT track | Can be toggled and localized | Text can be replaced without rendering the video again; support depends on player and container | YouTube, web players, HLS, and multilingual publishing |
| Burned-in captions | Always visible; cannot be turned off | Text is rendered into video frames, so edits require a new render | Social feeds or exports where separate text tracks may be stripped |
Choose a selectable track when viewers need language choice, caption controls, or future text corrections without re-encoding the picture. Choose burned-in text when the destination will not preserve text tracks or when captions must appear in every player. Burned-in text remains visible but does not provide separate-track controls. If you may localize or revise captions, retain a master video without burned-in captions.

Attach a subtitle track with FFmpeg
For a local MP4 delivery, FFmpeg can add an SRT stream as a subtitle stream. Container and player support vary, so test the resulting file in the actual destination player. The following example keeps video and audio streams and adds SRT as a subtitle stream:
ffmpeg -i with-audio.mp4 -i captions.srt \
-map 0:v:0 -map 0:a? -map 1:0 \
-c:v copy -c:a copy -c:s mov_text \
-metadata:s:s:0 language=eng with-subtitles.mp4
To burn captions into the picture, render the subtitles with FFmpeg’s subtitles filter. The exact filter setup can depend on the installed FFmpeg build and subtitle file path; quote paths that contain spaces or special characters:
ffmpeg -i with-audio.mp4 \
-vf "subtitles=captions.srt" \
-c:v libx264 -c:a copy burned-captions.mp4
Burn-in requires video encoding, so it takes longer than simply copying the video stream. Check the output’s visual quality, caption placement, and readability before publishing.
Serve subtitles with HLS
For HLS, FFmpeg documents a pattern that maps video, audio, and WebVTT subtitles and assigns subtitle groups in the master playlist. Treat the playlist as part of the deliverable and test it with the target HLS player; a WebVTT file alone is not enough if the playlist does not expose the subtitle rendition. See the FFmpeg HLS documentation for the documented stream mapping and subtitle group options.
6. Upload captions to YouTube
In YouTube Studio, open the video’s Subtitles section, choose the caption language, then add or upload the subtitle file. YouTube requires text and timestamps and can also use position and style information. Follow the current controls in Studio and preview captions on the published player; a correctly formatted file can still have timing or display issues in context.
7. Review the final video
Review the actual published or encoded result, not just the source files. Confirm that the opening audio is present, speech and music are balanced, captions match the final words, and the chosen player displays the track. If captions are burned in, check safe placement and contrast across bright and dark scenes. If using separate tracks, verify language selection and toggling in the target player.
8. Or skip the browser setup
If you need screenshots of the documentation, settings pages, or generated-video examples while building a workflow, ScreenshotNeo can capture a URL with one API request. It is a website screenshot API and MCP server; the API returns PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://stripe.com \
-o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({
access_key: 'YOUR_API_KEY',
url: 'https://stripe.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);
Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, no card required.
Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| No sound after export | The output command mapped only video, or selected the wrong input stream | Inspect streams with ffprobe; map the intended audio stream explicitly and render again. |
| Two copies of narration or music | The generated video already had audio and the new track was added without replacing or mixing intentionally | Inspect the source audio streams, then choose one track or mix them with deliberate levels. |
| Audio ends too early or the video is cut off | Input durations differ; -shortest ends output when the shortest mapped stream ends |
Check durations and remove -shortest if you need the video to continue; define how audio should end. |
| Captions appear out of sync | Timing was made against a different edit, or the audio/video was trimmed or shifted | Retiming cues against the final mixed export; inspect beginning, middle, and end. |
| SRT is rejected or displays incorrectly | Malformed timestamps, encoding issues, or unsupported container/player behavior | Check cue numbering and timestamp separators, save as UTF-8, and test another supported delivery format or player. |
| Subtitle track is missing in HLS | The playlist does not expose the subtitle rendition or associate it with the right group | Check the master playlist and subtitle group configuration against FFmpeg’s HLS documentation. |
| Burned text is unreadable or clipped | Placement, contrast, or line length does not fit the scenes or player crop | Review a rendered sample across varied scenes, adjust style and safe placement, and re-render. |
| Automatic transcript misses words | Speech is unclear, masked by music/noise, or separated by long silence | Improve the audio where possible and manually correct the transcript and timing. |
Performance, reliability, and cost
Copying compatible video and audio streams avoids unnecessary video encoding; burning subtitles requires rendering the frames again and therefore takes more processing time. Separate tracks make text corrections and localization cheaper in editing effort because you can replace the captions without rebuilding the picture. Burn-in can simplify delivery to players that ignore tracks, but each caption change means another render.
For reliability, preserve the original generated export, a clean master, audio source files, and versioned caption files. Keep the caption language and the final edit associated so that an SRT from an earlier cut is not accidentally shipped with a newer video. Preview the actual destination player because support for text tracks varies by container and player.
There is no single render cost or processing time: it depends on resolution, duration, codec settings, hardware, and whether video must be re-encoded. A stream-copy mux is generally less work than a burn-in render, while hosted video APIs may price according to their own plans and usage. Check the provider’s current terms before choosing a production workflow.
FAQ
Can I add subtitles after generating the video?
Yes. Make a timestamped SRT or WebVTT file and either attach it as a selectable track or render it into the picture.
Should I use SRT or WebVTT?
Use the format accepted by your target platform or player. Both are supported by FFmpeg, and Mux accepts either for text tracks.
Can I fix captions without exporting the video again?
Yes, if captions are a separate text track and your delivery platform lets you replace that track. Burned-in captions require another video render.
Do captions need sound descriptions?
For accessibility captions, include meaningful non-speech audio cues where they help convey what is happening. Translation subtitles may focus on spoken dialogue.
Why keep a clean master?
A master without burned-in captions can be reused for corrections, new languages, and destinations with different caption controls.


