What VideoRouter exposes for MiniMax speech, and how to pair audio with generated video
Updated 2026-10-02
If you are generating video with MiniMax, the next question is usually sound: narration, music or both. This page sticks to what VideoRouter's documentation and catalog actually list, and it is deliberately modest because the documented surface is small. Where something is not documented, it says so. The later sections are an honest, tool-agnostic guide to combining generated audio with generated video, which is the part that is usually underestimated.
What is documented for MiniMax speech
VideoRouter's text-to-speech page lists two MiniMax models on POST /v1/audio/speech: speech-02-hd and speech-02-turbo. The endpoint follows the shape of OpenAI's audio speech API and returns raw audio bytes in one response, with no streaming playback. Points the documentation is explicit about:
- MiniMax models use MiniMax's own voice-id vocabulary, and the
voicefield is required; there is no default. OpenAI voice names such asalloyare not the right vocabulary here. response_formatacceptsmp3(the default),opus,aac,flac,wavandpcm.- Most speech models bill per input character; see the model's page for its exact rate.
import os, requests
resp = requests.post(
"https://videorouter.sh/api/v1/audio/speech",
headers={"Authorization": f"Bearer {os.environ['VIDEOROUTER_KEY']}",
"Content-Type": "application/json"},
json={"model": "speech-02-turbo",
"voice": "<a MiniMax voice id>",
"input": "Your order has shipped.",
"response_format": "wav"},
timeout=120,
)
resp.raise_for_status()
open("narration.wav", "wb").write(resp.content)
The documentation does not list MiniMax voice ids on that page, so take them from MiniMax's own voice documentation or the model page, and test one before building around it.
What is not documented
VideoRouter's music generation page (POST /v1/audio/music) lists Google Lyria, Mureka, ElevenLabs, StepFun and Suno-via-Atlas models. MiniMax is not among them, so this site does not claim MiniMax music generation is available. Likewise, nothing in the documentation reviewed here describes voice cloning or a speech-to-video mode for MiniMax. If you need music for a video, choose from the models the music page actually lists, and check that page for current models and billing, which differ by provider (some flat per song, one per minute of requested length).
Separately, the H3 video model pages list reference-audio as an accepted input on some hosts. That is an input for conditioning generation, not a speech or music API; check the model page for the host you use before relying on it.
A pipeline: video first, narration second
The simplest reliable approach is to treat audio as a separate asset and combine them yourself:
- Write the narration script and decide the target length.
- Generate the speech and measure its duration with
ffprobe. - Generate video clips whose total length covers the narration. You are billed from the requested
duration_secs, so pick durations deliberately. - Mux audio onto video and trim to the shorter stream.
ffprobe -v error -show_entries format=duration -of csv=p=0 narration.wav
# attach narration to the video, keep the video stream as-is
ffmpeg -i clip.mp4 -i narration.wav -map 0:v -map 1:a \
-c:v copy -c:a aac -shortest final.mp4
If a clip has its own generated audio that you do not want, -map 0:v -map 1:a discards it by selecting only the video stream from the first input. If you want to keep both, mix them with -filter_complex "[0:a][1:a]amix=inputs=2" and lower the original with a volume filter.
Timing problems to plan for
- Length mismatch. Speech length depends on the text and voice, while video duration snaps to values the model supports. Measure both and adjust the script or concatenate several clips rather than stretching either one.
- Multiple clips, one voice-over. Generate speech for the whole script once, then cut the video to follow it. Per-sentence speech files joined later can have uneven pacing at the seams.
- Loudness. Normalise before delivery:
ffmpeg -i final.mp4 -af loudnorm -c:v copy normalised.mp4. - Lip movement. Narration over an off-screen speaker avoids the sync problem entirely. If you need on-screen speech, that is a different capability, and VideoRouter's documentation has a separate lip-sync page for it; read it rather than assuming video models do this.
Choosing where audio comes from
There are really three sources of sound for a generated video, and they have different failure modes. Pick one per project rather than mixing by accident.
| Source | What you control | Watch for |
|---|---|---|
| Audio produced by the video model itself | Mostly the prompt | Whether the model and host you chose produce audio at all; check the model page, and do not assume it from the family name |
Separate speech from /v1/audio/speech | Exact script, voice id and format | Length mismatch against the clip, which you resolve in editing |
Music from /v1/audio/music or your own library | Mood and length, depending on the model | Licence terms for anything that is not generated, and loudness balance against narration |
For a narrated explainer, separate speech gives you the most control, because the script is final before the picture is generated and you can write prompts that match each sentence. For a mood piece with no narration, leave the clip's own audio alone or lay music under it. In both cases, keep the audio file as an independent asset in your storage, so that re-rendering the picture later does not mean re-paying for the voice, and so a failed video job never costs you the narration.
A practical habit is to store a small manifest per finished video: the speech model and voice id, the character count, each clip's model id and requested duration, and the ffmpeg command used to combine them. When someone asks to change one line of narration, you regenerate one speech file and re-run the final command instead of starting over.
Cost and verification
Speech and video are billed separately, each by its own unit. Keep both in your per-job log: model ids, durations, job ids and character counts. Verify the finished file with ffprobe for stream count and duration before publishing, and listen to the first and last seconds, where sync problems surface first.
For the video side, see the call guide and the use-case overview; the speech documentation is on VideoRouter's text-to-speech page. Create a key to test both endpoints with one balance.
Frequently asked questions
Does VideoRouter offer MiniMax text-to-speech?
Its documentation lists speech-02-hd and speech-02-turbo on the /v1/audio/speech endpoint. MiniMax models require an explicit voice id, and the response is raw audio bytes.
Is MiniMax music generation available through VideoRouter?
The music generation documentation does not list a MiniMax model, so this site does not claim it. It lists Google Lyria, Mureka, ElevenLabs, StepFun and Suno models.
How do I add narration to a generated video?
Generate the speech file, measure its duration, then mux it onto the video with ffmpeg, mapping the video stream from the clip and the audio from the narration, and trimming to the shorter stream.
Are speech and video billed together?
No. They are separate endpoints with their own units, so log both per job if you need the total cost of a finished video.
Keep reading
- MiniMax H3 vs H3 Max vs H3 Max Turbo — Which Variant for Which Job
- MiniMax Video API Use Cases: What to Build with the H3 Family
- MiniMax vs Kling vs Seedance API: A Decision Framework
- MiniMax API Failover and Reliability: Pinning, Timeouts, Retries
VideoRouter puts it next to dozens of other video and image models behind one API key, so you can compare providers, prices and fail over automatically. Compare providers on VideoRouter →