Home › Guides › Voice and music with video

What VideoRouter exposes for MiniMax speech, and how to pair audio with generated video

Updated 2026-10-02

If you are generating video with MiniMax, the next question is usually sound: narration, music or both. This page sticks to what VideoRouter's documentation and catalog actually list, and it is deliberately modest because the documented surface is small. Where something is not documented, it says so. The later sections are an honest, tool-agnostic guide to combining generated audio with generated video, which is the part that is usually underestimated.

What is documented for MiniMax speech

VideoRouter's text-to-speech page lists two MiniMax models on POST /v1/audio/speech: speech-02-hd and speech-02-turbo. The endpoint follows the shape of OpenAI's audio speech API and returns raw audio bytes in one response, with no streaming playback. Points the documentation is explicit about:

import os, requests

resp = requests.post(
    "https://videorouter.sh/api/v1/audio/speech",
    headers={"Authorization": f"Bearer {os.environ['VIDEOROUTER_KEY']}",
             "Content-Type": "application/json"},
    json={"model": "speech-02-turbo",
          "voice": "<a MiniMax voice id>",
          "input": "Your order has shipped.",
          "response_format": "wav"},
    timeout=120,
)
resp.raise_for_status()
open("narration.wav", "wb").write(resp.content)

The documentation does not list MiniMax voice ids on that page, so take them from MiniMax's own voice documentation or the model page, and test one before building around it.

What is not documented

VideoRouter's music generation page (POST /v1/audio/music) lists Google Lyria, Mureka, ElevenLabs, StepFun and Suno-via-Atlas models. MiniMax is not among them, so this site does not claim MiniMax music generation is available. Likewise, nothing in the documentation reviewed here describes voice cloning or a speech-to-video mode for MiniMax. If you need music for a video, choose from the models the music page actually lists, and check that page for current models and billing, which differ by provider (some flat per song, one per minute of requested length).

Separately, the H3 video model pages list reference-audio as an accepted input on some hosts. That is an input for conditioning generation, not a speech or music API; check the model page for the host you use before relying on it.

A pipeline: video first, narration second

The simplest reliable approach is to treat audio as a separate asset and combine them yourself:

  1. Write the narration script and decide the target length.
  2. Generate the speech and measure its duration with ffprobe.
  3. Generate video clips whose total length covers the narration. You are billed from the requested duration_secs, so pick durations deliberately.
  4. Mux audio onto video and trim to the shorter stream.
ffprobe -v error -show_entries format=duration -of csv=p=0 narration.wav

# attach narration to the video, keep the video stream as-is
ffmpeg -i clip.mp4 -i narration.wav -map 0:v -map 1:a \
  -c:v copy -c:a aac -shortest final.mp4

If a clip has its own generated audio that you do not want, -map 0:v -map 1:a discards it by selecting only the video stream from the first input. If you want to keep both, mix them with -filter_complex "[0:a][1:a]amix=inputs=2" and lower the original with a volume filter.

Timing problems to plan for

Choosing where audio comes from

There are really three sources of sound for a generated video, and they have different failure modes. Pick one per project rather than mixing by accident.

SourceWhat you controlWatch for
Audio produced by the video model itselfMostly the promptWhether the model and host you chose produce audio at all; check the model page, and do not assume it from the family name
Separate speech from /v1/audio/speechExact script, voice id and formatLength mismatch against the clip, which you resolve in editing
Music from /v1/audio/music or your own libraryMood and length, depending on the modelLicence terms for anything that is not generated, and loudness balance against narration

For a narrated explainer, separate speech gives you the most control, because the script is final before the picture is generated and you can write prompts that match each sentence. For a mood piece with no narration, leave the clip's own audio alone or lay music under it. In both cases, keep the audio file as an independent asset in your storage, so that re-rendering the picture later does not mean re-paying for the voice, and so a failed video job never costs you the narration.

A practical habit is to store a small manifest per finished video: the speech model and voice id, the character count, each clip's model id and requested duration, and the ffmpeg command used to combine them. When someone asks to change one line of narration, you regenerate one speech file and re-run the final command instead of starting over.

Cost and verification

Speech and video are billed separately, each by its own unit. Keep both in your per-job log: model ids, durations, job ids and character counts. Verify the finished file with ffprobe for stream count and duration before publishing, and listen to the first and last seconds, where sync problems surface first.

For the video side, see the call guide and the use-case overview; the speech documentation is on VideoRouter's text-to-speech page. Create a key to test both endpoints with one balance.

Frequently asked questions

Does VideoRouter offer MiniMax text-to-speech?

Its documentation lists speech-02-hd and speech-02-turbo on the /v1/audio/speech endpoint. MiniMax models require an explicit voice id, and the response is raw audio bytes.

Is MiniMax music generation available through VideoRouter?

The music generation documentation does not list a MiniMax model, so this site does not claim it. It lists Google Lyria, Mureka, ElevenLabs, StepFun and Suno models.

How do I add narration to a generated video?

Generate the speech file, measure its duration, then mux it onto the video with ffmpeg, mapping the video stream from the clip and the audio from the narration, and trimming to the shorter stream.

Are speech and video billed together?

No. They are separate endpoints with their own units, so log both per job if you need the total cost of a finished video.

Keep reading

Using MiniMax is one part of the job.

VideoRouter puts it next to dozens of other video and image models behind one API key, so you can compare providers, prices and fail over automatically. Compare providers on VideoRouter →