Hailuo audio, voice and sound capabilities
MiniMax H3 generates video with native stereo audio at 32 kHz. The V2 API also accepts reference audio for voice, music, or soundscape guidance. MiniMax TTS, voice cloning, and music APIs are separate platform services and are not treated here as automatic Hailuo app workflows.
Audio paths
H3 generation
MiniMax H3 outputs native 32 kHz stereo audio with video. Dialogue support is documented for 11 stable languages.
Reference audio
Reference-to-video requests can use WAV or MP3 to guide voice timbre, music, or soundscape. Input limits are listed below.
Separate MiniMax audio APIs
MiniMax also documents TTS, voice cloning, timestamp return, and music services. They are useful companions, but not proof of an in-product Hailuo workflow.
What is verified
Every positive claim below is tied to an official source. Unconfirmed is not the same as unsupported.
| Question | Answer | Status | Source |
|---|---|---|---|
| Does MiniMax H3 generate sound with video? | Yes. MiniMax H3 generates video with native 32 kHz stereo audio. | Verified | MiniMax H3 model card Checked 2026-09-19 |
| Can audio guide generation? | Yes. H3 and H3 Max reference-to-video requests may include reference audio. WAV and MP3 are supported, with up to 3 clips, 2-15 seconds per clip, total audio duration of 15 seconds or less, and up to 15 MB per file. | Verified | MiniMax V2 video generation API Checked 2026-09-19 MiniMax video generation guide Checked 2026-09-19 |
| Can a prompt follow a reference voice? | Yes. The official V2 API reference example says, 'Voice timbre follows reference audio 1.' Reference audio can guide dialogue delivery, but this is not the same as the separate MiniMax voice-cloning API. | Verified | MiniMax V2 video generation API Checked 2026-09-19 |
| Which dialogue languages are documented for H3? | The H3 model card lists stable dialogue support for 11 languages: Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, and Spanish. | Verified | MiniMax H3 model card Checked 2026-09-19 |
| Does MiniMax provide standalone text-to-speech? | Yes, as a separate MiniMax platform API. Async long TTS supports up to 1 million characters, 100+ system voices, custom cloned voices, 40 listed languages, and sentence-level timestamps for subtitle alignment. | Verified | MiniMax Async Long TTS guide Checked 2026-09-19 |
| Does MiniMax provide standalone voice cloning? | Yes, as a separate MiniMax platform API. Source audio can be MP3, M4A, or WAV; it must be 10 seconds to 5 minutes long and no larger than 20 MB. The resulting voice ID is used with MiniMax speech synthesis. | Verified | MiniMax Voice Clone guide Checked 2026-09-19 |
| Does MiniMax provide standalone music generation? | MiniMax documents music generation, but its paid Music Generation and Lyrics Generation APIs stopped accepting new users on 20 August 2026. Existing paying users could continue. MiniMax directs new users to MiniMax Audio or the open-source MiniMax Music 3 model. | Verified | MiniMax Music Generation guide Checked 2026-09-19 |
| Can subtitles be generated? | MiniMax Async TTS can return sentence-level timestamps. Those timestamps can support subtitle alignment, but this does not prove that the Hailuo video app renders burned-in subtitles. | Verified | MiniMax Async Long TTS guide Checked 2026-09-19 |
| Does the Hailuo hosted app expose all of these audio controls? | Not verified here. API documentation does not establish the hosted app's in-product workflow, plan limits, or regional availability. Check the official Hailuo app for current controls. | Unconfirmed | No source claim |
| Does H3 Max output native audio? | Unconfirmed. The checked V2 API documents reference-audio input for H3 Max, while the model card's native 32 kHz stereo output statement is scoped to MiniMax H3. | Unconfirmed | MiniMax V2 video generation API Checked 2026-09-19 MiniMax H3 model card Checked 2026-09-19 |
Choose the right audio workflow
Sound inside the generated video
Use MiniMax H3 when the target is synchronized video plus native stereo audio. Include the desired soundscape and dialogue in the prompt.
Follow an existing voice or sound
Use reference-to-video with WAV or MP3 reference audio. Keep each clip between 2 and 15 seconds and total reference audio within 15 seconds.
Standalone voiceover or subtitles
Use MiniMax TTS or voice cloning separately, then combine the audio and video in an editor. Sentence-level timestamps can help align subtitles.
Audio FAQ
- Does Hailuo AI video have sound?
- For MiniMax H3, yes: the official model card says H3 generates video with native stereo audio at 32 kHz. H3 Max accepts reference audio as input, but this page does not infer native stereo output for H3 Max from that input capability.
- Can Hailuo follow a reference voice?
- The MiniMax V2 API supports reference audio in reference-to-video requests, and an official example says 'Voice timbre follows reference audio 1.' This is distinct from MiniMax's separate voice-cloning API.
- Does Hailuo include text-to-speech?
- MiniMax separately provides text-to-speech APIs, including async long-form TTS with custom cloned voices and sentence-level timestamps. Those platform APIs are not automatically the same as an in-product Hailuo video workflow.
- Can Hailuo generate subtitles?
- MiniMax Async TTS can return sentence-level timestamps, which can help align subtitles in an editor. This site does not claim that the Hailuo hosted app automatically creates or burns subtitles into video.