Hailuo 02 generates video from text, a first frame, or paired first and last frames.
Brands
MiniMax
Hailuo video generation plus MiniMax speech synthesis and voice cloning.
Models
Video
Hailuo Video
View details →Hailuo 02 and Hailuo 2.3 support text-to-video and image-to-video workflows, while 2.3 Fast is intended for faster image-driven iteration. The models accept prompts for action and camera direction; Hailuo's API also supports bracketed camera commands for compatible image-to-video variants.
Hailuo 2.3 generates video from text or a first-frame image, with controls for camera direction, duration, and resolution.
Hailuo 2.3 Fast for quicker image-to-video iteration from a supplied first frame.
Models
Audio
MiniMax Speech
View details →MiniMax Speech 2.8 HD turns text into speech using a selected voice ID. Omnipix exposes emotion, speed, volume, pitch, language assistance, channel layout, sample rate, bitrate, file format, and optional subtitles. It is suited to workflows that need more explicit control over both performance and the delivered audio file.
Synthesize controlled multilingual speech with preset or compatible cloned voices and production-oriented audio settings.
MiniMax Voice Cloning
View details →MiniMax Voice Cloning analyzes a reference audio file and returns a voice ID for compatible MiniMax speech models. Omnipix exposes a clone name, an accuracy control, and optional noise reduction and volume normalization. Clean, single-speaker source audio makes evaluation easier; use only recordings you are authorized to submit and reproduce.
Create a named MiniMax voice ID from an authorized reference recording for later speech synthesis.