MiniMax

MiniMax Speech

MiniMax Speech 2.8 HD turns text into speech using a selected voice ID. Omnipix exposes emotion, speed, volume, pitch, language assistance, channel layout, sample rate, bitrate, file format, and optional subtitles. It is suited to workflows that need more explicit control over both performance and the delivered audio file.

How to use this model

  1. 1

    Prepare the final narration text and add supported pause or interjection markers only where needed.

  2. 2

    Choose a preset voice ID or a compatible voice ID created with MiniMax Voice Cloning.

  3. 3

    Set emotion, speed, volume, and pitch; start near neutral values and change one dimension at a time.

  4. 4

    Select language assistance and normalization only when they match the script.

  5. 5

    Choose format, sample rate, bitrate, channel layout, and subtitles for the target workflow.

  6. 6

    Generate and review intelligibility, pronunciation, emotional consistency, and technical compatibility.

2.8 HDAvailable

Speech 2.8 HD

MiniMax Speech 2.8 HD accepts a script and voice ID, then exposes emotion, speed, volume, pitch, normalization, and language-assistance controls. Output settings include format, sample rate, bitrate, mono or stereo channels, and optional subtitle generation. A compatible cloned voice ID can be supplied in place of a preset voice.

Properties

  • Accepts preset MiniMax voices and compatible cloned voice IDs.
  • Controls emotion, speed, volume, pitch, normalization, and language assistance.
  • Supports MP3, WAV, FLAC, and PCM with selectable technical output settings.
  • Can request mono or stereo audio and optional subtitles.

Best for

  • Voiceovers
  • Audiobooks
  • Multilingual narration

Tips

  • Use <#0.5#> markers to insert timed pauses.

How to use this model

  1. 1

    Enter the narration text and select a preset or compatible cloned voice ID.

  2. 2

    Choose the emotion, then tune speed, volume, and pitch around neutral values.

  3. 3

    Set language assistance and English normalization only when appropriate for the script.

  4. 4

    Select audio format, sample rate, bitrate, mono or stereo, and subtitle output.

  5. 5

    Generate and verify speech clarity, emotional fit, pronunciation, subtitles, and file compatibility.

Parameters

Text

Required

Voice

Required
  • Deep Voice Man

  • Imposing Manner

  • Elegant Man

  • Casual Guy

  • Friendly Person

  • Decent Boy

  • Lively Girl

  • Exuberant Girl

  • Inspirational Girl

  • Young Knight

  • Abbess

  • Wise Woman

Speed

Required

Volume

Required

Pitch

Required

Emotion

Required
  • auto

  • happy

  • sad

  • angry

  • fearful

  • disgusted

  • surprised

  • calm

  • fluent

  • neutral

English Normalization

Required

Sample Rate

Required
  • 8000

  • 16000

  • 22050

  • 24000

  • 32000

  • 44100

Bitrate

Required
  • 32000

  • 64000

  • 128000

  • 256000

Audio Format

Required
  • mp3

  • wav

  • flac

  • pcm

Channel

Required
  • mono

  • stereo

Subtitle Enable

Required

Language Boost

Required
  • None

  • Automatic

  • Chinese

  • Chinese,Yue

  • Cantonese

  • English

  • Arabic

  • Russian

  • Spanish

  • French

  • Portuguese

  • German

  • Turkish

  • Dutch

  • Ukrainian

  • Vietnamese

  • Indonesian

  • Japanese

  • Italian

  • Korean

  • Thai

  • Polish

  • Romanian

  • Greek

  • Czech

  • Finnish

  • Hindi

  • Bulgarian

  • Danish

  • Hebrew

  • Malay

  • Persian

  • Slovak

  • Swedish

  • Croatian

  • Filipino

  • Hungarian

  • Norwegian

  • Slovenian

  • Catalan

  • Nynorsk

  • Tamil

  • Afrikaans