HeyGen

HeyGen Video

Video Agent develops a finished video from one brief, while Avatar Video renders a completed HeyGen look. Video Translation localizes footage into another language. Lipsync keeps the supplied video performance and synchronizes a separate audio track in Speed or Precision mode.

How to use this model

  1. 1

    Choose Video Agent for a complete video, Avatar Video for a selected presenter, Translation for localization, or Lipsync for replacing speech with supplied audio.

  2. 2

    For an avatar clip, select the avatar and provide either a script with a voice or a prerecorded audio track—not both.

  3. 3

    Set the aspect ratio, 720p or 1080p resolution, and output format for the intended placement.

  4. 4

    Review pronunciation, lip synchronization, framing, and speaker identity before publishing the rendered file.

3Available

HeyGen Avatar Video

Avatar Video requires a completed look that supports the selected Avatar IV or Avatar III engine and exactly one speech source. Choose either a script with a voice or a prerecorded audio track. Photo Avatars can use motion prompting and expressiveness with Avatar IV. Omnipix exposes 720p and 1080p output and uses 1080p by default.

Properties

  • Requires a completed avatar look compatible with the selected Avatar IV or Avatar III engine.
  • Script and audio inputs are mutually exclusive; a voice ID applies to script input.
  • Motion prompt and expressiveness are limited to Photo Avatars rendered with Avatar IV.

How to use this model

  1. 1

    Choose an engine, then select one of the completed avatar looks compatible with it.

  2. 2

    If using a script, select a voice and proofread pronunciation and pauses.

  3. 3

    For an Avatar IV Photo Avatar, optionally describe the motion and choose expressiveness.

  4. 4

    Set the aspect ratio, 720p or 1080p resolution, and output format.

  5. 5

    Render and review lip synchronization and framing before publication.

Examples

Parameters

Avatar

Required

Choose one of your saved avatars.

Avatar Engine

Optional

Avatar IV is recommended for motion controls; Avatar III provides a dedicated photo-to-video pipeline. The selected look must support the engine.

  • Avatar IV

  • Avatar III

Script

Optional

What the avatar should say. Use either Script or Audio.

Audio

Optional

Audio for the avatar to lip-sync. Use either Audio or Script.

Voice

Optional

Selected favorite voice to use when generating from script.

Title

Optional

Optional name for this generated video.

Resolution

Optional

Final video resolution.

  • 720p

  • 1080p

Aspect Ratio

Optional

Shape of the final video frame.

  • 16:9

  • 9:16

  • 4:5

  • 5:4

  • 1:1

  • auto

Remove Background

Optional

Remove the background when the selected avatar supports it.

Output Format

Optional

Download format for the final video.

  • mp4

  • webm

Motion Prompt

Optional

Natural-language direction for body motion and hand gestures. Available only for Photo Avatars using Avatar IV.

Expressiveness

Optional

Energy and range of movement. Available only for Photo Avatars using Avatar IV.

  • Low

  • Medium

  • High

3Available

HeyGen Lipsync

HeyGen Lipsync accepts one video and one audio file. Speed mode costs less and targets faster processing; Precision mode uses the higher-quality lip-sync path. Optional controls preserve the source format, generate captions, enhance speech, remove the music track, restrict the processed time range, and choose frame-rate handling. Omnipix stores the completed video and any returned caption file, then settles the final charge from HeyGen's factual output duration.

Properties

  • Requires one stored video and one stored audio file.
  • Speed is billed at $0.0333 per output second; Precision at $0.0667 per output second.
  • Final pricing uses the authoritative duration returned by HeyGen.

How to use this model

  1. 1

    Upload a video with one clearly visible speaking face and a clean replacement audio track.

  2. 2

    Choose Speed for lower-cost processing or Precision when lip-sync quality is more important.

  3. 3

    Set captions, source-format preservation, music removal, speech enhancement, time range, and frame-rate mode only when needed.

  4. 4

    Review the full output for mouth timing, audio continuity, frame pacing, and any returned captions.

Parameters

Video

Required

Audio

Required

Title

Optional

Mode

Optional
  • Speed

  • Precision

Generate Captions

Optional

Keep Source Format

Optional

Dynamic Duration

Optional

Disable Music Track

Optional

Enhance Speech

Optional

Start Time

Optional

End Time

Optional

Frame Rate Mode

Optional
  • Variable

  • Constant

  • Passthrough

3Available

HeyGen Video Agent

Video Agent turns a creative brief into a finished video. HeyGen handles scripting, avatar selection, scene composition, and rendering. Omnipix uses one-shot generate mode: each request produces one video and does not open an interactive revision session.

Properties

  • One-shot generate mode only; chat and revision messages are not exposed.
  • Billed from the authoritative completed video duration.

How to use this model

  1. 1

    State the purpose, audience, key message, desired tone, and approximate structure in one prompt.

  2. 2

    Include any wording, facts, or calls to action that must appear in the narration.

  3. 3

    Choose landscape or portrait only when the destination format requires it; otherwise let HeyGen infer orientation.

  4. 4

    Review the generated script, presenter, scene pacing, factual accuracy, and final call to action.

Parameters

Prompt

Required

Describe the video goal, audience, message, tone, visual direction, and any required narration.

Orientation

Optional

Optional output orientation. Leave empty to let HeyGen infer it from the prompt.

  • Landscape

  • Portrait

3Retired

HeyGen Image to Video

Image to Video sends the source portrait directly to HeyGen's v3 video API. It accepts exactly one speech source: a script paired with a voice, or an uploaded audio track. The model remains unavailable until its provider behavior and wallet pricing are verified in staging.

Properties

  • Input: one image and exactly one speech source.
  • Uses the v3 image video type without avatar-only motion controls.

How to use this model

  1. 1

    Upload a clear portrait with one visible subject.

  2. 2

    Provide either a script and voice or a prerecorded audio track.

  3. 3

    Choose the target aspect ratio and resolution.

  4. 4

    Review facial motion, lip synchronization, and framing before publishing.

Parameters

Portrait

Required

A clear, front-facing portrait image.

Script

Optional

Text spoken by the person in the image. Use either Script or Audio.

Audio

Optional

Audio to lip-sync. Use either Audio or Script.

Voice

Optional

Voice used to speak the script.

Title

Optional

Display title in HeyGen.

Resolution

Optional

Final video resolution.

  • 720p

  • 1080p

  • 4K

Aspect Ratio

Optional

Auto preserves the source image shape when possible.

  • 16:9

  • 9:16

  • 4:5

  • 5:4

  • 1:1

  • Auto

Remove Background

Optional

Remove the image background from the generated video.

3Retired

HeyGen Video Translation

Video Translation preserves the source performance while producing one localized result per target language. Use a clean source with intelligible speech, select the translation mode deliberately, and review each language independently.

Properties

  • Input: one source video; output: one translation job per target language.
  • Supports voice cloning and lip-synchronized localization.

How to use this model

  1. 1

    Upload a source video with clear speech and minimal overlapping voices.

  2. 2

    Select every target language and the available speed or precision mode.

  3. 3

    Set any available timing range or audio-only option before starting.

  4. 4

    Review terminology, speaker voice, timing, and lip synchronization for each result.

Examples

Parameters

Video

Required

Video to translate

Output Languages

Required

List of output languages for the translation

  • Afrikaans

  • Albanian

  • Amharic

  • Arabic

  • Armenian

  • Azerbaijani

  • Basque

  • Belarusian

  • Bengali

  • Bosnian

  • Bulgarian

  • Burmese

  • Catalan

  • Chinese (Simplified)

  • Chinese (Traditional)

  • Croatian

  • Czech

  • Danish

  • Dutch

  • English (UK)

  • English (US)

  • Estonian

  • Filipino

  • Finnish

  • French

  • Galician

  • Georgian

  • German

  • Greek

  • Gujarati

  • Haitian Creole

  • Hebrew

  • Hindi

  • Hungarian

  • Icelandic

  • Indonesian

  • Irish

  • Italian

  • Japanese

  • Javanese

  • Kannada

  • Kazakh

  • Khmer

  • Konkani

  • Korean

  • Lao

  • Latin

  • Latvian

  • Lithuanian

  • Malay

  • Malayalam

  • Maltese

  • Marathi

  • Mongolian

  • Nepali

  • Norwegian

  • Odia

  • Pashto

  • Persian

  • Polish

  • Portuguese

  • Punjabi

  • Romanian

  • Russian

  • Serbian

  • Sinhala

  • Slovak

  • Slovenian

  • Somali

  • Spanish

  • Sundanese

  • Swahili

  • Swedish

  • Tamil

  • Telugu

  • Thai

  • Turkish

  • Ukrainian

  • Urdu

  • Uzbek

  • Vietnamese

  • Welsh

  • Zulu

Audio

Optional

Audio to use for the translation

Translate Audio Only

Optional

Only translate audio, keep original video

Speaker Count

Optional

Number of speakers (improves speaker separation)

Mode

Optional

Translation quality mode: 'speed' (faster) or 'precision' (higher quality, uses avatar inference)

  • fast

  • quality

Enable Caption

Optional

Generate captions for translated video

Keep the Same Format

Optional

Preserve the source video's encoding specs (resolution, bitrate).

Enable Dynamic Duration

Optional

Allow dynamic duration adjustment

Disable Music Track

Optional

Remove background music

Enable Speech Enhancement

Optional

Enhance speech quality

Start Time

Optional

Start time in seconds for partial translation

End Time

Optional

End time in seconds for partial translation

Subtitle File

Optional

Provide subtitle file for the translated video