Avatar v2
This version requires one portrait and one speech recording. It synchronizes facial motion to the audio, accepts optional expression and movement guidance, and offers Standard and Professional quality modes. The audio file determines the generated video duration.
Properties
- Required inputs: one portrait image and one speech-audio file.
- Output duration matches the audio; Standard and Professional quality modes are available.
Best for
- Talking portraits, presenters, character dialogue, and localized voice tracks.
Avoid for
- Full-body motion, multi-person conversations, or scenes that require camera movement.
Tips
- Use a clear, front-facing portrait and clean speech audio for the strongest lip synchronization.
How to use this model
- 1
Upload a clear JPEG or PNG portrait with one front-facing subject.
- 2
Upload clean MP3, WAV, M4A, or AAC speech audio; the result will use the same duration.
- 3
Optionally describe subtle expression or movement that agrees with the portrait and speech.
- 4
Select Standard or Professional, then inspect lip synchronization and facial consistency across the full clip.
Parameters
Portrait
imagePortrait to animate. JPEG or PNG, up to 10 MB, at least 300 px on each side, with an aspect ratio from 1:2.5 to 2.5:1.
Audio
audioSpeech audio that drives the avatar. MP3, WAV, M4A, or AAC, up to 5 MB. The output video has the same duration as this audio.
Prompt
promptOptional instructions for the avatar's expression and movement.
Mode
modeStandard mode is faster and more economical; Professional mode prioritizes output quality.
Standard
Professional