Kling AI

Kling Avatar v2

Kling Avatar v2 animates a single portrait from an uploaded audio track. The audio determines the output duration and drives lip movement, while an optional prompt can guide expression and motion. Standard and Professional modes trade generation cost and speed against output quality; neither mode is intended for full-body action, multiple speakers, or camera-led scenes.

How to use this model

  1. 1

    Upload a sharp, front-facing portrait with one clearly visible person and no obstruction across the face.

  2. 2

    Add clean speech audio in a supported format; its duration determines the length of the generated video.

  3. 3

    Use the optional prompt only for compatible expression or movement guidance, without contradicting the portrait or speech.

  4. 4

    Choose Standard for faster, lower-cost iterations or Professional when output quality is the priority.

  5. 5

    Review lip synchronization, facial identity, blinking, teeth, and the beginning and end of the audio before publishing.

Avatar v2Available

Avatar v2

This version requires one portrait and one speech recording. It synchronizes facial motion to the audio, accepts optional expression and movement guidance, and offers Standard and Professional quality modes. The audio file determines the generated video duration.

Properties

  • Required inputs: one portrait image and one speech-audio file.
  • Output duration matches the audio; Standard and Professional quality modes are available.

Best for

  • Talking portraits, presenters, character dialogue, and localized voice tracks.

Avoid for

  • Full-body motion, multi-person conversations, or scenes that require camera movement.

Tips

  • Use a clear, front-facing portrait and clean speech audio for the strongest lip synchronization.

How to use this model

  1. 1

    Upload a clear JPEG or PNG portrait with one front-facing subject.

  2. 2

    Upload clean MP3, WAV, M4A, or AAC speech audio; the result will use the same duration.

  3. 3

    Optionally describe subtle expression or movement that agrees with the portrait and speech.

  4. 4

    Select Standard or Professional, then inspect lip synchronization and facial consistency across the full clip.

Parameters

Portrait

Required

Portrait to animate. JPEG or PNG, up to 10 MB, at least 300 px on each side, with an aspect ratio from 1:2.5 to 2.5:1.

Audio

Required

Speech audio that drives the avatar. MP3, WAV, M4A, or AAC, up to 5 MB. The output video has the same duration as this audio.

Prompt

Optional

Optional instructions for the avatar's expression and movement.

Mode

Optional

Standard mode is faster and more economical; Professional mode prioritizes output quality.

  • Standard

  • Professional