1.5 Pro
This version combines visual direction and sound design in one generation. It can start from an image, optionally target a last frame, and generate dialogue, effects, ambience, or music described in the prompt.
Properties
- Joint audio-video generation with text-to-video and image-driven inputs.
- Optional first/last-frame guidance; the last frame requires a first frame.
How to use this model
- 1
Describe visual action and sound as coordinated events on one timeline.
- 2
Add a first frame for composition control; add a last frame only with the first.
- 3
Select duration, aspect ratio, and whether audio should be generated.
- 4
Review action timing and audio synchronization together.
Examples
Parameters
FPS
fpsFrame rate (frames per second)
24
Seed
seedRandom seed. Set for reproducible generation
Image
imageInput image for image-to-video generation
Prompt
promptText prompt for video generation
Duration
durationVideo duration in seconds: min: 2, max: 12, default: 5
Aspect Ratio
aspect_ratioVideo aspect ratio. Ignored if an image is used.
21:9
16:9
4:3
1:1
3:4
9:16
9:21
Camera Fixed
camera_fixedWhether to fix camera position
Generate Audio
generate_audioGenerate audio synchronized with the video. When enabled, the model outputs a video with audio that matches the visuals.
Last Frame Image
last_frame_imageInput image for last frame generation. This only works if an image start frame is given too.