2.5
Seedance 2.5 uses one multimodal workflow for text-only and reference-guided generation. Add up to nine images, three videos, and three audio references, then address them by position in the prompt. Output is available at 480p or 720p for 4–15 seconds; a supplied image determines the aspect ratio.
Properties
- Single task type for text and mixed image, video, and audio guidance.
- Supports up to nine images, three videos, and three audio references.
Best for
- High-quality text-to-video generation
- Image-guided subject and style control
- Mixed image, video, and audio references
Avoid for
- 1080p output, output longer than 15 seconds, or audio-only reference generation
Tips
- Add only the image, video, and audio references needed for the result.
- Use @imageN, @videoN, and @audioN in the prompt to address references by position.
How to use this model
- 1
Start with a clear ordered description of the scene, motion, camera, and audio.
- 2
Attach only the image, video, and audio references needed for the result.
- 3
Use @imageN, @videoN, and @audioN to identify each reference in the prompt.
- 4
Choose 480p or 720p and a duration from 4 to 15 seconds.
Parameters
Prompt
promptText prompt describing the desired scene, motion, and action.
Resolution
resolutionOutput video resolution. This tier does not support 1080p.
480p
720p
Duration
durationGenerated video duration in seconds.
Aspect Ratio
aspect_ratioOutput aspect ratio. A reference image takes precedence when supplied.
21:9
16:9
4:3
1:1
3:4
9:16
auto
Reference Images
image_urlsUp to nine images for subject appearance, content, or style guidance.
Reference Videos
video_urlsUp to three motion, content, or style references totaling at most 15.4 seconds.
Reference Audio
audio_urlsUp to three MP3, OGG, WAV, M4A, or AAC references; audio requires an image or video reference.