Browse by modality

Video

Generate video from text, images, and references.

What you can do

Video models can generate a scene from text, animate a still image, follow reference materials, or transform an existing clip. They differ in motion quality, control over characters and shots, native sound, duration, speed, and moderation.

Generate video from text

Describe the scene, action, camera, lighting, and pacing when you want the model to build the entire shot.

Animate a photo or illustration

Use a starting image to retain the composition while adding camera movement, character motion, or environmental effects.

Control a scene with references

Provide first and last frames, character or style references, audio, or a motion clip when the selected model supports them.

Edit, translate, or add a presenter

Specialized workflows can revise a clip from instructions, translate speech with lip sync, or create a talking presenter.

Inputs

  • A prompt describing the shot, action, camera, and sound
  • A starting image, and in supported workflows a final frame
  • Optional image, video, motion, or audio references supported by the selected model
  • Duration, aspect ratio, resolution, and other model-specific controls

Outputs

  • A generated or transformed video clip
  • Native sound only in models that explicitly support audio generation
  • Model-specific duration and resolution rather than one shared format for the entire catalog

How to choose a model

Choose by workflow first: text, a starting image, several references, an existing video, or native audio. Then compare speed, quality, supported duration, and cost inside the relevant family.

Reviewed examples

The examples below are reviewed Omnipix generations. Open the family pages to see more outputs and the controls available for each version.

Limitations and trade-offs

  • Character identity, object shape, and fine details can drift during motion or between shots.
  • Native audio, reference video, first and last frames, and video editing are model-specific capabilities.
  • Longer, higher-resolution, and reference-heavy generations generally take more time and credits.
  • Moderation rules differ between providers and may reject a prompt that another model accepts.

All available model families

Black Forest Labs

View details

ByteDance

View details

Kling AI

View details