Google

Gemini Omni Video

Gemini Omni Flash is Google's preview model for conversational video generation and editing through the Interactions API. It accepts multimodal context, produces 3–10 second video at 720p and 24 FPS, and can refine a previously generated clip through natural-language follow-up instructions.

How to use this model

  1. 1

    Start with a prompt that names the scene, action, camera movement, lighting, mood, and required sound.

  2. 2

    Attach images or a short source video when the result should follow existing subjects, composition, or motion.

  3. 3

    Choose landscape or portrait output and a duration within the model's supported short-clip range.

  4. 4

    For an edit, state precisely what must change and what must remain untouched, including audio where relevant.

  5. 5

    Review the whole clip for continuity between shots; simplify the prompt if the model introduces unwanted scene changes.

Omni Flash PreviewAvailable

Omni Flash

Omni Flash accepts text, images, and short video context and returns a 3–10 second 720p clip with native audio. Follow-up edit instructions should identify both the requested change and the elements to preserve.

Properties

  • Inputs: text, image, and video context; output: 720p video at 24 FPS.
  • Supported output duration: 3–10 seconds.

How to use this model

  1. 1

    Provide the prompt and any image or video context required for the shot.

  2. 2

    Describe camera, action, lighting, mood, and audio explicitly.

  3. 3

    Select portrait or landscape output and a 3–10 second duration.

  4. 4

    For edits, name what changes and what remains unchanged.

Gemini Omni Flash selected in the Omnipix video workspace

Examples

Parameters

Prompt

Required

Text prompt describing the generated video.

Task

Required

Generation mode for Gemini Omni Flash.

  • Text to Video

  • Image to Video

  • Reference to Video

  • Video Editing

Duration

Optional

Requested output duration in seconds.

Resolution

Optional

Resolution of the output video.

  • 720p

Aspect Ratio

Optional

Aspect ratio of the output video.

  • 16:9

  • 9:16

First Frame Image

Optional

Input image to use as the starting frame when task is image_to_video.

Reference Images

Optional

Optional reference images for generation or editing. Refer to them in the prompt as <IMAGE_REF_0>, <IMAGE_REF_1>, etc. Maximum 6 images.

Source Video

Optional

Video to edit with the prompt.