Google has released Gemini Omni 1.1 Flash, a production-ready update to its native multimodal video generation and editing model. The update introduces a suite of creative controls and generative capabilities designed to move the tool from a general generator to a directable production asset for developers and professional creators.
Enhanced Scene Extension and Continuity
A primary feature of the update is an improved scene extension tool that allows users to seamlessly continue existing video footage. According to Google, the model now analyzes up to 10 seconds of prior context to maintain visual consistency and narrative adherence, a significant increase from previous models that referenced only the final second.
Users can extend videos in 10-second increments up to a cumulative total of 40 seconds. To ensure a continuous seam, the model edits some final frames of the input. However, specific constraints apply: extensions can only be appended to the end of a clip, with no support for prepending or mid-clip insertion. Additionally, uploaded input videos must be 10 seconds or shorter unless the user is extending a model-generated video in multi-turn. While spoken dialogue is supported in multi-turn extensions via a previous_interaction_id, users cannot add new dialogue when extending an uploaded video featuring a speaking subject.
Precision Camera and Character Control
The update introduces a first-and-last-frame specification feature, enabling creators to define the opening and closing frames of a shot while the model generates the continuous motion between them. This functionality is intended for looping videos, rapid scene transitions, zooms, and orbit shots.
To further maintain character consistency and visual context, the model allows video clips to be used as references. Creators can provide up to three seconds of a separate video as a reference when creating a scene. For example, motion from multiple dance videos can be loaded as references and applied to different backgrounds or characters.
Production Workflow and Tiered Pricing
To streamline iteration, Google introduced a 360p draft mode for rapid prototyping. These low-resolution drafts generate up to 60% faster than standard 720p output and cost roughly one-third as much. Once a concept is finalized, the model supports upscaling to 1080p and 4K resolutions.

The API pricing is structured per second of generated video as follows:
- 360p: $0.03
- 720p: $0.10
- 1080p: $0.15
- 4K: $0.30
Billing for 720p video runs at 5,792 tokens per second. To ensure provenance, every generated video includes SynthID watermarking, which is programmatically detectable but invisible to viewers.
Availability and Performance
Gemini Omni 1.1 Flash is available through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. It is also live in Google Flow for subscribers of Google AI Plus, Pro, and Ultra, with scene extension available within the Gemini app. Production users already utilizing the model include Runway, GMI Cloud, Figma Weave, and Adobe (Firefly).

According to the AI video evaluation platform Arena, the model has seen strong early results, ranking first in the Text-to-Video Arena—scoring 20 points above FLUX 3 Video—and second in the Image-to-Video Arena, where it trails only MiniMax-H3.
Certain limitations remain in the current release. The model does not support voice editing, audio references, or the use of YouTube URLs as sources. Additionally, it lacks support for system instructions, stop sequences, temperature, top_p, or negative prompts, though the latter can be included within the prompt text. While English is fully supported, other languages remain unevaluated.
Продолжение темы

