Google DeepMind released Gemini Omni 1.1 Flash on August 27, 2026, as a production-ready update for developers working with generative video. The model, available through Google AI Studio and the Gemini Enterprise Agent Platform, emphasizes enhanced control over video creation and editing. This release moves the focus from solely improving image quality to refining multimodal reference, conversational editing, and output specification control.
A central addition is scene extension, allowing developers to create longer, more consistent video sequences. Omni 1.1 Flash can analyze up to 10 seconds of prior footage when generating continuations, a notable increase from earlier models that relied on only the final second. This expanded context window aims to improve visual consistency and narrative flow across extended clips. Developers can add footage in 10-second increments, with a cumulative scene length reaching up to 40 seconds.
The update also introduces first-and-last-frame interpolation, enabling smoother transitions and precise camera movements. Users can specify both a starting and an ending image, and the model will generate the continuous video between them, facilitating effects like camera orbits, zoom transitions, and looping clips. Additionally, developers can provide up to three seconds of video reference material as multimodal input to guide visual context and maintain character consistency.
For efficient prototyping, Gemini Omni 1.1 Flash offers a 360p draft mode. Google states these previews generate up to 60% faster and cost one-third as much as standard 720p output, based on system throughput. This allows developers to iterate on concepts at a lower cost before committing to higher-resolution rendering. Final projects can be upscaled to 1080p or 4K resolution, providing options for polished, professional output. The pricing structure reflects this workflow, with per-second costs of $0.03 for 360p, $0.10 for 720p, $0.15 for 1080p, and $0.30 for 4K.
Gemini Omni 1.1 Flash supports conversational editing, which allows iterative refinement of videos through natural language instructions. The model maintains video state across turns, applying changes to specified elements while preserving other parts of the video without requiring re-uploads. The model processes text, images, audio, and video simultaneously, aiming for cohesive and controllable output. It also incorporates world knowledge from the broader Gemini family, integrating an understanding of physics, history, science, and cultural context.
The model is available under the stable Gemini API identifier `gemini-omni-1.1-flash`. Google's release notes classify it as generally available, with the earlier preview endpoint scheduled for deprecation on September 30, 2026. The model generates individual video outputs ranging from 3 to 10 seconds at 24 frames per second. Gemini Omni 1.1 Flash is accessible to developers via Google AI Studio and the Gemini Enterprise Agent Platform. It is also available globally in Flow for Google AI Plus, Pro, and Ultra subscribers, with scene extension features exposed within the Gemini app for these subscribers.
