Gemini Omni 1.1 Flash 40-Second Extension Measured
The 40-second scene extension claim on Gemini Omni 1.1 Flash comes from specific API tests; here is what the number covers and what changes for different
The announcement originates from product documentation written by two Google DeepMind product managers.
The text names Anish Nangia and Alisa Fortin as the authors of the post that presents the scene extension capability. It states that the model now analyzes up to 10 seconds of prior context instead of the single final second used in earlier versions. The same section lists the 10-second increment rule and the 40-second cumulative cap as the concrete limits developers must observe when calling the API with a previous_interaction_id. The post opens by describing the release as a production-ready update that offers improved control over generative video through Google AI Studio or the Gemini Enterprise Agent Platform.
Access routes through two documented paths. Developers begin in Google AI Studio for direct trials. Enterprise teams route calls through the Gemini Enterprise Agent Platform. Google AI Plus, Pro and Ultra subscribers receive scene extension inside the Gemini app at global availability starting on the release date. The post also lists separate documentation, a cookbook, and prompting guides that cover scene extensions, video references, and upscaling.
Scene extension records only the length of segments added after the initial clip.
The headline total therefore excludes the duration of the starting video supplied by the user. A developer who begins with a 15-second shot can produce an additional 40 seconds of extension and end with 55 seconds of final output while remaining inside the published limit. The measurement also excludes the time spent on the first and last frame specification calls that control camera movement between keyframes.
Scene extension allows you to take an existing video and continue generating footage seamlessly from where it left off.
The code example in the announcement shows an interaction.create call that passes previous_interaction_id and a text instruction such as "Continue the scene," then requests 360p output. Each successful response adds another 10-second block to the running total until the 40-second ceiling is reached. The same section supplies three prompt examples that demonstrate narrative continuation across separate calls while preserving visual style. One prompt set continues a conversation between two characters while adding dramatic music; another executes a cinematic optical dolly-zoom followed by a 360-degree orbital rotation around a frozen subject.
The 10-second context window and consistency improvements sit outside the extension total.
The announcement credits the longer context window for better visual consistency and narrative adherence across extension steps, yet those frames do not add to the 40-second count. The same paragraph notes that developers can still reference up to three seconds of additional video input for character or style consistency without affecting the extension budget. Up to three seconds of video references can be supplied in multimodal input to maintain continuity across separate generations. The feature description lists three concrete use cases: replacing dancers from reference clips with new characters while keeping the same large open space, maintaining a microscope lens effect through an entire scientific visualization, and preserving the exact size of a shocked face during a dolly-zoom.
First and last frame specification operates independently of the extension counter. The feature generates continuous video between two supplied keyframes, supporting camera orbits, zoom transitions and seamless loops. These calls consume their own generation quota and do not count toward the 40-second scene extension total. The post illustrates the capability with a whip-pan transition from a drummer to a saxophonist and ballet dancer, and with a continuous pull-back that reveals an entire harbor scene.
360p drafting changes the effective cost and speed for users who iterate before final 4K output.
The announcement states that 360p output runs up to 60 percent faster and at one-third the cost of the standard 720p resolution. It positions this mode for rapid prototyping and storyboard iteration, after which the same workflow can request 4K upscaling on the selected take. Access to the full feature set, including scene extension, requires a Google AI Plus, Pro or Ultra subscription in the Gemini app, while API usage routes through Google AI Studio or the Gemini Enterprise Agent Platform.
Adobe has integrated the model into Firefly, and Figma Weave uses it for branching versions and attaching references. GMI Cloud and Runway also list production deployments that rely on the new controls. One quoted Figma director states that extensions and 4K resolution move teams from generating videos to directing them. A GMI Cloud executive notes that detail accuracy holds up under scrutiny for educational content where reliability matters more than any single feature. Runway’s chief creative officer observes that the model fits naturally into existing prompt-plus-reference workflows.
Outbound documentation lives at the Google DeepMind blog post and Ars Technica coverage of related standards work. The 40-second limit would turn out to be different if the service changed the increment size or context window in a future update.
Keep reading
News
AI21 Labs Cuts 60% of Staff, Bets on Maestro
AI21 Labs slashes over 60% of staff, drops foundation models, and pivots to its Maestro agent optimization platform after Nebius acquisition talks collapse.
News
Alibaba Bans Claude Code Over Security Concerns
Alibaba told staff to remove Anthropic's Claude Code by July 10 over security concerns. Here's what triggered the ban and what it signals.
News
Anthropic Acquires Stainless: What It Means for AI
Anthropic bought Stainless, the SDK generator behind OpenAI and Cloudflare's client libraries. Here's the strategic play for AI agents.