Google DeepMind

Gemini Omni

Google's multimodal creation model — where Gemini's reasoning meets the ability to create. Generate and edit video from text, images, video, or audio with natural language. Every edit builds on the one before. Try free with Nano Banana Pro.

About

About Gemini Omni

Gemini Omni is Google DeepMind's multimodal creation model, announced at Google I/O 2025. It brings Gemini's reasoning ability together with generative media systems, enabling video generation and editing that goes beyond simple prompt-to-video output. The model understands scenes, actions, environments, physical behavior, and real-world context — producing results that feel intentional rather than random. Gemini Omni Flash is the first model in the Omni family, built for practical video creation and editing workflows where users can transform footage, guide results with references, and refine scenes through natural language conversation.

About Gemini Omni
Core Features Overview

Key Capabilities

Multimodal input, conversational editing, style transformation, and real-world knowledge — all in one model

Multi-Turn Conversational Editing

01

Edit videos through natural language conversation — each edit builds on the previous

Gemini Omni introduces a fundamentally different approach to video editing. Instead of starting from scratch with each generation, you can refine your video through a series of natural language instructions. Change the background, adjust the action, replace objects, shift the camera angle, or add visual effects — all while keeping the rest of the video stable. This conversational workflow means you can iterate toward your vision step by step, just like editing a document with tracked changes.

Prompt

Edit over multiple turns with consistency — change camera angle while maintaining scene coherence across sequential modifications

Multi-turn editing preserves scene coherence across sequential modifications

First establish the scene with a person in a room, then change the lighting to golden hour, then add rain on the window — each edit builds on the last

Sequential environment changes demonstrate conversational refinement

Output (Example)
2

Real-Time Style Transformation

02

Transform video aesthetics into metal, sketch, puppet, hologram, voxel, and more

Gemini Omni can transform the visual style of any input video while preserving the underlying motion, structure, and scene composition. Describe the target aesthetic — metallic surfaces, hand-drawn sketches, felt puppets, holographic projections, voxel art — and the model applies the transformation coherently across every frame. The original camera movement, character actions, and spatial relationships remain intact, creating a seamless style transfer that goes far beyond simple filters.

Prompt

When the person touches the mirror, make the mirror ripple beautifully like liquid, and the person's arm turns into reflective mirror material

Style transformation preserves motion while completely changing visual aesthetics to metallic

When the person touches the mirror, the entire environment turns into 3D voxel art with blocky geometric shapes

Complete environment transformation to voxel art while preserving spatial structure

Output (Example)
2

True Multimodal Input

03

Accept text, images, video, and audio as creative references in a single generation

Unlike models that only accept text or a single image, Gemini Omni can process multiple input types simultaneously. Provide text for direction, images for visual reference, video for motion guidance, and audio for speech or sound synchronization. The model synthesizes all inputs into a single cohesive video output. This makes it practical for real creative workflows where inspiration comes from multiple sources — a storyboard sketch, a reference clip, a voice recording, and a written description can all contribute to the final result.

Prompt

Add harp sounds synchronized to when I touch each fern leaf. Change the leaf structure to bioluminescent plant life with fireflies flying around

Combining video input with text instructions and audio reference for synchronized output

Visualize protein folding process using real-world scientific knowledge, rendered in claymation style with accurate molecular behavior

Real-world knowledge applied to scientific visualization with creative style

Output (Example)
2
FAQ

Frequently Asked Questions

Gemini Omni FAQ

What Creators Say About Gemini Omni

What Creators Say About Gemini Omni

Read All Stories
J
Jordan MitchellThe multi-turn editing on Nano Banana Pro changed how I approach video production. I can direct a scene through multiple rounds of refinement without losing continuity — it's the closest thing to having an AI cinematographer on set.
S
Samantha ColeWe use Gemini Omni's style transformation to repurpose a single shoot into dozens of variations — metal, sketch, hologram — all while keeping the original motion intact. Our content output tripled without additional filming.
D
Derek HuangThe real-world knowledge sets Gemini Omni apart. When I asked for a protein folding visualization, the molecular behavior was scientifically accurate — not just visually impressive. That's a first for any AI video tool I've used.

Explore More AI Video Models

Start Creating with Gemini Omni

Experience the power of Gemini Omni — free online