gemini-omni
Compare original and translation side by side
🇺🇸
Original
English🇨🇳
Translation
ChineseGemini Omni Flash (Google DeepMind)
Gemini Omni Flash (Google DeepMind)
Gemini Omni is Google DeepMind's video generation and editing model family, announced at I/O 2026. The first model, Gemini Omni Flash (, developer access since June 30, 2026), generates 3-10 second clips at 720p/24fps with synthesized audio via the Gemini Interactions API. Its differentiator in the OpenMontage fleet is stateful conversational editing: each generation returns an , and a follow-up call with edits that video in place — no other wrapped provider can refine a clip without regenerating it.
gemini-omni-flash-previewinteraction_idprevious_interaction_idOpenMontage wraps it as (native Gemini API, no gateway). It shares / with and — one key, three capabilities. Paid tier only: ~$0.10 per second of output video (billed as 5,792 output tokens/sec at $17.50/1M).
gemini_omni_videoGOOGLE_API_KEYGEMINI_API_KEYgoogle_imagengoogle_ttsGemini Omni是Google DeepMind于2026年I/O大会上发布的视频生成及编辑模型系列。首款模型Gemini Omni Flash(,自2026年6月30日起开放开发者访问)可通过Gemini Interactions API生成3-10秒、720p/24fps的剪辑,并附带合成音频。它在OpenMontage工具集中的差异化优势是有状态对话式编辑:每次生成都会返回一个,后续调用时传入即可原地编辑该视频——其他封装提供商无法在不重新生成的情况下优化剪辑。
gemini-omni-flash-previewinteraction_idprevious_interaction_idOpenMontage将其封装为(原生Gemini API,无需网关)。它与和共享/——一个密钥即可实现三项功能。仅支持付费 tier:每输出一秒视频约0.10美元(按每秒5792个输出令牌计费,单价为17.50美元/百万令牌)。
gemini_omni_videogoogle_imagengoogle_ttsGOOGLE_API_KEYGEMINI_API_KEYWhen to pick it (and when not)
适用场景与不适用场景
| Use it for | Prefer another provider for |
|---|---|
| Iterative refinement — generate, review, then edit the same clip in layers | One-shot cinematic hero clips (→ Seedance 2.0, see |
| Editing an existing/uploaded clip (restyle, add/remove objects, change text) | Clips longer than 10s or above 720p |
| On-screen rendered text and word-by-word text beats | Seed-reproducible generations (no seed support) |
| Reference-image-bound subjects/styles via prompt tags | First/last-frame interpolation (→ |
| Timecode-scheduled multi-beat clips from one prompt | Non-English narration (English only fully supported) |
Route through for generation operations. Editing () is a direct-tool operation — call from the registry, because the multi-turn interaction state lives outside the selector's model.
video_selectoredit_videogemini_omni_video| 适用场景 | 更适合其他提供商的场景 |
|---|---|
| 迭代优化——生成、审阅后分层编辑同一剪辑 | 一次性生成电影级主剪辑(→ Seedance 2.0,详见 |
| 编辑现有/已上传的剪辑(重新风格化、添加/移除对象、修改文本) | 时长超过10秒或分辨率高于720p的剪辑 |
| 屏幕渲染文本及逐字文本节拍 | 可通过种子复现的生成内容(不支持种子功能) |
| 通过提示标签绑定参考图像的主体/风格 | 首尾帧插值(→ |
| 通过单个提示生成带时间码调度的多节拍剪辑 | 非英语旁白(仅全面支持英语) |
生成操作需通过路由。编辑()是直接工具操作——需从注册表调用,因为多轮交互状态存储在选择器模型之外。
video_selectoredit_videogemini_omni_videoGeneration prompting
生成提示
Describe scene + camera + lighting + motion + audio. Official example:
Continuous, unbroken handheld shot of a fluffy tabby cat sitting on a sunny windowsill, looking out into a leafy garden. The cat's tail twitches slowly, and its ears rotate slightly toward ambient noises. Sunbeams illuminate dust motes in the air.
- Force a single shot explicitly: "In a single continuous shot," / "No scene cuts." Otherwise the model may cut between scenes.
- Negatives go in prose — there is no parameter: "No dialogue," "No extra sound effects."
negative_prompt - No sampler controls: system instructions, temperature, top_p, and seeds are all unsupported. The prompt is the only lever.
- Meta-prompt for quality: "Consider micro-detail, expression and timing to create a very rich, detailed but entirely natural scene."
描述场景+镜头+灯光+运动+音频。官方示例:
手持设备拍摄的连续无间断镜头,展示一只蓬松的虎斑猫坐在洒满阳光的窗台上,望向绿植繁茂的花园。猫的尾巴缓慢抽动,耳朵随环境噪音轻微转动。阳光照亮空气中的尘埃。
- 明确强制单镜头:“采用单个连续镜头拍摄,” / “无场景切换。”否则模型可能会切换场景。
- 否定描述用散文形式——没有参数:“无对话,”“无额外音效。”
negative_prompt - 无采样器控制:不支持系统指令、temperature、top_p和种子。提示是唯一的调节手段。
- 提升质量的元提示:“考虑微细节、表情和时机,打造一个丰富、细致且完全自然的场景。”
Timecode syntax
时间码语法
Schedule beats with bracketed ranges or natural language — this maps directly onto OpenMontage scene-plan timings:
[0-3s] A person is walking [3-6s] They stop and turn around"After 3 seconds, a woman enters the scene." / "At 5s the chorus starts in the background audio."
用括号范围或自然语言调度节拍——这直接映射到OpenMontage的场景计划时间:
[0-3s] 一个人正在行走 [3-6s] 他们停下并转身“3秒后,一名女子进入画面。” / “5秒时,背景音频开始播放副歌。”
Audio and on-screen text
音频与屏幕文本
Audio is synthesized automatically; direct it in the prompt: "Include calm background music," "The audio is a low tinny radio broadcast in the background." Rendered text works and can be timed:
One word on the screen at a time: 'did, you, know, that, Omni, can, do, awesome, text?' Each word appears for 1s.
音频会自动合成;可在提示中指定:“加入舒缓的背景音乐,”“音频为背景中微弱的收音机广播声。”渲染文本功能可用且可定时:
屏幕上一次显示一个单词:'did, you, know, that, Omni, can, do, awesome, text?' 每个单词显示1秒。
Reference images (<FIRST_FRAME>
/ <IMAGE_REF_N>
tags)
<FIRST_FRAME><IMAGE_REF_N>参考图像(<FIRST_FRAME>
/ <IMAGE_REF_N>
标签)
<FIRST_FRAME><IMAGE_REF_N>Pass local images via (they are sent in order), then bind them to roles inside the prompt with tags. indexes from 0 in the order supplied:
reference_image_paths<IMAGE_REF_N>in the style of <IMAGE_REF_0> a woman <IMAGE_REF_1> is walking[0-3s] A studio fashion sequence. Starting with woman <IMAGE_REF_0>, she is
holding <IMAGE_REF_1> [3-6s] Then we see the man <IMAGE_REF_2> holding <IMAGE_REF_3>- makes an image the opening frame:
<FIRST_FRAME>.<FIRST_FRAME> a woman is walking - Use high-resolution images; describe the intended motion specifically rather than "make it move."
- Say what each image is (product / character / style / background reference) — the model decides usage from context.
通过传入本地图像(按顺序发送),然后在提示内用标签将它们绑定到角色。按提供顺序从0开始索引:
reference_image_paths<IMAGE_REF_N>采用<IMAGE_REF_0>的风格,一名女子<IMAGE_REF_1>正在行走[0-3s] 一段工作室时尚序列。以女子<IMAGE_REF_0>开场,她手持<IMAGE_REF_1> [3-6s] 随后出现男子<IMAGE_REF_2>,他手持<IMAGE_REF_3>- 将图像设为开场帧:
<FIRST_FRAME>。<FIRST_FRAME> 一名女子正在行走 - 使用高分辨率图像;具体描述预期动作,而非“让它动起来”。
- 说明每张图像的用途(产品/角色/风格/背景参考)——模型会根据上下文决定如何使用。
Conversational editing (the differentiator)
对话式编辑(差异化优势)
Editing prompts are the opposite of generation prompts: short and surgical. Overly descriptive edit prompts cause unintended changes.
- Generate the base clip (subject + scene + motion). The tool returns in its result data.
interaction_id - Pass it back as with
previous_interaction_idand describe only the delta.operation="edit_video" - Append "Keep everything else the same." to pin unmentioned elements.
- Refine in layers — one turn for lighting, one for camera, one for action, one for audio.
Official good/bad pairs:
| Avoid | Instead |
|---|---|
| "In the video of the man sitting on the sofa, please add a small black cat..." | "Add a cat that jumps onto his lap, he begins to pet it. Keep everything else the same." |
| "Please remove the cell phone... and fill in the background so it looks like..." | "Make the phone invisible. Keep everything else the same." |
Other working edit prompts: "Make this video anime" / "Put a fashionable hat on this person" / "Change the lighting to be more dramatic" / "Change the text on the sign to say 'Omni Flash'".
Gotcha — : editing via only works if the prior call kept the interaction server-side ( defaults to true in ). Set only for one-shot generations you will never edit.
storeprevious_interaction_idstoregemini_omni_videostore=falseEditing uploaded videos: pass instead of ; the tool uploads it via the Files API. Unavailable in the EEA, Switzerland, and the UK (editing generated videos works everywhere).
input_video_pathprevious_interaction_id**编辑提示与生成提示相反:简短且精准。**过于描述性的编辑提示会导致意外更改。
- 生成基础剪辑(主体+场景+动作)。工具会在结果数据中返回。
interaction_id - 将其作为传入,设置
previous_interaction_id,并仅描述变更部分。operation="edit_video" - 追加**“其他内容保持不变。”**以固定未提及的元素。
- 分层优化——一轮调整灯光,一轮调整镜头,一轮调整动作,一轮调整音频。
官方正反示例对比:
| 避免写法 | 推荐写法 |
|---|---|
| “在男子坐在沙发上的视频中,请添加一只小黑猫...” | “添加一只跳到他腿上的猫,他开始抚摸它。其他内容保持不变。” |
| “请移除手机...并填充背景使其看起来像...” | “让手机消失。其他内容保持不变。” |
其他可行的编辑提示:“将此视频改为动漫风格” / “给这个人戴一顶时尚的帽子” / “将灯光调整得更具戏剧性” / “将标识上的文字改为'Omni Flash'”。
注意——参数:通过进行编辑仅当上一次调用在服务器端保留了交互状态时才有效(中默认值为true)。仅当你永远不会编辑一次性生成的内容时,才设置。
storeprevious_interaction_idgemini_omni_videostorestore=false编辑上传的视频:传入而非;工具会通过Files API上传视频。该功能在欧洲经济区、瑞士和英国不可用(编辑生成的视频在所有地区都可用)。
input_video_pathprevious_interaction_idHard limitations (preview)
硬性限制(预览版)
- Output: 3-10s, 720p, 24fps, MP4 with audio; aspect ratio or
16:9. All output carries an invisible SynthID watermark.9:16 - No seed, negative prompt, temperature, top_p, or system instructions.
- No video extension or first/last-frame interpolation; no voice editing.
- Audio reference inputs unsupported. Video references ≤3s are accepted by the schema but not processed correctly — don't rely on them.
- Multi-video prompting unsupported; may degrade output.
- English fully supported; other languages untested.
- Images of minors (EEA/CH/UK) and certain recognizable people are blocked for upload/editing.
- 输出:3-10秒,720p,24fps,带音频的MP4;宽高比为或
16:9。所有输出都带有不可见的SynthID水印。9:16 - 不支持种子、否定提示、temperature、top_p或系统指令。
- 不支持视频扩展或首尾帧插值;不支持语音编辑。
- 不支持音频参考输入。视频参考≤3秒虽能通过 schema 验证,但无法正确处理——请勿依赖此功能。
- 不支持多视频提示;可能会降低输出质量。
- 全面支持英语;其他语言未测试。
- 禁止上传/编辑未成年人图像(欧洲经济区/瑞士/英国)和某些可识别人物的图像。
Sources
参考资料
- Generation & editing guide: https://ai.google.dev/gemini-api/docs/omni
- Model card: https://ai.google.dev/gemini-api/docs/models/gemini-omni-flash
- Pricing: https://ai.google.dev/gemini-api/docs/pricing
- Announcement: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni/