Loading...
Loading...
Generate and conversationally edit short videos with Google Gemini Omni Flash (`gemini-omni-flash-preview`). Use when: (1) iterating on a clip with natural-language edits instead of regenerating ("make the phone invisible, keep everything else the same"), (2) generating 3-10s 720p clips with synthesized audio, rendered on-screen text, or timecoded beats, (3) binding reference images to roles with <FIRST_FRAME>/<IMAGE_REF_N> prompt tags, (4) editing an existing uploaded video. Accessed via the `gemini_omni_video` tool using the project's GEMINI_API_KEY/GOOGLE_API_KEY — the same key as Imagen and Google TTS.
npx skill4agent add calesthio/openmontage gemini-omnigemini-omni-flash-previewinteraction_idprevious_interaction_idgemini_omni_videoGOOGLE_API_KEYGEMINI_API_KEYgoogle_imagengoogle_tts| Use it for | Prefer another provider for |
|---|---|
| Iterative refinement — generate, review, then edit the same clip in layers | One-shot cinematic hero clips (→ Seedance 2.0, see |
| Editing an existing/uploaded clip (restyle, add/remove objects, change text) | Clips longer than 10s or above 720p |
| On-screen rendered text and word-by-word text beats | Seed-reproducible generations (no seed support) |
| Reference-image-bound subjects/styles via prompt tags | First/last-frame interpolation (→ |
| Timecode-scheduled multi-beat clips from one prompt | Non-English narration (English only fully supported) |
video_selectoredit_videogemini_omni_videoContinuous, unbroken handheld shot of a fluffy tabby cat sitting on a sunny windowsill, looking out into a leafy garden. The cat's tail twitches slowly, and its ears rotate slightly toward ambient noises. Sunbeams illuminate dust motes in the air.
negative_prompt[0-3s] A person is walking [3-6s] They stop and turn around"After 3 seconds, a woman enters the scene." / "At 5s the chorus starts in the background audio."
One word on the screen at a time: 'did, you, know, that, Omni, can, do, awesome, text?' Each word appears for 1s.
<FIRST_FRAME><IMAGE_REF_N>reference_image_paths<IMAGE_REF_N>in the style of <IMAGE_REF_0> a woman <IMAGE_REF_1> is walking[0-3s] A studio fashion sequence. Starting with woman <IMAGE_REF_0>, she is
holding <IMAGE_REF_1> [3-6s] Then we see the man <IMAGE_REF_2> holding <IMAGE_REF_3><FIRST_FRAME><FIRST_FRAME> a woman is walkinginteraction_idprevious_interaction_idoperation="edit_video"| Avoid | Instead |
|---|---|
| "In the video of the man sitting on the sofa, please add a small black cat..." | "Add a cat that jumps onto his lap, he begins to pet it. Keep everything else the same." |
| "Please remove the cell phone... and fill in the background so it looks like..." | "Make the phone invisible. Keep everything else the same." |
storeprevious_interaction_idstoregemini_omni_videostore=falseinput_video_pathprevious_interaction_id16:99:16