← Files Creative ClawARCHIVED FILE

skills/creativeclaw-plan-video/references/video/gemini-omni.md

9.85 KB · Oct 6, 2026 · 18:04 UTC

↓ Download file

See the change to this file →

# Creative Claw : Gemini Omni

Read [input modes and reference production](reference-production.md) before preparing media and [Review/Auto handling](review.md) before submission. These shared contracts take precedence over any recipe below. Never combine literal frames with reference arrays on standard routes. Load only this selected model guide, not every guide in the package.

Use the outcome skill's execution guidance for authorization, imports and job recovery.

Use `video/gemini-omni-flash` as Creative Claw's default video model. It is the primary recommendation for fast multimodal video generation and source-video editing with native audio.

## Core workflow

1. Define one clip: purpose, duration, aspect ratio, subject, action, camera, audio, and protected visual details.
2. Search for reusable assets and import any ChatGPT attachments into Creative Claw before passing them to URL fields.
3. Follow the reference-first pipeline in [reference production](reference-production.md): one clean keyframe per shot, made at 16:9 or 9:16 (Omni's only ratios), then its per-shot mode rule. Skip keyframes only on an explicit direct-generation request, a ready shot image, or a source edit.
4. Call `get_model_params({ model: "video/gemini-omni-flash" })` before generation when not already fetched for this task. Treat its current schema as authoritative.
5. Choose `resolution` from the current schema when output size matters. The direct Google route supports `360p`, `720p` (default), `1080p`, and `4k`; 1080p and 4K are upscaled outputs. Explain the selected duration, ratio, resolution, references, and audio plan when useful. Reuse existing authorization, including an explicitly requested batch; ask only when a material choice is unresolved or the proposed work expands the requested scope.
6. Call `generate_video` with `model: "video/gemini-omni-flash"`.
7. Let the inline viewer monitor the job. Call `check_job` only when a later tool needs the completed URL or no viewer is monitoring.
8. Inspect motion, identity, physics, framing, audio, dialogue, and text artifacts before describing the clip as complete.

## Choose the input mode

| Intent | Inputs | Prompt emphasis |
| --- | --- | --- |
| Direct text-to-video (explicit request only) | `prompt` only | Describe the complete visible scene and sound. |
| Animate a still | `image_url` | Describe what begins moving after the supplied first frame. |
| Reference-guided video | `image_urls` | Bind every reference to a role with `<IMAGE_REF_N>`. |
| Edit a source clip | one item in `video_urls` | Give one short change followed by “Keep everything else the same.” |

Do not combine modes casually. Use `image_url` when an image must be the literal first frame. Use `image_urls` when images should guide identity, product appearance, wardrobe, environment, or style without becoming the opening frame.

## Keyframes and continuity

For ads, branded content, character work, and multi-clip sequences:

1. Break the concept into short shots with one main action each.
2. Make each keyframe separately from the shared anchors (same image model, 16:9 or 9:16). Never pass a labeled grid, contact sheet, panels, captions, or prompt text to the video model.
3. People, several subjects or big motion: pass `image_urls` = [keyframe, identity anchor, product anchor] and cite them with `<IMAGE_REF_N>`. An exact opening such as a product hero: pass the keyframe as `image_url` with no `image_urls` or `character_id`.
4. For the next clip, return to the same anchors and look line. The last frame of clip N is only an extra composition cue for clip N+1, never a replacement for the anchors.

This reduces visual drift and makes revisions local to one shot.

## First and last frames

### Start frame

Pass the approved opening image as `image_url`. Treat it as frame zero. Describe the motion that follows rather than restating every visible detail.

Good:

```text
The woman turns toward the window as rain begins to trace the glass. Her coat,
face, and the room remain unchanged. Slow dolly-in, one continuous shot.
```

Weak:

```text
A woman wearing a red coat stands in a room by a window.
```

The weak version invites the model to reinterpret the already-approved frame.

### End frame

Google's underlying Omni model supports first-to-last interpolation, but Creative Claw's exposed fields can change. Pass `last_frame_url` only when `get_model_params` returns an end-frame field for the current route. Otherwise use Seedance 2.5 or MiniMax H3 Max for controlled first-to-last generation.

When supported, use two frames with the same ratio, subject identity, and plausible spatial continuity. Describe the transition, not two separate scenes:

```text
Begin exactly from the first frame. In one continuous orbital move, the camera
travels clockwise while the product lid opens and blue light grows from inside.
Arrive exactly at the supplied final frame. No cuts or teleporting objects.
```

## Reference syntax

Pass reference images in `image_urls`. Cite them with zero-based tokens:

```text
<IMAGE_REF_0> = exact product identity and geometry.
<IMAGE_REF_1> = lighting and material reference only.
<IMAGE_REF_2> = wardrobe reference for the actor.
```

Then direct the shot:

```text
Preserve the product from <IMAGE_REF_0> exactly. Borrow only the cool rim
lighting and glossy black environment from <IMAGE_REF_1>. The actor wears the
outfit from <IMAGE_REF_2>. She places the product on the pedestal as the camera
makes a slow 30-degree orbit. No cuts, no on-screen text, no geometry changes.
```

Preserve exact quoted copy, reference labels, dialogue, timecodes, colors, and approved layout or edit constraints in the prompt.

## Prompt formula

Order information by what the model must protect:

```text
Purpose: [ad, cinematic insert, product reveal, social clip].
References: [token → exact role; attributes to preserve].
Scene: [subject, environment, composition].
Action: [one filmable action with visible motion].
Camera: [framing + one intentional movement].
Look: [lighting, lens/medium, palette, texture].
Audio: [dialogue, ambience, effects, music, or explicit silence].
Timing: [optional natural beats or time ranges].
Constraints: [identity, product geometry, no cuts/text/subtitles/watermarks].
```

For reusable B-roll, use one subject, one action, and one camera idea. Say “single continuous shot, no scene cuts” when cuts would make the clip unusable.

## Timing and audio

Use the current runtime range discovered by `get_model_params`; the established Creative Claw route commonly supports 3-10 seconds, `16:9` or `9:16`, and `360p`, `720p`, `1080p`, or `4k` output resolution.

Natural beats work well:

```text
[0-3s] Slow push toward the unopened bottle.
[3-6s] The cap lifts and cold vapor spills across the table.
[6-8s] Hold on the clean hero angle.
```

Omni accepts no audio input (`audio_urls`). Its dialogue is native: the model invents the voice, and the voice changes between clips. For an exact or recurring voice, route the shot as described in [voice in video](voice-in-video.md) instead of making speech first for Omni.

Describe native audio explicitly:

- “No dialogue. Quiet studio ambience and a soft mechanical click.”
- “Dialogue, exact line: ‘Ready when you are.’ Natural room tone, no music.”
- “At five seconds, the percussion enters as the product locks into place.”
- “Generate a silent clip; no music, speech, ambience, or sound effects.”

Keep spoken copy short enough to fit naturally. Quote exact lines and preserve the approved wording.

## Source-video editing

Pass one source clip in `video_urls`. Use a concise delta:

```text
Replace the overcast sky with a warm sunset. Keep the people, timing, camera
motion, buildings, and every other detail unchanged.
```

Avoid redescribing the source. A long prompt increases unintended changes. Omni editing is best for one clear transformation per pass.

Verify the source duration before submission. The current Creative Claw Omni edit route accepts at most 10 seconds of uploaded source video. If the source is longer, do not submit it unchanged and do not regenerate the whole timeline in segments by default. Trim and edit only the interval that must change, then merge it between untouched source spans. For a request to keep the original and add a new ending, use `video/minimax-h3-max-extend` and keep the original untouched. Use Seedance 2.5 `extend` only for heavy references or a long continuation.

## Prompt examples

Product reveal:

```text
Premium ten-second product film. The matte-black headphones remain identical to
the approved first frame. They rotate slowly above a dark reflective plinth as a
thin ribbon of amber light travels across the ear cups. Macro commercial lens,
slow clockwise orbit, deep black background, crisp highlights. Sound design:
low electronic pulse and a soft magnetic click. Single continuous shot. No
people, text, captions, logo changes, or extra objects.
```

Character scene with references:

```text
<IMAGE_REF_0> is the exact character identity and face. <IMAGE_REF_1> is the
exact wardrobe. Preserve both. In a rain-soaked train station, she looks over
her shoulder and takes one step toward the arriving train. Medium close-up,
slow handheld push-in, cyan and amber practical lights. Natural rain, distant
train brakes, no dialogue. One shot; no face drift, wardrobe changes, text, or
extra people near camera.
```

Source edit:

```text
Transform the source video into a premium hand-painted anime look. Preserve the
exact motion, timing, people, composition, and camera path. Keep everything else
the same. No subtitles or added text.
```

## Quality and feedback

Reject static-subject pans when subject motion was requested, identity drift, product deformation, unexpected cuts, lip-sync mismatch, duplicate limbs, embedded text, and audio contradicting the prompt. Revise one failure at a time.

Send feedback only when the user requests it; include concrete model-specific failures without exposing private media.

SHA-256: 3cf52bebbd729671ba10fc3f607b866fa15dae2f6f379775824c1e128488504b