← Files Creative ClawARCHIVED FILE

skills/creativeclaw/references/workflows/video-gen.md

6.67 KB · Oct 5, 2026 · 12:04 UTC

↓ Download file

See the change to this file →

# Video generation and transformation

Match the model to the shot, reference structure, duration, resolution, audio needs, and transformation operation.

## Workflow

1. Define the deliverable: single shot or sequence, duration, ratio, subject, action, camera, audio, and continuity requirements.
2. Follow the reference-first pipeline in [reference production](../video/reference-production.md): reuse anchors (`search_assets`, `list_characters`, `get_theme`), write one look line, and make one keyframe per shot with `generate_image` at the video's aspect ratio. Skip keyframes only when the user asks for direct text-to-video, supplies a ready shot image, or is editing footage.
3. Pick one input mode per shot: people, several subjects or big motion → `image_urls` (keyframe + 2–3 anchors); exact opening → `image_url` = keyframe, no `image_urls` or `character_id`. For speech, pick a path from [voice in video](../video/voice-in-video.md).
4. Use `list_models({ category: "video" })` when discovery is needed and `get_model_params` for missing settings on the selected model; reuse schemas already fetched in this task.
5. State consequential settings briefly and proceed within existing authorization. Estimate with `operation: "video"` only for user-requested cost/budget help; do not add a routine approval question.
6. Call `generate_video`. Preserve literal reference tokens and timecodes.
7. Resolve the job only when needed, inspect the result, and report false motion, identity drift, broken physics, unwanted cuts, text artifacts, or bad audio. Deliver the result; inspection does not authorize another generation. Ask before another take unless the user explicitly requested that additional attempt.
8. Use focused processing tools for trim, scale, subtitles, frames, merging, isolation, or upscaling.

## Model picker

Runtime discovery is authoritative. Start here:

| Need | Model | Current specialty |
| --- | --- | --- |
| General generation, references, or source edit | `video/gemini-omni-flash` | Default; fast multimodal 3–10s generation/edit with native audio. |
| Premium cinematic, reference-rich or long video | `video/seedance-2.5` | 4–30s, native audio, optional first/last frames, and large mixed-reference sets. A 480p request is a usable draft; `extras.draft_job_id` finalizes the same take at 1080p (720p starts a new take). |
| Fast cinematic generation with strong adherence | `video/minimax-h3-max` | 5–15s, 480p/768p/1080p, native audio, optional first/last frames, and multimodal references. |
| Cheap, fast drafts | `video/minimax-h3-max-turbo` | Present as H3 Max Fast; text or start-frame drafts. Reference requests use the shared H3 Max reference route. |
| Native-audio single pass, or a document or webpage brief | `video/wan-3.0` | 2–30s, native audio, first/last frames, mixed references, or one public document or webpage as the brief. |

Recommend these models first. Use another runtime-listed model, including Seedance Mini, only when the user explicitly requests it or the five recommended choices cannot perform the operation.

For source-video work, use the edit-specific ranking in `edit-video.md` instead of the table above: `video/minimax-h3-max-extend` to extend a clip (Seedance 2.5 `extend` only for heavy references or a long continuation), `video/minimax-h3-max-insert` to replace an interval with new footage, Gemini Omni for targeted edits up to 10 seconds, and Seedance 2.5 for 4–30 second edits. Preserve untouched source spans and generate only the interval or continuation that needs new pixels. Never recommend or proactively route to an LTX or DreamActor model.

## Reference rules

- There is no universal reference-count requirement or cap. A request may use zero, one, or many references according to the selected model and operation.
- `image_url` is only the literal start/source image for image-to-video. If an image is a soft reference and should not become frame zero, use `image_urls`, even for one image.
- `last_frame_url` or the model's discovered boundary-frame field controls the end only on compatible models.
- `image_urls`, `video_urls`, and `audio_urls` are top-level video-tool reference arrays when supported. Do not move them into `extras` unless `get_model_params` explicitly says so.
- `character_id` appends the saved image to `image_urls` as a reference, never a start frame. Use it in reference mode; omit it with `image_url`/`last_frame_url` (the server rejects that mix).
- Use the exact token syntax required by the model. Examples include Seedance `@Image1`/`@Video1`/`@Audio1`, Kling `@Element1`, HappyHorse `@character1`, and Grok `<IMAGE_0>`. Verify the current schema and pass the prompt verbatim.
- Do not feed a labeled storyboard grid to a video model. Use clean, full-bleed generation frames.

## Model-specific guardrails

### Seedance 2.5

- Duration: whole seconds from 4 through 30, or `auto`.
- References: up to 30 images, 10 videos, 10 audio clips, 50 total.
- Audio references require at least one image or video reference.
- Use the named `@` reference tokens and disable prompt rewriting.

### MiniMax H3 Max

- Duration: 5–15s; the current hosted route accepts prompts up to 50,000 characters.
- Reference limits are model- and route-specific. Inspect `get_model_params({ model: "video/minimax-h3-max" })` instead of applying a global reference rule.
- Audio references require at least one image or video reference.

## Prompt structure

```text
[0s–Xs] Subject, action, environment, and framing.
Camera: one intentional movement.
Look: lighting, lens/medium, palette, texture.
Audio: dialogue in quotes, ambience, effects, music direction, or explicit silence.
Continuity: identity, wardrobe, product geometry, and protected references.
Constraints: single shot or named cuts; no unwanted text, captions, or watermark.
```

Give each short shot one main action and one camera idea. Use time blocks for multiple beats. If the output only pans across a still when subject motion was required, propose revised action verbs or a model better suited to physical motion. Obtain explicit authorization for another generation before executing the proposal.

## Multi-clip strategies

- **Parallel montage:** independent shots, generated together after approval.
- **Shared anchors (default):** every shot's keyframe and video request reuse the same Character/product/style anchors and look line.
- **Serial continuity:** a previous clip's last frame (`extract_frames`) is only an extra composition cue, never a replacement for the anchors.
- **Single-call multi-shot:** use a runtime model with native multi-shot structure only when its discovered schema fits the sequence.

Use `merge_media` only after individual clips are approved. Use `generate_speech` before timing narration-driven shots.

SHA-256: 4a4587d8e12460fac58b6902537d84915174e36dc65ac2b16bcdf8f2949d0ea1