← Creative ClawCONTENT HISTORY

Update to Creative Claw

Snapshot Oct 3, 2026 · 06:03 UTC · version 5.3.0

Collection source: downloaded plugin package. These snapshots do not have a confirmed matching collection source. Differences in file lists alone do not establish changes to the package.

WHAT CHANGED · RULE-BASED ANALYSIS

Instructions updated for creativeclaw-generate-video

Instruction wording changed from “[reference production](references/video/reference-production.md) before preparing media and ” to “”. 28 additional added or edited lines are in the evidence.

Observed in instructions or declared skills. Runtime behavior has not been tested.

Skill instructions

Before

[reference production](references/video/reference-production.md) before preparing media and [Review/Auto handling](references/video/review.md) before submission. Create one controlled video clip from text, a start frame, an optional end ...

After

[Review/Auto handling](references/video/review.md) before submission. Create one controlled video clip from text, a start frame, an optional end frame, or other model-supported references. Use the planning, UGC, or film skills when the d...

Supporting files

Before

[{"relative_path":"agents/openai.yaml","size_in_bytes":539},{"relative_path":"references/images/gpt-image-2.md","size_in_bytes":10996},{"relative_path":"references/images/index.md","size_in_bytes":456},{"relative_path":"references/images...

After

[{"relative_path":"agents/openai.yaml","size_in_bytes":539},{"relative_path":"references/images/gpt-image-2.md","size_in_bytes":11212},{"relative_path":"references/images/index.md","size_in_bytes":456},{"relative_path":"references/images...

Compare saved observations

Download comparison JSON
Full technical diff · 2 changed fields

changed /included_files

BEFORE
[
  {
    "relative_path": "agents/openai.yaml",
    "size_in_bytes": 539
  },
  {
    "relative_path": "references/images/gpt-image-2.md",
    "size_in_bytes": 10996
  },
  {
    "relative_path": "references/images/index.md",
    "size_in_bytes": 456
  },
  {
    "relative_path": "references/images/nano-banana-2.md",
    "size_in_bytes": 8817
  },
  {
    "relative_path": "references/images/nano-banana-pro.md",
    "size_in_bytes": 8779
  },
  {
    "relative_path": "references/images/seedream-5-pro.md",
    "size_in_bytes": 9556
  },
  {
    "relative_path": "references/job-recovery.md",
    "size_in_bytes": 3786
  },
  {
    "relative_path": "references/media-assembly.md",
    "size_in_bytes": 3876
  },
  {
    "relative_path": "references/platform-upload.md",
    "size_in_bytes": 1416
  },
  {
    "relative_path": "references/video/gemini-omni.md",
    "size_in_bytes": 9538
  },
  {
    "relative_path": "references/video/index.md",
    "size_in_bytes": 1823
  },
  {
    "relative_path": "references/video/minimax-h3-max.md",
    "size_in_bytes": 9097
  },
  {
    "relative_path": "references/video/other-models.md",
    "size_in_bytes": 4257
  },
  {
    "relative_path": "references/video/recipe-character-scene.md",
    "size_in_bytes": 2966
  },
  {
    "relative_path": "references/video/recipe-product-ad.md",
    "size_in_bytes": 2998
  },
  {
    "relative_path": "references/video/recipe-source-edit.md",
    "size_in_bytes": 2620
  },
  {
    "relative_path": "references/video/reference-production.md",
    "size_in_bytes": 7381
  },
  {
    "relative_path": "references/video/review.md",
    "size_in_bytes": 1909
  },
  {
    "relative_path": "references/video/seedance-2-5.md",
    "size_in_bytes": 12711
  },
  {
    "relative_path": "references/video/seedance-mini.md",
    "size_in_bytes": 1480
  },
  {
    "relative_path": "references/video/ugc-paths.md",
    "size_in_bytes": 2923
  },
  {
    "relative_path": "references/video/wan-3.md",
    "size_in_bytes": 8665
  },
  {
    "relative_path": "references/workflow-basics.md",
    "size_in_bytes": 7159
  }
]
AFTER
[
  {
    "relative_path": "agents/openai.yaml",
    "size_in_bytes": 539
  },
  {
    "relative_path": "references/images/gpt-image-2.md",
    "size_in_bytes": 11212
  },
  {
    "relative_path": "references/images/index.md",
    "size_in_bytes": 456
  },
  {
    "relative_path": "references/images/nano-banana-2.md",
    "size_in_bytes": 8855
  },
  {
    "relative_path": "references/images/nano-banana-pro.md",
    "size_in_bytes": 8824
  },
  {
    "relative_path": "references/images/seedream-5-pro.md",
    "size_in_bytes": 9556
  },
  {
    "relative_path": "references/job-recovery.md",
    "size_in_bytes": 3786
  },
  {
    "relative_path": "references/media-assembly.md",
    "size_in_bytes": 8654
  },
  {
    "relative_path": "references/platform-upload.md",
    "size_in_bytes": 1416
  },
  {
    "relative_path": "references/video/gemini-omni.md",
    "size_in_bytes": 10085
  },
  {
    "relative_path": "references/video/index.md",
    "size_in_bytes": 2176
  },
  {
    "relative_path": "references/video/minimax-h3-max.md",
    "size_in_bytes": 11220
  },
  {
    "relative_path": "references/video/other-models.md",
    "size_in_bytes": 4653
  },
  {
    "relative_path": "references/video/recipe-character-scene.md",
    "size_in_bytes": 2941
  },
  {
    "relative_path": "references/video/recipe-product-ad.md",
    "size_in_bytes": 3265
  },
  {
    "relative_path": "references/video/recipe-source-edit.md",
    "size_in_bytes": 3357
  },
  {
    "relative_path": "references/video/reference-production.md",
    "size_in_bytes": 5865
  },
  {
    "relative_path": "references/video/review.md",
    "size_in_bytes": 1909
  },
  {
    "relative_path": "references/video/seedance-2-5.md",
    "size_in_bytes": 15084
  },
  {
    "relative_path": "references/video/seedance-mini.md",
    "size_in_bytes": 1579
  },
  {
    "relative_path": "references/video/ugc-paths.md",
    "size_in_bytes": 2988
  },
  {
    "relative_path": "references/video/voice-in-video.md",
    "size_in_bytes": 3328
  },
  {
    "relative_path": "references/video/wan-3.md",
    "size_in_bytes": 9107
  },
  {
    "relative_path": "references/workflow-basics.md",
    "size_in_bytes": 7260
  }
]

changed /skill_md_contents

BEFORE
"---\nname: creativeclaw-generate-video\ndescription: \"Generate, animate, extend, reframe, or transform one video clip with Creative Claw. Use for a clear single-clip request when the user has not asked for a storyboard, UGC ad, or complete multi-shot film.\"\n---\n\n# Generate Video\n\nRead [video model selection](references/video/index.md), then only the selected model's guide. Model families are covered locally, with live-schema guidance for additional models; do not load every guide or require a sibling model skill. Read [reference production](references/video/reference-production.md) before preparing media and [Review/Auto handling](references/video/review.md) before submission.\n\nRead [shared execution guidance](references/workflow-basics.md) once per task before using tools. It covers existing authorization, model discovery, optional cost checks, imports, and recovery.\n\nCreate one controlled video clip from text, a start frame, an optional end frame, or other model-supported references. Use the planning, UGC, or film skills when the deliverable is a larger production. Use `creativeclaw-render-html-video` only when the user explicitly requests HTML/HyperFrames/code-driven rendering; use `creativeclaw-add-video-intro-outro` for video bookends.\n\nFor worked production flows, read only the relevant recipe: [product ad](references/video/recipe-product-ad.md), [consistent Character scene](references/video/recipe-character-scene.md), or [source edit and extension](references/video/recipe-source-edit.md). These explain asset preparation, shot prompting, assembly and output checks, without authorizing extra paid drafts.\n\n## Reference-first production\n\nStrongly recommend image preparation before video: reuse the user's media, develop shot-appropriate references with creativeclaw-generate-image, and prefer at least three complementary images where the video model supports them. Read [image model selection](references/images/index.md) and only the chosen image guide, plus [the reference-first workflow](references/video/reference-production.md) for approval, Character sheets, audio-first control and cross-shot continuity. Existing references count; do not force extra paid assets, exceed model limits or ignore an explicit direct-generation request.\n\nUse the video's Review card as the single approval of its exact references and settings, without a duplicate chat approval. Honor separately requested earlier checkpoints and consent requirements. Review does not gate earlier image/audio charges. For connected clips, reuse stable visual/audio anchors and request dialogue, ambience and effects without independently generated music per shot.\n\n## Workflow\n\n1. Define the clip's purpose, aspect ratio, duration, subject, one primary action, camera move, visual continuity, dialogue or sound, and required end state.\n2. Search and inspect the media the user referenced. Recommend a bounded reference-image preparation stage, using those sources to generate appropriate shot references when needed. Prefer reference conditioning; a still is a start frame only when it should define the exact opening composition.\n3. When the user asks for examples, references, styles, or similar concepts—or an open brief would materially benefit from choosing among concrete directions—use `creativeclaw-find-examples`. Filter by `output_type: \"video\"`, then load only the selected result. Do not search before every clip.\n4. Use `list_models({ category: \"video\" })` when choosing a model; for a known selection, use `get_model_params` directly. Reuse its current-task durations, resolutions, operations, and reference contract.\n5. Use `estimate_generation` with `operation: \"video\"` only when the user asks about cost, balance, affordability, or sets a budget. Treat returned alternatives as options; preserve explicitly chosen models, durations, and quality. Estimate-only requests do not authorize generation.\n6. State consequential settings briefly and proceed within the requested scope. Do not ask again when the user already requested the generation or approved that production stage.\n7. Write one chronological prompt: opening frame, subject action, camera behavior, environmental motion, audio or dialogue, ending frame, and exclusions.\n8. Call `generate_video`; use `check_job` only when another tool needs the completed URL or no inline viewer is monitoring the job.\n9. Inspect identity, anatomy, product fidelity, timing, camera motion, dialogue sync, and ending continuity. Deliver the result and describe any shortcomings. Suggest a focused revision, but do not generate another take without an explicit user request for that additional generation.\n\n## Authorization for additional videos\n\nA request for one video authorizes one generation attempt, not repeated attempts until the agent considers it good enough. For an explicitly requested batch or approved film shot list, generate only the requested number of clips, once each.\n\nBefore another take, replacement, model comparison, extension, or generative repair, require an explicit request for that additional video. A complaint, a request to inspect or diagnose a problem, or the agent noticing a defect does not authorize generation. If \"fix it\" could mean editing existing footage or generating again, explain the proposed repair and ask before generating again. An unused budget, a failed job, a refund, or `retryable: true` does not grant permission for another attempt.\n\nAn explicit \"generate another version\" or \"retry once\" is sufficient authorization for that scope; do not ask redundantly. For \"keep trying until perfect,\" agree on a finite attempt limit before starting further generations. Stop when the requested attempts finish and let the user decide what comes next.\n\nContinue status checks, retrieval, inspection, and drafting revised prompts without creating another video. Correct and resubmit a rejected input only when it is confirmed that no generation job was accepted or started and no credits were charged. For timeouts or uncertain submissions, follow [job recovery](references/job-recovery.md) before considering any replacement.\n\n## Model routing\n\n- Default to `video/gemini-omni-flash` for the best general balance of speed, quality, native audio, and reference-aware generation.\n- Use `video/seedance-2.5` for premium cinematic or reference-rich clips and longer shots where supported.\n- Use `video/seedance-2.0-mini`, presented to users as **Seedance Mini**, for economical drafts.\n- Use `video/minimax-h3-max` for fast cinematic motion and native-audio work.\n- Use `video/minimax-h3-max-turbo`, presented as **H3 Max Fast**, when speed and iteration cost matter most.\n- Honor an explicit model request, and load its local guide from the model-selection index for exact prompt and reference syntax.\n\nFor existing footage, do not apply the general generation ranking blindly:\n\n- Up to 10 seconds, prefer `video/gemini-omni-flash` for a targeted source edit.\n- From 4–30 seconds, prefer `video/seedance-2.5` for a full source edit or continuation.\n- When the user wants the original left unchanged with new footage added, generate only the new continuation from a short boundary segment and merge it with the untouched original.\n- Use Seedance Mini only for an explicitly cost-sensitive draft, not for exact source preservation.\n- Never recommend or proactively route to an LTX or DreamActor model.\n\n## Reference rules\n\n- `image_url` is only the literal start frame. It selects image-to-video and makes the supplied image frame zero. `last_frame_url` is the desired end frame when the selected model exposes it.\n- `image_urls`, `video_urls`, and `audio_urls` are model-specific reference arrays. If a supplied image should guide identity, style, character, product, or composition instead of becoming frame zero, use `image_urls`, even for exactly one image. Prefer at least three complementary images when supported, reusing existing assets and respecting model limits; this is a quality recommendation, not a minimum enforced by every model.\n- A saved `character_id` supplies the Character image as the start frame only when no explicit `image_url` is provided. For literal-frame animation, build identity into the approved frame and omit reference arrays. For reference-guided composition, use ordered references and omit literal frames and `character_id`. Do not mix these modes unless the current model contract explicitly supports it.\n- Preserve exact quoted copy, reference labels, dialogue, timecodes, colors, and approved layout or edit constraints in the prompt.\n- Discover transformation support on the selected model and connected tool schema. Do not assume a generic top-level `operation` selector exists; use only currently exposed fields and model-supported `extras` controls. Never silently switch an explicitly chosen model to obtain a transformation.\n\n## Prompt shape\n\nPrefer one subject action and one camera idea per clip. Describe what happens over time, not a pile of adjectives. Include exact spoken words only when needed, and specify what must not change. For multi-shot continuity, first create a storyboard and clean reference frames with `creativeclaw-plan-video`.\n\nConduct the workflow in the user's language and preserve quoted dialogue exactly. Confirm the chosen model supports the requested spoken language before relying on native audio.\n"
AFTER
"---\nname: creativeclaw-generate-video\ndescription: \"Generate, animate, extend, reframe, or transform one video clip with Creative Claw. Use for a clear single-clip request when the user has not asked for a storyboard, UGC ad, or complete multi-shot film.\"\n---\n\n# Generate Video\n\nRead [video model selection](references/video/index.md), then only the selected model's guide. Model families are covered locally, with live-schema guidance for additional models; do not load every guide or require a sibling model skill. Read [Review/Auto handling](references/video/review.md) before submission.\n\nRead [shared execution guidance](references/workflow-basics.md) once per task before using tools. It covers existing authorization, model discovery, optional cost checks, imports, and recovery.\n\nCreate one controlled video clip from text, a start frame, an optional end frame, or other model-supported references. Use the planning, UGC, or film skills when the deliverable is a larger production. Use `creativeclaw-render-html` only when the user explicitly requests HTML/HyperFrames/code-driven rendering; use `creativeclaw-edit-media` for video bookends.\n\nFor a permanent watermark on a finished video or a sequence assembled from existing images, video clips, and optional audio, use `merge_media` through `creativeclaw-edit-media`. Read [assembly guidance](references/media-assembly.md) for `overlay_images` and `compose_video`. Preserve the existing media instead of generating replacement footage.\n\nFor worked production flows, read only the relevant recipe: [product ad](references/video/recipe-product-ad.md), [consistent Character scene](references/video/recipe-character-scene.md), or [source edit and extension](references/video/recipe-source-edit.md). These explain asset preparation, shot prompting, assembly and output checks, without authorizing extra paid drafts.\n\n## Reference-first pipeline\n\nUnless the user asked for direct text-to-video, supplied a ready shot image, or is editing footage, follow this before `generate_video`. A video request authorizes one keyframe per shot: say so, don't ask.\n\n1. Anchors, reuse first: `search_assets`, `list_characters`, `get_theme`. Person: Character sheet + face portrait (real person: also their best original photo). Product: real photo or packshot, plus a label/logo close-up when text matters. A recurring person or product with no anchor: create it first (creativeclaw-create-avatar, creativeclaw-product-photoshoot).\n2. Look line: one sentence (palette, light, lens, medium), pasted into every keyframe and video prompt.\n3. Keyframe per shot: `generate_image` with the same image model all project (default `image/nano-banana-2`), `aspect_ratio` = the video's ratio, main anchor in `image_url`, others in `extras.image_urls`, roles named. One clean full-bleed frame; no text, grid or labels.\n4. Compare it to the anchors (face, label, logo, colors); fix with one targeted edit.\n5. Show keyframes and the plan (model, duration, ratio) in one message. Review mode: call `generate_video` now; the card is the approval. Auto: ask once unless the user said go.\n6. One mode per shot. People, several subjects or big motion: `image_urls` = [keyframe, identity anchor, product anchor], 2–4 total; `character_id` is fine here. Exact opening (product hero, logo reveal): `image_url` = keyframe, no `image_urls` or `character_id`.\n7. Next shot: same anchors and look line. A previous clip's last frame is only an extra composition cue.\n\nDetails: [reference production](references/video/reference-production.md). Image prompting: [image model index](references/images/index.md) and only the chosen guide.\n\n## Voices\n\nWhen a Character speaks and the voice is unknown, ask one question: design a new voice from a description (`design_voice`, three auditions), use your own voice (recording plus consent, creativeclaw-clone-voice), or pick a stock voice. If the voice doesn't matter, pick a stock voice and name it. Design auditions use the Character's real lines. In ChatGPT the Voice Studio card saves the choice; read it back from `list_characters` rather than saving again.\n\nThen pick one path per speaking shot from [voice in video](references/video/voice-in-video.md): native dialogue for a one-off, or speech first for an exact or recurring voice (an audio-capable model, or `video/sync-3` on a finished clip). Gemini Omni, the default, accepts no audio: don't make speech first for an Omni shot. Keep lines to at most 2.5 words per clip second.\n\n## Workflow\n\n1. Define the clip's purpose, aspect ratio, duration, subject, one primary action, camera move, visual continuity, dialogue or sound, and required end state.\n2. Unless the user asked for direct text-to-video, supplied a ready shot image, or is editing footage, follow the reference-first pipeline above before `generate_video`. Import media the user referred to first.\n3. When the user asks for examples, styles, or similar concepts, or an open brief would benefit from concrete directions, call `search_examples` with `output_type: \"video\"`, then load only the chosen result with `search_examples({ id })`. Do not search before every clip.\n4. Use `list_models({ category: \"video\" })` when choosing a model; for a known selection, use `get_model_params` directly. Reuse its current-task durations, resolutions, operations, and reference contract.\n5. Use `estimate_generation` with `operation: \"video\"` only when the user asks about cost, balance, affordability, or sets a budget. Treat returned alternatives as options; preserve explicitly chosen models, durations, and quality. Estimate-only requests do not authorize generation.\n6. State consequential settings briefly and proceed within the requested scope. Do not ask again when the user already requested the generation or approved that production stage.\n7. Write one chronological prompt: opening frame, subject action, camera behavior, environmental motion, audio or dialogue, ending frame, and exclusions.\n8. Call `generate_video`; use `check_job` only when another tool needs the completed URL or no inline viewer is monitoring the job.\n9. Inspect identity, anatomy, product fidelity, timing, camera motion, dialogue sync, and ending continuity. Deliver the result and describe any shortcomings. Suggest a focused revision, but do not generate another take without an explicit user request for that additional generation.\n\n## Authorization for additional videos\n\nA request for one video authorizes one generation attempt, not repeated attempts until the agent considers it good enough. For an explicitly requested batch or approved film shot list, generate only the requested number of clips, once each.\n\nBefore another take, replacement, model comparison, extension, or generative repair, require an explicit request for that additional video. A complaint, a request to inspect or diagnose a problem, or the agent noticing a defect does not authorize generation. If \"fix it\" could mean editing existing footage or generating again, explain the proposed repair and ask before generating again. An unused budget, a failed job, a refund, or `retryable: true` does not grant permission for another attempt.\n\nAn explicit \"generate another version\" or \"retry once\" is sufficient authorization for that scope; do not ask redundantly. For \"keep trying until perfect,\" agree on a finite attempt limit before starting further generations. Stop when the requested attempts finish and let the user decide what comes next.\n\nContinue status checks, retrieval, inspection, and drafting revised prompts without creating another video. Correct and resubmit a rejected input only when it is confirmed that no generation job was accepted or started and no credits were charged. For timeouts or uncertain submissions, follow [job recovery](references/job-recovery.md) before considering any replacement.\n\n## Model routing\n\n- Default to `video/gemini-omni-flash` for the best general balance of speed, quality, native audio, and reference-aware generation.\n- Use `video/seedance-2.5`, the premium cinematic model, for high-end, reference-rich or long clips (up to 30 s). Review mode stages an initial 1080p request as a 480p draft; Auto renders 1080p directly. Explicit 480p creates a draft in either mode. Only 1080p can finalize the same draft through `extras.draft_job_id`; 720p starts a new take.\n- Use `video/minimax-h3-max` for fast cinematic motion and native-audio work.\n- Use `video/minimax-h3-max-turbo`, presented as **H3 Max Fast**, for cheap, fast drafts and iteration.\n- Use `video/wan-3.0` for native-audio clips of 2–30 s in one pass, or when a document or webpage drives the video.\n- Use Seedance Mini only when the user asks for it.\n- Honor an explicit model request, and load its local guide from the model-selection index for exact prompt and reference syntax.\n\nFor existing footage, do not apply the general generation ranking blindly:\n\n- To extend a clip, default to `video/minimax-h3-max-extend`; it keeps the source's characters, setting, motion, and look. Use one source of 1.625–60 seconds in `video_urls`, a `duration` of 5–15 seconds for the new footage, keep `aspect_ratio: \"auto\"` unless cropping was requested, and describe what happens next. Its default `extras.output: \"extended\"` returns the source plus new footage; `\"continuation\"` returns only the new segment. Use `video/seedance-2.5` `extend` only for heavy references or a long continuation.\n- To replace an interval inside a clip with new footage, use `video/minimax-h3-max-insert`: one source in `video_urls`, `extras.start_time` where the new scene begins and `extras.resume_time` where the original resumes, both on the source timeline. `duration` sets the new scene's length.\n- Up to 10 seconds, prefer `video/gemini-omni-flash` for a targeted source edit.\n- From 4–30 seconds, prefer `video/seedance-2.5` for a full source edit.\n- When the user wants the original left unchanged with new footage added, generate only the new continuation (with H3 Max Extend, `extras.output: \"continuation\"`) and merge it with the untouched original.\n- Never recommend or proactively route to an LTX or DreamActor model.\n\n## Reference rules\n\n- `image_url` is only the literal start frame. It selects image-to-video and makes the supplied image frame zero. `last_frame_url` is the desired end frame when the selected model exposes it.\n- `image_urls`, `video_urls`, and `audio_urls` are model-specific reference arrays. If a supplied image should guide identity, style, character, product, or composition instead of becoming frame zero, use `image_urls`, even for exactly one image. Use 2–4 strong references; more is not better.\n- `character_id` appends the saved image to `image_urls` as a reference, never a start frame. Use it in reference mode; omit it with `image_url`/`last_frame_url` (the server rejects that mix). For literal-frame animation, build identity into the approved frame and omit reference arrays.\n- Preserve exact quoted copy, reference labels, dialogue, timecodes, colors, and approved layout or edit constraints in the prompt.\n- Discover transformation support on the selected model and connected tool schema. Do not assume a generic top-level `operation` selector exists; use only currently exposed fields and model-supported `extras` controls. Never silently switch an explicitly chosen model to obtain a transformation.\n\n## Prompt shape\n\nPrefer one subject action and one camera idea per clip. Describe what happens over time, not a pile of adjectives. Include exact spoken words only when needed, and specify what must not change. For a multi-shot piece, plan it with `creativeclaw-plan-video`.\n\nConduct the workflow in the user's language and preserve quoted dialogue exactly. Confirm the chosen model supports the requested spoken language before relying on native audio.\n"

SKILL.md line diff

--- before
+++ after
@@ -5,25 +5,41 @@
 
 # Generate Video
 
-Read [video model selection](references/video/index.md), then only the selected model's guide. Model families are covered locally, with live-schema guidance for additional models; do not load every guide or require a sibling model skill. Read [reference production](references/video/reference-production.md) before preparing media and [Review/Auto handling](references/video/review.md) before submission.
+Read [video model selection](references/video/index.md), then only the selected model's guide. Model families are covered locally, with live-schema guidance for additional models; do not load every guide or require a sibling model skill. Read [Review/Auto handling](references/video/review.md) before submission.
 
 Read [shared execution guidance](references/workflow-basics.md) once per task before using tools. It covers existing authorization, model discovery, optional cost checks, imports, and recovery.
 
-Create one controlled video clip from text, a start frame, an optional end frame, or other model-supported references. Use the planning, UGC, or film skills when the deliverable is a larger production. Use `creativeclaw-render-html-video` only when the user explicitly requests HTML/HyperFrames/code-driven rendering; use `creativeclaw-add-video-intro-outro` for video bookends.
+Create one controlled video clip from text, a start frame, an optional end frame, or other model-supported references. Use the planning, UGC, or film skills when the deliverable is a larger production. Use `creativeclaw-render-html` only when the user explicitly requests HTML/HyperFrames/code-driven rendering; use `creativeclaw-edit-media` for video bookends.
+
+For a permanent watermark on a finished video or a sequence assembled from existing images, video clips, and optional audio, use `merge_media` through `creativeclaw-edit-media`. Read [assembly guidance](references/media-assembly.md) for `overlay_images` and `compose_video`. Preserve the existing media instead of generating replacement footage.
 
 For worked production flows, read only the relevant recipe: [product ad](references/video/recipe-product-ad.md), [consistent Character scene](references/video/recipe-character-scene.md), or [source edit and extension](references/video/recipe-source-edit.md). These explain asset preparation, shot prompting, assembly and output checks, without authorizing extra paid drafts.
 
-## Reference-first production
+## Reference-first pipeline
+
+Unless the user asked for direct text-to-video, supplied a ready shot image, or is editing footage, follow this before `generate_video`. A video request authorizes one keyframe per shot: say so, don't ask.
+
+1. Anchors, reuse first: `search_assets`, `list_characters`, `get_theme`. Person: Character sheet + face portrait (real person: also their best original photo). Product: real photo or packshot, plus a label/logo close-up when text matters. A recurring person or product with no anchor: create it first (creativeclaw-create-avatar, creativeclaw-product-photoshoot).
+2. Look line: one sentence (palette, light, lens, medium), pasted into every keyframe and video prompt.
+3. Keyframe per shot: `generate_image` with the same image model all project (default `image/nano-banana-2`), `aspect_ratio` = the video's ratio, main anchor in `image_url`, others in `extras.image_urls`, roles named. One clean full-bleed frame; no text, grid or labels.
+4. Compare it to the anchors (face, label, logo, colors); fix with one targeted edit.
+5. Show keyframes and the plan (model, duration, ratio) in one message. Review mode: call `generate_video` now; the card is the approval. Auto: ask once unless the user said go.
+6. One mode per shot. People, several subjects or big motion: `image_urls` = [keyframe, identity anchor, product anchor], 2–4 total; `character_id` is fine here. Exact opening (product hero, logo reveal): `image_url` = keyframe, no `image_urls` or `character_id`.
+7. Next shot: same anchors and look line. A previous clip's last frame is only an extra composition cue.
+
+Details: [reference production](references/video/reference-production.md). Image prompting: [image model index](references/images/index.md) and only the chosen guide.
+
+## Voices
 
-Strongly recommend image preparation before video: reuse the user's media, develop shot-appropriate references with creativeclaw-generate-image, and prefer at least three complementary images where the video model supports them. Read [image model selection](references/images/index.md) and only the chosen image guide, plus [the reference-first workflow](references/video/reference-production.md) for approval, Character sheets, audio-first control and cross-shot continuity. Existing references count; do not force extra paid assets, exceed model limits or ignore an explicit direct-generation request.
+When a Character speaks and the voice is unknown, ask one question: design a new voice from a description (`design_voice`, three auditions), use your own voice (recording plus consent, creativeclaw-clone-voice), or pick a stock voice. If the voice doesn't matter, pick a stock voice and name it. Design auditions use the Character's real lines. In ChatGPT the Voice Studio card saves the choice; read it back from `list_characters` rather than saving again.
 
-Use the video's Review card as the single approval of its exact references and settings, without a duplicate chat approval. Honor separately requested earlier checkpoints and consent requirements. Review does not gate earlier image/audio charges. For connected clips, reuse stable visual/audio anchors and request dialogue, ambience and effects without independently generated music per shot.
+Then pick one path per speaking shot from [voice in video](references/video/voice-in-video.md): native dialogue for a one-off, or speech first for an exact or recurring voice (an audio-capable model, or `video/sync-3` on a finished clip). Gemini Omni, the default, accepts no audio: don't make speech first for an Omni shot. Keep lines to at most 2.5 words per clip second.
 
 ## Workflow
 
 1. Define the clip's purpose, aspect ratio, duration, subject, one primary action, camera move, visual continuity, dialogue or sound, and required end state.
-2. Search and inspect the media the user referenced. Recommend a bounded reference-image preparation stage, using those sources to generate appropriate shot references when needed. Prefer reference conditioning; a still is a start frame only when it should define the exact opening composition.
-3. When the user asks for examples, references, styles, or similar concepts—or an open brief would materially benefit from choosing among concrete directions—use `creativeclaw-find-examples`. Filter by `output_type: "video"`, then load only the selected result. Do not search before every clip.
+2. Unless the user asked for direct text-to-video, supplied a ready shot image, or is editing footage, follow the reference-first pipeline above before `generate_video`. Import media the user referred to first.
+3. When the user asks for examples, styles, or similar concepts, or an open brief would benefit from concrete directions, call `search_examples` with `output_type: "video"`, then load only the chosen result with `search_examples({ id })`. Do not search before every clip.
 4. Use `list_models({ category: "video" })` when choosing a model; for a known selection, use `get_model_params` directly. Reuse its current-task durations, resolutions, operations, and reference contract.
 5. Use `estimate_generation` with `operation: "video"` only when the user asks about cost, balance, affordability, or sets a budget. Treat returned alternatives as options; preserve explicitly chosen models, durations, and quality. Estimate-only requests do not authorize generation.
 6. State consequential settings briefly and proceed within the requested scope. Do not ask again when the user already requested the generation or approved that production stage.
@@ -44,30 +60,32 @@
 ## Model routing
 
 - Default to `video/gemini-omni-flash` for the best general balance of speed, quality, native audio, and reference-aware generation.
-- Use `video/seedance-2.5` for premium cinematic or reference-rich clips and longer shots where supported.
-- Use `video/seedance-2.0-mini`, presented to users as **Seedance Mini**, for economical drafts.
+- Use `video/seedance-2.5`, the premium cinematic model, for high-end, reference-rich or long clips (up to 30 s). Review mode stages an initial 1080p request as a 480p draft; Auto renders 1080p directly. Explicit 480p creates a draft in either mode. Only 1080p can finalize the same draft through `extras.draft_job_id`; 720p starts a new take.
 - Use `video/minimax-h3-max` for fast cinematic motion and native-audio work.
-- Use `video/minimax-h3-max-turbo`, presented as **H3 Max Fast**, when speed and iteration cost matter most.
+- Use `video/minimax-h3-max-turbo`, presented as **H3 Max Fast**, for cheap, fast drafts and iteration.
+- Use `video/wan-3.0` for native-audio clips of 2–30 s in one pass, or when a document or webpage drives the video.
+- Use Seedance Mini only when the user asks for it.
 - Honor an explicit model request, and load its local guide from the model-selection index for exact prompt and reference syntax.
 
 For existing footage, do not apply the general generation ranking blindly:
 
+- To extend a clip, default to `video/minimax-h3-max-extend`; it keeps the source's characters, setting, motion, and look. Use one source of 1.625–60 seconds in `video_urls`, a `duration` of 5–15 seconds for the new footage, keep `aspect_ratio: "auto"` unless cropping was requested, and describe what happens next. Its default `extras.output: "extended"` returns the source plus new footage; `"continuation"` returns only the new segment. Use `video/seedance-2.5` `extend` only for heavy references or a long continuation.
+- To replace an interval inside a clip with new footage, use `video/minimax-h3-max-insert`: one source in `video_urls`, `extras.start_time` where the new scene begins and `extras.resume_time` where the original resumes, both on the source timeline. `duration` sets the new scene's length.
 - Up to 10 seconds, prefer `video/gemini-omni-flash` for a targeted source edit.
-- From 4–30 seconds, prefer `video/seedance-2.5` for a full source edit or continuation.
-- When the user wants the original left unchanged with new footage added, generate only the new continuation from a short boundary segment and merge it with the untouched original.
-- Use Seedance Mini only for an explicitly cost-sensitive draft, not for exact source preservation.
+- From 4–30 seconds, prefer `video/seedance-2.5` for a full source edit.
+- When the user wants the original left unchanged with new footage added, generate only the new continuation (with H3 Max Extend, `extras.output: "continuation"`) and merge it with the untouched original.
 - Never recommend or proactively route to an LTX or DreamActor model.
 
 ## Reference rules
 
 - `image_url` is only the literal start frame. It selects image-to-video and makes the supplied image frame zero. `last_frame_url` is the desired end frame when the selected model exposes it.
-- `image_urls`, `video_urls`, and `audio_urls` are model-specific reference arrays. If a supplied image should guide identity, style, character, product, or composition instead of becoming frame zero, use `image_urls`, even for exactly one image. Prefer at least three complementary images when supported, reusing existing assets and respecting model limits; this is a quality recommendation, not a minimum enforced by every model.
-- A saved `character_id` supplies the Character image as the start frame only when no explicit `image_url` is provided. For literal-frame animation, build identity into the approved frame and omit reference arrays. For reference-guided composition, use ordered references and omit literal frames and `character_id`. Do not mix these modes unless the current model contract explicitly supports it.
+- `image_urls`, `video_urls`, and `audio_urls` are model-specific reference arrays. If a supplied image should guide identity, style, character, product, or composition instead of becoming frame zero, use `image_urls`, even for exactly one image. Use 2–4 strong references; more is not better.
+- `character_id` appends the saved image to `image_urls` as a reference, never a start frame. Use it in reference mode; omit it with `image_url`/`last_frame_url` (the server rejects that mix). For literal-frame animation, build identity into the approved frame and omit reference arrays.
 - Preserve exact quoted copy, reference labels, dialogue, timecodes, colors, and approved layout or edit constraints in the prompt.
 - Discover transformation support on the selected model and connected tool schema. Do not assume a generic top-level `operation` selector exists; use only currently exposed fields and model-supported `extras` controls. Never silently switch an explicitly chosen model to obtain a transformation.
 
 ## Prompt shape
 
-Prefer one subject action and one camera idea per clip. Describe what happens over time, not a pile of adjectives. Include exact spoken words only when needed, and specify what must not change. For multi-shot continuity, first create a storyboard and clean reference frames with `creativeclaw-plan-video`.
+Prefer one subject action and one camera idea per clip. Describe what happens over time, not a pile of adjectives. Include exact spoken words only when needed, and specify what must not change. For a multi-shot piece, plan it with `creativeclaw-plan-video`.
 
 Conduct the workflow in the user's language and preserve quoted dialogue exactly. Confirm the chosen model supports the requested spoken language before relying on native audio.
Full snapshot data
{
  "description": "Generate, animate, extend, reframe, or transform one video clip with Creative Claw. Use for a clear single-clip request when the user has not asked for a storyboard, UGC ad, or complete multi-shot film.",
  "included_files": [
    {
      "relative_path": "agents/openai.yaml",
      "size_in_bytes": 539
    },
    {
      "relative_path": "references/images/gpt-image-2.md",
      "size_in_bytes": 11212
    },
    {
      "relative_path": "references/images/index.md",
      "size_in_bytes": 456
    },
    {
      "relative_path": "references/images/nano-banana-2.md",
      "size_in_bytes": 8855
    },
    {
      "relative_path": "references/images/nano-banana-pro.md",
      "size_in_bytes": 8824
    },
    {
      "relative_path": "references/images/seedream-5-pro.md",
      "size_in_bytes": 9556
    },
    {
      "relative_path": "references/job-recovery.md",
      "size_in_bytes": 3786
    },
    {
      "relative_path": "references/media-assembly.md",
      "size_in_bytes": 8654
    },
    {
      "relative_path": "references/platform-upload.md",
      "size_in_bytes": 1416
    },
    {
      "relative_path": "references/video/gemini-omni.md",
      "size_in_bytes": 10085
    },
    {
      "relative_path": "references/video/index.md",
      "size_in_bytes": 2176
    },
    {
      "relative_path": "references/video/minimax-h3-max.md",
      "size_in_bytes": 11220
    },
    {
      "relative_path": "references/video/other-models.md",
      "size_in_bytes": 4653
    },
    {
      "relative_path": "references/video/recipe-character-scene.md",
      "size_in_bytes": 2941
    },
    {
      "relative_path": "references/video/recipe-product-ad.md",
      "size_in_bytes": 3265
    },
    {
      "relative_path": "references/video/recipe-source-edit.md",
      "size_in_bytes": 3357
    },
    {
      "relative_path": "references/video/reference-production.md",
      "size_in_bytes": 5865
    },
    {
      "relative_path": "references/video/review.md",
      "size_in_bytes": 1909
    },
    {
      "relative_path": "references/video/seedance-2-5.md",
      "size_in_bytes": 15084
    },
    {
      "relative_path": "references/video/seedance-mini.md",
      "size_in_bytes": 1579
    },
    {
      "relative_path": "references/video/ugc-paths.md",
      "size_in_bytes": 2988
    },
    {
      "relative_path": "references/video/voice-in-video.md",
      "size_in_bytes": 3328
    },
    {
      "relative_path": "references/video/wan-3.md",
      "size_in_bytes": 9107
    },
    {
      "relative_path": "references/workflow-basics.md",
      "size_in_bytes": 7260
    }
  ],
  "name": "creativeclaw-generate-video",
  "skill_md_contents": "---\nname: creativeclaw-generate-video\ndescription: \"Generate, animate, extend, reframe, or transform one video clip with Creative Claw. Use for a clear single-clip request when the user has not asked for a storyboard, UGC ad, or complete multi-shot film.\"\n---\n\n# Generate Video\n\nRead [video model selection](references/video/index.md), then only the selected model's guide. Model families are covered locally, with live-schema guidance for additional models; do not load every guide or require a sibling model skill. Read [Review/Auto handling](references/video/review.md) before submission.\n\nRead [shared execution guidance](references/workflow-basics.md) once per task before using tools. It covers existing authorization, model discovery, optional cost checks, imports, and recovery.\n\nCreate one controlled video clip from text, a start frame, an optional end frame, or other model-supported references. Use the planning, UGC, or film skills when the deliverable is a larger production. Use `creativeclaw-render-html` only when the user explicitly requests HTML/HyperFrames/code-driven rendering; use `creativeclaw-edit-media` for video bookends.\n\nFor a permanent watermark on a finished video or a sequence assembled from existing images, video clips, and optional audio, use `merge_media` through `creativeclaw-edit-media`. Read [assembly guidance](references/media-assembly.md) for `overlay_images` and `compose_video`. Preserve the existing media instead of generating replacement footage.\n\nFor worked production flows, read only the relevant recipe: [product ad](references/video/recipe-product-ad.md), [consistent Character scene](references/video/recipe-character-scene.md), or [source edit and extension](references/video/recipe-source-edit.md). These explain asset preparation, shot prompting, assembly and output checks, without authorizing extra paid drafts.\n\n## Reference-first pipeline\n\nUnless the user asked for direct text-to-video, supplied a ready shot image, or is editing footage, follow this before `generate_video`. A video request authorizes one keyframe per shot: say so, don't ask.\n\n1. Anchors, reuse first: `search_assets`, `list_characters`, `get_theme`. Person: Character sheet + face portrait (real person: also their best original photo). Product: real photo or packshot, plus a label/logo close-up when text matters. A recurring person or product with no anchor: create it first (creativeclaw-create-avatar, creativeclaw-product-photoshoot).\n2. Look line: one sentence (palette, light, lens, medium), pasted into every keyframe and video prompt.\n3. Keyframe per shot: `generate_image` with the same image model all project (default `image/nano-banana-2`), `aspect_ratio` = the video's ratio, main anchor in `image_url`, others in `extras.image_urls`, roles named. One clean full-bleed frame; no text, grid or labels.\n4. Compare it to the anchors (face, label, logo, colors); fix with one targeted edit.\n5. Show keyframes and the plan (model, duration, ratio) in one message. Review mode: call `generate_video` now; the card is the approval. Auto: ask once unless the user said go.\n6. One mode per shot. People, several subjects or big motion: `image_urls` = [keyframe, identity anchor, product anchor], 2–4 total; `character_id` is fine here. Exact opening (product hero, logo reveal): `image_url` = keyframe, no `image_urls` or `character_id`.\n7. Next shot: same anchors and look line. A previous clip's last frame is only an extra composition cue.\n\nDetails: [reference production](references/video/reference-production.md). Image prompting: [image model index](references/images/index.md) and only the chosen guide.\n\n## Voices\n\nWhen a Character speaks and the voice is unknown, ask one question: design a new voice from a description (`design_voice`, three auditions), use your own voice (recording plus consent, creativeclaw-clone-voice), or pick a stock voice. If the voice doesn't matter, pick a stock voice and name it. Design auditions use the Character's real lines. In ChatGPT the Voice Studio card saves the choice; read it back from `list_characters` rather than saving again.\n\nThen pick one path per speaking shot from [voice in video](references/video/voice-in-video.md): native dialogue for a one-off, or speech first for an exact or recurring voice (an audio-capable model, or `video/sync-3` on a finished clip). Gemini Omni, the default, accepts no audio: don't make speech first for an Omni shot. Keep lines to at most 2.5 words per clip second.\n\n## Workflow\n\n1. Define the clip's purpose, aspect ratio, duration, subject, one primary action, camera move, visual continuity, dialogue or sound, and required end state.\n2. Unless the user asked for direct text-to-video, supplied a ready shot image, or is editing footage, follow the reference-first pipeline above before `generate_video`. Import media the user referred to first.\n3. When the user asks for examples, styles, or similar concepts, or an open brief would benefit from concrete directions, call `search_examples` with `output_type: \"video\"`, then load only the chosen result with `search_examples({ id })`. Do not search before every clip.\n4. Use `list_models({ category: \"video\" })` when choosing a model; for a known selection, use `get_model_params` directly. Reuse its current-task durations, resolutions, operations, and reference contract.\n5. Use `estimate_generation` with `operation: \"video\"` only when the user asks about cost, balance, affordability, or sets a budget. Treat returned alternatives as options; preserve explicitly chosen models, durations, and quality. Estimate-only requests do not authorize generation.\n6. State consequential settings briefly and proceed within the requested scope. Do not ask again when the user already requested the generation or approved that production stage.\n7. Write one chronological prompt: opening frame, subject action, camera behavior, environmental motion, audio or dialogue, ending frame, and exclusions.\n8. Call `generate_video`; use `check_job` only when another tool needs the completed URL or no inline viewer is monitoring the job.\n9. Inspect identity, anatomy, product fidelity, timing, camera motion, dialogue sync, and ending continuity. Deliver the result and describe any shortcomings. Suggest a focused revision, but do not generate another take without an explicit user request for that additional generation.\n\n## Authorization for additional videos\n\nA request for one video authorizes one generation attempt, not repeated attempts until the agent considers it good enough. For an explicitly requested batch or approved film shot list, generate only the requested number of clips, once each.\n\nBefore another take, replacement, model comparison, extension, or generative repair, require an explicit request for that additional video. A complaint, a request to inspect or diagnose a problem, or the agent noticing a defect does not authorize generation. If \"fix it\" could mean editing existing footage or generating again, explain the proposed repair and ask before generating again. An unused budget, a failed job, a refund, or `retryable: true` does not grant permission for another attempt.\n\nAn explicit \"generate another version\" or \"retry once\" is sufficient authorization for that scope; do not ask redundantly. For \"keep trying until perfect,\" agree on a finite attempt limit before starting further generations. Stop when the requested attempts finish and let the user decide what comes next.\n\nContinue status checks, retrieval, inspection, and drafting revised prompts without creating another video. Correct and resubmit a rejected input only when it is confirmed that no generation job was accepted or started and no credits were charged. For timeouts or uncertain submissions, follow [job recovery](references/job-recovery.md) before considering any replacement.\n\n## Model routing\n\n- Default to `video/gemini-omni-flash` for the best general balance of speed, quality, native audio, and reference-aware generation.\n- Use `video/seedance-2.5`, the premium cinematic model, for high-end, reference-rich or long clips (up to 30 s). Review mode stages an initial 1080p request as a 480p draft; Auto renders 1080p directly. Explicit 480p creates a draft in either mode. Only 1080p can finalize the same draft through `extras.draft_job_id`; 720p starts a new take.\n- Use `video/minimax-h3-max` for fast cinematic motion and native-audio work.\n- Use `video/minimax-h3-max-turbo`, presented as **H3 Max Fast**, for cheap, fast drafts and iteration.\n- Use `video/wan-3.0` for native-audio clips of 2–30 s in one pass, or when a document or webpage drives the video.\n- Use Seedance Mini only when the user asks for it.\n- Honor an explicit model request, and load its local guide from the model-selection index for exact prompt and reference syntax.\n\nFor existing footage, do not apply the general generation ranking blindly:\n\n- To extend a clip, default to `video/minimax-h3-max-extend`; it keeps the source's characters, setting, motion, and look. Use one source of 1.625–60 seconds in `video_urls`, a `duration` of 5–15 seconds for the new footage, keep `aspect_ratio: \"auto\"` unless cropping was requested, and describe what happens next. Its default `extras.output: \"extended\"` returns the source plus new footage; `\"continuation\"` returns only the new segment. Use `video/seedance-2.5` `extend` only for heavy references or a long continuation.\n- To replace an interval inside a clip with new footage, use `video/minimax-h3-max-insert`: one source in `video_urls`, `extras.start_time` where the new scene begins and `extras.resume_time` where the original resumes, both on the source timeline. `duration` sets the new scene's length.\n- Up to 10 seconds, prefer `video/gemini-omni-flash` for a targeted source edit.\n- From 4–30 seconds, prefer `video/seedance-2.5` for a full source edit.\n- When the user wants the original left unchanged with new footage added, generate only the new continuation (with H3 Max Extend, `extras.output: \"continuation\"`) and merge it with the untouched original.\n- Never recommend or proactively route to an LTX or DreamActor model.\n\n## Reference rules\n\n- `image_url` is only the literal start frame. It selects image-to-video and makes the supplied image frame zero. `last_frame_url` is the desired end frame when the selected model exposes it.\n- `image_urls`, `video_urls`, and `audio_urls` are model-specific reference arrays. If a supplied image should guide identity, style, character, product, or composition instead of becoming frame zero, use `image_urls`, even for exactly one image. Use 2–4 strong references; more is not better.\n- `character_id` appends the saved image to `image_urls` as a reference, never a start frame. Use it in reference mode; omit it with `image_url`/`last_frame_url` (the server rejects that mix). For literal-frame animation, build identity into the approved frame and omit reference arrays.\n- Preserve exact quoted copy, reference labels, dialogue, timecodes, colors, and approved layout or edit constraints in the prompt.\n- Discover transformation support on the selected model and connected tool schema. Do not assume a generic top-level `operation` selector exists; use only currently exposed fields and model-supported `extras` controls. Never silently switch an explicitly chosen model to obtain a transformation.\n\n## Prompt shape\n\nPrefer one subject action and one camera idea per clip. Describe what happens over time, not a pile of adjectives. Include exact spoken words only when needed, and specify what must not change. For a multi-shot piece, plan it with `creativeclaw-plan-video`.\n\nConduct the workflow in the user's language and preserve quoted dialogue exactly. Confirm the chosen model supports the requested spoken language before relying on native audio.\n"
}

SHA-256 of public snapshot: 3a79dd876fbad2750a81e18237c0568e3435d7977cbbb04c6c375fda6331f2f7