Update to Creative Claw
Snapshot Oct 3, 2026 · 06:03 UTC · version 5.3.0
Collection source: downloaded plugin package. These snapshots do not have a confirmed matching collection source. Differences in file lists alone do not establish changes to the package.
Instructions updated for creativeclaw-edit-media
Instruction wording changed from “Claw: trim clips, resize for social formats, add automatic captions, transcribe, clean speech, extract frames, or combine finished media. Use when the source footage or recording should be preserved; route invented scenes and generative ...” to “Claw without regenerating it: trim, cut and reorder chosen moments, resize or reframe, caption, transcribe, clean speech, extract frames, add or mix audio, burn a logo watermark, compose images and clips, or add an intro/outro. Use when ...”. 63 additional added or edited lines are in the evidence.
Observed in instructions or declared skills. Runtime behavior has not been tested.
Product description
Claw: trim clips, resize for social formats, add automatic captions, transcribe, clean speech, extract frames, or combine finished media. Use when the source footage or recording should be preserved; route invented scenes and generative ...
Claw without regenerating it: trim, cut and reorder chosen moments, resize or reframe, caption, transcribe, clean speech, extract frames, add or mix audio, burn a logo watermark, compose images and clips, or add an intro/outro. Use when ...
Skill instructions
Claw: trim clips, resize for social formats, add automatic captions, transcribe, clean speech, extract frames, or combine finished media. Use when the source footage or recording should be preserved; route invented scenes and generative ...
Claw without regenerating it: trim, cut and reorder chosen moments, resize or reframe, caption, transcribe, clean speech, extract frames, add or mix audio, burn a logo watermark, compose images and clips, or add an intro/outro. Use when ...
Supporting files
[{"relative_path":"agents/openai.yaml","size_in_bytes":546},{"relative_path":"references/job-recovery.md","size_in_bytes":3786},{"relative_path":"references/media-assembly.md","size_in_bytes":3876},{"relative_path":"references/platform-u...
[{"relative_path":"agents/openai.yaml","size_in_bytes":603},{"relative_path":"references/edit-contract.md","size_in_bytes":4098},{"relative_path":"references/job-recovery.md","size_in_bytes":3786},{"relative_path":"references/media-assem...
Compare saved observations
Download comparison JSONFull technical diff · 3 changed fields
changed /description
"Edit existing video or audio with Creative Claw: trim clips, resize for social formats, add automatic captions, transcribe, clean speech, extract frames, or combine finished media. Use when the source footage or recording should be preserved; route invented scenes and generative transformations to generation skills."
"Edit existing video or audio with Creative Claw without regenerating it: trim, cut and reorder chosen moments, resize or reframe, caption, transcribe, clean speech, extract frames, add or mix audio, burn a logo watermark, compose images and clips, or add an intro/outro. Use when the source footage should be kept; use generation skills for invented or transformed footage and create-reels to pick highlights."
changed /included_files
[
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 546
},
{
"relative_path": "references/job-recovery.md",
"size_in_bytes": 3786
},
{
"relative_path": "references/media-assembly.md",
"size_in_bytes": 3876
},
{
"relative_path": "references/platform-upload.md",
"size_in_bytes": 1416
},
{
"relative_path": "references/workflow-basics.md",
"size_in_bytes": 7159
}
][
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 603
},
{
"relative_path": "references/edit-contract.md",
"size_in_bytes": 4098
},
{
"relative_path": "references/job-recovery.md",
"size_in_bytes": 3786
},
{
"relative_path": "references/media-assembly.md",
"size_in_bytes": 8654
},
{
"relative_path": "references/platform-upload.md",
"size_in_bytes": 1416
},
{
"relative_path": "references/workflow-basics.md",
"size_in_bytes": 7260
}
]changed /skill_md_contents
"---\nname: creativeclaw-edit-media\ndescription: \"Edit existing video or audio with Creative Claw: trim clips, resize for social formats, add automatic captions, transcribe, clean speech, extract frames, or combine finished media. Use when the source footage or recording should be preserved; route invented scenes and generative transformations to generation skills.\"\n---\n\n# Edit Existing Media\n\nRead [shared execution guidance](references/workflow-basics.md) once per task before using tools. It covers existing authorization, model discovery, optional cost checks, imports, and recovery.\n\nTurn supplied footage or audio into a finished derivative. Use the user's requested edits and sensible defaults without a settings questionnaire. Preserve the original asset.\n\n## Select only necessary operations\n\n| Requested change | Tool and important boundary |\n| --- | --- |\n| Shorten a video | `trim_video`: set `start_time` explicitly, including zero, and either `end_time` or `duration`. |\n| Resize or make a vertical version | `scale_video`: set even dimensions and an explicit `mode`. `crop` is center crop; `pad` preserves the full frame; `stretch` distorts it. This is not subject tracking or AI enhancement. |\n| Burn automatic captions | `add_subtitles`: operates on the current video and transcribes automatically. Set the known spoken `language`; do not translate unless requested. |\n| Get a transcript or find a spoken passage | `transcribe`: pass exactly one of `audio_url` or `video_url`; reuse existing timing when available. |\n| Clean a noisy recording | `isolate_audio`: preserve and compare the original; cleanup can alter speech. |\n| Get still frames | `extract_frames`: discover the current tool schema for first/last/sampling controls. |\n| Combine clips or add one finished audio track | `merge_media`: read [assembly guidance](references/media-assembly.md). |\n| Enhance resolution or remove a background | `upscale_media` or `remove_background`, only for the requested media type and supported parameters. |\n\nUse `creativeclaw-generate-video` for invented footage or a generative source-video transformation. Use `creativeclaw-add-video-intro-outro` for bookends. HTML rendering requires the user's explicit choice; a captions request alone does not select that route.\n\nUse `creativeclaw-create-reels` when selecting highlights from long footage, or `creativeclaw-cut-and-reframe-video` for explicit multi-cut edits, moving crop paths and source-timed word captions. Both require `cut_and_reframe_video` to be available and configured.\n\n## Workflow\n\n1. Reuse the known asset or import the missing source through [platform-upload.md](references/platform-upload.md). Use available metadata/playback to establish duration, aspect ratio, audio, and protected content; do not invent measurements.\n2. Apply the smallest requested edit chain. For a user-specified range, trim directly. Transcribe first only if text or timing is needed to choose cuts. For a clip selection request, use the transcript to choose a coherent section and verify visual relevance when possible.\n3. Finish cuts and merges before captions. Resize to the final framing before burning text. For vertical delivery, use center crop only when the subject remains visible; otherwise use padding or ask about a material framing tradeoff. Do not claim automatic face tracking.\n4. Resolve each queued operation with `check_job` before passing its result into the next. On failure, follow [job-recovery.md](references/job-recovery.md) and resume from the latest completed derivative.\n5. Add automatic subtitles to the final edited video if requested. The current tool accepts style/language settings, not an edited transcript or SRT payload. Do not promise verbatim corrected or translated captions through unsupported fields.\n6. Inspect final length, framing, spoken content, caption sync/safe areas, and audio using available capabilities. Name/tag the result with the project and edit role; provide the finished asset and any material inspection limitation.\n\nFor “make this a captioned vertical clip,” reuse the supplied footage, trim only if a shorter clip was requested, resize with a suitable explicit mode, then add captions. Skip model discovery and generation when all requested changes are supported processing operations.\n\n## Limits that change the plan\n\n- A public YouTube URL can go directly to `transcribe` through ElevenLabs Scribe. It is not a directly downloadable source for trimming, resizing, or subtitle rendering; obtain the actual video asset for those steps.\n- Automatic subtitles already transcribe. Do not add a separate transcription job unless the user also needs a transcript or it guides clip selection.\n- Mixing narration with music, ducking, crossfades, or exact caption-file rendering require additional capabilities. Explicit moving crops and source-timed word captions are supported by `cut_and_reframe_video` when available; automatic face tracking is not. State the specific missing operation before creating assets that cannot be assembled as requested.\n- `estimate_generation` does not estimate these processing operations. Only discuss cost when relevant to the user's request, and do not label a partial generation estimate as the total edit cost.\n"
"---\nname: creativeclaw-edit-media\ndescription: \"Edit existing video or audio with Creative Claw without regenerating it: trim, cut and reorder chosen moments, resize or reframe, caption, transcribe, clean speech, extract frames, add or mix audio, burn a logo watermark, compose images and clips, or add an intro/outro. Use when the source footage should be kept; use generation skills for invented or transformed footage and create-reels to pick highlights.\"\n---\n\n# Edit Existing Media\n\nRead [shared execution guidance](references/workflow-basics.md) once per task before using tools. It covers existing authorization, model discovery, optional cost checks, imports, and recovery.\n\nTurn supplied footage or audio into a finished derivative. Apply the requested edits with sensible defaults and no settings questionnaire. Keep the original asset unchanged.\n\n## Pick only the operations you need\n\n| Requested change | Tool and key limits |\n| --- | --- |\n| Shorten a video to one range | `trim_video`: set `start_time` explicitly, including zero, and either `end_time` or `duration`. |\n| Cut, reorder, or reframe several chosen moments | `cut_and_reframe_video`: see [Multi-cut edits](#multi-cut-edits). |\n| Resize or make a vertical version | `scale_video`: even dimensions and an explicit `mode`. `crop` is a center crop, `pad` keeps the full frame, `stretch` distorts. No subject tracking. |\n| Burn automatic captions | `add_subtitles`: transcribes the current video itself. Set the spoken `language`; do not translate unless asked. |\n| Get a transcript or find a passage | `transcribe`: exactly one of `audio_url` or `video_url`. |\n| Clean a noisy recording | `isolate_audio`: keep and compare the original; cleanup can change speech. |\n| Get still frames | `extract_frames`: check the schema for first/last/sampling controls. |\n| Put narration or music on a clip | `merge_media` `merge_audio_video`: see [Audio](#audio). |\n| Join clips or audio files in order | `merge_media` `merge_videos` or `merge_audios`. |\n| Burn a logo or copyright mark onto a finished video | `merge_media` `overlay_images` with a transparent image. `remove_background` can prepare a logo. |\n| Build a video from timed images, clips, and optional audio | `merge_media` `compose_video`. |\n| Add an intro, outro, or both | See [Intros and outros](#intros-and-outros). |\n| Upscale or remove a background | `upscale_media` or `remove_background`, for the media type and parameters they support. |\n\nRead [assembly guidance](references/media-assembly.md) before any `merge_media` call.\n\nRoute elsewhere: `creativeclaw-generate-video` for invented footage or a generative change to the source; `creativeclaw-create-reels` when the agent must pick highlights from long footage; `creativeclaw-render-html` only when the user explicitly wants HTML or HyperFrames rendering. A caption request alone does not select HTML. For an exact text watermark, `render_html_image` with `transparent_background: true` can make the PNG for `overlay_images`.\n\n## Workflow\n\n1. Reuse the known asset or import the source through [platform-upload.md](references/platform-upload.md). Establish duration, aspect ratio, audio, and burned-in text from metadata and playback; do not invent measurements.\n2. Run the smallest edit chain. Trim a user-given range directly. Transcribe first only if text or timing is needed to choose cuts.\n3. Finish cuts and merges first, then resize to the final framing, then burn captions. Use a center crop only when the subject stays visible; otherwise pad or ask about the framing tradeoff.\n4. Resolve each queued job with `check_job` before passing its result on. On failure, follow [job-recovery.md](references/job-recovery.md) and resume from the last completed derivative.\n5. `add_subtitles` takes style and language settings, not an edited transcript or SRT. Do not promise corrected or translated captions through fields it lacks.\n6. Check final length, framing, speech, caption sync and safe areas, and audio. Name and tag the result, and state any check you could not do.\n\nFor \"make this a captioned vertical clip\": reuse the footage, trim only if a shorter clip was asked for, resize with an explicit mode, then add captions. No model discovery or generation is needed.\n\n## Multi-cut edits\n\nUse `cut_and_reframe_video` when the user has chosen the moments (or `creativeclaw-create-reels` has) and wants them cut, reordered, or reframed in one render. It does not pick highlights, transcribe, or track faces. If the tool is missing or returns \"not configured\", say so. Do not fake an exact multi-cut edit with a chain of lossy trims.\n\nRead [the cut-and-reframe contract](references/edit-contract.md) before building input.\n\n1. The source must be a video asset in this workspace. Keep an unchanged source URL and a source-timed transcript. Inspect metadata and frames; do not invent sizes, word times, or coordinates.\n2. Choose safe boundaries. Listen around each edge when you can. Avoid clipped consonants, breaths, and cut-off reactions. Prefer a natural gap, keep enough lead-in and tail, and end on a short release without entering the next word. The renderer does not extend ranges. Do not remove every pause or change what someone meant.\n3. Set the requested width and height (1080×1920 is only the default). Use `pad` when position is uncertain. Use `center_crop` only after checking the subject stays in frame. For `crop`, give an even rectangle at the output aspect ratio, with static coordinates or keyframes set from real observations. If a speaker leaves the crop, revise the path or pad; never stretch or guess.\n4. Captions are optional. Pass real source-timed words; the tool remaps them after the cuts and burns them last. With only sentence timing, render the cuts first and run `add_subtitles` on the result. Never fabricate word alignment. Use plain captions for overlapping speakers or right-to-left text unless you can review karaoke visually.\n5. Submit one call per output. Save the job ID and the exact edit plan. After a timeout, poll the existing job; do not resubmit blindly.\n6. Check every splice with about 1.5 seconds of context on each side, the first and last words, crop continuity, caption spelling and sync, and audio. Technical checks do not prove the edit reads well. Fix only identified faults; do not rerender speculatively.\n\n## Intros and outros\n\nGet the bookend segments, then concatenate them around the main video.\n\n- **Supplied clips:** import and use them.\n- **HTML title cards:** only when the user asks for HTML, HyperFrames, or code-rendered cards, or accepts that option when you offer it for exact text, fonts, and logos. Render them with `creativeclaw-render-html`. An intro/outro request alone does not authorize `render_html_video`.\n- **Generated cinematic bookend:** use `creativeclaw-generate-video`. To match the look, `extract_frames` from the main video and pass the frame as a style reference.\n\n1. Establish the main video's URL, size, aspect ratio, frame rate, audio, and platform.\n2. Use the requested copy, logo, duration, and audio. Ask only about missing copy; do not invent a slogan or call to action. Keep bookends short.\n3. Make segments at the main video's size and frame rate. Resolve and inspect each one.\n4. Merge in playback order. Set `canvas_video_index` to the main video's zero-based index and `video_fit: \"pad\"` to keep mismatched frames whole, unless cropping is authorized.\n\n```text\nmerge_media({\n operation: \"merge_videos\",\n video_urls: [\"<intro-url>\", \"<main-video-url>\", \"<outro-url>\"],\n canvas_video_index: 1,\n video_fit: \"pad\",\n pad_color: \"black\"\n})\n```\n\nOmit a bookend that was not requested. `merge_videos` is a hard cut: no dissolves, crossfades, or audio carried across the join. If the main audio must continue under a bookend or fade across a cut, say so before making segments.\n\n## Audio\n\n- `merge_audio_video` replaces the clip's audio by default and ends at the shorter input. Check both durations first; do not silently shorten the video to fit a voiceover.\n- `audio_mode: \"mix\"` keeps the clip's own sound and layers the new track over it, at the video's full length. Set `original_volume` and `added_volume` (0–1); about 0.3 keeps music under speech.\n- `compose_video` layers one audio track over the clips at full level, and the output can run past the last clip. Use it when the added audio should extend beyond the picture.\n- `merge_audios` joins files end to end; it does not layer them.\n- There is no ducking, volume automation, or crossfade. Say which operation is missing before creating assets that cannot be assembled as asked.\n\n## Limits\n\n- A public YouTube, Google Drive, or social video page can go straight to `transcribe` as `video_url`. It is not a source for trimming, resizing, or subtitles; import the actual file for those.\n- `add_subtitles` already transcribes. Add a separate `transcribe` job only for a transcript the user wants or for choosing cuts.\n- `estimate_generation` does not cover these processing operations. Discuss cost only when it matters to the request, and do not present a partial estimate as the total.\n"SKILL.md line diff
--- before +++ after @@ -1,45 +1,95 @@ --- name: creativeclaw-edit-media -description: "Edit existing video or audio with Creative Claw: trim clips, resize for social formats, add automatic captions, transcribe, clean speech, extract frames, or combine finished media. Use when the source footage or recording should be preserved; route invented scenes and generative transformations to generation skills." +description: "Edit existing video or audio with Creative Claw without regenerating it: trim, cut and reorder chosen moments, resize or reframe, caption, transcribe, clean speech, extract frames, add or mix audio, burn a logo watermark, compose images and clips, or add an intro/outro. Use when the source footage should be kept; use generation skills for invented or transformed footage and create-reels to pick highlights." --- # Edit Existing Media Read [shared execution guidance](references/workflow-basics.md) once per task before using tools. It covers existing authorization, model discovery, optional cost checks, imports, and recovery. -Turn supplied footage or audio into a finished derivative. Use the user's requested edits and sensible defaults without a settings questionnaire. Preserve the original asset. +Turn supplied footage or audio into a finished derivative. Apply the requested edits with sensible defaults and no settings questionnaire. Keep the original asset unchanged. -## Select only necessary operations +## Pick only the operations you need -| Requested change | Tool and important boundary | +| Requested change | Tool and key limits | | --- | --- | -| Shorten a video | `trim_video`: set `start_time` explicitly, including zero, and either `end_time` or `duration`. | -| Resize or make a vertical version | `scale_video`: set even dimensions and an explicit `mode`. `crop` is center crop; `pad` preserves the full frame; `stretch` distorts it. This is not subject tracking or AI enhancement. | -| Burn automatic captions | `add_subtitles`: operates on the current video and transcribes automatically. Set the known spoken `language`; do not translate unless requested. | -| Get a transcript or find a spoken passage | `transcribe`: pass exactly one of `audio_url` or `video_url`; reuse existing timing when available. | -| Clean a noisy recording | `isolate_audio`: preserve and compare the original; cleanup can alter speech. | -| Get still frames | `extract_frames`: discover the current tool schema for first/last/sampling controls. | -| Combine clips or add one finished audio track | `merge_media`: read [assembly guidance](references/media-assembly.md). | -| Enhance resolution or remove a background | `upscale_media` or `remove_background`, only for the requested media type and supported parameters. | +| Shorten a video to one range | `trim_video`: set `start_time` explicitly, including zero, and either `end_time` or `duration`. | +| Cut, reorder, or reframe several chosen moments | `cut_and_reframe_video`: see [Multi-cut edits](#multi-cut-edits). | +| Resize or make a vertical version | `scale_video`: even dimensions and an explicit `mode`. `crop` is a center crop, `pad` keeps the full frame, `stretch` distorts. No subject tracking. | +| Burn automatic captions | `add_subtitles`: transcribes the current video itself. Set the spoken `language`; do not translate unless asked. | +| Get a transcript or find a passage | `transcribe`: exactly one of `audio_url` or `video_url`. | +| Clean a noisy recording | `isolate_audio`: keep and compare the original; cleanup can change speech. | +| Get still frames | `extract_frames`: check the schema for first/last/sampling controls. | +| Put narration or music on a clip | `merge_media` `merge_audio_video`: see [Audio](#audio). | +| Join clips or audio files in order | `merge_media` `merge_videos` or `merge_audios`. | +| Burn a logo or copyright mark onto a finished video | `merge_media` `overlay_images` with a transparent image. `remove_background` can prepare a logo. | +| Build a video from timed images, clips, and optional audio | `merge_media` `compose_video`. | +| Add an intro, outro, or both | See [Intros and outros](#intros-and-outros). | +| Upscale or remove a background | `upscale_media` or `remove_background`, for the media type and parameters they support. | -Use `creativeclaw-generate-video` for invented footage or a generative source-video transformation. Use `creativeclaw-add-video-intro-outro` for bookends. HTML rendering requires the user's explicit choice; a captions request alone does not select that route. +Read [assembly guidance](references/media-assembly.md) before any `merge_media` call. -Use `creativeclaw-create-reels` when selecting highlights from long footage, or `creativeclaw-cut-and-reframe-video` for explicit multi-cut edits, moving crop paths and source-timed word captions. Both require `cut_and_reframe_video` to be available and configured. +Route elsewhere: `creativeclaw-generate-video` for invented footage or a generative change to the source; `creativeclaw-create-reels` when the agent must pick highlights from long footage; `creativeclaw-render-html` only when the user explicitly wants HTML or HyperFrames rendering. A caption request alone does not select HTML. For an exact text watermark, `render_html_image` with `transparent_background: true` can make the PNG for `overlay_images`. ## Workflow -1. Reuse the known asset or import the missing source through [platform-upload.md](references/platform-upload.md). Use available metadata/playback to establish duration, aspect ratio, audio, and protected content; do not invent measurements. -2. Apply the smallest requested edit chain. For a user-specified range, trim directly. Transcribe first only if text or timing is needed to choose cuts. For a clip selection request, use the transcript to choose a coherent section and verify visual relevance when possible. -3. Finish cuts and merges before captions. Resize to the final framing before burning text. For vertical delivery, use center crop only when the subject remains visible; otherwise use padding or ask about a material framing tradeoff. Do not claim automatic face tracking. -4. Resolve each queued operation with `check_job` before passing its result into the next. On failure, follow [job-recovery.md](references/job-recovery.md) and resume from the latest completed derivative. -5. Add automatic subtitles to the final edited video if requested. The current tool accepts style/language settings, not an edited transcript or SRT payload. Do not promise verbatim corrected or translated captions through unsupported fields. -6. Inspect final length, framing, spoken content, caption sync/safe areas, and audio using available capabilities. Name/tag the result with the project and edit role; provide the finished asset and any material inspection limitation. - -For “make this a captioned vertical clip,” reuse the supplied footage, trim only if a shorter clip was requested, resize with a suitable explicit mode, then add captions. Skip model discovery and generation when all requested changes are supported processing operations. - -## Limits that change the plan - -- A public YouTube URL can go directly to `transcribe` through ElevenLabs Scribe. It is not a directly downloadable source for trimming, resizing, or subtitle rendering; obtain the actual video asset for those steps. -- Automatic subtitles already transcribe. Do not add a separate transcription job unless the user also needs a transcript or it guides clip selection. -- Mixing narration with music, ducking, crossfades, or exact caption-file rendering require additional capabilities. Explicit moving crops and source-timed word captions are supported by `cut_and_reframe_video` when available; automatic face tracking is not. State the specific missing operation before creating assets that cannot be assembled as requested. -- `estimate_generation` does not estimate these processing operations. Only discuss cost when relevant to the user's request, and do not label a partial generation estimate as the total edit cost. +1. Reuse the known asset or import the source through [platform-upload.md](references/platform-upload.md). Establish duration, aspect ratio, audio, and burned-in text from metadata and playback; do not invent measurements. +2. Run the smallest edit chain. Trim a user-given range directly. Transcribe first only if text or timing is needed to choose cuts. +3. Finish cuts and merges first, then resize to the final framing, then burn captions. Use a center crop only when the subject stays visible; otherwise pad or ask about the framing tradeoff. +4. Resolve each queued job with `check_job` before passing its result on. On failure, follow [job-recovery.md](references/job-recovery.md) and resume from the last completed derivative. +5. `add_subtitles` takes style and language settings, not an edited transcript or SRT. Do not promise corrected or translated captions through fields it lacks. +6. Check final length, framing, speech, caption sync and safe areas, and audio. Name and tag the result, and state any check you could not do. + +For "make this a captioned vertical clip": reuse the footage, trim only if a shorter clip was asked for, resize with an explicit mode, then add captions. No model discovery or generation is needed. + +## Multi-cut edits + +Use `cut_and_reframe_video` when the user has chosen the moments (or `creativeclaw-create-reels` has) and wants them cut, reordered, or reframed in one render. It does not pick highlights, transcribe, or track faces. If the tool is missing or returns "not configured", say so. Do not fake an exact multi-cut edit with a chain of lossy trims. + +Read [the cut-and-reframe contract](references/edit-contract.md) before building input. + +1. The source must be a video asset in this workspace. Keep an unchanged source URL and a source-timed transcript. Inspect metadata and frames; do not invent sizes, word times, or coordinates. +2. Choose safe boundaries. Listen around each edge when you can. Avoid clipped consonants, breaths, and cut-off reactions. Prefer a natural gap, keep enough lead-in and tail, and end on a short release without entering the next word. The renderer does not extend ranges. Do not remove every pause or change what someone meant. +3. Set the requested width and height (1080×1920 is only the default). Use `pad` when position is uncertain. Use `center_crop` only after checking the subject stays in frame. For `crop`, give an even rectangle at the output aspect ratio, with static coordinates or keyframes set from real observations. If a speaker leaves the crop, revise the path or pad; never stretch or guess. +4. Captions are optional. Pass real source-timed words; the tool remaps them after the cuts and burns them last. With only sentence timing, render the cuts first and run `add_subtitles` on the result. Never fabricate word alignment. Use plain captions for overlapping speakers or right-to-left text unless you can review karaoke visually. +5. Submit one call per output. Save the job ID and the exact edit plan. After a timeout, poll the existing job; do not resubmit blindly. +6. Check every splice with about 1.5 seconds of context on each side, the first and last words, crop continuity, caption spelling and sync, and audio. Technical checks do not prove the edit reads well. Fix only identified faults; do not rerender speculatively. + +## Intros and outros + +Get the bookend segments, then concatenate them around the main video. + +- **Supplied clips:** import and use them. +- **HTML title cards:** only when the user asks for HTML, HyperFrames, or code-rendered cards, or accepts that option when you offer it for exact text, fonts, and logos. Render them with `creativeclaw-render-html`. An intro/outro request alone does not authorize `render_html_video`. +- **Generated cinematic bookend:** use `creativeclaw-generate-video`. To match the look, `extract_frames` from the main video and pass the frame as a style reference. + +1. Establish the main video's URL, size, aspect ratio, frame rate, audio, and platform. +2. Use the requested copy, logo, duration, and audio. Ask only about missing copy; do not invent a slogan or call to action. Keep bookends short. +3. Make segments at the main video's size and frame rate. Resolve and inspect each one. +4. Merge in playback order. Set `canvas_video_index` to the main video's zero-based index and `video_fit: "pad"` to keep mismatched frames whole, unless cropping is authorized. + +```text +merge_media({ + operation: "merge_videos", + video_urls: ["<intro-url>", "<main-video-url>", "<outro-url>"], + canvas_video_index: 1, + video_fit: "pad", + pad_color: "black" +}) +``` + +Omit a bookend that was not requested. `merge_videos` is a hard cut: no dissolves, crossfades, or audio carried across the join. If the main audio must continue under a bookend or fade across a cut, say so before making segments. + +## Audio + +- `merge_audio_video` replaces the clip's audio by default and ends at the shorter input. Check both durations first; do not silently shorten the video to fit a voiceover. +- `audio_mode: "mix"` keeps the clip's own sound and layers the new track over it, at the video's full length. Set `original_volume` and `added_volume` (0–1); about 0.3 keeps music under speech. +- `compose_video` layers one audio track over the clips at full level, and the output can run past the last clip. Use it when the added audio should extend beyond the picture. +- `merge_audios` joins files end to end; it does not layer them. +- There is no ducking, volume automation, or crossfade. Say which operation is missing before creating assets that cannot be assembled as asked. + +## Limits + +- A public YouTube, Google Drive, or social video page can go straight to `transcribe` as `video_url`. It is not a source for trimming, resizing, or subtitles; import the actual file for those. +- `add_subtitles` already transcribes. Add a separate `transcribe` job only for a transcript the user wants or for choosing cuts. +- `estimate_generation` does not cover these processing operations. Discuss cost only when it matters to the request, and do not present a partial estimate as the total.
Full snapshot data
{
"description": "Edit existing video or audio with Creative Claw without regenerating it: trim, cut and reorder chosen moments, resize or reframe, caption, transcribe, clean speech, extract frames, add or mix audio, burn a logo watermark, compose images and clips, or add an intro/outro. Use when the source footage should be kept; use generation skills for invented or transformed footage and create-reels to pick highlights.",
"included_files": [
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 603
},
{
"relative_path": "references/edit-contract.md",
"size_in_bytes": 4098
},
{
"relative_path": "references/job-recovery.md",
"size_in_bytes": 3786
},
{
"relative_path": "references/media-assembly.md",
"size_in_bytes": 8654
},
{
"relative_path": "references/platform-upload.md",
"size_in_bytes": 1416
},
{
"relative_path": "references/workflow-basics.md",
"size_in_bytes": 7260
}
],
"name": "creativeclaw-edit-media",
"skill_md_contents": "---\nname: creativeclaw-edit-media\ndescription: \"Edit existing video or audio with Creative Claw without regenerating it: trim, cut and reorder chosen moments, resize or reframe, caption, transcribe, clean speech, extract frames, add or mix audio, burn a logo watermark, compose images and clips, or add an intro/outro. Use when the source footage should be kept; use generation skills for invented or transformed footage and create-reels to pick highlights.\"\n---\n\n# Edit Existing Media\n\nRead [shared execution guidance](references/workflow-basics.md) once per task before using tools. It covers existing authorization, model discovery, optional cost checks, imports, and recovery.\n\nTurn supplied footage or audio into a finished derivative. Apply the requested edits with sensible defaults and no settings questionnaire. Keep the original asset unchanged.\n\n## Pick only the operations you need\n\n| Requested change | Tool and key limits |\n| --- | --- |\n| Shorten a video to one range | `trim_video`: set `start_time` explicitly, including zero, and either `end_time` or `duration`. |\n| Cut, reorder, or reframe several chosen moments | `cut_and_reframe_video`: see [Multi-cut edits](#multi-cut-edits). |\n| Resize or make a vertical version | `scale_video`: even dimensions and an explicit `mode`. `crop` is a center crop, `pad` keeps the full frame, `stretch` distorts. No subject tracking. |\n| Burn automatic captions | `add_subtitles`: transcribes the current video itself. Set the spoken `language`; do not translate unless asked. |\n| Get a transcript or find a passage | `transcribe`: exactly one of `audio_url` or `video_url`. |\n| Clean a noisy recording | `isolate_audio`: keep and compare the original; cleanup can change speech. |\n| Get still frames | `extract_frames`: check the schema for first/last/sampling controls. |\n| Put narration or music on a clip | `merge_media` `merge_audio_video`: see [Audio](#audio). |\n| Join clips or audio files in order | `merge_media` `merge_videos` or `merge_audios`. |\n| Burn a logo or copyright mark onto a finished video | `merge_media` `overlay_images` with a transparent image. `remove_background` can prepare a logo. |\n| Build a video from timed images, clips, and optional audio | `merge_media` `compose_video`. |\n| Add an intro, outro, or both | See [Intros and outros](#intros-and-outros). |\n| Upscale or remove a background | `upscale_media` or `remove_background`, for the media type and parameters they support. |\n\nRead [assembly guidance](references/media-assembly.md) before any `merge_media` call.\n\nRoute elsewhere: `creativeclaw-generate-video` for invented footage or a generative change to the source; `creativeclaw-create-reels` when the agent must pick highlights from long footage; `creativeclaw-render-html` only when the user explicitly wants HTML or HyperFrames rendering. A caption request alone does not select HTML. For an exact text watermark, `render_html_image` with `transparent_background: true` can make the PNG for `overlay_images`.\n\n## Workflow\n\n1. Reuse the known asset or import the source through [platform-upload.md](references/platform-upload.md). Establish duration, aspect ratio, audio, and burned-in text from metadata and playback; do not invent measurements.\n2. Run the smallest edit chain. Trim a user-given range directly. Transcribe first only if text or timing is needed to choose cuts.\n3. Finish cuts and merges first, then resize to the final framing, then burn captions. Use a center crop only when the subject stays visible; otherwise pad or ask about the framing tradeoff.\n4. Resolve each queued job with `check_job` before passing its result on. On failure, follow [job-recovery.md](references/job-recovery.md) and resume from the last completed derivative.\n5. `add_subtitles` takes style and language settings, not an edited transcript or SRT. Do not promise corrected or translated captions through fields it lacks.\n6. Check final length, framing, speech, caption sync and safe areas, and audio. Name and tag the result, and state any check you could not do.\n\nFor \"make this a captioned vertical clip\": reuse the footage, trim only if a shorter clip was asked for, resize with an explicit mode, then add captions. No model discovery or generation is needed.\n\n## Multi-cut edits\n\nUse `cut_and_reframe_video` when the user has chosen the moments (or `creativeclaw-create-reels` has) and wants them cut, reordered, or reframed in one render. It does not pick highlights, transcribe, or track faces. If the tool is missing or returns \"not configured\", say so. Do not fake an exact multi-cut edit with a chain of lossy trims.\n\nRead [the cut-and-reframe contract](references/edit-contract.md) before building input.\n\n1. The source must be a video asset in this workspace. Keep an unchanged source URL and a source-timed transcript. Inspect metadata and frames; do not invent sizes, word times, or coordinates.\n2. Choose safe boundaries. Listen around each edge when you can. Avoid clipped consonants, breaths, and cut-off reactions. Prefer a natural gap, keep enough lead-in and tail, and end on a short release without entering the next word. The renderer does not extend ranges. Do not remove every pause or change what someone meant.\n3. Set the requested width and height (1080×1920 is only the default). Use `pad` when position is uncertain. Use `center_crop` only after checking the subject stays in frame. For `crop`, give an even rectangle at the output aspect ratio, with static coordinates or keyframes set from real observations. If a speaker leaves the crop, revise the path or pad; never stretch or guess.\n4. Captions are optional. Pass real source-timed words; the tool remaps them after the cuts and burns them last. With only sentence timing, render the cuts first and run `add_subtitles` on the result. Never fabricate word alignment. Use plain captions for overlapping speakers or right-to-left text unless you can review karaoke visually.\n5. Submit one call per output. Save the job ID and the exact edit plan. After a timeout, poll the existing job; do not resubmit blindly.\n6. Check every splice with about 1.5 seconds of context on each side, the first and last words, crop continuity, caption spelling and sync, and audio. Technical checks do not prove the edit reads well. Fix only identified faults; do not rerender speculatively.\n\n## Intros and outros\n\nGet the bookend segments, then concatenate them around the main video.\n\n- **Supplied clips:** import and use them.\n- **HTML title cards:** only when the user asks for HTML, HyperFrames, or code-rendered cards, or accepts that option when you offer it for exact text, fonts, and logos. Render them with `creativeclaw-render-html`. An intro/outro request alone does not authorize `render_html_video`.\n- **Generated cinematic bookend:** use `creativeclaw-generate-video`. To match the look, `extract_frames` from the main video and pass the frame as a style reference.\n\n1. Establish the main video's URL, size, aspect ratio, frame rate, audio, and platform.\n2. Use the requested copy, logo, duration, and audio. Ask only about missing copy; do not invent a slogan or call to action. Keep bookends short.\n3. Make segments at the main video's size and frame rate. Resolve and inspect each one.\n4. Merge in playback order. Set `canvas_video_index` to the main video's zero-based index and `video_fit: \"pad\"` to keep mismatched frames whole, unless cropping is authorized.\n\n```text\nmerge_media({\n operation: \"merge_videos\",\n video_urls: [\"<intro-url>\", \"<main-video-url>\", \"<outro-url>\"],\n canvas_video_index: 1,\n video_fit: \"pad\",\n pad_color: \"black\"\n})\n```\n\nOmit a bookend that was not requested. `merge_videos` is a hard cut: no dissolves, crossfades, or audio carried across the join. If the main audio must continue under a bookend or fade across a cut, say so before making segments.\n\n## Audio\n\n- `merge_audio_video` replaces the clip's audio by default and ends at the shorter input. Check both durations first; do not silently shorten the video to fit a voiceover.\n- `audio_mode: \"mix\"` keeps the clip's own sound and layers the new track over it, at the video's full length. Set `original_volume` and `added_volume` (0–1); about 0.3 keeps music under speech.\n- `compose_video` layers one audio track over the clips at full level, and the output can run past the last clip. Use it when the added audio should extend beyond the picture.\n- `merge_audios` joins files end to end; it does not layer them.\n- There is no ducking, volume automation, or crossfade. Say which operation is missing before creating assets that cannot be assembled as asked.\n\n## Limits\n\n- A public YouTube, Google Drive, or social video page can go straight to `transcribe` as `video_url`. It is not a source for trimming, resizing, or subtitles; import the actual file for those.\n- `add_subtitles` already transcribes. Add a separate `transcribe` job only for a transcript the user wants or for choosing cuts.\n- `estimate_generation` does not cover these processing operations. Discuss cost only when it matters to the request, and do not present a partial estimate as the total.\n"
}SHA-256 of public snapshot: cf16446f1c9b2269e59aba5ae059bf5f022923c1507bb22c175411dbbb88a036