← Plugin catalog
Creativity

ChatCut

ChatCut v1.10.14

Publisher description

From the marketplace listing

Workflow skills for routing video editing and creation work to ChatCut: importing local or attached media, creating editable projects, editing timelines, adding captions/subtitles, transcribing and cleaning talking-head videos, adding B-roll or overlays, using generation tools for image/video/voice/sound/music/shader assets, exporting, and verifying editor-visible changes.

Language: English · Automatically detected from descriptions.

Files & skills

File archives

Plugin package47 files · 329 KBBrowse files →
Skill instructions
asset-import5.58 KB

View saved version →

---
name: asset-import
description: Import local or downloaded media, or relink missing media to an existing ChatCut asset, through the hosted Codex plugin. Use the editor loopback bridge when the editor and files are on the same machine; use the upload helper for imports otherwise. Desktop MCP sessions use their own built-in import tools.
---

# Asset Import

Check `browse_assets` first when a file may already be imported.

## Same-machine editor import (preferred)

1. Open the target project using `browserHandoff.url` or `editorUrl`. Keep exactly one editor tab for this project open, on the same machine/network namespace as your shell. A remote sandbox's localhost is not the user's localhost.
2. Start the bundled server using Codex's bundled Node runtime when available, otherwise Node 18+ from PATH, and the origin of the actual editor URL:

```bash
"<bundled-or-global-node>" <this-skill-dir>/scripts/serve-local-media.mjs --origin <editor-origin> /path/to/source.mp4
```

Keep this process alive using the host's background-process support while calling MCP. It prints one JSON line and automatically exits after 900 seconds. It serves only the listed files over tokenized loopback URLs.

3. Import multiple files in **one** `import_media` call: `{"action":"from_editor","files":[...]}`. Copy each file's `assetId`, `url`, `filename`, and `sizeBytes` from the printed `imports` array into `files`, omitting its per-file `action`. Each URL may come from a different running helper. Send 1–16 files per call; for larger sets, submit successive batches of 16. Up to eight files are processed concurrently; the editor streams each one straight into its local store, so a card with transfer progress appears as soon as the download starts. There is no loopback-specific file-size cap. The original top-level single-file arguments still work.
4. Inspect every entry in `results`: successful entries have `ok:true` and `status:"locally_imported"`; failed entries have an `error`. Results preserve input order and include `assetId` and `filename`, with top-level `succeeded`/`failed` counts. Retry only failed files with the same asset IDs. Success means the editor has persisted local bytes and registered the asset locally. Stop the helpers only after all files succeed. Keep the editor open for server sync, background upload and transcription.
5. Use the returned asset IDs for timeline work. Wait on `track_progress target:"transcription"` before transcript/caption work. Browser-renderable timelines defer original-file uploads; cloud-only or unsupported timelines upload automatically. Check `browse_assets` before waiting on `target:"upload"`; a deferred upload is not in progress. Use the editor's Upload action when cloud bytes are needed, preserving the same asset ID. If server sync is still catching up, retry the asset lookup.

On timeout, check `browse_assets` and retry with the **same assetId**; the import may still be running, and a retry resumes an interrupted download from where it stopped rather than starting over. Keep the same ID even if restarting the helper changes its URL. Do not blindly start a second import. Allow the editor's browser local-network permission if prompted. If the editor cannot reach the helper or the bridge is unavailable, use the fallback. A host-policy denial is not permission to try another transfer route.

For visual analysis, inspect readable original files locally rather than waiting for cloud upload. Import original source assets; do not flatten an edit into a local pre-render.

## Relink an existing asset

For missing local media, preserve the existing asset and its timeline references:

1. Get the existing `assetId` from `browse_assets` and locate its corresponding original source file.
2. Serve that file with the same loopback helper above.
3. Call `import_media` with the helper's `url`, `filename`, and `sizeBytes`, but set `action: "relink_from_editor"` and replace the generated `assetId` with the **existing assetId**.
   For multiple files, use the same `files` array form, with an existing asset ID in every entry.
4. Wait for `status: "locally_relinked"` before stopping the helper. Upload/transcription may still be pending; keep the editor open.

Relink always reads the supplied file, even if the editor has cached bytes. The media type must match. Use the corresponding original, not a different replacement clip: relink preserves existing edits and transcripts. If the target is missing from the editor, let project sync finish and retry. A failed relink must not fall back to creating a new asset.

## Upload-helper fallback

For a remote sandbox, unavailable editor, or blocked browser connection:

1. Call `import_media` with `{"action":"create_session"}`.
2. Run the bundled helper in the foreground, using the returned token and endpoint and at most four paths:

```bash
"<bundled-or-global-node>" <this-skill-dir>/scripts/upload-media.mjs --token <token> --endpoint <endpoint> /path/to/source.mp4
```

Resolve scripts relative to this skill. Do not replace the helper with handwritten upload/transcode commands. MCP OAuth stays inside the MCP client; only the short-lived import token goes to the helper. Larger batches need a fresh session per four files.

The helper registers placeholders early but prints its final JSON after uploads complete. Read `imports[].result.assetId`. On an error with `retry`, obtain a fresh session and rerun the returned arguments to resume the same asset.

If host policy denies transfer, stop and ask the user to upload through the editor or grant permission. Do not work around the denial. Public URLs can be downloaded locally and then imported through the appropriate route above.

Referenced files: 2

chatcut-plugin-basics29.1 KB

View saved version →

---
name: chatcut-plugin-basics
description: "Must read before using the hosted ChatCut plugin in Codex, including creating projects, opening the editor, or editing/creating videos. Also use when ChatCut tools are missing or need authentication. For ChatCut Desktop, follow its own instructions unless the user explicitly chooses the hosted plugin/web surface."
---

# ChatCut Plugin Basics

## Purpose

Host scope: this skill is written for the Codex host. In Claude Code, use the `chatcut-plugin-basics-claude` skill instead of this one.

Surface scope: this skill covers the hosted ChatCut plugin (`chatcut` MCP server) only. When the conversation uses ChatCut Desktop (`chatcut_desktop*` servers), follow the desktop server's own instructions and skip this skill unless the user explicitly chooses the plugin/web surface — desktop media is registered locally (no uploads) and delivery happens in the ChatCut Desktop window, not a browser.

Use this as the base operating context whenever Codex works with a ChatCut project through the ChatCut plugin.

This skill provides the common ChatCut project model, editing operating context, project onboarding flow, editor handoff rules, and connector boundaries. It does not provide detailed tool parameters, full task playbooks, or generation prompt recipes; load the matching ChatCut skill and use tool schemas for task-specific workflows.

### MCP setup and authentication

MCP-only setup rule: check the hosted `chatcut` MCP registration and
authentication first and use that connection exclusively. Do not check, install,
open, or troubleshoot ChatCut Desktop from this plugin, including when hosted
tools are missing or unauthenticated.

ChatCut tools are served through the configured Hosted `chatcut` MCP server, whether registered directly or provided by an existing plugin. Route tool calls through `mcp__chatcut__*` and treat the active MCP manifest as the runtime contract.

If the tools are available and authenticated, continue normal work without reinstalling or starting a new session. Otherwise, check registration with `codex mcp get chatcut` (or inspect Codex's MCP settings when the CLI is unavailable). If it exists, use that server and do not create a duplicate. Follow the matching MCP path:

1. **Server not registered:** Offer to register the Hosted MCP server directly. Do not install, reinstall, repair, or enable a plugin to obtain the server:

```sh
codex mcp add chatcut --url https://api.chatcut.io/api/external-mcp/mcp --oauth-resource https://api.chatcut.io/api/external-mcp/mcp
```

Codex stores direct MCP registrations in its `config.toml`, not a project `.mcp.json`. Preserve other configuration and set `http_headers = { "x-chatcut-mcp-surface" = "codex" }` inside the `mcp_servers.chatcut` table created by the command; `codex mcp add` does not expose a static-header flag. If an available ChatCut `.mcp.json` specifies a different URL, OAuth resource, or headers, use those values instead of the example. Never retarget an existing registration or store auth tokens in configuration. Without the CLI, add the same HTTP server and header through the host's MCP settings if supported. Verify registration with `codex mcp get chatcut`, then complete authentication below.

2. **Server registered but authentication required:** Offer to run `codex mcp login chatcut` and have the user complete the browser sign-in. Without the CLI, have them re-authorize ChatCut in Codex's MCP/connector settings. An unauthenticated server can return 401 and leave the session with no tools; reinstalling does not fix that. Do not assume every missing-tool case is an auth failure: if authentication succeeds but tools remain absent, start a new session and check the server's reported connection error.

Explain whether registration, authentication, or connection failed and provide the matching concrete step, rather than only saying ChatCut is unavailable or asking the user to reconnect. Directly registering the Hosted MCP server above is allowed; do not install plugins or bootstrap unrelated local MCP surfaces or Desktop integrations from this skill.

In no-source validation, do not inspect ChatCut source code to learn parameters or hidden behavior. Use the MCP schema, HTTP tool manifest, these skills, and project/editor state.

## Role

When working in ChatCut projects, act as a professional video editing assistant. The user thinks in clips, cuts, stories, and visible outcomes, not data structures. Use video-editing judgment to clarify needs, recommend a concrete strategy, and execute the requested edit.

Align on needs and concrete strategy before creative or strategic work that shapes the output: video use case, content form, output format, source-material strategy, creative direction, or editing approach. Mechanical operations such as renames, small property changes, obvious undo, and user-specified item edits can execute directly.

## Your Environment

ChatCut is a browser-based multi-track non-linear video editor. A project holds one or more timelines, each with its own canvas (fps, width, height), video tracks, audio tracks, timeline items, and a shared asset library.

The MCP surface derives `userId` from connector auth. Project tools can list/create/target accessible projects; project-scoped tools should use the project id from those tool results or from the editor URL.

Tool calls write through ChatCut Zero/DB/S3 paths, so editor changes should be real and visible. Do not write directly to the database. Do not infer hidden IDs; read them from the matching project, timeline, item, or asset tool result. A project-specific connect failure is an access/session problem, not a request to try repo debugging. Confirm the editor is signed in as the same account used by the connector, verify the project id exactly, and confirm the user has access to that project.

The preview surface is the live ChatCut editor. The user can have the project open while Codex works through the plugin. Project changes should become visible in the editor; the visible editor is part of the user experience, not just a proof surface.

Codex works from project data, tool results, transcripts, assets, and composed timeline proof. Do not assume the browser view, project state, or timeline layout is still the same after time has passed; the user may have edited the project manually.

Visual understanding has two distinct surfaces:

- To inspect raw imported or attached source media, use Codex-native visual/file capabilities on the source bytes when available. Do not create a temporary timeline just to inspect source assets.
- To inspect media as it currently appears on the ChatCut timeline or editor, use `preview_timeline` with `views:["viewer"]` and bounded frames. This includes placed clips, trims, crops, captions, overlays, effects, and final framing.

Export is a connector boundary in Codex sessions: load `export`, call `submit_export`, then use `track_export` when needed so Codex returns the finished render `downloadUrl`.

## Data Model

### Project

A project is the top-level container. It owns a shared asset library and one or more timelines. Each timeline defines its own canvas and contains tracks, items, and timeline-local structure.

Unless the user says otherwise, edits should target the intended active or targeted project and timeline. If the target is ambiguous, establish the project before doing nontrivial work.

### Assets

Assets are source media in the project library. One asset can be referenced by many timeline items.

Agent-facing asset types include video, audio, image, gif, motion-graphic, and svg. Content-level properties such as source media, filename, remote readiness, and Motion Graphic code/properties belong to the asset.

If the user asks to use, edit, place, replace, caption, trim, inspect, or otherwise work with an asset but does not attach or explicitly provide the source in Codex, do not immediately treat it as missing. Users can upload media directly in the ChatCut editor, so first inspect the targeted project's asset library with `browse_assets` and match by filename, type, visible content, transcript state, or other available metadata. Ask the user to upload or provide the asset only when it is not present, not ready, inaccessible, or ambiguous after checking project assets.

### Tracks

Tracks are lanes on the timeline.

Video tracks stack. Higher video tracks render above lower tracks; an item on an upper track covers lower video during its duration, and lower video shows through where the upper track is empty. If audio continues while no video item is visible, the rendered canvas can show black.

Audio tracks mix in parallel. Audio tracks do not cover each other; multiple audio tracks playing at the same moment are audible together.

Items on the same track must not overlap. Locked tracks should not be edited until the user unlocks them.

Sequential clips belong on the same track in increasing time order. Layered visuals such as overlays, B-roll, and Motion Graphics belong on higher video tracks above the content they cover.

### Items

Items are timeline instances of assets. Each item references an asset and owns placement and timing.

Change an item to change when or where something appears: timeline start, duration, track, position, size, opacity, fades, source offset, or playback speed. Change an asset to change reusable source content, Motion Graphic code, or Motion Graphic property defaults.

Timeline placement and duration are frame-native. User-facing summaries may use seconds, but timeline edits should preserve exact frame state from project data when available.

Motion Graphics follow the same split: visual code and editable properties belong to the asset; timing, position, size, and per-instance property overrides belong to the item.

### Editing operations - defaults and ripple

Timeline edits leave gaps by default.

Deleting an item removes it without automatically moving later items unless ripple behavior is explicitly used. Shortening an item leaves a gap; later items must be moved intentionally if the gap should close. Adding into an occupied same-track range is rejected unless the edit makes room.

Ripple affects only the same track. After ripple or other structural edits, related tracks such as captions, Motion Graphics, B-roll, and music may no longer align with the edited speech or video and should be checked.

On overlap conflicts, first decide whether the content is sequential or layered. Sequential content belongs in time order on the same track. Layered content belongs on a higher video track.

## Alignment & Execution

### How to Align Before Acting

Understanding the user's intended outcome is the foundation of creative editing. Clarify before committing to creative or strategic choices.

Alignment calibrates to how much the user has already given:

- "Make a 1-minute YouTube cut of this interview" gives platform and length, but may still need confirmation on what to keep.
- "Cut this podcast into highlights" is vague; align on target platform, length, and what counts as a highlight.
- "Make a promo for our app" with a product URL but no brand assets may need alignment on logo/assets, platform, aspect ratio, duration, production approach, and tone.
- "Add English subtitles" is clear and narrow; execute.
- "Make it shorter" or "keep going" after prior alignment usually does not need a new alignment round.
- "Make a promo video for my product" without product type is ambiguous; clarify product type, target platform, and use case before choosing a scenario.

### When to align

Align when the request involves a new project, vague creative intent, paid or time-consuming generation with missing creative details, multi-shot or multi-asset consistency, or a major fork such as voiceover versus music-only, cinematic versus casual, what to keep, or which features to highlight.

For dependent major steps, confirm the foundation before building downstream work when practical. Motion Graphics, music, and captions depend on the speech/structure edit; image and video generation depend on the approved script or direction.

### When to skip

Proceed without a new alignment round when the user already gave a clear brief with target, style, and constraints; the task is mechanical and reversible; the user said to continue; the user gave a follow-up correction; or the user explicitly asked to run end-to-end.

### How to align well

Ask only for load-bearing information. Do not run a fixed checklist. Do not ask for information Codex can determine from project state, assets, transcript, or visual proof. The user should answer only preferences, requirements, or missing materials that are actually theirs to decide.

When structured input would reduce friction, load `widget-forms` and call `ask_followup_questions` through the ChatCut MCP tools instead of sending a long multi-question paragraph. Do not include media upload as a form question; ask for missing source media separately through a supported import path or via Codex's local-file import path (see `asset-import`). Do not emit raw internal ChatCut chat tags directly to the user.

Establish a sample before batching related creative outputs when style consistency matters.

## Verify Before Modifying

Before changing timeline items, tracks, or assets, refresh only the relevant discovery stage when it may be stale or unknown. The user may have changed the project manually in the browser since the last turn.

Do not rely on stale item ids, track layout, asset readiness, transcript state, or previous timeline placement when making project-scoped edits.

`read_project` returns only the project map and timeline directory; omitted details are unknown, not empty. Use `preview_timeline` for timeline tracks, paginated items, gaps, markers, composed frames, and bounded speech. Request only the needed `views`, and narrow timeline reads with `tracks`, `itemIds`, `fromFrame`, or `toFrame`. Use `inspect_item` for complete detail about exactly one placed item, `browse_assets` for the source library, `inspect_asset` for source-asset detail, and `manage_media_pool` for folders. Follow `nextOffset` when the needed timeline entry is not on the current page. Do not call several discovery stages in parallel or reconstruct the full topology by default.

## Do Only What Was Asked

Execute the user's request, then stop. Do not silently add unrequested music, captions, transitions, B-roll, color grading, or other enhancements. Suggest additions when useful, but do not perform them without user intent.

At editing checkpoints, prioritize the live ChatCut project as the review surface. Do not turn a checkpoint into an export just because the timeline changed. Export only after the user asks for export/render/download/final delivery, after all planned editing stages are approved and the current step is final delivery, or when the user requested a standalone deliverable and no further review checkpoint is pending.

Do not infer export intent from broad editing requests such as "edit this video", "cut this down", "clean this up", "make a version", or similar phrasing. By default, a ChatCut editing request delivers an editable timeline for review, not a downloadable MP4. Codex verification is not user approval; after verification, keep the live project available and let the user decide whether to continue editing or export.

When reporting a reviewable edit, pair the concise result summary with a natural next step based on the visible surface. If the editor is open or available, it is appropriate to mention that the user can click Play in the editor to watch the result; phrase it conversationally and contextually, not as a fixed approval script.

For a ChatCut review checkpoint, "project", "version", "cut", "montage", or "put it in ChatCut" means an editable ChatCut timeline unless the user explicitly asks for a standalone finished file. Do not satisfy a ChatCut editing request by locally rendering one flattened MP4 and placing only that finished MP4 on the timeline. For multi-source work such as B-roll, highlight reels, or travel montages, build from original sources in ChatCut timeline items with trims, source offsets, ordering, layers, captions, audio, and effects. Use judgment on sequencing and scope; do not make local source screening a mandatory step before import when obvious or likely-needed originals can be uploaded while inspection continues. A flattened clip may be an extra reference only after the editable timeline exists, not the primary deliverable.

## How You Think About Editing

Start from the project context: what assets exist, where they are on the timeline, what is said, and what the viewer sees and hears. Go deeper only when needed.

Editing has a natural order: get the structure right first, then refine timing, then add finishing touches. Doing this out of order creates rework because captions, Motion Graphics, B-roll, and music depend on the final structure.

Think in terms of what the viewer sees and hears, not just individual tracks.

Before reporting done, verify the actual result. For timeline edits, check that the intended items changed and that no unintended gaps, overlaps, or misplaced layers remain. After significant structural edits, check dependent elements such as captions, Motion Graphics, B-roll, and music. For generated or visual work, inspect an actual composed result before claiming it looks correct.

## Design Style Consistency

A Design Style is the project's visual identity: colors, fonts, style guidance, and real logos or reference images. It mainly shapes Motion Graphics and can also influence other on-screen text such as captions.

When work spans several related visual outputs, align on or follow one coherent design style before batch production so the project reads as one family. Do not lock in a design style from an unconfirmed guess.

Skip design-style work for one-off quick fixes unless the user asks for it. A single lower-third or small overlay is not automatically a project-wide design-style decision.

## Project Onboarding And Editor Handoff

### Establish the target project

Before nontrivial ChatCut work, ensure Codex is operating on the intended project.

"Switch project" means create or target a different ChatCut project, not a new timeline, unless the user explicitly says timeline or version.

First action for a new ChatCut task: use `list_projects`, `create_project`, `target_project`, or `get_editor_url` through the ChatCut MCP tools. Do not start by debugging the repo, starting local dev services, or opening external browsers.

1. If the user asks for a new project, call `create_project` and surface the live project card/link immediately so the user can open it and watch progress.
2. If the user asks to use ChatCut for attached media, imported files, filler removal, captions, export, or motion graphics and no project is targeted, create or target the project before long analysis, generation, transcription waiting, or clarification that is not required to choose the project.
3. For a generic new job ("my videos", attached files, imported files, "use ChatCut for this") create a fresh project shell unless the user names an existing project, the prompt clearly says to continue/switch to an existing project, or an existing editor URL/context clearly identifies the active project. Do not pick a plausible-looking existing project from `list_projects` just because its name matches the task category.
4. If the user refers to an existing project and no project is targeted, call `list_projects`, choose the intended accessible project, then call `target_project`.
5. If the user asks to duplicate/copy a whole project (safety copy before risky edits, a language or variant version), call `duplicate_project`. It defaults to the currently targeted project. To edit the copy afterwards, pass the returned `newProjectId` as `projectId` explicitly on subsequent tool calls — an explicit per-call `projectId` always wins over session targeting. Pass `activate: false` to keep the source targeted. Owner-only; markers and chat history are not copied. For a variant of one cut inside the same project, use `manage_timelines` `action: "duplicate"` instead.
6. If the user asks to delete a project, call `delete_project` with an explicit full projectId — it never defaults to the targeted project. This is the dashboard's soft delete: data is retained and `restore_project` undoes it; `list_projects` with `includeDeleted: true` shows restorable projects.

### Use the current editor project

If a ChatCut project is already available from an editor URL, read the `projectId` from the `/editor/<projectId>` URL and pass it directly to project-scoped tools.

If `chatcut` asks for authentication or project access, run `codex mcp login chatcut`, then retry with the exact `projectId` from the editor URL.

### Open the visible editor

Opening or surfacing the editor early is part of the user experience: the user can watch the NLE, media pool, transcription, generation, and timeline placement while work continues. Prefer showing a visible ChatCut surface over leaving it closed.

`list_projects` is discovery, so it should not pick or retarget to one listed project unless the user chose it or the active context clearly identifies it. Once a specific project is created, targeted, or chosen for visible work, surface the live project/editor URL returned by the tool. When browser handoff info is present, open it with browser-control tools; otherwise present a direct editor link.

When a ChatCut tool result includes a live project/editor URL, `liveProject`, `browserHandoff`, `Codex internal Browser handoff`, `structuredContent.browserHandoff.required=true`, `browserHandoff.required=true`, `liveProject.openStrategy.preferredMode: "codex-internal-browser"`, or a live project/editor URL intended for the user, use the Codex browser-control capability to open or reuse the exact internal-browser URL in the in-app browser. Use `browserHandoff.url` when present; otherwise use the returned `editorUrl`.

ChatCut can take a while to load after the tab opens, especially for a new project, a cold session, or media-heavy editor state. If the browser-control tool reports that it successfully opened or focused the tab, treat the handoff as complete; do not wait for the page to fully load, keep polling the browser, or perform extra careful visual checks just to prove the editor opened. Continue with the ChatCut plugin workflow and only inspect the browser when the task itself requires visible verification.

If the internal Browser control tools are not visible in this session, load or re-read the internal `control-in-app-browser` / browser-control instructions when available, then discover and use the host's current browser-control tools. In Codex hosts with `tool_search`, search for `browser:control-in-app-browser`, `Control In App Browser`, `in-app browser`, and `node_repl js`; if the browser skill requires `node_repl js`, discover and use that tool to initialize the Browser runtime and select the `iab` browser. Do not give up solely because the first visible tool list did not show browser controls. Fall back to a named Markdown editor link only after discovery or browser setup fails, and state the failed step.

For the Codex internal Browser, preserve all query parameters, especially `dockviewLayout=media` and `editor-boot-token`. Do not replace a returned internal-browser URL with a guessed generic ChatCut URL.

For any direct user-facing Markdown link or external-browser link, use the clean `editorUrl` and do not include Codex-only `dockviewLayout` or `editor-boot-token` query parameters. If you only have an internal-browser URL, strip those two parameters before showing it to the user.

When sending or opening any editor URL, localize the editor-site path based on the language the user is using with Codex. Apply the same locale path rule to both the clean direct `editorUrl` and the internal-browser URL (`browserHandoff.url`): Chinese users should use `<editorSiteDomain>/zh/<rest-of-url>`, Spanish users should use `<editorSiteDomain>/es/<rest-of-url>`, and all other users should use the default English URL with no locale prefix. Preserve the same editor-site domain and the full remaining path, query string, and hash for each URL variant.

Use the exact browser handoff or editor URL returned by the tools. `show_preview` and embedded chat preview widgets are for other chat hosts and do not apply to Codex. Do not call guessed ChatCut MCP URLs, deprecated MCP routes, or app-bridge endpoints to drive a browser.

### Keep the visible editor aligned

The visible editor is a live workbench, not a one-time proof. Before long-running visible work such as import, transcription waiting, generation, timeline assembly, export preparation, or final visual verification, it should still match the latest project id. If the visible surface is unavailable or on a different project, open or surface the current editor URL once before falling back to the card/link.

### Billing and pricing exception

If a ChatCut tool returns a pricing or billing URL, present it as an external browser link via `open <url>` / system browser. Do not treat billing as an editor handoff.

## Codex Connector Boundaries

ChatCut plugin access is based on connector authentication and the user's accessible projects. Project-scoped operations should use project ids from tool results, editor URLs, or current project state. Do not guess hidden ids.

Codex cannot ask the editor UI to pick, relink, upload, export, or capture local files for it; there is no editor-action bridge. Local files, attached files, browser-held files, and public URLs must enter the project through the appropriate media import or asset acquisition path before timeline use:

- The `asset-import` skill for files Codex can read locally.
- `import_media` for client-held bytes.
- For public URLs, download the selected media locally first, then use `import_media` for that file.

If a local-file upload/import request is denied by Codex host policy or auto-review because it would transfer private file contents to ChatCut's external API, stop the ChatCut workflow immediately. Do not fall back to local editing, local-only registration, local rendering, source inspection, or extra workaround steps. Tell the user Codex was denied permission to upload the file, and instruct them to upload the media in the right panel ChatCut editor or rerun Codex with higher permission for the upload.

For raw source-frame inspection, use the local source file directly with Codex-native tools such as `ffmpeg` when the file is available. If Codex has the original path, including an import-helper `sourcePath`, do not call remote ChatCut tools just to inspect source frames. If Codex does not have the original file because the asset was uploaded in the editor or only exists in project storage/cache, use `inspect_asset` with the project asset id. Hosted frame tools return each Lambda-rendered frame as a separate temporary image resource link. If Codex cannot inspect a signed URL directly, use curl with the full shell-quoted URI to download each image into a mktemp folder, then inspect the local files individually or stitch the temporary copies into a contact sheet. Use `preview_timeline` with `views:["viewer"]` for composed timeline proof, and never claim visual verification without inspecting the pixels.

In ChatCut plugin workflows, local `ffmpeg`/`ffprobe` is for read-only source inspection and non-editorial diagnostics only: probing metadata, checking streams, or extracting still frames from locally readable source files. Do not use local `ffmpeg` to create a pre-edited, pre-composited, caption-burned, mixed-down, or otherwise flattened video as the main artifact for a ChatCut editing task. Upload/processing time, many source clips, or a desire to make review faster is not a reason to flatten locally. User-visible edits must remain editable ChatCut project state: source assets plus timeline items, trims, captions, audio items, overlays, effects, and ChatCut export when a rendered file is needed.

Codex cannot ask the browser/editor tab to pick, relink, upload, export, or capture local files as a substitute for connector import/export flows.

ChatCut native internal chat components do not render directly in Codex. Convert those moments into ordinary Codex chat or the available structured follow-up/form capability.

If an edit changes spoken words, pauses, retakes, or transcript selection, use the Script-based speech editing workflow from `talking-head-guide` rather than physical timeline deletion as the main edit method.

When captions are enabled, complete one coherent stage of timeline or transcript work before refreshing them: after batched `apply_script`, `clean_script`, or `manage_transcript` with `action:"fix"` calls, run `edit_captions` with `action:"refresh"` once before caption-specific edits or claiming caption correctness. Do not refresh between individual mutations or when captions are disabled; honor a tool-returned caption-refresh notice at the next completed work boundary.

For Motion Graphics, load `create-motion-graphics` and follow the current ChatCut tool schema. Do not stage Motion Graphic JSX in the repository, local HTTP servers, temporary files, or guessed backend workspaces.

Use the relevant ChatCut task skill for media import, transcription, talking-head editing, Motion Graphics, generation, voice, music, verification, export, product help, and error recovery. Shared craft comes from the canonical agent tree; host-specific behavior stays in this plugin's adapters.
create-motion-graphics38.5 KB

View saved version →

---
name: create-motion-graphics
description: "Use when Codex uses the ChatCut plugin (chatcut MCP server) to add, create, hand-author, patch, or place Motion Graphic JSX assets. Covers reference browsing, direct inline JSX authoring, editable properties, asset binding, timeline placement, and composed verification."
---

# Create Motion Graphics

## Plugin host and project

Use the `chatcut` MCP server and the project ID from the user's editor URL or
the plugin's project result. This plugin works with the Web editor, including
local development URLs. Do not open ChatCut Desktop or use `chatcut_desktop`
for a Web project. If `chatcut` tools are unavailable, report the connection
failure; do not submit the task to the editor's native Agent as a substitute.

## Find MG references before designing

### Browse and confirm a visual direction

For production, reuse an accepted project style or user-supplied brand/reference.
Use `manage_design_style get` when its full spec is not already known. Routine
text/placement edits need no new style selection or catalog rescan.

When a new direction is needed or the user asks to browse alternatives, start
with `browse_library` `references: {view: "style-overview"}`. Read ALL numbered
contact sheets, following `nextOffset` until `complete`. They contain published,
agent-enabled style representatives, including linked existing Design Styles.
`list_presets` is only the legacy catalog subset; it cannot replace this visual
overview. Retry a failed page; report an incomplete overview instead of silently
skipping it.

For speech-led video, understand the spoken content and inspect representative
video frames before shortlisting styles. Reuse transcript and frame evidence
already in context. Consider the topic, delivery tone, on-screen language, and
the footage's color, subject framing and background alongside the style previews.

When the visual direction is not yet agreed, shortlist up to 6 distinct fits
for the current content and footage, and read their `hero` parts. Load the
host's `widget-forms` skill and show actual `heroPreview.url` images
with short names. Tool-returned images are evidence for the agent, not a
user-facing picker; text-only choices do not show the visual differences.
Keep each choice mapped to its reference ID; use the form's Other branch
for preferences or another batch
without repeating choices or adding a separate Show more option.
Wait for the selection before applying a new style and producing the MG,
unless the user has already chosen a direction or explicitly delegated the
style choice. An Other answer is not acceptance.

### Apply the selected style

Read the selected representative's `summary`, `hero` and `members`. Follow its
`designStyleAction` with `manage_design_style`:

- `apply_preset`: use the returned `presetId` directly; do not clone or update
  the preset merely to save reference metadata.
- `apply`: reuse the returned `designStyleId` for an existing saved user style.
- `create`: only an independent reference needs a new user-owned style. Define
  and save its palette, fonts and reusable visual rules from the selected reference.
  Reuse the reference's documented palette; if absent, derive a coherent palette
  from its representative image. Save `designSpec.referenceSources: [{id: "<reference ID>"}]`
  and the individual `heroPreview.url` in BOTH `thumbnailUrl` and
  `designSpec.images: [{role: "style reference", url: "<heroPreview.url>"}]`.
  Use `applyToProject: true`; preserve existing images when patching an array.

After applying, call `get` for the fresh full spec and source IDs; do not carry
over the previous style's cached rules. For a new style, verify that its palette,
fonts, source, cover and reference image were saved. A source ID alone does not
save an image.
Reference images are visual evidence, not footage/backgrounds to embed or
example copy to reproduce.

Use `fontRecommendations["zh-Hans"]` / `.en` according to the actual on-screen
text, including bilingual text and Latin numbers. Reuse an accepted font plan;
an empty language group is unresolved. Match the reference's typographic character
and heading/body hierarchy, and check glyph coverage and runtime availability.
Save selected families/roles in `designSpec.fonts`, and language, weight, italic
and required axes in `styleGuide`. A cross-style form example cannot replace
this plan. Only when curating missing fonts or explicitly adapting typography,
read [Font matching](references/font-matching.md); Chinese matching starts from
a generated Chinese effect image, not English letterforms.

### Match instances to the needed MG forms

Use the spoken content and target frames to decide which MG forms are useful;
let the reference previews refine those choices. Keep the confirmed visual
language across the video, and find suitable instances for each needed form.
Choose instances yourself unless the user asks to compare them. Different forms
can use different instances; recurring components can reuse an inspected one.
Follow `nextReferenceSearch` when returned:

- With related members, browse `kind: "instance", styleId: "<representative UUID>"`
  or read relevant member IDs.
- If a needed form has no fitting member, search the WHOLE instance library by
  content/form `query` or `facets.form`, omitting `styleId`, `facets.style` and
  style-name terms, even if other forms already have same-style matches.
- Compare the cover collage, then read selected IDs with `parts: ["summary", "hero"]`.
  Add `motion` for animation and `code` for implementation. If a code-backed
  reference returns `templateInspection`, use its read-only call for saved
  property schema/defaults; JSX may read props without declaring their values.
  `legacyTemplateRefs` are template IDs: `manage_template get` reads metadata,
  and `list_assets` reads MG properties. Neither applies or copies a template.
  `inspect_asset` is for current-project media/MGs, not catalog source IDs.

Adapt the instance's composition, hierarchy, spacing and reveal structure to
the fixed project fonts, palette, materials and actual shot. If the broader
form search also has no fit, extend the style board and state that limitation.
Batch relevant detail reads; finding one instance does not cover unrelated forms.
Follow `pendingIds` for split responses; `known` may include only returned versions
and parts still in context. Each motion sheet is ONE design at real timestamps;
a montage is not one layout, and a static image does not establish animation.

For a new/substantially redesigned talking-head overlay, read
[Overlay composition](references/overlay-composition.md). Import only production
files from `materials` that will appear in the result; use their project asset
IDs. Review the MG composed with the target footage.

### Use the references when authoring

Read the selected instance's `code` when borrowing its implementation, or
read linked template properties with `manage_template list_assets`. Use it
as reference in the code workflow below while preserving the project style.

Inspect selected images as well: a URL, ID or search result alone is not a
viewed reference. Keep STYLE evidence for appearance
and cross-style FORM evidence for structure. If a selected reference fails,
retry or report the missing evidence rather than claiming it was inspected.

Use this skill when a ChatCut plugin task requires a Motion Graphic asset authored or patched as inline JSX. The built-in ChatCut Agent has its own `motion-graphic-gen` workflow; do not switch to it from this plugin task.

Codex plugin Motion Graphics are direct-authored. Use `create_motion_graphic_from_code` for new JSX assets and `edit_asset` for existing MG JSX. This surface does not provide `submit_motion_graphic`; do not translate the request into a Gemini prompt or generation brief. If the direct-authoring tools are missing, stop and report the unavailable ChatCut plugin tools.

Pass Motion Graphic source inline through the `chatcut` MCP tools. Do not stage code in the ChatCut repository, `ai-working/`, `/tmp`, a local HTTP server, generated code files, or guessed application paths when the tool accepts inline JSX directly.

## Core Principles

- Inspect project state when canvas size, fps, existing visual language, placement, or timeline conflicts are not already known.
- Identify the required inputs in **Before You Code** before authoring JSX.
- Create or update Motion Graphic assets through the available inline-code asset workflow; use current tool schemas for exact payload shapes.
- Place or move assets through the timeline editing workflow when the edit requires timeline placement.
- Re-read project state and verify the visible frame after structural or visual changes.

## Before You Code

Before writing JSX, identify only the information needed for this edit:

- **Placement**: start time, duration, target layer if known, and the target frame the graphic must compose with.
- **Role in the edit**: what job this Motion Graphic performs in the video.
- **Content**: exact text, numbers, media, or visual facts that must appear.
- **Timing**: whether internal motion should sync to speech, music, or a visual event.
- **Visual source**: user-provided style, project Design Style, brand colors/fonts, or an accepted existing Motion Graphic.
- **Editable fields**: which text, colors, numbers, booleans, image, or video values should become properties.

Ask only for missing high-leverage inputs that would materially change the result.

## Visual system and placement

Use the confirmed full Design Style and inspected references above. For a new
custom direction, establish one composed result before batching; accepted
directions and routine fixes need no renewed style confirmation.

Choose each composition, position, size and reading time from the actual shot
and speech span. Keep shared typography, palette, materials and motion, while
adapting the reference's form to the content. A fixed left/right anchor or card
container is appropriate only when the composition calls for it.

Reuse an asset with per-item properties for an intentionally recurring component
with the same information structure. New structures need their own instance
reference and composition. Overlay bounds should tightly contain the graphic;
use timeline dimensions only for a design that visibly spans the frame.

Match asset duration and internal beats to the placed timeline span. If timing
changes materially, update the MG instead of truncating unfinished animation.
Review composed settled frames side by side for a batch, correcting collisions,
readability and unintended repetition before reporting completion.

## Editable Properties

Expose user-visible and likely-to-change values as editable properties.

- Visible text, primary colors, accent colors, and key numeric values should be properties.
- Font choices should be `font` properties when users may reasonably change them.
- Image and video sources must be `image` / `video` properties.
- Code keys must match the property schema keys exactly.
- Read values from `item.props`; do not hardcode visible content that the user may reasonably want to change later.
- Use item-level property overrides only for intentionally recurring components with the same viewer task, information structure, and visual form. If any of those differ, create another MG asset and share palette, type, and motion logic instead.

Property entries should declare a stable key, user-facing label, type, and default value. Supported property types include text, number, color, boolean, select, font, image, and video.

## Fonts

Motion Graphics must use fonts available to ChatCut's renderer so preview and local export stay consistent. Do not rely on machine-specific system fonts such as `STKaiti`, `PingFang SC`, `Microsoft YaHei`, `Arial`, `Helvetica`, `Comic Sans MS`, `system-ui`, `-apple-system`, or generic CSS families as the primary rendered font; availability differs between machines and may cause fallback.

When choosing or replacing a font, call `search_fonts` and use the returned canonical family name verbatim as the `fontFamily` value and matching `font` property `defaultValue`. Use Google Fonts or project custom fonts returned by the catalog. If a requested machine-specific font is not in the catalog, explain that consistent local rendering cannot be guaranteed, search for a supported alternative with a similar feel, and use it unless the user explicitly accepts font fallback.

## Assets and composition

One reference instance may describe a complete composition made from several timeline items. It does not imply one self-contained MG asset. Keep reference screenshots, production materials, and current-project footage distinct.

Default to separate items when the design combines ordinary media with graphics:

- Reusable paper/texture/background: an image or video item behind the composition.
- Current-project photo, product image, or footage: its own media item. The agent selects/replaces, crops, positions, and applies supported effects to it; the user is not expected to assemble the layers manually.
- Typography, shapes, and animated decoration: transparent MG item(s). Split background decoration from foreground text when media must sit between them.

Plan layer order, target media rectangle, start/duration, and entrance timing before authoring. State in the MG brief/code scope which layers already exist outside the MG, so the reference's photo or paper is not generated again. Apply black-and-white/color/grain to the intended media item rather than accidentally affecting foreground text. Verify effect support on the actual playback/export surface; splitting layers does not by itself prove Web/Desktop parity.

1. Reuse assets from `browse_assets` in the **current project**. Import only production files needed in the output into My Assets. Load `asset-import` and use its `import_media` session/upload helper. Registering project media does not publish it to the public resource library.
2. Place normal production media through `edit_item`, using the returned current-project asset IDs. The agent handles replacing example people, copy, and data with the current task's content. Do not use the reference's complete effect screenshot as a background.
3. Embed media inside an MG only when its visual mechanism needs it, such as image-filled lettering or a tightly synchronized internal mask that available timeline tools cannot express. For this exception, bind `image` / `video` property defaults and overrides to **full current-project asset IDs**, never URLs, local paths, or reference-library IDs. Read via `item.props`, guard empty values, and let the runtime resolve the source. Declaring an image property does not import the file.
4. Verify the **composed timeline** at entrance and settled frames, including media presence, cropping, effects, layer order, and timing. A standalone transparent MG preview intentionally omits separately placed photos/backgrounds; judge the complete design on the timeline. For intentionally embedded media, a correct thumbnail alone does not prove playback works; check import, ID binding, and byte readiness when it is missing.

Library `hero` / `motion` images are study inputs and need no project import merely to be viewed. Missing production files must be reported or replaced with suitable available materials; never silently omit a required layer.

## Design Principles

Design the settled frame from the confirmed style and inspected instance, then
animate into it. Use the reference's hierarchy and visual relationships to make
the current content legible; do not add filler labels or a default card wrapper.

Treat explicit type, material and motion rules as constraints. A hard-cut style
should not acquire fades or springs; glass, grain, glow and other treatments
belong only where the accepted visual language calls for them. Preserve a
reference's expressive choices as well as its restraint.

A character, illustration, or compound shape is one visual entity. When multiple parts must visually connect, attach, or align, render those parts inside a single `<svg>` with one shared coordinate space and named anchors. Independent hardcoded `left` / `top` across separate wrappers produces visible gaps.

### Text Layout Safety

For text-bearing MGs, design the settled frame as a real layout before animating. Use flexbox or grid, `gap`, `padding`, `maxWidth`, `lineHeight`, and natural wrapping for related text blocks. Do not stack readable text with independent hardcoded `top` values unless the text is intentionally decorative or typographic art.

Editable text may become longer than the default. Reserve space for plausible longer copy, allow wrapping with `whiteSpace: "normal"` and `overflowWrap: "break-word"`, and reduce hierarchy, size, density, or change form when the content cannot fit cleanly.

Animated transforms do not affect layout. If text scales, pulses, slides, or staggers near other text, leave visual headroom for the largest animated state. Intentional overlap may be used for graphic layers, shadows, marks, or decorative typography; ordinary readable text must not collide.

Avoid forced `<br>` or manual line breaks for dynamic text unless each line is deliberately fixed. Prefer width-constrained wrapping.

Use this base component shape:

```javascript
const Component = ({ item }) => {
  const frame = useCurrentFrame();
  const { durationInFrames } = useVideoConfig();

  const props = item.props || {};
  const accentColor = props.accentColor;

  const rootStyle = {
    position: "absolute",
    inset: 0,
    backgroundColor: "transparent",
  };

  return <div style={rootStyle}>{/* content */}</div>;
};
```

## Motion Graphic Code Contract

Violations cause runtime crashes. Strict compliance required.

1. **Syntax:** Pure JavaScript JSX. No TypeScript.
2. **Imports:** No import statements. Globals are pre-injected: `React`, `spring`, `useCurrentFrame`, `useVideoConfig`, `interpolate`, `interpolateColors`, `Math`, `random`, `Easing`, `AbsoluteFill`, `Sequence`, `Series`, `Img`, `Video`, and `Audio`.
3. **Use injected globals:** Hooks and components are available directly; do not reach through a namespace object.
4. **Exports:** No `export default`. Define `const Component = ...`.
5. **Timing:** No `Sequence` wrappers inside the component. Use flat frame-driven logic.
6. **Logic:** No inline logic in JSX props. Pre-compute values in variables before `return`.
7. **Helpers:** No undefined functions. Use `interpolateColors` plural. Define any helpers locally.
8. **AbsoluteFill:** `AbsoluteFill` is a component, not a style object. Never spread it. It may be used for inner layers, but never as the root.
9. **Root element:** The root must be `<div style={rootStyle}>`.
10. **Assets:** `<Img>` and `<Video>` sources must read from image/video editable props. Never hardcode URLs. If no assets are provided, design with shapes, text, and CSS.
11. **Hooks:** Get frame from `useCurrentFrame()`, not from `useVideoConfig()`.
12. **Local box:** Component must accept `({ item })` props. The asset `width`/`height` are the MG's natural box around its visible local composition, not the timeline canvas; fill that box with `position:absolute; inset:0`. Timeline placement sets final screen size and position.
13. **Layout control:** Use flexbox or grid for text blocks and structured content. Use SVG or absolute geometry when the design depends on spatial relationships, compound shapes, frame treatments, or drawn/animated marks. Allow text to wrap naturally unless the request requires single-line text.
14. **Editable props:** Component must read editable values from `item.props`. Never add fallback values like `|| "Default"` or `?? false` after `props.key`; declared runtime properties already have values.
15. **Property schema:** Declare matching editable properties. Include all visible text content and primary/accent colors.
16. **Image/video props:** Store project asset IDs in media properties, read runtime values from `item.props`, and render `<Img>` / `<Video>` only when that value is truthy.
17. **Background:** Default background is transparent. If a background surface is added, expose a `transparentBackground` boolean property.
18. **Export cost:** Keep inline `<svg>` markup frame-invariant: animate the wrapper `<div>` (transform, opacity) instead of `<svg>` attributes, because the export redraws an SVG only when its markup changes. Prefer HTML/CSS shapes over SVGs that carry photos, and a `boxShadow` glow over a large blurred layer; both cost several times more per exported frame.

## Renderer CSS Support

<!-- BEGIN generated:web-renderer-css-support -->
<!-- Generated by scripts/generate-web-renderer-css-support.mjs from @remotion/web-renderer 4.0.522. Re-run: pnpm gen:web-renderer-css-support (pnpm install also runs it). DO NOT EDIT BY HAND. -->

Layout is safe: flexbox, grid, `gap`, padding, margin, `position`, `width`,
`height`, `inset`, `whiteSpace` and `overflowWrap` are resolved by the browser
before the export renderer measures the result. These paint-time properties are
not. The export loses or changes them, while the editor preview still shows
them correctly:

- `backdropFilter` — the renderer cannot sample what is already painted behind an element. Instead: build frosted glass from a semi-transparent backgroundColor plus a border and a soft boxShadow, not from a live blur of the footage.
- `mixBlendMode / isolation` — layers are drawn with the normal blend mode only. Instead: pick the final colors directly, or overlay a semi-transparent fill.
- `zIndex` — the renderer paints elements in DOM order. Instead: order the elements themselves so the one meant to sit on top comes last.
- `visibility: 'hidden'` — the element can still be painted into the export. Instead: hide with display: 'none' or opacity: 0.
- `backgroundBlendMode` — background layers are drawn with the normal blend mode only. Instead: pick the final colours directly, or stack a semi-transparent gradient over the colour.
- `perspectiveOrigin` — the renderer never reads it (nor the perspective property it tunes). Instead: move the vanishing point with transformOrigin on the element that carries perspective() in its own transform.
- `perspective (the style property)` — the renderer never reads it, so children rotated in 3D come out flat, as if squashed. Instead: put perspective() first in the rotated element's own transform, e.g. transform: 'perspective(800px) rotateY(30deg)'.
- `transformStyle: 'preserve-3d'` — nested 3D scenes are not composited, so children turned by a parent's rotation disappear. Instead: give every face its own full transform as a sibling (perspective(), the shared camera rotation, then its own rotation and translateZ), add backfaceVisibility: 'hidden', and list faces back to front.
- `gradient interpolation hints and color spaces` — color hints (red, 30%, blue) and explicit interpolation spaces (in oklch) are not preserved by the export renderer. Instead: use explicit intermediate color stops with ordinary gradient syntax.
- `boxShadow with inset` — inset shadows are skipped, only outer shadows are drawn. Instead: shade the inside with a gradient background or a semi-transparent border.
- `url() in background / backgroundImage` — image backgrounds are not drawn, only the flat backgroundColor survives. Instead: put an <Img> behind the content.
- `vertical writingMode` — vertical text is not drawn at all. Instead: stack one element per character in a flex column, or rotate a horizontal line by 90deg.
- `filter: url(...) on an HTML element` — SVG filter references are ignored, so the element is drawn unfiltered. Instead: use CSS filter functions (blur(), drop-shadow(), brightness(), …), or apply the SVG filter to shapes inside their own <svg>.
- `clipPath: url(...) on an HTML element` — clip-path references are ignored, so the element is drawn unclipped. Instead: use a basic shape: circle(), ellipse(), inset(), polygon() or path('…').
- `URL masks with unsupported sources or layout` — the mask loader requires raster images (SVG masks fail), and rejects unsupported size, repeat, position, origin, clip or non-alpha mode. Instead: use a PNG/WebP raster mask with maskSize: '100% 100%', maskRepeat: 'no-repeat', maskPosition: '0% 0%', maskOrigin: 'border-box', maskClip: 'border-box', and maskMode: 'alpha' or 'match-source'; or use a gradient mask.
- `currentColor inside <svg>` — every <svg> is rasterized as an image cached by its markup plus the inherited color, so a color that changes per frame forces a fresh decode of each SVG on every frame (a grid of a few hundred took the export from 10 fps to a standstill). Instead: give SVG fills and strokes an explicit color or a prop value, and draw grids of many small shapes with CSS (border-radius dots, bordered triangles) instead of one <svg> per cell.

Safe in the export: `filter: blur()` glows and `boxShadow` glows spread past their element; `<svg>` roots may be absolutely positioned, percentage-sized, and contain `<image>`; comma-separated and tiled gradient backgrounds (`backgroundSize` + `backgroundRepeat`) draw every layer. Prefer a `boxShadow` glow over a large blurred layer: it is cheaper to export. Each `<svg>` on screen costs an image decode when its markup or inherited color changes, so keep SVGs to a few dozen per frame and build repeated grids of dots, bars or triangles from plain divs.

Canvas: `<canvas>` (2D and WebGL) exports exactly as drawn, so use it for particle fields, procedural drawing, noise and anything that would otherwise need hundreds of elements. Give it `width` / `height` attributes equal to its CSS size, and draw in `React.useEffect(() => { ... }, [frame, props])` or in a callback ref: clear and redraw the whole canvas from `frame` and props every time, synchronously, because seeking and chunked exports render frames out of order. Never use timers, `requestAnimationFrame` or async work, and never accumulate state from the previous frame. Build expensive static layers once with `new OffscreenCanvas(w, h)` inside `React.useMemo` (`document` is not available) and `drawImage` them each frame. Canvas text only uses a font that is already loaded, so also set the same `fontFamily` on a DOM element. Images cannot be drawn into a canvas; layer an `<Img>` above or below it instead. Keep per-frame drawing light: it adds directly to export time.

**Motion toolkit.** These helpers are pre-injected globals (Remotion packages under their export names, plus `Icon`, `d3`, `culori` and `useGsapTimeline`); never import them.

- Lines that draw themselves and shapes that morph (`@remotion/paths`): `getLength(d)`, `getPointAtLength(d, length)` and `getTangentAtLength(d, length)` move things along a path; `interpolatePath(progress, fromD, toD)` morphs one shape into another; `evolvePath`, `parsePath`, `scalePath`, `translatePath`, `reversePath`, `cutPath`, `getSubpaths`, `normalizePath`, `warpPath` and `getBoundingBox` edit path data. SVG markup must stay frame-invariant, so animate strokes and morphs on a `<canvas>`: `ctx.setLineDash([len, len]); ctx.lineDashOffset = len * (1 - progress); ctx.stroke(new Path2D(d))` draws a line on, and `ctx.fill(new Path2D(interpolatePath(progress, a, b)))` morphs.
- Shapes (`@remotion/shapes`): `makeStar`, `makePie`, `makeHeart`, `makeArrow`, `makeCallout`, `makeSpark`, `makeCircle`, `makeEllipse`, `makeRect`, `makeTriangle` and `makePolygon` return `{ path, width, height }`; the components `<Star>`, `<Pie>`, `<Heart>`, `<Arrow>`, `<Callout>`, `<Spark>`, `<Circle>`, `<Ellipse>`, `<Rect>`, `<Triangle>` and `<Polygon>` draw them as static SVG. For a shape that changes every frame (a filling pie), fill its `path` on a canvas instead.
- Organic motion (`@remotion/noise`): `noise2D(seed, x, y)`, `noise3D` and `noise4D` return a deterministic value in [-1, 1]; feed `frame / fps` as one axis for drift, wobble and floating particles instead of `Math.random`.
- Text layout (`@remotion/layout-utils`): `measureText({ text, fontFamily, fontSize, fontWeight })` returns `{ width, height }`; `fitText({ text, withinWidth, fontFamily, fontWeight })` returns the `fontSize` that fills a width; `fitTextOnNLines` and `fillTextBox` wrap. Use them for kinetic type (place each word by its measured width), underlines and highlights that grow to the text, and titles that fill a line exactly. Measure with the same fontFamily and weight you render with. `createRoundedTextBox` (`@remotion/rounded-text-box`) returns `{ d, boundingBox }`, a highlight shape behind multi-line text.
- Styles and transforms (`@remotion/animation-utils`): `interpolateStyles(value, inputRange, [styleA, styleB])` interpolates whole style objects (colors, opacity, sizes); `makeTransform([Transform.translate(x, y), Transform.rotate(deg), Transform.scale(s)])` builds a transform string from the builders on `Transform` (translate, rotate, scale, skew and their X/Y/Z variants, perspective, matrix).
- Motion trails (`@remotion/motion-blur`): `<Trail layers={6} lagInFrames={0.6} trailOpacity={0.9}>…</Trail>` leaves echoes behind fast movement; its children must animate from `useCurrentFrame()`. There is no camera motion blur: the export cannot blend its samples.
- Icons: `<Icon name="rocket" size={56} color="#ffca28" strokeWidth={2} />` draws any Lucide icon by its kebab-case name ("arrow-right", "chart-line", "shield-check", "coffee", "users", "globe", "lightbulb", "brain", …). Always pass an explicit `color`; animate the wrapper (scale, opacity, position), not the icon's own props. An unknown name is rejected with the closest real names.
- Charts (`d3.*`): scales (`d3.scaleLinear`, `d3.scaleBand`, `d3.scaleTime`, `d3.scaleOrdinal`, …), shapes (`d3.line`, `d3.area`, `d3.arc`, `d3.pie`, `d3.stack`, curves such as `d3.curveMonotoneX`), arrays (`d3.max`, `d3.extent`, `d3.range`, `d3.bin`, …) and layouts (`d3.hierarchy` with `d3.treemap`, `d3.pack`, `d3.partition`). d3 only computes numbers and path strings: draw bars and tiles as positioned divs grown by `frame`, and stroke or fill the generated paths on a `<canvas>` with `new Path2D(d)` (a growing donut is `arcGen({ ...slice, endAngle: slice.startAngle + (slice.endAngle - slice.startAngle) * progress })`). Never use d3-selection, transitions or timers; they are not available.
- Colour (`culori.*`): `culori.interpolate([a, b], 'oklch')` returns a function of 0–1 whose result goes through `culori.formatHex`; blending in OKLCH stays vivid where RGB passes through grey. `culori.wcagContrast(a, b)` picks a readable text colour for a background; `culori.samples(n)` gives evenly spaced stops for a palette.
- GSAP timelines (`useGsapTimeline`): `const scope = useGsapTimeline(({ timeline, selector }) => { timeline.from(selector('.word'), { y: 40, opacity: 0, duration: 0.5, ease: 'back.out(1.7)', stagger: 0.08 }, 0); })`, then `<div ref={scope}>…</div>`. The timeline is paused and seeked to the current frame, so it is deterministic; author durations in seconds, target elements by className through `selector`, and never call play(), seek(), callbacks or random values (they throw). The DrawSVG (`{ drawSVG: '0%' }`), MorphSVG (`{ morphSVG: targetPathElement }`) and MotionPath (`{ motionPath: { path, align, alignOrigin: [0.5, 0.5] } }`) plugins are registered. DrawSVG and MorphSVG rewrite SVG markup every frame, the one exception to keeping SVGs frame-invariant: use them on at most two paths per graphic, and draw anything more on a canvas. Let GSAP own the properties it animates: do not also set that element's transform or opacity in its React style.
- `<FitText maxFontSize={56} minFontSize={32} lines={1} style={{ width: 360 }}>{props.value}</FitText>` shrinks its text until it fits its box on at most `lines` lines; its box needs a bounded width.
- 3D: the export projects `perspective()` only from an element's own transform, so write `transform: 'perspective(800px) rotateY(30deg)'` on the rotated element itself. For a solid (a cube, a flipping card with two sides), make each face a sibling with the full chain `perspective(800px) rotateX(a) rotateY(b) <face rotation> translateZ(d)`, set `backfaceVisibility: 'hidden'`, and list faces back to front.
- Call injected globals such as `useCurrentFrame`, `spring` and `interpolate` directly: destructuring them from `React` yields undefined.

**Shared web validation.** `create_motion_graphic_from_code` and source/schema changes through `edit_asset` use the same web validator. Syntax/API errors and known renderer incompatibilities must be fixed before granting `metadata.renderContract: "web-renderer-v1"`. Uncertain static findings are non-blocking warnings, not rendering failures: make at most one best-effort repair, then accept the stamped result instead of repeatedly rewriting it. Asset edits may apply reported autofixes. Generation retries known incompatibilities within its bounded attempt budget and may finish unstamped if they remain. `edit_item` placement and instance changes never revalidate source or write stamps. `validateOnly` never writes. `inspect_asset` and `preview_timeline` remain read-only and never grant stamps. Desktop uses its separate native policy.

<!-- END generated:web-renderer-css-support -->

## Placement And Review

Do not author JSX from timing alone. Inspect the target frame first: timing tells you when; the frame tells you form, placement, and background. For a batch of overlays, make one target-frame screenshot/contact sheet and decide each MG's settled frame, speech span, read time, and placement relationship before choosing final anchors, sizes, and durations.

Design the settled frame first: choose the moment when the MG is most readable, place the final layout there, then animate into that composition.

Before authoring JSX, make four linked editor decisions. They prepare the Motion Graphic asset and the later timeline placement.

| Decision               | Question                                                               | Output                                                           |
| ---------------------- | ---------------------------------------------------------------------- | ---------------------------------------------------------------- |
| **Content**            | What idea deserves a visual layer?                                     | The message or visual fact the MG expresses.                     |
| **Timing**             | When should it land and leave?                                         | Speech span, read time, duration, and any internal motion beats. |
| **Form and placement** | What kind of MG is it, and where does it belong in the composed frame? | MG form/size, then `edit_item` placement after asset creation.   |
| **Background**         | Is this an overlay on the footage, or its own moment?                  | Transparent or opaque background.                                |

Placement principles:

- Compose the footage and MG together from the frame you inspected: subject, camera framing, visual weight, captions/subtitles when present, and the MG's job.
- Place the MG where it makes the frame read best for that moment.
- Keep necessary information readable at video scale without zooming; if the MG feels detached, change form, timing, scale, or skip.
- Account for captions only when captions are present or planned.
- Treat full-frame MGs as intentional beats, not as a workaround for awkward overlay placement.

Default to a transparent overlay unless a full-frame beat is intended. A transparent overlay still uses a natural-box asset; do not use a transparent timeline-sized asset just for placement.

Place and review:

- Place with `edit_item` (adds/updates). Prefer an explicit rectangle once you know the frame: one horizontal anchor, one vertical anchor, width, and height.
- Verify with screenshots. Pass multiple frames in one tool call — settled state appears alongside any transient mid-animation frames. Compare frames before concluding: apparent truncation, missing elements, or "broken design" visible in only some of the batch is animation, not a real flaw. If unclear, re-capture more frames around the suspect one before adjusting anything. Judge from the settled frames.
- Check the full frame: necessary information is clear at video scale, captions remain readable when present, MG content is correct, text is legible, and the composition feels balanced and intentional.
- For text-heavy MGs, inspect the settled frame where the most text is visible. Check for text-on-text overlap, clipped lines, overflow outside the natural asset box, and readable content covered by animated scale or translate states.
- If it fails, first adjust position and size. If position/size cannot make it work, change the design form. Verify each recurring component form on a target frame before expanding it.

## Asset And Timeline Flow

### Create A New Asset

Create new Motion Graphic assets by passing inline JSX and editable property metadata through the current ChatCut asset-creation tool. Use the tool schema for the exact field names and accepted duration format.

Choose the MG's natural box, duration, asset name, description, and property schema from the edit requirements. The asset duration should match the intended placed span, including internal entrance, hold, and exit timing. The asset creation step only creates the asset; timeline placement is separate.

For `create_motion_graphic_from_code`, pass that natural box as `width`/`height`; use timeline dimensions only for intentional full-frame MGs. If the content occupies only part of the screen, place and scale the bounded asset with `edit_item` instead of baking screen coordinates into a full-canvas MG.

### Patch An Existing Asset

Before patching, inspect the existing asset code and property schema. Preserve unrelated behavior, property keys, and timeline timing unless the requested change requires otherwise.

Patch with full inline replacement source through the current asset-update tool.

### Place On The Timeline

Use the timeline editing workflow for placement, movement, trimming, and per-instance property overrides. Dry-run large or uncertain transactions when the tool surface supports validation.

## Verification

A successful tool call is not verification.

- Re-read asset state after asset creation or update.
- Re-read timeline state after placement, movement, trimming, or property overrides.
- For visible changes, inspect a real composed frame using the normal ChatCut visual verification path.
- For a batch, compare the composed settled frames side by side; verify that each visual job has a fitting form, repeated surfaces/anchors/rhythms are intentional, and each placement works for its own target frame.
- If the result is wrong, classify the failure before retrying: invalid tool shape, invalid JSX, missing/incorrect property key, timeline placement, async asset readiness, or canvas/export safety.
- An export contract error from `create_motion_graphic_from_code` or `edit_asset` lists styles the browser export cannot draw; rewrite them and retry. A legacy warning means `renderContract` was missing (see Renderer CSS Support).
- Fix placement with timeline edits and bad rendering with JSX/property changes.

Referenced files: 2

digital-human29.8 KB

View saved version →

---
name: digital-human
description: Create and manage script- or audio-driven AI avatar videos from an official or saved presenter, an imported portrait, or a representative still prepared from video. Use for 数字人、数字分身、虚拟人、照片开口说话、人像口播、改稿不重拍, AI avatar, AI avatar video, avatar video, talking avatar, talking photo, photo avatar, video avatar, AI presenter, virtual presenter, virtual spokesperson, digital human, or a person's digital twin, including choosing or creating the avatar and deciding its voice, script, aspect ratio, and output quality. Do not use for a static profile-picture avatar, game or 3D character creation, an industrial digital twin, ordinary B-roll, generic image-to-video, video translation, talking-head editing, lip-sync dubbing alone, or TTS-only requests.
user-invocable: true
---

# Digital Human

Create a reusable presenter identity or generate a lip-synced presenter video
without exposing the underlying provider. Treat avatar selection, speech, script,
format, consent, generation, and verification as one stateful conversation.

## When to Use

Use this skill when the user wants to:

- Make an official or saved AI avatar, talking avatar, or virtual presenter speak.
- Turn a consented portrait into a reusable digital-human identity.
- Use a representative still from an imported video as a photo avatar.
- Create a synthetic AI presenter from a description, when live capabilities allow it.
- Draft or revise the narration for a digital-human video.
- Regenerate, reorder, remove, retry, or accept segments from an avatar-video batch.

Treat **AI avatar** as the broad default English concept. **Digital twin** means a
reusable likeness of a specific real person; **photo avatar** or **talking photo**
means a still-image source; **AI presenter**, **virtual presenter**, and **virtual
spokesperson** describe the presenter's role. **Digital human** is valid but is often
used for broader enterprise or interactive experiences.

Route adjacent requests elsewhere:

- Speech, voice audition, or voice cloning without avatar video: use the Voice Skill.
- Ordinary footage cleanup or presenter editing: use the Talking Head Guide.
- Generic image-to-video or video generation without a speaking avatar: use the
  Video Generation Skill.
- Translation or dubbing of an existing video: use the dedicated translation or
  dubbing workflow when available.

## Required Inputs

Resolve these slots before submission, but do not ask for them in a fixed order:

1. **Target project** — use the active project when it is unambiguous.
2. **Avatar identity and ready revision** — an official avatar, a saved avatar, or
   a newly created photo or synthetic avatar.
3. **Speech source** — one exact available voice, one cloned voice, or one imported
   audio asset. A text-driven source also requires the final script.
4. **Final script** — preserve the user's wording and punctuation. Show any AI-written
   draft in full and let the user edit or approve it before generation.
5. **Aspect ratio** — `auto`, `1:1`, `4:5`, `5:4`, `9:16`, or `16:9`. Explicit user
   choice wins; otherwise omit it so the backend uses `auto` and preserves the
   selected avatar's original framing. Do not substitute the current project ratio.
6. **Resolution** — `1080p` by default; use `720p` when requested or when live
   capabilities require it, and explain any fallback.
7. **Optional controls** — motion guidance only when the live capability response
   says it is supported. For an official catalog avatar, use its returned
   `hasBackground` value: preserve the original scene when true and request a
   transparent background when false. For a saved/custom avatar, preserve its original
   background by default. Request a transparent background only when the user explicitly
   asks to remove the background or use transparency. Do not ask the user to choose a
   background mode; briefly state the resolved default before submission. An explicit
   user request to preserve or remove the background always wins.

For avatar creation also resolve:

- A non-empty name. Keep numeric validation limits out of the question label;
  only explain a limit if the submitted name actually fails validation.
- One final avatar-creation source: an imported JPEG/PNG portrait, a
  representative JPEG/PNG still prepared from an imported video, or a prompt
  for a synthetic identity when supported.
- Explicit consent when a real person's likeness is involved.

Do not pass a video directly to photo-avatar creation unless the live capability
contract explicitly supports it. For an imported video, use the host's prepared
representative still or help select a clear, unobstructed, front-facing frame, then
use that image asset as the source.

When a user attaches or references a video containing a person and says "this
avatar", "this person", "use her/him", `这个形象`, `这个人`, `用她`, `用他`, or an
equivalent phrase, resolve the person in that exact attachment as the intended
likeness source. Do not treat the attachment as the requested speech source merely
because it has audio or a transcript. If the user also supplies a new topic or an
approved script, that new topic or script wins; do not read or reuse the video's
spoken content unless the user explicitly asks for it.

The attached video is a prospective avatar source, not a ready `identityId`. Enter
the video-to-avatar creation flow: prepare a clean representative still, obtain the
required likeness consent, create the identity, and then reuse that identity for the
requested generation. Do not show official or saved avatar choices while this
prospective source remains valid. Offer other identities only when the user declines
to use the person in the video or no eligible face frame can be prepared.

If "this" could reasonably mean the person's likeness, the video's existing pixels,
or its spoken content, ask exactly one targeted clarification before choosing a
workflow: "Do you want to use the person in this video as the avatar and have them
speak your new script?" Do not replace this with a generic "What do you want to do
with this video?" question.

When the user is about to upload or pick a video for avatar creation, tell them up
front that the chosen frame must show the face clearly with nothing covering it — no
burned-in subtitles, captions, watermarks, stickers, or hands over the face — and
prefer a frame with a neutral, front-facing pose. Never upload an inspection contact
sheet or any preview that carries an overlaid time label as the avatar source; use a
clean frame with no added markings.

A video source also carries the person's voice. After the avatar is created from a
video, offer to clone that voice from the same video so the avatar speaks with it:

1. Run the Voice Skill's cloning flow (`manage_custom_voice action=clone` with the
   video's asset and explicit voice-cloning consent — likeness consent alone does
   not cover the voice).
2. When the cloned voice is ready, call
   `manage_avatar action=set-default-voice identityId=<avatar> voiceRef=<cloned voiceRef>`
   so it becomes this avatar's saved default voice.
3. Later generations with this avatar should use the returned `defaultVoiceRef` by
   default, while still letting the user pick another voice.

This is an offer, not an automatic step: skip it when the user declines or the video
has unusable audio (music, multiple speakers, heavy noise).

## Workflow

This is an intent-driven constraint resolver, not a linear wizard. Do not walk the
user through every section or ask for fields in a preset order.

On every turn:

1. Extract all facts already supplied by the user, attachments, active project,
   selected media, saved avatar state, and prior answers in this request.
2. Resolve deictic phrases such as "this", `这个`, `这个形象`, and `这个人` against
   the attached or selected media and the user's action words before inferring a
   generic media workflow.
3. Infer the user's immediate requested action: explore, choose, create, revise,
   generate, inspect progress, retry, reorder, remove, accept, place, or export.
4. Compute only the blockers for that action. Defaults and live saved bindings count
   as resolved unless they conflict with an explicit request.
5. If nothing blocks the action, execute it immediately.
6. Otherwise ask about the single most important missing or ambiguous constraint,
   preferably with structured choices, then recompute from the new state.

For example, a user who names a ready avatar, supplies approved text, and accepts its
saved voice should not be asked to choose them again. A user inspecting a running
batch does not need to resolve a new script or aspect ratio. Consent is requested only
for the creation or cloning action that needs it.

### Load live capabilities and state when relevant

Use `ToolSearch` to load `manage_avatar` and `submit_avatar_video` if they are not
already visible. Call `manage_avatar` for live capabilities before
offering creation or advanced options.

If the user has neither selected an exact identity nor supplied a prospective
likeness source, list the current official and saved identities before recommending
one. Do not infer availability from chat history, examples, or provider knowledge.
For official avatars, use the live catalog's `tags` as factual selection filters and
its `description` as recommendation guidance. Match those fields against the user's
requested presenter and use case, and never infer a missing age, gender, ethnicity,
role, setting, or capability from the preview's appearance.
For the empty saved-avatar recommendation state, call the official catalog with
`limit: 5`; do not page through the catalog or turn a full catalog page into one
recommendation form.

An avatar attached from ChatCut's Digital Humans library or native visual picker is
already an exact selection. Reuse its attached ChatCut `avatarId` as `identityId`
across later confirmation and retry turns. An official attached identity is directly
usable by `submit_avatar_video`; it does not need to be imported or rediscovered in
the first catalog page. Before resolving speech for an attached identity, call
`manage_avatar` with `action: get` and that exact ID unless a live result in the
current turn already supplies its `defaultVoiceRef`. Never reject or replace an
attached avatar just because it is absent from the five-item recommendation page.

Only use identities and revisions reported as ready. Pending, failed, deleting,
deleted, or orphaned identities are not valid generation inputs.

If the digital-human tools are unavailable, explain that this generation capability
is not connected. Do not silently substitute generic video generation or call the
backend HTTP routes directly.

### Resolve the avatar when it is missing or invalid

Respect an explicit user choice. Otherwise:

- A person-containing video explicitly or provisionally resolved as the likeness
  source: continue the video-to-avatar creation flow. Do not enter the generic avatar
  picker and do not ask the user to choose among saved or official identities.
- No ready saved avatar: show at most five suitable official avatars and include the
  reserved **create** card in the same choice surface. This is a strict empty-state
  contract: never render a sixth official recommendation, and never omit the
  creation option merely because official avatars are available.
- Ready saved avatars exist: show them first (up to five total cards), then the
  reserved create card.
- Exactly one ready saved avatar: offer it first and ask whether to use it.
- Multiple ready saved avatars: prefer higher live usage count; when counts tie, prefer
  the newest creation. If usage metadata is absent, rank by newest creation only;
  never invent a frequency.

Avatar identity selection is media-preview-only and is a strict exception to the generic
single-branch `<choices/>` rule. Always use visual cards through the Widget Forms
Skill, including when exactly one ready saved avatar exists. Never ask a text-only
category question such as "use the saved avatar" versus "browse official avatars",
and never use `<choices/>`, `<form-single>`, or label-only identity cards for avatar
selection. For an avatar card in ChatCut's native widget — official or the user's own
saved avatar — write only its exact `identityId` as `avatar-id`; the host resolves the
current name, preview video or fallback image, aspect ratio, and submitted value from the live avatar
identity even when the tool payload does not expose a preview URL. Never copy `value`,
`name`, `media`, or `aspect-ratio` onto an `avatar-id` option. If the host cannot
resolve preview media, omit that identity rather than falling back to text. Prefer
the live preview video when available and use the preview image only as its poster
or fallback. When
the user has ready saved avatars, place their cards before official recommendations.
Ask one most important missing question at a time and do not repeat answered questions.

The raw tags below are the embedded ChatCut protocol only. In a published Codex or
Claude Code plugin, the Widget Forms Skill's host adapter takes priority: never emit
raw ChatCut tags. Map each live identity to the host's visual-choice surface using
`identityId` as its stable value and `name` as its label. In Codex, pass
`previewVideoUrl` as `previewVideo` and `previewImageUrl` as `preview` so the image is
the video's poster and failure fallback. Keep the create action
as a separate `create_avatar` value; plugin hosts return that value to the workflow
instead of opening ChatCut's native dialog automatically.
For a saved identity whose `list` row has no preview URLs, call `manage_avatar
action=get` for that exact candidate before rendering it; use the selected ready
revision's returned preview media and do not invent or copy media from another avatar.

One reserved label-only option value opens a native editor surface instead of
submitting an answer. Write it with no `media`:

- `create_avatar` — opens the native creation dialog where the user uploads a photo
  and creates a personal avatar. Keep its card label a short "create your own
  avatar" phrase (e.g. 创建专属数字人); do not mention uploading a photo — the
  dialog explains that itself.

Clicking that card does not produce a form answer. When the user finishes in the
native dialog, their chosen or newly created avatar is sent back automatically as
this question's answer with the avatar attached; treat that reply as the avatar
selection and do not re-ask.

The selection shape is therefore one `<form-visual>` containing up to five avatar
preview cards (saved first, then official) followed by the creation card. Localize
the visible question and card label to the conversation language:

```text
<widget>
  <form-visual id="avatar" question="Please choose an avatar" required="true">
    <visual-option avatar-id="<saved-or-official-identity-id>"/>
    <visual-option value="create_avatar" name="Create your own avatar"/>
  </form-visual>
</widget>
```

Creating an identity is separate from generating a video:

1. Resolve the source or prompt and name.
2. Obtain required consent.
3. Create the identity through `manage_avatar`.
4. Poll until it is ready or terminal.
5. Save it under **My avatars**. Do not auto-generate a video merely because avatar
   creation succeeded.

When asking for a portrait in an embedded ChatCut Widget, always mark the upload as an avatar source so
the host uses the same validation, JPEG normalization, private transport storage, and
upload-readiness behavior as the native Digital Humans dialog:

```text
<widget>
  <form-files id="avatar_source" label="Please upload a portrait photo"
    accept="image/jpeg,image/png" purpose="avatar-source"
    multiple="false" required="true"/>
  <form-text id="avatar_name" label="Name this avatar" required="true"/>
</widget>
```

Do not use an unmarked generic `<form-files>` upload for avatar creation.
Published plugin forms do not provide this native upload control. Ask the user to
attach the portrait through their host, then use the Asset Import Skill and pass the
returned ChatCut image asset id to `manage_avatar`.

### Resolve speech only when the requested action needs it

For an official avatar, never silently use its paired voice and never immediately
open the broader voice catalog. When the live result provides `defaultVoiceRef` and
the user has not already made this choice, first ask one blocking text-only decision
in the conversation language: whether to use this avatar's official built-in voice
or choose another voice. In embedded ChatCut, render that decision with localized
`<choices options="Use official voice,Choose another voice"/>`, then stop and wait.
If the user accepts, resolve speech with that exact `defaultVoiceRef` and do not call
`manage_avatar action=voices`. Call `action=voices` only after the user chooses
another voice, or when the live identity has no `defaultVoiceRef`. Do not ask for
language or script and then pre-emptively call `action=voices` in the same turn.

For a saved or newly created avatar, reuse its saved voice binding when present and
confirm it; otherwise offer available voices. When the user settles on a voice they
want this avatar to keep, save it with `manage_avatar action=set-default-voice`;
`list`/`get` then return it as the avatar's `defaultVoiceRef`.

For voice discovery, audition, or cloning, follow the Voice Skill and reuse its live
voice list, consent gate, and generated voice asset. Do not duplicate voice-cloning
logic here.

When an exact voice still needs to be selected, never present voice names as prose
beside a free-text field. Use the exact `voiceRef` returned by `manage_avatar` as the
submitted value. Voice selection is audio-preview-only and is a strict exception to
the generic `<choices/>` rule:

- Never render `<choices/>`, `<form-single>`, prose voice names, or label-only voice
  options. A voice without a returned playable `previewUrl` must be omitted from the
  recommendation surface rather than presented as an unplayable choice.
- When the user wants to audition voices, prefer up to five suitable available
  official, preset, or custom voices that actually include `previewUrl`. Render them as one
  required `<form-visual id="voiceRef" media-kind="audio">`; `hasPreview: true`
  without `previewUrl` is not playable and must not be presented as an audition
  card. Always append one no-media `<visual-option value="clone_voice"
name="Clone Voice"/>` action, localized to the conversation language, as the
  final card. This clone option is required on every recommended-voice surface.
- An official avatar's paired default voice may still be used without listing the
  broader official voice catalog. Do not show that default voice as an audition
  card unless an actual audio `previewUrl` was returned for it. Never reuse the
  avatar image or video URL as a voice preview.
- Copy each returned `previewUrl` exactly. In particular, keep
  `/voice-samples/...` root-relative; do not expand it to `chatcut.com` or invent
  another host.
- When script text is also being confirmed, put the voice selector and script field
  in the same widget. Do not make the user type a voice name into the script field.

```text
<widget>
  <form-visual id="voiceRef" label="Please choose a voice and listen to its preview" media-kind="audio" required="true">
    <visual-option value="<voiceRef>" name="<voice name>"
      media="<previewUrl>" media-kind="audio" summary="<voice traits>"/>
    <visual-option value="clone_voice" name="Clone a voice"/>
  </form-visual>
  <form-textarea id="script" label="Please confirm the script" required="true" default="<script>"/>
</widget>
```

If the source began as a user video, ask whether the user wants that video's voice.
If yes, follow the Voice Skill's explicit voice-cloning authorization flow before
using it. An attachment alone is not permission to clone. If no, continue with the
normal voice picker.

An existing project audio asset may drive the avatar directly. When audio is the
speech source, do not invent or require text unless the user also wants a transcript
or script review.

### Resolve a script only for text-driven generation

The script may be pasted by the user, derived from selected project material, or
written collaboratively. When drafting:

1. Ask only for missing intent such as audience, goal, tone, facts, and duration.
2. Produce the complete proposed script.
3. Let the user edit it or approve it explicitly.
4. Submit exactly the confirmed content.

Each explicit text segment must contain 1–5000 characters. Omit explicit segments
for an ordinary script so the backend generates one continuous video up to 5000
characters. Use explicit segments only when the user requests meaningful paragraph
or shot boundaries. Longer text is split automatically and may create at most 10
segments. Never summarize, translate, drop, duplicate, or reorder content to fit a
limit. Preserve wording and
punctuation; non-semantic line whitespace may be normalized automatically.

Text-driven avatar generation is limited by the 5000-character segment boundary,
not by an estimated audio duration. ChatCut TTS and cloned voices are rendered to
audio before avatar-video submission. Long audio-backed text is prepared as smaller
ordered segments, and each generated audio file is validated from its real duration.
Existing project audio longer than 600 seconds is automatically split into ordered
parts before submission. Pass the original asset once; never split it manually or
submit one generation call per part. If a progress update is useful, say only that
ChatCut is splitting the long audio while preserving the complete content and order.
Do not add implementation details about where or how the work runs, and never expose
hidden derived audio assets. Never shorten or summarize approved content to satisfy
these limits.

### Resolve output settings from intent and available defaults

Preserve the selected avatar's original framing by default: omit `aspectRatio` and
let the backend resolve it to `auto`. Never use the current project ratio as the
avatar-video default. If the user explicitly asks for a different ratio, pass that
request instead.

For an official catalog avatar, resolve the default from the `hasBackground` field
returned by `manage_avatar action="catalog"`:

- `hasBackground: true` — pass `removeBackground: false` and preserve the original
  scene.
- `hasBackground: false` — pass `removeBackground: true` so the result is a
  transparent WebM when transparency is supported.

For a saved/custom avatar without catalog metadata, keep `removeBackground: false` as
the default and preserve the source image background. Pass `removeBackground: true`
only when the user explicitly asks for a transparent background or background removal.
Do not turn this into a blocking question. Before submission, briefly tell the user
whether the original scene will be preserved or the result will be transparent. An
explicit user request overrides the catalog default. If transparency is unsupported,
explain the limitation and keep the original background.

Do not promise crop, background removal, motion control, transparency, resolution,
or prompt-based identity creation until the live capability response confirms it.

### Submit as soon as the required constraints are resolved

Summarize the resolved avatar, voice or audio source, script state, aspect ratio, and
resolution before the first generation when any choice remains consequential. Then
call `submit_avatar_video` once.

Treat long-script output as an ordered batch of avatar-video segments, not as one
audio file or one opaque job. After submitting, tell the user that generation has
started and end the turn. Never call `track_progress`, `manage_avatar action=get-batch`,
or add a Bash sleep in the same turn merely to keep the turn open, including when a
queued follow-up step will eventually need the completed asset. The project library
shows the in-progress generation and receives the completed video assets in the
background.

Call `track_progress` in a later user turn only when the user asks for status or resumes
a follow-up step that depends on the completed asset, such as placing it on the timeline
or reviewing the result. When a job is still non-terminal, say that generation is still
running without quoting its numeric progress: provider progress values are coarse
internal stages, not user-facing completion estimates. Use `manage_avatar
action=get-batch` when ordered segment and attempt state is needed. Keep the stable
segment key and original order.

For a failed segment, offer a targeted retry. Do not resubmit successful segments or
the entire batch without the user's instruction. Use the management tool for supported
segment regeneration, removal, reordering, attempt selection, and acceptance.

Generated results become normal ChatCut video assets. Add them to the timeline only
when the user requested placement or it is clearly part of the active editing task.
Do not export unless asked.

## Provider Selection

Provider selection is an internal ChatCut backend detail. Always call the
ChatCut-native tools and let the backend construct provider requests, resolve private
IDs, enforce capabilities, and persist results. Never call an underlying provider API
or CLI directly from this skill.

Select only through ChatCut capabilities, catalog results, saved bindings, and the
native tools. User-facing language is limited to concepts such as:

- **Official avatars** and **My avatars**
- **Official voices**, **cloned voices**, and **project audio**
- **Photo avatar** and **synthetic avatar**

Provider names, engine or model names, provider-specific avatar or voice IDs, raw
provider URLs, request payloads, and provider error messages are internal. Never show
them in prose, option labels, progress updates, or failure messages. Translate failures
into actionable product language while preserving the real status.

Do not route around ChatCut's feature gates, quotas, billing checks, safety policy,
catalog visibility, or provider registry.

## Consent and Safety

A real-person photo avatar requires an explicit affirmative confirmation immediately
before creation. Use the user's language and keep the meaning exact. For Chinese:

> 我确认拥有该肖像,或已获得创建和使用该数字人形象的授权,并承诺不将其用于
> 冒充他人、欺诈或其他违法用途。

For English:

> I confirm that I own this likeness or have permission to create and use this
> digital avatar, and I will not use it for impersonation, fraud, or unlawful activity.

En español:

> Confirmo que soy titular de los derechos sobre esta imagen o que tengo permiso para crear y usar este avatar digital, y no lo utilizaré para suplantar identidades, cometer fraude ni realizar actividades ilícitas.

Do not infer consent from an upload, a previous unrelated confirmation, ownership of
the project, or a third party's instruction. Record consent only through the native
tool's consent fields. Synthetic prompt-created identities do not require real-person
likeness consent unless the prompt or references identify a real person.

Refuse non-consensual impersonation, deceptive identity use, fraud, harassment,
sexual exploitation, or attempts to evade safeguards. Do not create a real-person
identity from uncertain authorization. Voice cloning has a separate consent gate;
follow the Voice Skill even when avatar consent was already obtained.

## Verification

Before reporting success, verify through live tool results:

1. The chosen identity revision was ready when submitted.
2. The job or batch reached a completed terminal state and returned output asset IDs.
3. Every expected segment key exists exactly once and appears in the confirmed order.
4. Text-backed segment manifests preserve the confirmed wording and punctuation,
   allowing only documented non-semantic whitespace normalization.
5. The generated videos exist as usable project or media-library assets.
6. Any requested timeline placement is present and ordered correctly.

Preview the generated assets or composed timeline when visual verification is available.
Do not claim the face, lip sync, framing, or background looks correct from status fields
alone. If visual inspection is unavailable, say what was structurally verified.

## Hard Rules

- Never expose provider, engine, or model identity to the user.
- Never call an underlying provider directly or construct provider payloads in the
  conversation layer.
- Never invent an avatar, voice, capability, usage count, saved binding, or asset ID.
- Never generate from a non-ready identity revision.
- Never treat a video as a supported identity source without an explicit live capability;
  prepare a representative still for the current photo-avatar contract.
- Never pass an external source URL to avatar creation; import the source into the
  target ChatCut project first.
- Never infer likeness consent or voice-cloning consent from an attachment.
- Never clone a video's original voice without the Voice Skill's explicit authorization.
- Never silently rewrite, translate, truncate, reorder, or duplicate confirmed script text.
- Never ask for information already available from the project, attachment, or live tools.
- Never run a fixed wizard or ask for avatar, voice, script, ratio, and resolution in a
  preset sequence; resolve only what the current user action still needs.
- Ask one most important missing question at a time; prefer structured visual choices.
- Never show provider-private IDs, URLs, payloads, or raw errors in user-visible output.
- Never bypass ChatCut tools, feature gates, billing, quota, or safety checks.
- Never retry a whole successful batch because one segment failed.
- Never claim visual quality without inspecting the generated result.
- Never place media on the timeline or export it unless the task calls for that action.
- For official catalog avatars, obey the returned `hasBackground` default: preserve a
  configured scene background and remove a configured absent background. For
  saved/custom avatars, preserve the source background unless the user explicitly asks
  for transparency or background removal. Explicit user intent always overrides either
  default.
export1.89 KB

View saved version →

---
name: export
description: Hosted ChatCut plugin sessions only (the `chatcut` MCP server). If the conversation is driving ChatCut Desktop (a `chatcut_desktop*` MCP server), load this skill only when the user explicitly chooses the plugin/web surface — desktop sessions otherwise ship their own instructions and tools. Export or deliver a ChatCut project through the hosted connector, including video, audio, subtitles, XML, and render-status checks.
---

# Export

Export only when the user asks for a render, download, or final standalone deliverable. A review checkpoint normally stays as an editable ChatCut timeline.

Use `submit_export` with the current tool schema. Confirm the target timeline, range, resolution, codec, and format from project state and user intent; do not guess a non-active timeline. Record every returned `renderId`.

Use `track_export` for status. It is the render-job tracker; do not use generation/transcription `track_progress` for exports. For the latest project export, use the tool's latest-render option when available rather than guessing an id.

When a render completes, return its `downloadUrl`. In Claude Code, download the completed file to a collision-safe local path and provide the path, local-file preview link, and concise render metadata. In Codex, surface the completed download URL through the host's normal file/link delivery. Do not claim delivery from a queued or running render.

For subtitle or NLE XML export, use the corresponding `submit_export` format and report warnings about unsupported or dropped elements. For transparent Motion Graphic delivery, use the dedicated MG export tool exposed by the current manifest and track the returned render ids.

If cloud rendering is blocked by media that is not remotely readable, use `asset-import` when the user permits upload. Otherwise report the blocker; do not flatten or substitute the edit locally without user intent.
known-errors2.2 KB

View saved version →

---
name: known-errors
description: Diagnose ChatCut plugin tool failures, rejected mutations, unexpected response shapes, and blocked import, generation, or render operations.
---

# Known Errors

Treat the active MCP tool schema and returned structured error as authoritative. Do not retry with remembered legacy payloads when the current manifest differs.

## Mutation Rejections

- On same-track overlap, do not force the write or silently delete the conflicting item. Decide whether content is sequential or layered, choose an available track when appropriate, or ask which item should win.
- On locked-track or stale-id failures, refresh the affected timeline scope and retry only with current ids and state.
- On validation errors, change only the rejected field or transaction shape; preserve unrelated project state.

## Import And Media

- Hosted plugin surfaces use the `asset-import` adapter and `import_media`; do not substitute backend-only `push_asset` or `download_media` calls when they are absent.
- If media conversion fails, use the helper's structured retry when provided. Otherwise explain the unsupported source and ask for a compatible replacement; do not repeat the same failing editor import path.
- Wait for upload only before byte-dependent operations. A known asset id can be used for metadata and timeline placement while bytes continue uploading.

## Generation And Motion Graphics

- Use only generation tools visible on the current host. Codex and Claude Code plugins direct-author Motion Graphics with `create_motion_graphic_from_code`; the built-in ChatCut Agent may expose `submit_motion_graphic`.
- Preserve the original provider or content-policy failure. Do not spend credits on repeated identical retries or silently switch models.

## Verification And Export

- A successful mutation is not visual proof; load `verification` when the result must be seen.
- Use `track_export` for render jobs. If cloud render cannot read an asset, resolve remote readiness through `asset-import` only with user permission, or report the limitation.

On auth or project-access errors, verify the exact project id and that the editor and plugin use the same ChatCut account before attempting code or infrastructure debugging.
multicam-sync15.9 KB

View saved version →

---
name: multicam-sync
description: Synchronize footage from a multi-camera / multi-recorder shoot — several cameras plus separate audio recorders covering one session, imported as loose clips — and optionally turn all or part of it into speaker-follow footage for a larger edit. Use when a user drops in multiple clips from the same recording and wants them aligned, asks for multicam / 多机位 / multi-angle sync, wants a "cut to whoever is talking" edit, or refers to camera A/B, angles, or separate lav or field recordings that need to line up with picture.
user-invocable: true
---

# Multicam Sync

Parts 1–3 always run: they produce the synced master, a verifiable fact that every
later edit is rebuilt from. Part 4 cuts a draft from it, and runs only on request.
Transcripts do almost all of the work. AI locates; arithmetic decides frames.

Use ordinary timelines, tracks, and items only. Never create a compound or nested
multicam clip, and never depend on a camera-switcher item. If a main edit already
exists, leave it untouched while building the synced master and speaker-follow
cutting board on separate timelines.

## Part 1 — Discover the structure

Work out what you actually have before aligning anything. Never ask the user how many
cameras there are. The material already answers it.

### The two rules that do the work

**Same words at the same moment → cannot be sequential spans of one camera.** Overlapping
transcripts mean two devices recorded one event simultaneously, so they are different
sources or simultaneous angles.

**No shared words, and one ends where the next begins → one camera that stopped and
restarted.** These are sequential spans of a single capture source, not separate angles.

### Steps

1. Read every asset's transcript, build a pairwise overlap matrix — how much text
   each pair shares, and where — and split by the rules above.
2. Video assets are **angles**; audio-only assets are **recorders**.
3. Corroborate with filenames, folder structure, camera-model metadata, durations —
   tie-breakers only, never over transcript evidence. Real shoots ship three cameras
   all named `C0001.MP4`.

### Report before acting

State the structure in the user's terms and wait if anything is ambiguous:

> 2 camera angles (Pocket3, 3 spans · Pocket4, 2 spans), 2 audio recorders
> (MIC1, MIC2). Session runs 53 minutes.

When evidence is thin — a source with little speech, an overlap resting on a handful
of matched lines — say so. Do not fill the gap with a guess.

## Part 2 — Align and place

One number per clip: timeline time minus source time. Keep that number constant for
every piece cut from the clip. Do not time-stretch automatically: a changing offset
may be a bad match, a clock difference, or dropped frames, and each needs review.

### Steps

1. **Reference**: the source that covers the whole session with the richest transcript —
   usually a dedicated audio recorder, not a camera.
2. **Use the renderer when available**: on the separate master timeline, put the
   untrimmed clips on ordinary source tracks, then call `multicam_sync` with the
   same-take items and the reference. This is the preferred Web/Desktop path. It
   uses existing source timing when decisive, otherwise audio correlation; it does
   not create a multicam object.

   Accept only `applied`, `already_synced`, or an understood `partial` result. Read
   `alignmentEvidence` for method, match confidence, correlation, overlap, and
   relative offset. Read `placementEvidence` for the actual post-sync
   `timelineSourceOffsetSeconds`. Do not use a skipped or low-confidence item. If
   the tool is unavailable or finds no confident alignment, leave the master
   untouched and use the transcript fallback below.

3. **Transcript fallback**: run
   `scripts/transcript-offset.mjs <utterances.json>` from this skill. Do not
   improvise the math. It first lets shared phrases vote on a coarse offset, then
   takes the median of near-identical utterance-pair deltas. It also rejects too few
   pairs, inconsistent pairs, and early/late drift. Only place files whose output is
   `confident:true`; report every printed issue for the others. `--help` documents
   the input and sign convention.
4. **Continuity check** (free, run it): spans of one camera were recorded
   back-to-back, so span N+1's offset minus span N's must equal span N's duration.
   Each offset was measured independently, so agreement confirms both. Sub-frame
   agreement is the norm; a multi-frame gap means dropped frames or a mismatch —
   say which. **An overlap or gap between placed spans of one camera means an
   offset is wrong. Recompute it — never trim a span to make it fit.**
5. **Place transcript-fallback results**: shift everything so the earliest clip
   starts at 0. One video track per camera, one audio track per recorder.
   Convert seconds → frames once at the end, never accumulating, and **floor at
   an asset's tail, never round up** — a
   rounded-up final frame claims source the file doesn't have; renders tolerate it
   silently, Script editing later refuses it.

### Report

Per clip: actual timeline-source offset, method and supporting evidence (renderer
confidence/correlation/overlap or transcript pair count/spread/drift), and the
continuity-check result.

## Part 3 — Identify and label

Turn "track 1 / track 2" into who is actually on it. Everything here is evidence-first:
a wrong confident claim is a failure; an honest "can't tell" is not.

### Who is in the session

Speakers usually name themselves or each other — self-introductions, banter. Pull real
names from the transcript. If none appear, A/B is fine; never invent names.

### Which camera frames whom

From the transcript, pick 3–4 moments where only one person is talking, spread across
the session. View the synced frame from **every** camera at the same wall-clock moment:
the talking face identifies the person; clothing and seating anchor identity across
angles. Two-shots are identity anchors, not noise. Classify framing while you're
there: close-up / medium / two-shot; note reframing if the dominant framing changes.

Angle labels are summaries, not promises: framing can drift, fail, then recover.
Inspect the exact synced interval before every cut; never choose from the label alone.

Hard check before concluding: at each self-introduction, the face whose lips are
moving is that name's owner. "Both cameras frame the same person" contradicts a
two-camera, two-mic, two-voice structure — treat that conclusion as an error until
frames at both introductions prove it.

### Which mic belongs to whom

A lav is dramatically louder for its wearer — typically 15 dB or more. Measure, don't
infer:

- During one person's solo speech vs the other's, compare **the same track against
  itself**. The in-track contrast cancels recorder gain. Above ~4 dB it decides;
  below, say "indistinct" — that itself is a finding (ambient mic, shared mic).
- Judge the **distribution**, not one moment: consistent → owned mic; 50/50 → not a
  personal mic; flips mid-session → handheld passed around or seats changed.
- **Never use utterance counts or transcript volume as evidence of mic ownership.**
  Crosstalk transcribes fine; both speakers appear fully on both mics' transcripts.
  The overall loudness difference between two mics is not evidence either — only
  the in-track solo-vs-solo contrast is.
- Fragmented diarization heals mechanically: cluster speaker-ids by their median
  cross-track level difference. Never hand-reconcile speaker ids across assets —
  they are per-asset serials.

### Which source is the program audio

Decide what the viewer will hear. Default: dedicated recorders beat camera embedded
audio. The user's word beats everything — "camera A has the good audio" is a program
audio assignment; that camera's audio track then behaves exactly like a recorder
(its offset is already known from Part 2). Record the assignment in the report.

With no clean recorder, compare camera mics at the same solo-speech moments. Prefer
one continuous source that is good enough; switch for a speaker or passage only when
another is clearly better for a sustained stretch and the handoff is inaudible.
Judge intelligibility, noise, clipping, and reverb — not labels or utterance counts.

### Label

Rename tracks in place with the shortest evidenced label: `Cam · <subject or view>`
and `Audio · <speaker or source>`. Append a shot size only when it distinguishes
otherwise similar angles: `WS` (wide shot), `MS` (medium shot), or `CU`
(close-up), for example `Cam · Speaker A · MS` or `Cam · Two-shot · WS`.
Otherwise omit it. Fall back to numbered labels when identity is uncertain.

Deliver an evidence table alongside: claim | evidence (timestamp + what was seen or
measured) | confidence. Decline to label what the evidence doesn't support —
fragmented transcript speaker ids are usually in that category — and say so.

### Deliver the master, then offer the cut

The master is the deliverable — report structure, offsets, identities, and evidence.
Note that it stacks angles and is a reference, not something to watch: the top track
covers the others. If the user only asked to sync, stop here, but **offer** the
speaker-follow draft rather than leaving them with tracks and no next step.

## Part 4 — Cut to the speaker (on request)

Turn the synced master into a watchable draft: one video track that follows the
conversation, program audio continuous underneath.

### The master is never edited

All cutting happens on a **new timeline** (same fps/canvas). The synced master is the
source of truth every derived cut can be rebuilt from — if it changes, every offset
becomes unverifiable. When any instruction, taken literally, would break sync
(e.g. "start the audio at frame 0" when its synced position is not 0), keep sync and
say why. Preserving sync outranks literal wording.

Treat the speaker-follow timeline as a source or cutting board, not as an opaque clip.
When the user wants part of it in an existing edit, materialize only those ordinary
picture and audio ranges into the main timeline. To substitute another angle later,
take the same wall-clock interval from that camera using the master offset. The master
is the timing record; no hidden multicam data structure is required.

### The conversation drives the cut

Build one conversation document first: each person's speech taken from their
assigned program source (Part 3), crosstalk dropped, interleaved by wall-clock via
the Part 2 offsets, with real names and timestamps. Then write the angle plan —
start–end, angle, reason — **before placing anything**.

If the user also asks to remove fillers, shorten answers, or restructure the
conversation, use the talking-head workflow to decide which speech ranges remain.
Multicam owns sync and angle choice; talking-head owns content. Finalize the content
plan before materializing picture cuts.

### Editorial objective

Keep the viewer oriented, emotionally informed, and visually awake with the fewest
cuts that add something.

**Meaning chooses what to show. Rhythm chooses when to cut. Orientation chooses how
wide to go.** Stay while the frame is still revealing; move when another frame gives
more; return to the room when the relationship needs refreshing.

Angle rules (defaults, user's brief wins):

- **Carrier of the moment**: the active speaker is the baseline, not the law. Show
  the speaker when information originates there, the listener when the reaction is
  the meaning, a pair or subgroup when the relationship carries the beat, and the
  room when attention is divided.
- **Shot scale**: use the smallest grouping that preserves the beat. `CU` isolates
  thought or emotion; `MS` is the conversational default; a two-shot or group shot
  shows a relationship; `WS` restores geography. Use a wider view at a new
  question or topic, a participant change, overlap, shared laughter or silence,
  physical movement, or after a long run of isolated singles. Hold it long enough
  to read — usually 2–5s — then tighten when attention concentrates.
- **Duration pressure**: a normal shot needs about 2–3s to arrive. Treat 4–10s as
  a useful conversational range, not a metronome. After roughly 8–12s on an
  unchanged single, actively look for a motivated alternative. By 20–25s the hold
  should be deliberate. There is no maximum while performance, emotion, or visual
  information is still developing.
- **Speaker changes**: do not chase every sound. A <2s interjection usually stays
  on the current shot. Let an unanticipated new speaker begin for about a second
  before cutting; a direct question may motivate showing the respondent while they
  prepare to answer. A short important line — introduction, direct address,
  punchline — earns a shot; widen around it if necessary.
- **Reactions**: use a visible, truthful reaction when it changes how the line lands
  or refreshes a static hold. A reaction is usually 1.5–3s. Do not insert a generic
  nod merely for variety, and never borrow a reaction from another wall-clock moment.
- **Rhythm**: cut on a thought, breath, gesture, look, laugh, or relationship change,
  not on a timer or arbitrary word boundary. Sentence boundaries are safe, not
  mandatory.
- **Seams**: one clean speech seam at a sentence boundary needs nothing. Cover a
  cluster of visible jump cuts with one wall-clock-synced reaction or relationship
  shot spanning the cluster. Never break source sync to manufacture coverage.
- A camera restart under an unchanged angle is a same-angle join, not an editorial
  cut. Count and report it separately.

### Materialize flat

- **One video track**, alternating angle segments, no gaps. Each segment's source
  offset comes from the master placement math, never re-derived from transcripts.
- For a full-length speaker-follow draft, keep program audio continuous and untrimmed
  at its synced position. For a shortened or reordered cut, use identical kept
  source-time ranges on every program mic so picture and audio cannot drift.
- Mute every camera segment's embedded audio.
- If audio outruns picture (recorders stopped later), keep the audio and leave the
  tail dark. Never fabricate or freeze picture to cover it.

### Keep program audio coherent

Keep every isolated program mic open across each retained range. When program audio
comes from camera mics, use only the chosen source for that passage, crossfade at a
quiet boundary, and do not stack them. Word timestamps carry ±30–60ms of ASR noise,
so boundaries need natural handles; hard cuts on word
timestamps clip breaths and word onsets. Backchannel ("嗯", laughter) on an idle
mic is part of the conversation. After content editing, run the talking-head
workflow's audio smoothing step. If mic isolation is weak, flag it for an audio pass
rather than gating speakers independently.

### Verify last, and prove it with numbers

Verification is the **final action**, after every edit and every visual check.
Anything that touches the timeline — including dragging in a browser to look at a
frame — can move an item. **Inspection is not read-only.** If you interacted with
the editor UI at any point, re-run these checks afterwards; a pass from before that
interaction is void.

- **Wall-clock invariant**, every segment and every program-audio item: timeline
  time − source time must equal that source's master offset.
- **Angle spot check**: at a few sampled shots — include the longest — the dominant
  speaker in the window must match the angle's person.

Report the **actual values**, not the word "verified": `MIC1 at 111, MIC2 at 111,
offsets 3.700 / 3.700`. A claim you cannot print numbers for has not been checked.

The same rule governs recovery. If you disturb an item, **restore every field and
re-read the row to prove it** — position, duration, and source offset each fail
independently, and fixing the one you noticed is not a restore.

Also report: cut count, shot length min/avg/max, the conversation document, the
plan, and every place a rule conflicted with the material plus what you chose.

Referenced files: 1

music4.14 KB

View saved version →

---
name: music
description: Shared ChatCut entry point for newly generated music, soundtrack, theme music, 配乐, background music, an intro theme, a music bed, BGM, or vocal songs. If vocals are unspecified, ask whether the user wants a vocal song or instrumental music before choosing a generator. Use `submit_music` for either branch, with `generationType` selecting instrumental or song mode.
user-invocable: true
---

# Music

Use `submit_music` to create a new instrumental background-music or vocal-song audio asset. For an explicit sung song, lyrics-to-song request, or vocal MV song, read [`references/song.md`](references/song.md) for the lyric workflow, then call `submit_music` with `generationType:"song"`. When a request only says music, soundtrack, theme music, or 配乐 and does not specify vocals, ask the user to choose vocals or instrumental first instead of assuming.

The tool submits one generation job and returns a `jobId`. The generated audio asset is available after `track_progress` reports completion.

## Capability Boundary

Mureka creates a new original music asset; it does not edit, clean up, remix, separate, or adjust existing audio.

Mureka also cannot guarantee exact beat, drop, lyric, or timestamp alignment. Generate the music asset first, then use timeline tools for placement, trimming, looping, fades, and ducking. If the user requires precise beat-level sync, explain that this must be handled as a timeline/audio edit rather than guaranteed by the generation model.

## Workflow

1. Treat the vocal choice as a required gate before asking about style, mood,
   duration, instruments, or any other creative detail. Words such as music,
   soundtrack, theme music, 配乐, "for this video", or "for an MV" do not
   resolve that choice. If the user has not explicitly chosen vocals or
   instrumental music, reply only with the equivalent binary question in the
   user's conversation language. For Chinese use "要带人声的歌曲,还是纯音乐/无
   人声配乐?"; for English use "Would you like a vocal song or instrumental,
   no-vocals backing music?" Do not spend credits or ask other questions in
   that turn.
2. For explicit instrumental, BGM, music-bed, or no-vocals requests, write a
   concise prompt describing style, energy, instrumentation, mood, tempo, and
   edit role, then call `submit_music` with `generationType:"instrumental"`.
3. For explicit singing, lyrics-to-song, or vocal-MV requests, read
   [`references/song.md`](references/song.md), resolve lyrics and creative
   direction there, then call `submit_music` with `generationType:"song"`.
4. Provide a short descriptive `name` when useful.
5. Use `track_progress` if the next edit needs the completed asset.
6. Place, trim, loop, or duck the audio with timeline tools after the asset
   exists.

## Tool Input Shape

`submit_music` accepts:

- `generationType` — `instrumental` (default) for no-vocals BGM, or `song` for
  an original vocal song;
- `prompt` — required for instrumental mode and optional music/vocal direction
  for song mode; maximum 1,024 characters;
- `lyrics` — required for song mode and invalid for instrumental mode; maximum
  5,000 characters;
- `gender` — optional `female` or `male` preference for song mode only;
- `name` — optional descriptive asset name.

## Prompt Shape

For instrumental music, combine:

- genre or instrumentation: "minimal electronic", "warm acoustic guitar", "cinematic piano";
- energy: "upbeat", "calm", "tense", "confident";
- role: "under a product walkthrough", "intro sting", "background bed under speech";
- constraints: "not distracting", "no vocals", "short loop feel" when needed.

## Rules

- Never route an ordinary BGM request to vocal song generation, even if the
  prompt is stylistically ambiguous.
- Do not cover speech with loud music; lower volume or duck under narration.
- Do not call the generation path to remix, clean up, separate, or replace an
  existing copyrighted recording. The result is a new original asset.
- Submit one generation job by default. Do not create several variants in
  parallel unless the user explicitly asks for that.
- Do not promise exact lyric timing, beat alignment, or automatic MV assembly.

Referenced files: 1

product-help7.18 KB

View saved version →

---
name: product-help
description: Answer current ChatCut product questions using the latest official Docs, Releases, and Changelog. Use for how ChatCut works; UI layout, buttons, and feature instructions; troubleshooting; feature or fix availability on Web, Desktop, or Agent Plugin and minimum version requirements; credits including costs, balance, usage history, validity, or recent charges; plans, subscriptions, renewals, ChatCut Pro, pricing, card or Alipay (支付宝) payments, and billing; Desktop downloads; Agent Plugin installation or updates for Codex and Claude Code; and manual GUI guidance when an action cannot be completed directly. NOT for live project, asset, timeline, or item state; use the matching read tool instead.
user-invocable: false
---

# ChatCut Product Help

For a ChatCut product question, the first substantive action is to search the
exact topic within `https://chatcut.io/docs` and open the most relevant official
page. Do not answer from memory, a search snippet, or an unopened result. Use the
host's native search and page-reading capabilities; this Skill defines retrieval
and decision rules, not a second copy of product facts.

## Route the question

- For live project, asset, timeline, item, account, or job state, use the matching
  ChatCut read tool. If the question is only about live state, no Docs lookup is
  needed.
- For current behavior, visible UI labels, installation, credits, pricing,
  billing, and troubleshooting, use the owning product Docs page.
- For availability, regressions, or version requirements, use **Current versions**
  and the **Changelog** in addition to the owning product Docs page.
- For a mixed question, use tools for live state and Docs for what that state means
  or what the user should do next.

Treat this Skill only as stable routing guidance. Never answer a changeable
product fact from remembered Skill content when current Docs can be checked. Do
not use unofficial domains.

## Establish the relevant product surface

Before answering a surface- or version-dependent question, identify both the
surface hosting this agent and the surface the user is asking about.

- Prefer the explicit ChatCut runtime profile or the latest `[chatcut-runtime]`
  envelope. Map `chatcut-web-editor` to Web, `chatcut-desktop` to Desktop, and
  `chatcut-workbench-client` or `chatcut-chat-connector` to Agent Plugin.
- When Agent Plugin is hosted by ChatCut Desktop, treat its bundled Skills as part
  of the Desktop installation. A Desktop update updates that Plugin surface too;
  never ask the user to update the Plugin separately in this runtime.
- A user can ask an Agent Plugin about Desktop, or ask Desktop about a Web project.
  When the user names a target surface, answer for that target instead of assuming
  it is the same as the agent's host surface.
- Use an app URL, client name, or version supplied by runtime context as supporting
  evidence. Do not infer the current surface merely because the user mentions a
  product in an example.
- If no reliable signal exists and the surface changes the answer, ask which of
  Web, Desktop, or Agent Plugin they are using. If the answer is surface-independent,
  proceed without an unnecessary question.

## Find the smallest useful source

1. Extract the feature name, visible UI label, error text, product surface, and
   version from the user's question or runtime context when available.
2. Search the exact topic with a concise query containing only its essential terms.
   Prefer the most specific official result and open the smallest set of pages that
   establishes the answer.
3. For ordinary usage questions, start with the owning feature or troubleshooting
   page. For availability, regression, update, or version questions, directly open
   `https://chatcut.io/docs/releases` and `https://chatcut.io/docs/changelog`, then
   open the matching Changelog entry.
4. For one user question, issue at most two search queries total: the exact topic,
   then one narrower query if needed. After that, stop searching and use
   `https://chatcut.io/docs/llms.txt` only as an index, then open the target page.
   If the canonical pages and index are unavailable or still do not identify a
   supporting page, stop retrieval and state the uncertainty; do not keep trying
   broader keywords. Do not load the full Docs corpus or treat the index itself as
   the answer.
5. Treat search snippets as discovery only. Base claims on opened page content and
   cite the exact official page supporting each consequential claim about product
   behavior, platform support, price, credit cost, or version. Prefer a detail page
   over the Docs home or a Changelog listing page.
6. Do not infer an unsupported feature or version from a similar entry. If official
   pages do not establish the answer, state the uncertainty.

## Decide whether the user must update

- Web changes are deployed server-side. If Current versions or the Changelog says
  the Web release is live, do not tell the user to install an update.
- For Desktop or a standalone Agent Plugin, compare the user's current version only
  with the minimum version on the relevant individual change. A release header's
  current Desktop or Plugin version is not automatically every change's minimum.
  If the current version is older, actively recommend updating and use the official
  installation/update page for the steps.
- Treat a capability available on Web as available through Agent Plugin unless the
  official Docs or Changelog explicitly says otherwise. Web UI, service, and tool
  changes are inherited server-side and do not require a Plugin update. Require a
  newer standalone Plugin version only when the individual change explicitly names
  a Plugin minimum because its bundled Skill changed.
- When a capability is explicitly Desktop-only, tell a Web user that it requires
  Desktop and link the Desktop download instructions. Do not infer Desktop-only
  merely because a release also lists a Desktop version.
- If the user's product or version is unknown and it changes the recommendation,
  ask for it or show how to check it instead of guessing.

## Answering and fallback rules

- Try to complete supported editing actions with ChatCut tools first. Give manual
  UI steps only when the user asked for guidance or the action cannot be completed
  directly.
- Use localized visible UI names verified in the current Docs. Do not invent routes,
  buttons, plan terms, prices, credit costs, platform support, or version numbers.
- Current feature Docs describe present behavior. Changelog entries describe
  historical rollout and minimum versions. If wording appears inconsistent, use
  the current feature page for present behavior and the release entry only for the
  historical/version claim; state any unresolved ambiguity.
- Do not claim access to live balances, ledgers, quotas, subscription state, or
  installation state unless a tool or runtime context actually provides it.
- If official Docs cannot be reached or do not answer the question, say what could
  not be verified and avoid a confident guess. Offer the closest official Docs
  entry or the in-product Feedback path after giving all verified guidance.
- Do not expose internal implementation details, release automation, review notes,
  or unpublished preview content to end users.
shader-gen13.1 KB

View saved version →

---
name: shader-gen
description: AI shader generator for WebGL video effects, transitions, masks, and color grading (LUT / 调色 / 电影感 / film look). Use when the user wants a video effect (滤镜 / 特效), a transition (转场 / crossfade / wipe / cube / 3d), a mask (蒙版 / 遮罩 / reveal), a zoom / push-in (推近 / 推镜头), or a color grade — try the built-in effects (zoom, builtin LUTs) before generating a new shader.
user-invocable: true
---

# Shader Generator

Submit-only: creates a backend generation job, returns `jobId`. Use the `track_progress` tool for job lifecycle after submission.

**Always use `generate.ts` for new shaders.** Manual authoring is only for editing existing asset code — never as a fallback when generation fails.

## Catalog-first rule — try existing assets before generation

Before generating a shader, call `browse_library` unless the user names an exact asset id that is already visible in `browse_assets`.

`browse_library` is the source of truth for built-in effects, built-in transitions, and project effect/transition assets. Built-ins are stable global asset ids, not per-project DB assets, so they may not appear in `browse_assets`.

Apply catalog entries with `edit_item`, do **not** call `submit_shader`.

Good catalog searches:

```text
browse_library(query: "zoom")
browse_library(category: "transitions", query: "dissolve")
browse_library(category: "audio-fx")
```

Generate only when no catalog entry matches the user's intent closely enough.

### `builtin:zoom` uses the same track-bound placement as every effect

Effects always have timeline geometry. For a whole-clip effect, pass `targetItemId`; ChatCut resolves it to a clip-anchored range covering that clip. For an explicit timeline range, pass `trackId` + `trackBoundFrom` + `trackBoundDurationInFrames`.

```text
# Zoom on the entire video clip
edit_item(json: '{"adds":[{"type":"effect","assetId":"builtin:zoom","targetItemId":"<clip-id>","propertyOverrides":{"magnification":1.5,"shape":"hold"}}]}')

# Zoom on a sub-range of the clip (e.g. frames 90–150 only, a punch zoom on a beat)
edit_item(json: '{"adds":[{"type":"effect","assetId":"builtin:zoom","mode":"track-bound","trackId":"<trackId>","trackBoundFrom":90,"trackBoundDurationInFrames":60,"propertyOverrides":{"magnification":2,"shape":"punch"}}]}')
```

Use `preview_timeline({views:["timeline"],tracks:["V1"]})` to obtain track item ids and timeline-frame ranges. Use `inspect_item({itemId:"..."})` when exact item detail is needed.

| Key             | Type   | Range / values                             | Default | Notes                                          |
| --------------- | ------ | ------------------------------------------ | ------- | ---------------------------------------------- |
| `magnification` | number | 1–4                                        | `1.5`   | Zoom factor; 1 = no zoom, 2 = 2× in            |
| `focalPointX`   | number | 0–1                                        | `0.5`   | Horizontal focal point (0 = left, 1 = right)   |
| `focalPointY`   | number | 0–1                                        | `0.5`   | Vertical focal point (0 = top, 1 = bottom)     |
| `shape`         | select | `punch` / `hold` / `slow-push` / `instant` | `hold`  | Animation curve                                |
| `focalMode`     | select | `auto` / `manual`                          | `auto`  | `auto` picks subject; `manual` uses focalPoint |
| `easeInFrames`  | number | 0–60                                       | `8`     | Frames to ramp in                              |
| `easeOutFrames` | number | 0–60                                       | `8`     | Frames to ramp out                             |

Omit `propertyOverrides` entirely for default zoom. Send only the keys you want to change — patch semantics.

### Clip-anchored vs adjustment-track

Effect items have two placements, both with a concrete time range:

- **Clip-anchored**: pass `targetItemId` for the whole clip, or include an explicit range that intersects the clip. The stored range is local to that clip and follows it when it moves.
- **Adjustment-track**: pass `trackId` + `trackBoundFrom` + `trackBoundDurationInFrames` for a range over empty track space. The range is absolute on the timeline.

Use `targetItemId` for whole-clip effects. Use explicit track geometry only when the requested range differs from a clip's full duration.

### Built-in LUT properties

```text
edit_item(json: '{"adds":[{"type":"effect","targetItemId":"<clip-id>","assetId":"builtin:slog3-s709","propertyOverrides":{"intensity":1}}]}')
```

| Key         | Type   | Range | Default | Notes                          |
| ----------- | ------ | ----- | ------- | ------------------------------ |
| `intensity` | number | 0–1   | `1`     | LUT strength; 1 = full applied |

To swap: delete the effect and re-add with a different `assetId`. To remove: delete the effect item.

User-uploaded `.cube` LUT assets take this exact same shape — only `assetId` differs. See "Applying an Existing LUT Asset" below.

## Supported Targets

Effects and transitions apply to `video`, `image`, and `gif` items.

## Type Routing

Before generating anything, check two non-generation paths first:

1. **Catalog entry** — use `browse_library` for built-in and project effects/transitions.
2. **User-uploaded `.cube` LUT asset** that already exists in the project library — bind it instead of generating, see "Applying an Existing LUT Asset" below. It shows up in `browse_assets` as `type: effect` with a `lut`-typed entry in `editableProperties`.

| User wants                                                           | `--type`     |
| -------------------------------------------------------------------- | ------------ |
| Video appearance (color, blur, glow, grain, distortion)              | `effect`     |
| Color grade / look (teal-orange, cinematic, vintage, LUT-style)      | `effect`     |
| Visibility control (mask, reveal, wipe, shape cutout, gradient fade) | `effect`     |
| Blend between clips (crossfade, dissolve, slide, 3D cube/page flip)  | `transition` |

"LUT-style" in the table means **generating a fresh GLSL color grade that resembles a LUT** — only when the user wants something new. If they want to apply a `.cube` file already in the library, don't generate; bind the existing asset instead.

No separate LUT or mask generator for the generation path — those are all `effect`.

## Applying an Existing LUT Asset

**Default target is the timeline.** "Apply this LUT" means an `edit_item` effect on the clip. Only reach for `edit_asset sourceLut` when the user asks for the source itself everywhere it appears — every clip cut from that asset, or a log-to-Rec.709 normalization of the footage — because that changes every instance of the asset on every timeline.

`.cube` files uploaded by the user become **effect assets with `category: "lut"`**. Applying one to a clip is **not** generation — it is the same `edit_item` effect shape as a built-in LUT, with that asset's own id as `assetId`:

```text
edit_item(json: '{"adds":[{"type":"effect","targetItemId":"<clip-id>","assetId":"<lut-effect-asset-id>","propertyOverrides":{"intensity":1}}]}')
```

Key points:

- `assetId` is the LUT effect asset's real id. There is no literal `"lut"` assetId, and no LUT binding nested inside `propertyOverrides`.
- `propertyOverrides` carries only `intensity` (0–1, default 1). The `.cube` binding lives on the asset, not on the effect item.
- Find the id with `browse_assets type:"effect"`: a LUT asset is the one whose `editableProperties` contain a `lut`-typed key. Built-in LUTs come from `browse_library category:"luts"`.
- `targetItemType` defaults to `video`; also supports `image`, `gif`.
- To swap: delete the effect and re-add with the other LUT's `assetId`. To remove: delete the effect item.
- To grade a whole source clip everywhere it appears instead of one timeline item, use `edit_asset` update with `{"sourceLut":{"assetId":"<lut-effect-asset-id>"}}` on the video/image/gif asset; `{"sourceLut":null}` removes it.
- If the `.cube` is not in the project yet, you cannot get it in: this surface has no LUT import path, and `.cube` is not an accepted chat attachment either, so never ask the user to send you the file. Ask them to drag it into the editor's media pool instead — the editor registers it as a LUT asset and it then shows up in `browse_assets`. Do not tell them ChatCut cannot handle `.cube`; the editor can.

Do not call `submit_shader` for this path.

After applying, confirm with `preview_timeline` that the effect is listed on the target clip's track. Do not report success from the `edit_item` response alone.

## Usage

Before calling `submit_shader`, restate the user's intent in one concrete sentence, then proceed immediately. After `track_progress` returns, state what was produced in one line — do NOT ask "Keep it or regenerate it?".

```ts
submit_shader({
  type: "effect",
  prompt: "Chromatic aberration with RGB split",
  name: "Chromatic Aberration",
});

submit_shader({
  type: "transition",
  prompt: "Smooth crossfade with soft edge",
  name: "Crossfade",
});

submit_shader({
  type: "effect",
  prompt: "Cinematic teal-orange color grade",
});

submit_shader({
  type: "effect",
  prompt: "Stronger version",
  referenceAssetIds: ["effect_asset_id"],
});
```

## Strategy

- Submit, then stop. Tell user the job was created.
- Use the `track_progress` tool for status/wait after submission.
- Generation always produces a library asset — never refuse because the timeline isn't ready.
- **Apply is separate and optional.** Only apply when user explicitly asks ("加到视频", "apply", "用到第一段"). When ambiguous, default to library-only.

## Editing Existing Properties

Any time you're about to edit shader `asset.properties`, applied effect/transition `item.propertyOverrides`, or promote a hardcoded shader value, read [`references/property-changes.md`](references/property-changes.md) first.

It reinforces that shader `properties` is an array, but the allowed shader property types are only `number`, `boolean`, `color`, `select`, and `vec2`. Motion Graphic properties are also arrays, but use a different type set.

## Parameters

| Param               | Description                                                                                                                                                    | Default |
| ------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------- |
| `type`              | `"effect"` or `"transition"` (req'd)                                                                                                                           | —       |
| `prompt`            | Description of the shader (req'd)                                                                                                                              | —       |
| `name`              | Asset name shown in library                                                                                                                                    | —       |
| `referenceAssetIds` | Asset ids. Image id → model LOOKS AT it for visual inspiration. Effect/transition id → reuse its code as style anchor (≤1 per submit, kind must match `type`). | —       |

## Output

Returns `{ success, job: { jobId, status }, manage: { status, wait, watch } }`.

## Applying to Timeline

Only when user explicitly requests. Refresh the affected timeline with `view:"timeline"`, passing `track` to narrow the read when appropriate.

### Effect

```text
edit_item(json: '{"adds":[{"type":"effect","targetItemId":"<id>","assetId":"<id>","enabled":true,"propertyOverrides":{}}]}')
```

### Transition

Requires two adjacent same-track endpoints. `edit_item` validates live seam feasibility and refuses durations that would require freeze frames or overlapping neighboring transitions. If the add fails, retry with the suggested `durationInFrames`, trim the clips to expose handles, delete/shorten neighboring transitions, or keep a hard cut.

```text
edit_item(json: '{"adds":[{"type":"transition","assetId":"<id>","outgoingItemId":"<id1>","incomingItemId":"<id2>","durationInFrames":30}]}')
```

## Validation & Verification

### Backend Validation

When generating via `generate.ts`, the backend handles validation automatically (transpile, AST security, class structure, retry on failure).

### Manual Code Verification

**NEVER write shader code from scratch.** Always use `generate.ts` for new shaders. This section is ONLY for modifying existing shader code that was already generated.

When writing shader code manually, read `${CLAUDE_SKILL_DIR}/references/design-principles.md` first. If the change touches editable properties, also read `${CLAUDE_SKILL_DIR}/references/property-changes.md`.

Typical workflow:

1. `inspect_asset` with the shader `assetId` and `includeCode: true` — read the current source.
2. Edit the source in your own context.
3. `edit_asset` with `action=update`, the same `assetId`, and the full replacement source inline in `json.code`. Validation runs automatically on update — if code is invalid, the update is rejected with error details.

Referenced files: 5

talking-head-guide24.9 KB

View saved version →

---
name: talking-head-guide
description: Guide for editing speech-led videos where spoken delivery or conversation drives the cut — single-speaker talking-head / 口播, two- or multi-speaker interview / 访谈, video podcast, lecture, tutorial, course, and similar formats. Load before analyzing, planning, or performing non-trivial edits of those formats, including speech cleanup (剪口播 / 口播剪辑 / 去口癖 / clean up fillers / smooth speech), pause or repeated-take removal, motion graphics layered onto the footage (口播加 MG / 加动画), B-roll (加 B-roll / add B-roll), music, or captions. For motion graphics specifically, use this together with the active Motion Graphics skill/workflow available in the current ChatCut environment — this skill adds speech-specific guidance (rhythm-aware timing, frame-aware placement, subject/caption protection, placement verification).
user-invocable: true
---

# Speech-Led Video Editing (Talking Head, Interview, Podcast)

## What this skill covers

**Required input**: an existing speech-led source registered or imported into the project — for example a single-speaker talking-head / 口播, a two- or multi-speaker interview / 访谈, a video podcast, lecture, tutorial, or course. For transcript-based A-roll editing, follow `asset-import`, then start as soon as the transcript is ready. If the user wants to start without source footage (e.g., generate a fresh talking-head from scratch), this skill doesn't apply.

### When source media is missing

- **Inside ChatCut Web (including embedded agents):** Reuse source media already in the project or at a user-provided accessible path; otherwise, load `widget-forms` and include its native file-upload field (`<form-files>`) in the reply, together with any related edit preferences.
- **Standalone Claude Code:** The user may provide the source in the conversation or upload it directly in the open ChatCut editor. Use `asset-import` when a conversation attachment or local path still needs to be imported into ChatCut. For an editor upload, continue from the registered project asset without importing it again.
- **Standalone Codex:** Ask the user to attach or drop the source into the conversation, then follow `asset-import`. Do not use or redirect the user to the editor's Browser upload path: it can currently crash the Codex Browser runtime.

If this workflow is running through a Codex/connector host and the task creates, targets, or opens a ChatCut project for the user, satisfy any returned `browserHandoff.required=true`, `Codex internal Browser handoff`, or equivalent live project handoff before starting nontrivial edits and again before final delivery if the visible editor no longer matches the project.

Independent treatments that can be applied to speech-led videos. Pick the ones that match what the user wants — not all are needed every time.

- **A-roll editing** (Chinese product term: **语音剪辑**, including **去口癖、停顿、重复**) — transcript-based speech editing. Common operations include cleanup, highlight extraction, restructure, opening hook, and others as needed for the aligned outcome.
- **Motion graphics overlay** (write the full name **Motion Graphics** in English user-facing copy, not "MG"; the fixed Chinese product term is **MG 动画**, not alternatives such as "动效", "字幕条", or "动态字幕") — reinforce key information, structured content, and topic transitions with on-screen motion graphics
- **B-roll** (industry term — keep as "B-roll" in any language, do not translate) — cover jump cuts or visualize what's being said
- **Background music** (Chinese product term: **背景音乐**) — set mood and smooth micro-gaps
- **Captions** (Chinese product term: **字幕**) — on-screen text for accessibility
- **AI Voice Isolation** (Chinese product term: **AI 人声隔离**) — clean or isolate spoken human voice with DeepFilterNet3, picture untouched. See the `voice-isolation` skill.

> When the user's language is Chinese, **use the exact product terms in parentheses above** in widget options, choices options, and conversation copy. Do not retranslate them; that would conflict with terminology elsewhere in the product.

## What shapes the edit

Beyond picking treatments, a talking-head edit is shaped by several orthogonal variables. When the user's ask is vague, these are what's worth clarifying first:

- **Target** — platform (YouTube / TikTok / Shorts / ...), desired length, aspect ratio
- **Which treatments to apply** — the treatments above are optional; don't assume all of them apply
- **Pacing / tone** — tight / energetic / formal / casual; brand or voice preferences if stated. (For MG visual style, follow the active Motion Graphics skill/workflow.)

When more than one of these variables is missing, ask with one form after loading `widget-forms`. Do not ask markdown numbered questions and then append `<choices/>` for only one part of the same intake.

## Order of execution

When multiple treatments have been aligned with the user, they depend on each other and must be finalized in dependency order. This section is **only relevant after alignment** — it doesn't tell you what to start with on a fresh request.

The speech timing (set by A-roll editing) anchors everything downstream — MG placement, B-roll cut-covers, music duration, and caption sync all reference the final speech timeline.

So: finalize A-roll editing before committing any visual, audio, or text layer. Don't write captions against pre-edit speech, don't cut music to pre-edit length, don't place MG against timing that will shift.

Once the cut structure is final, run `smooth_audio` once as the last audio step — it micro-crossfades every hard audio cut and fades exposed edges so edits don't pop. Run it after `apply_script`/`clean_script` restructuring (reflow drops transitions); it's idempotent, so re-run it if the timeline changes again.

**You must confirm the result with the user after each major step before starting the next**, unless the user has explicitly asked to run end-to-end without stopping. Key checkpoints when multiple treatments apply: after A-roll editing finalizes the speech timing; before MG generation (confirm style and direction, and, when it isn't obvious, whether it sits over the video as an overlay or takes the whole frame); after MG generation; same pattern for B-roll, music, and captions. **Don't bundle multiple checkpoints into one response — confirm each step separately.** An upstream mistake forces redoing everything downstream (e.g., MG placed against pre-cleanup timing wastes generation credits when the timeline shifts).

---

## A-roll editing

### Scenario

In a talking-head workflow, the first step is usually A-roll editing: editing the original spoken footage.

A-roll edits are ultimately applied to the timeline and change what the viewer actually hears and sees. However, the editing decisions should usually start from the transcript, because the core question is: what spoken content should the viewer hear, and what should be removed, compressed, or reordered?

### Common A-roll tasks

A-roll editing is not only cleanup. First decide what spoken-content task the user is asking for, then choose the editing strategy and tools.

Common tasks:

- **Cleanup** — remove mistakes, repeated attempts, verbal habits, filler words, and meaningless pauses so the speech becomes clearer and more natural.
- **Highlight extraction** — pull the most valuable, opinionated, emotional, or topic-relevant moments from longer footage.
- **Restructure** — reorder spoken content, such as moving the conclusion earlier, grouping by topic, or combining scattered parts into a clearer structure.
- **Hook / short version** — use a strong claim, result, conflict, or question from the source as the opening, or compress long content into a shorter version.
- **Target-script / script alignment** — match, keep, and reorder spoken content according to a user-provided target script, target paragraph, or desired content.

Cleanup is the most common task and the one most likely to fail from bad boundary decisions. It is described in detail below. Other tasks get shorter rules, but still follow the shared A-roll principles: complete meaning, clear boundaries, and natural listening flow.

### Shared A-roll principles

These principles apply to all A-roll tasks, not only cleanup.

- **Decide the task before choosing the tool.** Do not let tool availability change the editing strategy.
- **Edit by complete semantic units.** Whenever possible, move/delete/keep complete sentences, complete ideas, complete answers, or complete steps. Do not cut out a half-sentence just because a few words match.
- **When the task names what to keep, trim to that boundary.** The inverse of the rule above, for any task that specifies which content to keep — restoring a specific sentence, matching a target script, pulling a named highlight, building a version: keep exactly the requested span. Trim the kept range to start and end at the requested words and drop the off-script head/tail of the source `[sN]` segment it sits in; keeping a whole segment for one requested sentence is over-keeping that drags in unrequested speech. This applies only when the task names what to keep — never to open-ended cleanup, where you keep complete units (above).
- **Do not stitch unfinished fragments across retakes.** Do not combine incomplete pieces from different attempts into one artificial sentence. This does not make the earlier attempt disposable: keep a complete useful lead-in, setup, contrast, category, evaluation, or context if it is not repeated later and can naturally connect to the later complete retake.
- **Preserve connective tissue.** List labels, contrast words, subjects, verbs, and adjacent source words are not filler when removing them makes a kept idea ungrammatical, abrupt, or misleading. Trim the smallest span that keeps the line speakable.
- **Keep listening flow natural.** The result should still have natural phrasing and breathing room. Do not make sentences feel glued together just to make them "clean."
- **Be conservative when boundaries are uncertain.** If unsure whether a cut harms meaning, logic, or listening flow, keep it or make a smaller cut.
- **Confirm complex changes first.** For complex restructuring, aggressive shortening, structural changes, or generated hooks, confirm target length, structure direction, and what to preserve with the user before editing.
- **Explain content, never indices.** You MUST NOT explain edits to the user with internal addresses such as `[sN]`, `[cN]`, `[gap]`, word indices, clip ids, or segment ids. The user cannot see those addresses and will not understand what they mean. Use the actual spoken content, a short quote, or a plain-language description of the edit.
- **Never name a screen position for a panel.** When you invite the user to review or fine-tune the result, call it "the Transcript panel" (Chinese product term: "文字稿面板") — never a direction (left / right / side / 左侧 / 右侧). The layout is rearrangeable and the panel does not sit in a fixed corner.

### Cleanup goals and decisions

#### What good cleanup means

Good cleanup does not mean making the video as short as possible, and it does not mean rewriting the speaker into a different script.

Good cleanup means:

- The logic stays coherent
- The expression becomes clearer
- The audio feels natural
- Obvious mistakes, repeated attempts, meaningless stalls, and filler are removed
- The speaker's intent, tone, and natural rhythm are preserved

Bad cleanup usually falls into two failure modes:

- Under-cleaning: obvious mistakes, repetition, long pauses, or filler remain.
- Over-cleaning: sentences are cut off, meaning is missing, rhythm becomes too hard, or the result sounds stitched together.

Default principle: remove defects without changing meaning; make speech smoother, not harder; prefer small local cuts over whole-sentence or whole-segment deletion; when unsure whether a cut harms meaning, keep it.

#### How to judge common cleanup cases

Below are the common cleanup categories and how to make editing decisions for each.

##### Meaningless filler words

Fillers fall into two categories.

The first category is clearly meaningless hesitation sounds. These are usually safe to remove:

- `um`
- `uh`
- `er`
- `ah`
- `呃`
- `额`
- `えー` / `えーと` / `えっと`
- `あー`
- `んー`

When they do not carry special meaning, use `clean_script` for bulk cleanup of its supported literal tokens: `um`, `uh`, `er`, `ah`, `呃`, and `额`. Japanese hesitation sounds require a Script edit (`read_script` → edit `timeline.md` → `apply_script`); `clean_script` does not currently remove them.

The second category depends on context and must not be removed by word list alone:

- `so`
- `like`
- `然后`
- `就是`
- `嗯`
- `啊`
- `那个`
- `那`
- `对`
- `所以`
- `但是`
- `あの` / `あのー`
- `その` / `そのー`
- `まあ`
- `なんか`
- `ちょっと`
- `で`
- `うーん`
- `やっぱり`

How to decide:

- If the word is only hesitation or padding, remove it.
- If it carries sequence, continuation, contrast, cause, reference, response, emphasis, or natural tone, keep it.
- If removing it makes the surrounding words sound hard-spliced, keep it or only compress the pause.
- If unsure, keep it.

Examples:

- `um, I think this solves the main problem` -> remove `um`.
- `It works like a checklist` -> keep `like`; it is a comparison.
- `The upload failed, so we retried it` -> keep `so`; it carries cause/result.
- `right after the call, send the recap` -> keep `right`; it modifies timing.
- `然后我们再看第二点` -> keep `然后`; it marks sequence.
- `えーと、次のポイントです` -> remove `えーと`; it is pure hesitation.
- `あの人が言っていた通りです` -> keep `あの`; it is a demonstrative, not hesitation.
- `あの、ちょっと待ってください` -> remove the leading `あの`; it is hesitation here.
- `まあ、悪くないと思います` -> keep `まあ`; it softens the judgement and carries tone.
- `なんか違和感がある` -> keep `なんか`; it means "somehow" here, not padding.

##### Retakes and repeated attempts

A retake is when the speaker retries the same intended idea because they misspoke, got stuck, forgot words, or restarted. Retake cleanup is not "delete repeated text." The goal is to keep one complete, natural, logically coherent version of the intended idea.

Use this decision path:

1. Decide whether it is really a retake.
   Treat it as a retake only when multiple attempts are trying to say the same intended idea. Do not treat it as a normal retake when the repetition is intentional emphasis, a rhetorical beat, a structural marker, or a second pass that adds new information or tone.
2. Define the complete version to keep.
   A complete version may include more than the main content sentence. It may need a lead-in, connector, section marker, topic setup, contrast, qualifier, subject, object, or conclusion. These are not filler when the kept content depends on them.
3. Cut only the failed or covered part.
   Remove only words that are wrong, dangling, abandoned, or fully covered by the kept version. The cut boundary starts at the repeated or failed idea, not automatically at the earlier transition, setup, or continuous speech. If earlier speech contains useful context that the kept version does not repeat, keep it.
4. Choose the best complete attempt.
   If several attempts are complete, usually prefer the later one because it is often closer to the speaker's intended take. But do not choose the last attempt mechanically. If the later attempt is missing needed context, structure, subject, object, or conclusion, keep the more complete version or preserve the missing lead-in from the earlier attempt.

A repeated lead-in is redundant only when another equivalent lead-in remains naturally connected to the kept content. If removing every copy makes the result lose structure or sound abrupt, keep one natural copy and remove only the extra restarts. Do not stitch unfinished fragments from different attempts into one artificial sentence.

Examples are patterns, not a closed list:

- Local false start inside a kept sentence:
  `There, there's no After Effects, no Premiere, no DaVinci Resolve learning.`
  Keep the complete sentence, but remove the abandoned restart:
  `There's no After Effects, no Premiere, no DaVinci Resolve learning.`
  Do not keep the stray first word just because the full sentence is otherwise useful.
- Repeated structural lead-in:
  `And secondly, ... and secondly, we're introducing a brand new UI.`
  Remove the extra restart, but keep one natural lead-in attached to the kept content:
  `And secondly, we're introducing a brand new UI.`
  Do not delete every structural marker and leave only:
  `We're introducing a brand new UI.`
- Useful setup before a failed ending:
  `Then the next one is different from comedy. It is popular on Disney Plus. It is called...`
  Later retake:
  `It is a popular Disney Plus show called Love Story.`
  Keep useful setup that the later retake does not repeat, and cut from the failure point:
  `Then the next one is different from comedy. It is a popular Disney Plus show called Love Story.`

##### False starts and unfinished fragments

Use `false starts / unfinished fragments` for this category. `False start` is the more natural editing/transcription term for a speaker beginning a phrase and then restarting or abandoning it; `unfinished fragment` makes the dangling half-sentence case explicit.

Only remove a fragment when it clearly does not form useful information.

Safe to remove:

- The speaker abandons the thought and a complete version appears later.
- The segment is only a dangling phrase, such as "this is actually..." with no completion.
- It is clearly the leftover beginning of a failed attempt.

Do not remove:

- A sentence that is imperfect but contains useful information.
- A lead-in that provides the subject, object, or context needed later.
- Content that provides setup, contrast, conclusion, emotion, or tone.

If only part of a sentence or segment is wrong, do not delete the useful content around it. Remove only the bad word, phrase, or pause; if a local cut cannot sound natural, keep the segment.

##### Pauses and breaths

Pause cleanup should default to compression, not zeroing out. Spoken video needs natural breathing room.

Default rules:

- Obvious long pauses over 0.8-1s: usually compress to about 0.3s.
- Between sentences: keep about 0.3-0.5s so listeners can hear natural phrasing.
- Around topic shifts, contrast, or emphasis: keep slightly longer pauses when needed; do not make the delivery too rushed.
- Short breaths inside one sentence: if they are normal breathing, do not remove them.
- Clear long pauses inside one sentence: compress them, but not so tightly that adjacent words sound glued together.
- Long pauses before a retake: if the failed attempts around it are removed, remove the pause with them.
- If the user provides explicit thresholds, follow them. For example, if the user says "only process pauses over 0.8s and keep at least 0.3s", do not process natural pauses under 0.8s.

### Other A-roll task guidance

Load only the reference for the branch the user selected. Do not preload references for unrequested treatments; default cleanup requires no reference.

For A-roll editing scenarios other than cleanup — including highlight extraction, restructure, hook / short version, target-script / script alignment, or building versions, highlights, and excerpts — read `.claude/skills/talking-head-guide/references/other-a-roll-editing-scenarios.md` completely before planning or editing that A-roll branch.

### A-roll / transcript-based editing workflow

Use this flow for any A-roll task driven by transcript meaning.

If the user asks to preview the whole requested edit before it is applied, do not run `clean_script` first because it writes immediately and has no preview parameter. Call `read_script({ showSilence: true })`, stage fixed-filler strikes, semantic edits, and pause edits together in `timeline.md`, then call `apply_script({ preview: true })`. Encode pause compression as `[silence=Xs→Ys]` with the Unicode `→`; a bare change to `[silence=Ys]` is not the cross-surface rewrite syntax and Web will not apply it as compression. When a struck filler has retained silence on both sides, deleting it joins those durations: a 250ms cap is their combined retained total, not 250ms per side. For example, left `0.48s→0.25s` plus an unchanged right `0.13s` yields `0.38s`; allocate the two sides to total `0.25s` instead of expecting the planner to infer a global cap. If preview validates and the request already authorizes application, immediately call `apply_script` again without preview; do not ask for another confirmation. If the user explicitly requested preview only, stop after the preview.

For edits that may be applied immediately, use the following flow:

1. Start with orientation. Call `read_script`, then read `timeline.md` once to understand the user's goal, the content structure, and whether fixed fillers or long pauses are present. If you will run `clean_script`, do not build the full semantic edit from this pre-clean read.
2. Outside the preview-first path above, run the mechanical cleanup pass before semantic editing when fixed fillers or long pauses are present. Do not use this step for context-dependent fillers, retakes, repeated sentences, or anything that needs meaning.
3. After `clean_script`, always read the refreshed clean `timeline.md` before semantic editing. Use this refreshed file as the source of truth; previously read text may be stale. Then edit with semantic judgment: choose the best retake, clean false starts, remove repeated or failed attempts, preserve useful setup and context, reorder content when needed, and keep the speech natural. Use the current filesystem tool schema: batch related exact replacements if `Edit` exposes an `edits` array; otherwise use its supported exact-replacement fields and check validation after each edit. Inline strikethrough removes spoken words; never remove unstruck source words or correct ASR in `timeline.md`.
4. Validation runs after each filesystem edit. If diagnostics are returned, the changed `timeline.md` remains current: read it and fix forward with `Edit`. Only after validation passes, immediately apply the semantic edit with `apply_script` before starting another global pass.
5. Read the regenerated clean result and check what the viewer will actually hear: broken logic, missing context, over-deletion, missed cleanup, wrong order, or pauses that feel too tight or too long. Fix clear problems only.

`[sN]` rows are ASR segments, not semantic units. A complete sentence, idea, retake, or transition may span several `[sN]` rows, and one `[sN]` row may contain only part of a sentence. Before deciding what to delete or keep, mentally reconstruct the complete spoken sentence or idea across adjacent rows.

---

## MG Overlay

### Goal

Motion graphics layered into A-roll reinforce what the speaker is conveying — deepening the audience's impression of the key points and helping them grasp content that's hard to land through speech alone. Complete A-roll editing first; MG timing is based on the post-edit timeline.

This section only adds talking-head timing, frame-composition, subject/caption protection, and review constraints. For visual style alignment, MG creation or authoring, implementation constraints, editable properties, asset sizing, and verification, use the active Motion Graphics skill/workflow available in the current ChatCut environment.

### MG workflow

For talking-head MG work, treat the video as one edited piece, not as isolated graphics.

When talking-head Motion Graphics is selected, including a plan-only or handoff-analysis request, first activate the Motion Graphics Skill available on the current surface: call `Skill` with `skill: "motion-graphic-gen"` when that generator Skill is available; on Codex, where it is intentionally absent, call `Skill` with `skill: "create-motion-graphics"` instead. Then read `.claude/skills/talking-head-guide/references/motion-graphics.md` completely before planning, generating, placing, or reviewing MG. This talking-head guide supplies the scene-specific editorial context but does not replace the active Motion Graphics Skill.

---

## B-roll

For B-roll, read `.claude/skills/talking-head-guide/references/b-roll.md` completely before sourcing, choosing, placing, or reviewing it.

---

## Multicam (multiple camera angles of the same take)

For Multicam, read `.claude/skills/talking-head-guide/references/multicam.md` completely before aligning or switching multiple camera angles of the same take.

---

## Track roles (turn on auto-ducking)

For Track roles, read `.claude/skills/talking-head-guide/references/audio-and-music.md` completely before assigning audio roles or changing ducking behavior.

---

## Background Music

For Background Music, read `.claude/skills/talking-head-guide/references/audio-and-music.md` completely before adding or fitting it.

---

## Captions

For Captions, read `.claude/skills/talking-head-guide/references/captions.md` completely before editing them.

Referenced files: 6

transcription2.67 KB

View saved version →

---
name: transcription
description: Hosted ChatCut plugin sessions only (the `chatcut` MCP server). If the conversation is driving ChatCut Desktop (a `chatcut_desktop*` MCP server), load this skill only when the user explicitly chooses the plugin/web surface — desktop sessions otherwise ship their own instructions and tools. Use for ChatCut transcription, transcript readiness, captions, subtitles, transcript repair, filler removal, and speech-led editing setup.
---

# Transcription

1. Use `browse_assets` to identify the video/audio asset and its transcription state.
2. For newly imported local media, complete the `asset-import` workflow first.
3. Check `track_progress` with target `transcription`. It returns current status; follow the returned check-back guidance and do not busy-loop.
4. Use `find_transcript` for timestamped text lookup.
5. Use the current caption tools to enable, inspect, translate, or style captions only after transcription is ready.

Do not call a transcription stuck from one pending status. Treat an explicit failed terminal state immediately; otherwise allow at least `max(5 minutes, min(60 minutes, 2 x asset duration))`, or at least 10 minutes across multiple checks when duration is unknown.

`no_audio` / no-speech is a successful terminal analysis result: the media is usable, but there is no transcript to read or caption. Do not report it as a transcription failure or retry it unless the user says the media contains speech that should have been detected.

Use `trigger_transcript` when the user explicitly wants transcription started or restarted. It is idempotent: only `idle` and `error` start work; ready, no-audio, and in-progress assets are left unchanged. If prepared transcription audio exists it is reused; otherwise the open Web/Desktop editor is asked to prepare and upload audio through its native pipeline. Then inspect readiness with `track_progress`.

For an `idle` or explicitly failed (`error`) run, call `trigger_transcript` with the asset id, then check transcription progress again. If an in-progress run has genuinely exceeded the stuck threshold, use `manage_transcript` action `retry_transcription` because `trigger_transcript` intentionally leaves in-progress work unchanged. The manage tool also remains available for low-level recovery where an external MCP host must supply prepared `audioBase64`. Repair source words with the transcript-fix action instead of rewriting visible captions when the source transcript itself is wrong.

For semantic speech edits, load `talking-head-guide` and use the Script workflow. Use mechanical cleanup only for fixed fillers and pauses; do not replace transcript-aware editing with destructive physical timeline cuts.
verification1.72 KB

View saved version →

---
name: verification
description: Hosted ChatCut plugin sessions only (the `chatcut` MCP server). If the conversation is driving ChatCut Desktop (a `chatcut_desktop*` MCP server), load this skill only when the user explicitly chooses the plugin/web surface — desktop sessions otherwise ship their own instructions and tools. Verify that ChatCut plugin edits are reflected in project structure and visible timeline output before reporting success.
---

# Verification

Use both structural and visual evidence when the requested result is visible:

1. Re-read the affected scope with `read_project`, `preview_timeline`, `inspect_item`, `browse_assets`, or `inspect_asset` as appropriate. Use the latest response for ids, tracks, timing, and readiness.
2. Inspect composed timeline frames through the visual tool exposed by the current manifest. Successful rendering or mutation metadata is not visual proof until the returned pixels are inspected.

Use local source bytes for raw source understanding when the original path is available. Use `inspect_asset` when the original is not locally available. Source frames help select moments but do not prove timeline composition; use composed timeline frames for trims, layers, captions, effects, crops, transitions, and Motion Graphics.

If signed frame URLs are returned and the host cannot inspect them directly, download the complete shell-quoted URLs into a temporary directory and inspect those files. Do not report visual success from URLs alone.

If visual proof is blocked by upload/readiness or renderer state, report the blocker explicitly and ask the user to inspect the live editor when appropriate. Never convert a failed verification into an unsupported tool fallback or a false completion claim.
video-gen22.3 KB

View saved version →

---
name: video-gen
description: AI video generation via Seedance, Kling, Gemini Omni, MiniMax H3, and MiniMax H3 Max. Use when the user wants to generate a video clip — text-to-video, image-to-video, first/last-frame transitions, reference-guided generation — or wants to modify / edit / extend an existing generated clip.
user-invocable: true
---

# Video Gen

Submits one video generation job per call and returns a `jobId`. Job status and later check-backs belong to `track_progress`; this skill does **not** place videos on the timeline automatically.

## When to Use

Any time the user wants to generate a video clip — text-to-video, image-to-video, first-last-frame transition, reference-based generation, or generatively editing / extending an existing video (producing new generated footage based on a source clip; not timeline trimming).

## Models

| Model            | Reference                                                    | Strengths                                                                                                                                                                                                            |
| ---------------- | ------------------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `seedance-2-5`   | [references/seedance25.md](references/seedance25.md)         | Default. 4-30s, up to 50 multimodal references, audio-only reference, timestamp control, and mp4/mov output. Strongest Seedance choice for new generation, edit, and extend.                                         |
| `seedance2`      | [references/seedance2.md](references/seedance2.md)           | Seedance 2.0 full model. 4-15s, rich multimodal references, and up to 1080p output.                                                                                                                                  |
| `seedance2fast`  | [references/seedance2.md](references/seedance2.md)           | Seedance 2.0 Fast — **same inputs and 4-15s range as `seedance2`, returns in minutes instead of ~ten, ~20% cheaper, capped at 720p**. Use when the user is iterating and turnaround matters more than peak fidelity. |
| `seedance2mini`  | [references/seedance2.md](references/seedance2.md)           | Seedance 2.0 mini — **same inputs and 4-15s range as `seedance2`, ~50% cheaper, capped at 720p**. Use for cheap drafts and high-volume batches where the full model's fidelity isn't required.                       |
| `kling`          | [references/kling.md](references/kling.md)                   | Camera-control language via prompt; strong on fine emotional / performance control for character shots. 1080p via `mode:"pro"`.                                                                                      |
| `minimax-h3`     | [references/minimax-h3.md](references/minimax-h3.md)         | MiniMax H3. 4-15s, 768p/2k, first/last frames, and up to 9 image + 3 video + 3 audio references (12 total). Strong multimodal reference following.                                                                   |
| `minimax-h3-max` | [references/minimax-h3-max.md](references/minimax-h3-max.md) | MiniMax H3 Max. 5-15s, 480p/768p, text or first/last-frame generation. No reference mode or 2K output.                                                                                                               |
| `omni`           | [references/omni.md](references/omni.md)                     | Gemini Omni 1.1 Flash. Targeted edits/extensions via `continueFrom`, first/last frames, and cheap 360p drafts. 3-10s; 720p default, upscaled 1080p/4k.                                                               |

**IMPORTANT:** Before generating, READ the chosen model's reference for capabilities, input channels, modes, prompt structure, and model-specific behavior.

## Model Selection

**For a new generation**, `seedance-2-5` is the **default**. Only switch when one of:

- **User explicitly named a model** ("用 Kling", "use Kling", "用 Omni", "用 mini", "/kling") — switch, no need to re-ask.
- **User explicitly asks for MiniMax H3 or needs combined image/video/audio reference guidance with optional 2K output** — use `minimax-h3`.
- **User explicitly asks for MiniMax H3 Max or wants its highest-fidelity text/frame generation without multimodal references** — use `minimax-h3-max`.
- **Seedance clearly can't or won't do it well** — when you hit a known case where Seedance struggles, propose switching to `kling` and confirm before submitting.
- **User explicitly asked for a fast / cheap draft they expect to revise** — propose `omni` (360p drafts, ≤10s, supports targeted edits) and confirm before submitting.
- **User needs Seedance output at 1080p** — prefer `seedance-2-5`; it and full `seedance2` support 1080p.
- **User wants Seedance's 2.0 inputs cheaply, or many clips at once, and is fine with 720p** — propose `seedance2mini` (4-15s, ~50% of the 2.0 full-model cost) and confirm before submitting.
- **User is iterating on 2.0-style output and cares about turnaround** — propose `seedance2fast` (480p or 720p, 4-15s, ~80% of the 2.0 full-model cost, returns much sooner) and confirm before submitting.

Otherwise stay on `seedance-2-5`.

**For revising a clip that already exists**, the choice is made in Step 4 — not here. A revision request is not a reason to change the default for future fresh generations. Omni 1.1 replaces the previous Omni choice; keep the `omni` alias and do not invent a second legacy choice.

Access note: all eight models require ChatCut paid video-generation entitlement (subscription or paid credits). Never present another model as a free workaround for a subscription gate; offer free Motion Graphic animation instead when the user asks for a free path.

## Tool Params

| Param             | Values                                                                                                         | Default                           |
| ----------------- | -------------------------------------------------------------------------------------------------------------- | --------------------------------- |
| `prompt`          | video description (required)                                                                                   | —                                 |
| `model`           | `seedance-2-5`, `seedance2`, `seedance2fast`, `seedance2mini`, `kling`, `minimax-h3`, `minimax-h3-max`, `omni` | seedance-2-5                      |
| `durationSeconds` | seconds                                                                                                        | 5                                 |
| `ratio`           | see model docs                                                                                                 | 16:9                              |
| `resolution`      | model-specific; H3: 768p/2k; H3 Max: 480p/768p; Seedance: 480p/720p/1080p; Omni: 360p/720p/1080p/4k            | 720p (H3 models: 768p)            |
| `outputFormat`    | `mp4`, `mov` (`seedance-2-5` only)                                                                             | mp4 (generate); mov (edit/extend) |
| `taskMode`        | `generate`, `edit`, `extend` (`seedance-2-5` or `omni`)                                                        | generate                          |
| `name`            | descriptive asset name (required)                                                                              | —                                 |
| `firstFrame`      | project asset ref                                                                                              | —                                 |
| `lastFrame`       | project asset ref                                                                                              | —                                 |
| `continueFrom`    | video asset ref — **`omni` only**, source to edit/extend                                                       | —                                 |

Model-specific params (e.g., Kling `mode`, Seedance `refImages` / `refVideos` / `refAudios`) — see the model's reference.

## Input Resolution

`firstFrame` / `lastFrame` / `refImages` / `refVideos` / `refAudios` all take a project asset reference. Prefer a full UUID or short prefix from `browse_assets`; `asset://<id>` and same-project asset URLs returned by asset tools are also accepted. Per-slot type: frame slots and `refImages` → image; `refVideos` → video; `refAudios` → audio.

External URLs and base64 are not accepted. If the source is a public URL, download it into the sandbox workspace, import the local file through `asset-import` + `push_asset`, and pass the resulting asset id.

## Workflow

Four-step loop. For each new generation, restart from Step 1 if the user's intent has shifted.

### Step 1 — Align scope with the user

Before writing any prompt, align on three dimensions:

1. **Duration & segments** — total length, how many shots, and whether they live in one clip or several.

   If the user has already stated a direction ("做一段", "in one video", "分别生成", "split into N shots", etc.), follow it — don't second-guess.

   Otherwise, surface the two paths and let the user pick:
   - **Multi-shot within one clip** (see model ref) — single inference, subject / lighting / style physically consistent across sub-shots; fits a coherent narrative within the per-clip duration cap.
   - **Multiple clips** — each clip is independently controllable and re-rollable, but identity and style continuity have to be carried by anchors; fits durations beyond the cap or hard scene breaks.

   Offer the trade-off; do not pick for the user.

2. **Content** — what each clip depicts. Summarize back what you understood, segment by segment. When content is vague (e.g. "generate a video of a girl dancing"), the user typically hasn't specified one or more of:
   - **Subject**: who / what is the main subject (appearance, outfit, defining features)?
   - **Action**: what are they doing? (For talking / emotional shots, what micro-expression?)
   - **Scene**: where — setting, time of day, environmental details?
   - **Lighting / color mood**: what atmosphere?
   - **Camera**: any shot-size / angle / movement preference?
   - **Style**: visual style or reference (cinematic / anime / documentary / ...).

   Focus on the items that matter for this specific request and can't be safely inferred — don't turn this into a blank-filling exercise. Summarize the understood parts back to the user before proceeding.

3. **Consistency anchors** — only when multiple shots reuse a character, object, or scene: identify which anchor (reference image or video) to pin across shots. For sourcing rules, see §Visual consistency across shots below.

For each dimension, check the user's words:

- **Clear** — proceed.
- **Ambiguous or missing** — ASK the user. Do not guess, do not default to your own interpretation. A round-trip confirmation is cheaper than a wasted generation.

#### What NOT to do

- **Do not "tell then submit"** — announcing "I'll make this as 2 clips" and immediately submitting is not alignment, it's a unilateral decision with announcement.
- **Do not default to splitting a single-video request into multiple clips.** A single clip can carry multiple sub-shots (see model ref), with subject / lighting / style physically consistent across them. Surface the trade-off, then let the user choose.
- **Do not skip the ask** because you think the answer is obvious.

#### Hard overrides (user's explicit word wins)

- "one clip / single clip / 一条 / 一个镜头 / in 1 clip" → never split, even if the description is objectively long.
- "N shots / N 段 / N 个镜头" → generate exactly N.
- "use this image / 用这张图" → use as reference, don't substitute.

### Step 2 — Write the prompt

See the chosen model's reference for prompt structure and param combinations (e.g., Seedance's 8-element structure and modes; Kling's prompt tips). Before submitting, check:

- `name` is a **descriptive** asset name — descriptive enough for the user (and you in later turns) to recognize this asset in the project library. Avoid vague names like "Untitled" or "clip 1".
- Param combination matches the user's intent — see the **Modes** section in the model's reference.
- Generated video audio is not a tool parameter. Seedance and Kling are submitted with audio enabled by the backend.
- On validation failure, read the error and fix the inputs — **do not blindly retry the same invalid arguments**.

### Step 3 — Submit one, wait, confirm

**Submit one generation job at a time.** Unless the user explicitly asked for multiple clips in parallel, do not submit the next clip until the current one completes and the user has reviewed it. Parallel submission hides problems: if the first shot has drift or wrong framing, the user would rather redo it once than have several misaligned shots to discard.

- `submit_video.ratio` controls the generated asset only; it does not change the project timeline canvas. If the user requested a final output aspect ratio (for example "9:16 vertical" or "16:9 landscape"), set the timeline canvas to the same ratio with `manage_timelines` action=update (e.g. ratio:"9:16") before placing the completed asset. If the user asked for no black bars / full-bleed, pass `fit:"cover"` when setting the canvas or updating/adding the visual item.
- Do not use this skill for job management — use `track_progress` for an immediate status read and follow its later check-back guidance.
- After submitting, end your turn (tell the user the job was created) unless a follow-up task is already queued.
- When the job finishes, surface the result to the user for review before proceeding to the next shot.
- For Omni source edits, follow the first/middle/end frame review in its reference before claiming the requested change succeeded.
- Model-specific failure handling — see the model's reference.

### Step 4 — Iterate

When the user wants a next clip, a revision, or a continuation, first decide what the existing clip **is** in the next call:

| The existing clip is…                                                                          | Path                                                         |
| ---------------------------------------------------------------------------------------------- | ------------------------------------------------------------ |
| **The thing being modified** — a targeted change; preservation is not pixel-exact              | `omni` + `continueFrom`                                      |
| **A source to re-generate from** — new motion / camera, whole-clip style shift, extend, bridge | `seedance-2-5` + `refVideos`; set `taskMode` for edit/extend |
| **Not reusable** — the intent changed                                                          | regenerate, restart at Step 1                                |

**A localized change defaults to `omni` + `continueFrom`.** For an explicitly chosen Omni continuation use `taskMode:"extend"` + `continueFrom`; for transitions between two images use `firstFrame` + `lastFrame`. Keep Seedance 2.5 as the general edit/extend path otherwise. Do not wait for the user to say "keep everything else" — "把气球改成黄色" / "remove the text in the corner" already means it. Route to `seedance-2-5` only when the change genuinely needs re-inference (motion, camera, duration, whole-clip style), not because Seedance is more familiar.

`omni` source edits/extensions require a project video of up to 10 seconds. For longer sources use `seedance-2-5 + refVideos` or ask to trim first. This integration is stateless: do not promise the upstream 40-second multi-turn workflow.

**No generative edit is lossless.** Omni 1.1 supports upscaled 1080p/4k, but higher resolution does not guarantee detail or unchanged pixels. Review the result before replacing a timeline item; the original asset remains available.

Then apply the standing rules:

- **If it's the next shot in a multi-shot sequence** — reuse the established anchor (see §Visual consistency across shots below for principles, model ref for flag-level details).
- **If the user's feedback is ambiguous** ("it doesn't feel right") — ask what specifically to change before regenerating.
- **If the same text-prompt adjustment has failed twice** — stop adjusting text. Switch to reference images, or switch to an edit path (`omni` `continueFrom` for targeted changes, seedance `refVideos` for re-generation).
- Each new generation restarts the loop at Step 1 — realign if scope shifted.
- `submit_video` reports the omni edit round; when it warns about chain depth, relay the suggestion to the user instead of silently continuing.

## Visual consistency across shots

Text alone cannot reliably maintain visual identity across shots; visual references constrain output far more precisely than words.

### Anchors: the cornerstone of consistency

An **anchor** is a reference image or video pinned across every shot that shares the same character, object, or style. Any multi-shot sequence with recurring visual elements needs an anchor — don't try to reproduce them from text.

### Sourcing an anchor

Have reference awareness. When the user's request involves a recurring character / object / scene, think about what anchor to use **before** writing prompts:

- **Check the project first.** What has the user already provided or approved? Uploaded images, previously generated and approved shots, or earlier project assets can all serve as anchors.
- **Match the user's intent.** If the user pointed to a specific asset ("use this photo", "像上一段那样"), use that. If they described a character only in words, no anchor exists yet and one must be established.
- **When in doubt, ask the user.** Don't guess which asset to pin, and don't silently generate a new anchor when the user may already have one in mind.

### Establishing a new anchor (with user consent)

When no existing asset fits and one must be generated, propose it to the user first — it costs credits and shapes every downstream shot. Model-specific paths — see the chosen model's ref.

### Using the anchor

- Pass the anchor in **every shot** that shares the character / object / style. The specific flag(s) to use depend on the model — see the model's ref.
- Describe the anchor by appearance in the prompt, not by name: "The BLACK RACING CAR with chrome exhaust" constrains far more than "Fleetmaster". When role confusion is likely, add explicit negations: "The motorcycle does NOT transform."
- Refer to the anchor with `@Image1` / `@Video1` in the prompt — not vague phrases like "the same car as before".
- When a shot depends on a previous generation, check it with `track_progress` in a later turn until it is terminal, then use its `outputAssetId` as the anchor reference. `track_progress` is an immediate status read; `action=wait` is only a compatibility alias and does not block. Do not submit dependent shots in parallel.

### Multi-character projects

When a project has multiple named characters with distinct attributes (e.g. Faz with fire energy, Kev with ice energy), treat each character as a **separate anchor** — one reference asset per character. In every prompt:

- Name the **active** character and attach their distinctive attributes ("Kev has **blue ice** electric energy").
- Add explicit negations for the others to prevent attribute leakage ("NOT red fire energy, NOT Faz's look").
- Pin the correct character's anchor (model-specific flag — see model ref). Do not reuse another character's anchor by accident.

Missing either explicit attribution or negation causes cross-character attribute mixing.

**Multiple characters in the same frame.** For shots where multiple characters appear together (especially facing the camera), the model is prone to face-swap or body-clipping. Add **strong positional + outfit anchors** to each character and prefer a **fixed camera** for that shot:

- "the character on the LEFT wears a grey-blue tactical jacket, short beard, silver earring"
- "the character on the RIGHT wears a red cape with gold trim, long braided hair"
- "fixed camera, medium shot, both characters clearly separated"

Positional words (left / right / foreground / background) + distinctive outfit colors give the model enough signal to keep the characters apart.

### Escalate when text adjustments fail

If a visual-identity issue (wrong character, drift, color mismatch) persists after **two text-prompt adjustments** on the same shot, stop adjusting text. Text is not a substitute for an anchor. Escalate to:

- Adding or switching the anchor.
- Edit mode where the model supports it (see model ref for how to invoke).

Do **not** submit a third text-only retry on the same consistency issue.

### When to skip anchoring

Simple, one-off, or exploratory requests do not need anchors — generate directly.

## Run

```ts
// Text-to-video (Seedance 2.5 default)
submit_video({
  model: "seedance-2-5",
  prompt: "A cat walks across a sunny windowsill",
  name: "Cat on windowsill",
});

// Image-to-video with Seedance 2.5 — pass the project asset id directly
submit_video({
  model: "seedance-2-5",
  prompt: "The scene comes to life, gentle breeze rustles the curtains",
  firstFrame: "abc12345",
  name: "Living room animation",
});

// Kling text-to-video — only after Model Selection check
submit_video({
  model: "kling",
  prompt: "A sports car drifts around a wet corner",
  name: "Car drift shot",
});
```

After submission, return the `jobId` and end the turn by default. Call `track_progress` only when the user later asks for status or a dependent step needs the completed asset. Use `action=status`; it returns immediately, so follow its later check-back guidance and do not loop or expect `action=wait` to block.

## Config Mode

For complex multimodal jobs, build the full args object up front and pass it in a single call:

```ts
submit_video({
  model: "seedance-2-5",
  prompt: "...",
  name: "...",
  refImages: ["def67890", "ghi24680"],
  refVideos: ["abc99999"],
  refAudios: ["jkl55555"],
  durationSeconds: 8,
  ratio: "9:16",
});
```

## Rules

- Always provide `name` with a descriptive asset name.
- Default to submit-only. End your turn after submitting unless a follow-up task is queued.
- Do not try to manage jobs through `submit_video` — status and later check-backs belong to `track_progress`.
- Generation costs credits. Before submitting, briefly tell the user what you're about to generate.

Referenced files: 6

video-translation12.6 KB

View saved version →

---
name: video-translation
description: Translate, dub, and localize speech in an existing video while optionally preserving speaker voices, generating translated captions, or synchronizing the speaker's lip movements. Use when the user asks for 视频译制、多语言配音、 把视频里的中文变成英文、让视频里的人说另一种语言、保留原音色、翻译声音、 口型同步、translated video, video dubbing, voice translation, or lip-synced localization, as well as traducción de video, doblaje de video, traducir un video, or sincronización labial. This Skill MUST be loaded before composing treatment choices for an ambiguous video-translation request in any language, including a turn that only asks the user to choose a treatment. Do not use when the user only wants subtitles translated, wants an SRT/VTT file, wants text or a script translated before generating a new avatar video, or wants ordinary TTS with a newly selected voice.
user-invocable: true
---

# Video Translation

Create a new localized video asset from an existing project video. Treat
speech translation, dubbing, voice preservation, lip synchronization, captions,
quality, authorization, generation, and verification as one workflow.

Lip-synced translation starts with source-transcript review in an editable
ChatCut Widget. The recognized source wording is never treated as final until
the user has edited or approved that form and submitted it. Audio-only
translation keeps the existing direct-submit flow and skips this review.

## When to Use

Use this skill when the user wants an existing video to:

- Make its speakers speak another language.
- Produce a dubbed or localized version.
- Preserve the original speakers' vocal identity across languages.
- Synchronize visible mouth movements with translated speech.
- Translate only the spoken audio while keeping the original picture.
- Generate a translated video with optional translated captions.
- Create a language-specific version for another market.

Typical triggering requests include:

- "让视频里这个人说英语。"
- "把中文口播做成英文版,口型也对上。"
- "把这段采访译制成日语,保留每个人的声音。"
- "做一个西班牙语配音版。"
- "只把声音翻译成英语,画面不要变。"
- "把这个中文数字人成片改成英文。"

Route adjacent requests elsewhere:

- Only translate, add, edit, or export captions: use caption translation.
- Translate a script before generating a new avatar video: translate the text,
  then use the Digital Human Skill.
- Generate speech with a chosen replacement voice: use the Voice Skill.
- Transcribe spoken content without changing it: use the Transcription Skill.
- Edit pauses, mistakes, framing, or B-roll in talking-head footage: use the
  Talking Head Guide.
- Generate new footage rather than localize an existing video: use Video Gen.

Treat these requests as ambiguous:

- "把这个视频翻译成英文。"
- "帮我做一个英文版。"
- "中文改成英语。"
- "做一个海外版。"
- "把这个视频国际化一下。"

For an ambiguous request, ask one treatment question with text choices. Keep
the wording natural for the conversation, and include all applicable paths:

- Lip-sync translation: continue through the editable source-transcript review
  and submit with `audioOnly: false` or omit `audioOnly`.
- Audio-only translation: keep the existing direct-submit path and submit with
  `audioOnly: true`.
- Subtitle-only translation: use caption translation and do not submit a video
  translation job.

Keep user-facing terminology in one language. These are localized concept names,
not fixed full option sentences; descriptions may stay natural for the context:

| Conversation language | Lip-sync path                        | Audio-only path          | Subtitle-only path            |
| --------------------- | ------------------------------------ | ------------------------ | ----------------------------- |
| Chinese               | 口型同步翻译                         | 只翻译声音               | 只翻译字幕                    |
| English               | Lip-synced translation               | Audio-only translation   | Subtitle-only translation     |
| Spanish               | Traducción con sincronización labial | Traducción solo de audio | Traducción solo de subtítulos |

The target translation language does not control this copy; the user's
conversation language does. Never append a second-language gloss such as
`(lip-sync)` to a localized label.

Do not omit lip-sync when the source has a visible speaker. Do not relabel
audio-only translation as generic "AI voice replacement"; choosing a new
synthetic voice is a separate Voice Skill workflow. If the user already stated
one treatment, skip this question and follow that route directly.

Do not submit a paid video-translation job until changing the spoken audio is
explicitly requested or confirmed.

## Workflow

1. **Resolve the source video.** It must be an imported project video asset;
   use the exact asset id from the attachment, selection, or `browse_assets`.
   Local or external media must be imported first.
2. **Resolve the target language from the live catalog.** Call
   `submit_video_translation` with `action: "list_languages"` and no other
   arguments. This only reads the catalog; it requires no project, transcript
   review, confirmation, or credits. Then copy the matching returned value
   verbatim into `targetLanguage`, including any dialect, region,
   or script qualifier. Do not invent a language name or send an ISO code.
   Match the user's requested language and variant against the returned catalog;
   confirm the language when the user only implied a market ("海外版"). If a
   submission reports an unsupported language, use its suggested candidates or
   refresh the catalog with `action: "list_languages"` instead of retrying
   guessed names. The `action` argument is required on every call.
3. **Pick the treatment:**
   - Lip-sync: `mode: "speed"` — translated dubbing that preserves
     the speakers' vocal identity, plus synchronized mouth movements. Continue
     through the source-transcript review below.
   - `audioOnly: true` — translate the audio only and keep the picture
     untouched. Use for screen recordings, voice-over footage, or when the user
     says the picture must not change. Prefer this for transparent WebM when
     preserving the alpha channel matters. **Skip steps 4–6 and keep the existing
     direct-submit flow; do not require or pass `reviewedSourceTranscript`.**
4. **Prepare the source transcript for lip-sync only.** If word-level
   transcription is not complete, call `trigger_transcript`, wait for it with
   `track_progress`, and retry only when it is ready. Call `read_script`, then
   read the matching `library/<filename>.md` source transcript. Do not use the
   editable `timeline.md` cut as the translation source. **Do not substitute
   `inspect_asset` transcript ranges for `read_script`; the range result may be
   partial and is not the canonical complete source transcript.**
5. **Render the editable source-text review for lip-sync only.** Load the
   `widget-forms` Skill and reuse the same `<form-textarea>` confirmation pattern
   used by the Digital Human Skill. Strip the library's `[sN]` addresses and
   speaker-rendering rows, but keep **exactly one recognized transcript segment
   per line**, in source order. Put that complete text in the textarea's
   `default`; use a localized label that asks the user to check and correct the
   recognized original text. Do not expose segment ids, word indices, file
   syntax, or timestamps.

   Hard preflight before emitting the Widget:
   - The field id is exactly `reviewedSourceTranscript`.
   - The prefill attribute is exactly `default`, never `default-value`,
     `defaultValue`, `value`, or `placeholder`.
   - `default` contains the actual complete non-empty recognized transcript,
     not a placeholder or an omitted value. If the transcript text has not been
     read successfully, do not render the Widget; read the canonical library
     document first.
   - Preserve one recognized source segment per line. Verify that the first and
     last non-empty source segments are both present before sending the form.
   - In the embedded raw-tag route, apply the `widget-forms` XML attribute
     escaping rule to the complete transcript before placing it in `default`.
     Never put unescaped recognized text inside the tag.

   Embedded ChatCut example (localize visible copy):

   ```text
   Please review the recognized source text and correct any mistakes:

   <widget>
     <form-textarea id="reviewedSourceTranscript" label="Please review and correct the source text" rows="12" required="true" default="<recognized source text; one segment per line>"/>
   </widget>
   ```

   Follow `widget-forms` for the active host rather than emitting raw tags in a
   host that does not support the embedded protocol. Do not add a separate
   yes/no question or a handwritten submit button. **Stop here and wait. Never
   submit a translation in the same turn that first presents the Widget.**

6. **Use the submitted revision exactly.** The Widget submission is explicit
   confirmation. Read the complete `reviewedSourceTranscript` textarea answer
   from the user's next message and preserve its wording, punctuation, and
   order exactly. Pass through the returned line breaks when present, but do
   not reject or rewrite normal line-break edits made inside the textarea; the
   tool realigns them to the source timing. Do not paraphrase the text and do
   not ask the user to confirm the same text again. If the user replies outside
   the Widget with further corrections, reopen the same editable Widget with
   those corrections applied rather than reverting to a prose transcript.
7. **Duration behavior.** By default the output may run slightly longer or
   shorter than the source so the translated speech keeps a natural pace. Pass
   `keepDuration: true` only when the user needs the exact original length
   (for example to swap it into an existing timeline slot), and mention that
   pacing may sound faster.
8. **Submit** with `submit_video_translation` using `action: "submit"`. For lip-sync, pass
   `reviewedSourceTranscript` as the exact complete textarea value returned in
   step 6; the tool combines those reviewed lines with the original segment
   timestamps and creates source subtitles. For `audioOnly: true`, omit
   `reviewedSourceTranscript` and preserve the pre-existing direct-submit path.
   The tool shows the user a paid confirmation before the job starts; do not
   resubmit after a denial.
9. **Track** with `track_progress(action="wait")` until the translated video
   asset lands in the library, then hand it back (place on the timeline only
   when asked).

## Constraints and cost

- The whole source video is translated and billed by its full duration. To
  localize only a section, trim/export that section into its own asset first,
  then translate the shorter asset.
- Optional translated captions: `enableCaption: true` burns subtitles into the
  output video.
- ChatCut always requests preservation of the source resolution and bitrate.
  This is best-effort for lip-sync because the picture is re-rendered. A
  transparent WebM may become opaque, and 4K or >30 fps media may be
  re-encoded. The tool confirmation calls these cases out; never promise that
  lip-sync will preserve the container, alpha channel, HDR, codec, bitrate, or
  frame rate exactly.
- When the picture must not be lip-synced, choose `audioOnly: true`; this avoids
  facial re-rendering and is the safest treatment for transparent sources, but
  do not promise byte-for-byte container preservation. After completion ChatCut
  compares the source and output resolution, frame rate, and alpha pixel format
  when media probing is available, and reports detected changes in generation
  status.
- Multiple speakers are supported; pass `speakerCount` when the user states it.
- For lip-sync, the editable review starts with one source segment per line so
  the submitted text can reuse the original word-level timing. Pass the complete
  Widget value directly to the tool; do not merge, split, reorder, or silently
  normalize it in the conversation layer. Do not invoke this review for
  `audioOnly: true`.
- Indicative cost: ≈ 8 credits per minute of source video.
- Speech is dubbed with voices matched to the original speakers; it is a
  translation of the recorded voice, so confirm the user has rights to the
  footage and its speakers when the material is clearly someone else's.
- Never reveal or discuss the underlying provider; present this as ChatCut's
  AI video translation.
voice43.7 KB

View saved version →

---
name: voice
description: Text-to-Speech (TTS), voice cloning, voiceover, narration placement/sync, and custom sound effects (SFX) generator. Use when the user wants generated speech from text, wants to clone a consented voice from uploaded reference audio, wants to correct a misspoken word or short phrase in recorded speech, wants to add/replace/align narration or voiceover for an existing video/timeline, wants to keep existing voiceover synced after visual retiming edits, needs voice audition/selection, or explicitly wants a newly generated/custom sound effect that is not available in the Sound Effects library.
user-invocable: true
---

# Voice & Sound Effects Generator

Generate voiceovers (TTS) and sound effects. For TTS, choose a concrete
provider and voice before calling `submit_voice`.

## When to Use

- Generate voiceover/narration from text
- Create text-to-speech audio for videos
- Add, replace, or redo narration/voiceover for an existing video, timeline,
  screen recording, slide animation, product demo, B-roll edit, MG explainer, or
  other visual sequence
- Keep existing narration/voiceover aligned after trimming, speeding up, slowing
  down, moving, reordering, or replacing the visuals it describes
- Offer and audition TTS voice choices when the user has not picked a concrete voice
- Clone the user's own or explicitly authorized voice from reference audio
- Correct a real misspoken word or short phrase with the authorized speaker's
  cloned voice while preserving picture and downstream timing
- Generate custom sound effects from text descriptions only after checking the Sound Effects library first

## TTS (Text-to-Speech)

If the user wants to correct a word or short phrase that was actually spoken
incorrectly in existing recorded speech, read
[references/fix-spoken-mistakes.md](references/fix-spoken-mistakes.md) before
editing or generating. That workflow preserves the picture and protects later
timing while replacing only the faulty sound. Do not use synthesized speech
when the recording is already correct and only its transcript is wrong; fix
the transcript instead. Use the ordinary narration path for a full rewrite or
complete new voiceover.

If the current request has an existing visual target and the user wants
narration, voiceover, dubbing, or replacement speech for that target, read
[references/video-sync.md](references/video-sync.md) before drafting new
narration, using existing narration text to generate TTS, or placing audio. Do
this even when the user did not explicitly say "sync" or "match the visuals";
the existence of a visual target means narration timing and meaning may need to
follow on-screen content. Use the normal standalone TTS path only when there is
no visual target or the user just wants an audio asset from text.

Also read [references/video-sync.md](references/video-sync.md) when the timeline
already has narration/voiceover and the user asks to change the visuals while
keeping that voiceover aligned. This is a sync maintenance task even if no new
TTS is needed.

Use `manage_voice` as the single catalog and management entry, and
`submit_voice` for final TTS. `manage_voice action="list" voiceType="all"`
returns official and cloned candidates in one normalized list. Only
`voiceType="custom"` supports create-access, apply, clone, preview, rename, or
delete actions. The current contracts are:

- `provider` is required. Use `doubao` for Chinese-optimized narration and
  `elevenlabs` for English or multilingual narration, or `fish-audio` for a
  ready ChatCut custom voice.
- `voiceId` is required and provider-specific. Do not mix catalogs.
- `submit_voice` creates an audio asset only. Timeline placement, replacement,
  trimming, and alignment happen later with timeline tools.
- For long narration, multiple `submit_voice` calls can be useful: split at
  natural pauses, sentence groups, or script beat boundaries when the workflow
  benefits from separately timed or placed voice clips, such as storyboard beats,
  scene-level ad segments, or a user request for separate assets.
- For Doubao, `speedRatio`, `loudnessRatio`, `pitch`, `emotion`,
  `emotionScale`, `performancePrompt`, and `explicitDialect` are supported
  knobs, but not every voice supports every expressive control. Use the
  selected official entry returned by `manage_voice action="list"` as the
  authority.
- For ElevenLabs, `modelId`, `speed`, and `stability` are the supported voice
  knobs. For `eleven_v3`, inline audio tags are available for expressive
  delivery such as emotion, tone, nonverbal cues, accent hints, pauses, or
  local pacing.
- For `fish-audio`, pass the ready voice's ChatCut `customVoiceId` as
  `voiceId`. `submit_voice` performs the authoritative Pro/readiness check,
  creates a generation job, and consumes standard TTS credits after success.
  Never pass or request the provider's internal Fish model id. The current
  Fish S2.1 model supports inline square-bracket cues in `text` for local
  emotion, delivery, and paralinguistic control.

### Keep voice providers private

Doubao, ElevenLabs, Fish Audio, their model names, and other provider identity
are internal implementation details. Never expose or attribute them in
user-facing replies, progress updates, voice recommendations, audition cards,
clone instructions, success summaries, or errors. Use ChatCut product language
instead:

- Call curated catalog voices `official voices` / `官方音色`.
- Call saved custom voices `cloned voices` / `克隆音色`, `My Voices` /
  `我的音色`, or use the voice's user-visible saved name.
- Describe status and failures at the ChatCut feature level. If a tool or
  provider error contains a provider name, model name, provider voice id, or
  provider URL, preserve the actionable meaning but remove those details
  before replying.

Provider names and provider-specific ids remain valid only in internal tool
arguments, tool-result interpretation, and these implementation instructions.
Do not copy them from tool output into visible UI metadata or prose.

Official voice controls come from `manage_voice action="list"`:

- Call it for the target narration language and the user's visible locale.
  Treat each returned `(provider, voiceId, name, summary, sampleUrl,
capabilities)` tuple as atomic.
- Use only fields in that entry's `capabilities.supportedControls`; never infer
  capability support from a remembered voice name or family.
- Controls are not guarantees of a specific acting style. Use the returned
  tags and sample to pick a naturally suitable voice, then use supported
  controls for moderate delivery changes.
- For entries that advertise `audioTags`, inline audio tags are available when
  the user asks for expressive delivery such as emotion, tone, nonverbal cues,
  accent hints, or local pacing. Official examples fit these useful TTS
  categories:
- Emotion/tone tags include `[happy]`, `[sad]`, `[angry]`, `[excited]`,
  `[curious]`, `[sarcastic]`, `[crying]`, `[annoyed]`, `[appalled]`,
  `[thoughtful]`, `[surprised]`, and `[mischievously]`; vocal delivery and
  nonverbal cue tags such as `[whispers]`, `[laughs]`, `[sighs]`, `[exhales]`,
  `[inhales deeply]`, `[clears throat]`, `[snorts]`, `[swallows]`,
  `[wheezing]`, and `[coughs]`;
  pacing/pause/local speed tags such as `[slowly]`, `[pause]`,
  `[short pause]`, `[long pause]`, `[rushed]`, and `[drawn out]`; and
  accent/special-performance tags such as
  `[strong X accent]`, for example `[strong French accent]`, plus `[sings]`,
  `[singing]`, `[woo]`, and `[pirate voice]`. Official examples are
  non-exhaustive; similar auditory tags can be tried when the user explicitly
  asks for that delivery and the tag describes how the voice should sound, not
  a visual action. Write tags directly in `text`, close to the short phrase
  they should affect. Treat tags as local guidance, not paragraph-wide controls.
- For pauses and pacing, use punctuation, text structure,
  shorter generated segments, or local audio tags such as `[short pause]` and
  `[slowly]` when needed.

Fish Audio control support for cloned voices:

- The current `fish-audio` route uses Fish S2.1. Put concise natural-language
  cues in square brackets directly in `submit_voice.text`, for example
  `[happy]`, `[calm]`, `[angry]`, `[excited]`, `[whisper]`, `[laugh]`,
  `[sigh]`, `[gasp]`, `[pause]`, `[emphasis]`, `[inhale]`, or `[exhale]`.
  S2.1 is not limited to a fixed tag list, so a specific auditory description
  such as `[whispers sweetly]` or `[laughing nervously]` is also valid.
- Place a cue immediately before the phrase or moment it should affect. Fish
  S2.1 accepts cues anywhere in the text, so multiple short cues can create
  local transitions, for example
  `[calm] 先别着急。[excited] 好消息是,我们已经找到解决办法了!` or
  `I thought it was over [gasp] but then the lights came back [relieved].`
- Use cues sparingly and only when the requested delivery benefits from them.
  Prefer one clear instruction at a transition over stacking conflicting
  directions. Treat the result as model guidance rather than a deterministic
  editing boundary, and split into separate `submit_voice` calls when exact
  clip-level timing or independent retries matter.
- Do not use Fish S1's legacy `(parenthesis)` emotion syntax on this route, and
  do not add a separate `emotion` argument: the S2.1 cue belongs inside
  `text`. Keep the default voice-cloning preview line untagged. Do not add cues
  to a custom preview on the user's behalf, but preserve cues the user
  intentionally includes within the 100-character preview limit.

```ts
// English / multilingual via ElevenLabs
mcp__skill__submit_voice({
  provider: "elevenlabs",
  text: "Hello world",
  voiceId: "peter",
});

// Chinese via Doubao
mcp__skill__submit_voice({
  provider: "doubao",
  text: "你好世界",
  voiceId: "liuchang",
});

// With speed adjustment (Doubao only)
mcp__skill__submit_voice({
  provider: "doubao",
  text: "这是一段稍快的中文旁白。",
  voiceId: "liuchang",
  speedRatio: 1.5,
});

// With expressive Doubao controls
mcp__skill__submit_voice({
  provider: "doubao",
  text: "这次事故提醒我们,安全永远不能侥幸。",
  voiceId: "liuchang",
  emotion: "sad",
  emotionScale: 3,
  performancePrompt: "痛心但克制,语速稍慢,像新闻专题旁白",
  pitch: -1,
  speedRatio: 0.92,
});

// With ElevenLabs delivery controls
mcp__skill__submit_voice({
  provider: "elevenlabs",
  text: "The launch changed how teams plan their daily work.",
  voiceId: "peter",
  speed: 0.95,
  stability: 0.4,
});

// With local Fish Audio S2.1 delivery cues on a ready cloned voice
mcp__skill__submit_voice({
  provider: "fish-audio",
  voiceId: "<confirmed custom voice id>",
  text: "[calm] 先别着急。[excited] 好消息是,我们已经找到解决办法了!",
  name: "Expressive custom voiceover",
});
```

## Voice Audition Before Generation

### Emit the native clone entry only as a final action

`<clone-voice/>` is an executable editor action, not prose, code, or an
internal process label. Never quote it, wrap it in backticks, describe it as
the next step, or emit it in a progress update, pre-tool explanation, plan, or
other intermediate assistant message.

Finish every required inspection and tool call first. If the native dialog is
still the correct route afterward, follow the **dialog-entry sequence** under
**Native ChatCut editor** below. That sequence must be the final user-facing
assistant message for the run: do not call another tool, add anything after
its completion reminder, or emit the tag again. A single Agent run may render
at most one standalone native clone entry.

### Route cloning by host and available reference

Decide the host path before loading `widget-forms` or asking for clone inputs:

- In the native ChatCut editor, first check whether the user has already
  supplied a readable audio attachment or explicitly identified an accessible
  audio-bearing ChatCut asset for this cloning request, including a timeline
  audio or video item. If so, do not make the user choose the same source again
  and do not emit `<clone-voice/>`. Resolve or derive its audio asset, run the
  creation preflight, collect only the missing name, preview text, and explicit
  authorization, then use the Agent-driven cloning flow below.
- In the native ChatCut editor, use the editor dialog only when a usable
  reference has not already been supplied. Follow the native dialog path under
  **Custom Voice Cloning** below.
- In external Codex / Claude hosts, never emit `<clone-voice/>`; those hosts do
  not render or dispatch the native editor action. Use the Agent-driven
  attachment flow below.

An arbitrary voice already present in the project is not consent or a cloning
reference. Use the direct path only when the user supplied or identified the
audio for the current cloning request and later gives the full authorization
required below.

When the user explicitly chooses an audio-bearing timeline item, resolve it and
its source range with `preview_timeline`. Reuse a 10-second-to-3-minute audio
asset directly. For video, `pull_asset`, extract the chosen range as supported
audio with ffmpeg, then load `asset-import` and `push_asset` the result. For any
source over 3 minutes, use a clean 30–60-second excerpt; under 10 seconds, ask
for a longer reference. Continue with the resulting audio asset id without
emitting `<clone-voice/>`; the source choice is not authorization.

### Choose the interaction from live state

Treat the unified voice lookup as a mandatory gate. Whenever the user requests
TTS without naming a concrete official or custom voice, and has not already
explicitly chosen voice cloning:

1. If `manage_voice` is not already loaded, use `ToolSearch` to load it.
2. Call `manage_voice action="list" voiceType="all"` with the target narration
   language, conversation locale, and any useful official-voice filters.
3. Use the returned normalized `voices` list as the only candidate directory.
   It already puts cloned voices before matching official voices.

Do this before recommending or rendering any voice-selection UI. Never infer
the user's saved or official voices from chat history or a static Skill file.

An explicit request to clone a voice is already a concrete source choice. Do
not call the combined list merely to choose between official and cloned voices
in that case. First resolve whether the current request already supplies a
usable reference; perform `manage_voice voiceType="custom"
action="check-create-access"`, then follow the selected host route. The
standalone native entry may appear only after that work is complete, under the
final-action discipline above.

The Agent may infer useful tone, mood, delivery, or use-case suggestions from
the narration text. Treat these as recommendation signals, not as the user's
choice of voice source. A reasonable inference must never exclude a ready,
playable custom voice from the combined audition surface below.

Classify the request before rendering anything:

- **Concrete custom voice:** when the user names a ready custom voice or chooses
  the only named custom-voice pill, use that exact custom voice. When the user
  chooses `Choose my cloned voice` and several are ready, show only those ready
  custom voices as playable choices before continuing.
- **Concrete official preset:** use or confirm that exact preset; do not reopen
  the source-choice branch.
- **Clone action selected:** only now enter the host-specific cloning route
  above in the very next reply. Do not wait for the user to ask again, tell
  them to open a menu or panel, or leave the choice as an unhandled pill. In
  the native editor, finish any required inspection or preflight before the
  dialog entry; in external Codex / Claude hosts, begin the attachment flow.
- **At least one ready custom voice, but no concrete voice named:** render one
  combined playable audition surface. Put ready custom voices that have a real
  `previewUrl` first, add 2-4 suitable official voices with playable samples,
  and finish with `Clone another voice`. Do not ask a text-only saved-versus-
  official source question and do not use `<choices/>`.
- **No ready custom voice and no voice preference:** infer a small, varied set
  of 2-4 generally suitable official voices, render them as playable cards,
  and finish with `Clone my voice`. Do not use text-only source choices.
- **No ready custom voice and voice requirements available:** use explicit
  requirements such as "middle-aged male", "warm female", or "professional"
  first; otherwise reasonable traits inferred from the narration may guide the
  recommendation. Show 2-4 matching official playable cards and one clone
  action card in the same selection surface. Do not append a standalone
  `<clone-voice/>` button to this reply.

Localize the branch labels and adjust them to live state. Use these meanings:

- Chinese:
  - Named custom: `使用「<name>」`
  - Custom list: `选择我的克隆音色`
  - Curated catalog: `选择官方音色`
  - First clone: `克隆我的音色`
  - Additional clone: `克隆新音色`
- English:
  - Named custom: `Use “<name>”`
  - Custom list: `Choose my cloned voice`
  - Curated catalog: `Choose an official voice`
  - First clone: `Clone my voice`
  - Additional clone: `Clone another voice`
- Spanish:
  - Named custom: `Usar «<name>»`
  - Custom list: `Elegir mi voz clonada`
  - Curated catalog: `Elegir una voz oficial`
  - First clone: `Clonar mi voz`
  - Additional clone: `Clonar otra voz`

Native Chinese example when no usable reference has been supplied:

```text
<widget>
  <form-visual id="voiceId" label="请选择并试听音色" media-kind="audio" required="true">
    <visual-option value="<returned voiceId>" name="<returned localized name>" summary="<returned summary>" media="<returned sampleUrl>" media-kind="audio"/>
    <visual-option value="clone_voice" name="克隆我的音色"/>
  </form-visual>
</widget>
```

When no usable reference has been supplied, the dialog action belongs inside
the playable grid:

```html
<visual-option value="clone_voice" name="克隆我的音色" />
```

Always expose one appropriate path to cloning during ordinary voice selection,
but never show both a clone action card and the standalone native
`<clone-voice/>` entry in the same reply. Provider availability, plan, and
quota govern what happens after the user chooses cloning, not whether the
choice is visible. Offering the choice does not require consent; explicit
permission and a supported reference are required only before the clone tool
call. A scenario reference may intentionally use a narrower audition without
the clone action and explain its direct-from-project fallback in prose.

Before recommending, rendering, or submitting a TTS voice option, use the
current `manage_voice action="list" voiceType="all"` result. It is the live
catalog shared with ChatCut Web, Desktop, and published plugins. Do not create
voice options from memory, translated names, filenames, or broad descriptions.

First determine two separate languages:

- User conversation language: the language the user used to talk to you. Use
  this for surrounding copy, option names, and summaries.
- Target narration language: the language of the text being synthesized. Use
  this only to choose provider and voice catalog.

Load `widget-forms` before collecting input or rendering a choice. That skill
owns the current host's form, attachment, and media-card behavior. This skill
owns the voice candidates, required fields, safety rules, and the asset ids
passed to voice tools. Do not embed host-specific UI instructions here.

"help me generate ... voice over in Chinese" is an English conversation asking
for Chinese narration, so the audition widget copy stays in English while the
voice candidates come from Doubao.

For an official audition:

1. Call `manage_voice action="list" voiceType="all"` with the target narration
   `language`, user conversation `locale`, and any explicit gender, age, tone,
   or use-case requirements. If none were explicit, a concise inferred
   tone/use-case query may guide the official shortlist. Present inferred traits
   as a suggestion, not as a stated user preference. If nothing official
   matches, broaden only optional filters and clearly describe the closest
   supported choices; never drop returned playable custom voices because they
   lack official catalog tags.
2. Use returned ready custom voices with previews first, followed by 2-4
   returned official voices.
3. Load `widget-forms` and request one required `playable_single_choice`. Give
   every official option the `voiceId`, localized `name`, `summary`, and
   matching `sampleUrl` from the same returned tuple. Pass `sampleUrl`
   verbatim; never prepend an inferred S3, CDN, editor, localhost, or production
   base URL, and never reconstruct it from the filename pattern. This differs
   from a custom voice's `previewUrl`, which must be used exactly as returned by
   `manage_voice`.
4. Always add one no-media action option with stable value `clone_voice` as the
   final card on every ordinary recommendation/selection surface. Label it
   `Clone my voice` when no ready custom voice exists, or `Clone another voice`
   when one does; localize it using the meanings above. In native ChatCut this
   is a compact `<visual-option>` without `media`, not a separate button. In
   external hosts it is the equivalent label-only option. A scenario-specific
   reference may explicitly replace this ordinary surface with a narrower one.
5. Wait for the user to choose. If the answer maps to `clone_voice`, render the
   host-specific cloning entry/workflow on the next turn; do not start cloning
   from the audition reply itself.
6. For an official voice, call `submit_voice` with the selected tuple's
   `provider` and `voiceId` verbatim.
7. For an existing custom voice, call `manage_voice voiceType="custom"
action="apply"` with its full mapped `customVoiceId`. When final speech is
   requested, pass that same id to `submit_voice` as `voiceId` with
   `provider: "fish-audio"`.

For a custom-voice-only audition, include every relevant `ready` custom voice
that has a `previewUrl`. Use the full `customVoiceId` as its stable value, the
stored name as its display label, `previewText` as its summary when present,
and the exact tool-returned `previewUrl` as audio media. Do not copy, download,
rewrite, validate by hostname, or invent that URL. If a ready voice lacks a
preview, omit it from the audition surface rather than making an unplayable
text-only card.

Keep every option's type, submission id, display label, preview, summary, and
capabilities tied to the same `manage_voice` tuple. The target narration
language filters official entries; the conversation language controls all
visible copy. Keep the label-to-tuple map in context so a host that returns
visible labels can still map the answer without another confirmation.

## Custom Voice Cloning

Voice cloning is a separate consented flow. Proactively offering it during
voice selection is required and is not the same as starting a clone. Never
clone a third party's voice merely because a clip is present in the project.

### Native ChatCut editor

When the user has not already supplied a usable reference audio attachment or
ChatCut audio asset for this cloning request, delegate creation to the existing
editor dialog and follow the **dialog-entry sequence** below. Do not first ask
for the voice name, language, reference upload, or consent in an Agent widget.
The dialog owns fresh entitlement and slot checks, recording/upload,
authorization, durable source storage, Fish Audio registration, preview, and
retry.

When the user has already supplied a usable reference for this cloning request,
do not send them back through the dialog and do not ask them to reselect the
file. Follow **Agent-driven cloning from an available reference** below. After
the creation preflight succeeds, collect only missing fields: explicit
authorization, voice name, and preview text. Then call
`manage_voice voiceType="custom" action="clone"` with the resolved ChatCut
audio asset id.

For the dialog path, the Agent owns presenting the entry. Immediately after the
user chooses the clone branch:

1. Send one short, natural sentence that tells the user they can click the
   button below to start cloning.
2. Render exactly `<clone-voice/>` on its own line.
3. Add one short, localized sentence asking the user to send a message when
   cloning is complete so the Agent can continue the request.

Adapt the instruction and reminder to the conversation and language. This
three-part sequence must be the final assistant message for the run after all
required tools have finished; never place it in a pre-tool or intermediate
message. Examples:

- Chinese: `可以点击下方按钮开始克隆你的音色啦。\n<clone-voice/>\n克隆完成后告诉我一声,我再继续。`
- English: `Click the button below to start cloning your voice.\n<clone-voice/>\nLet me know when cloning is complete, and I'll continue.`
- Spanish: `Haz clic en el botón de abajo para empezar a clonar tu voz.\n<clone-voice/>\nAvísame cuando termine la clonación y continuaré.`

These are examples, not fixed copy. Do not say “follow these steps” or imply
that the Agent will collect clone inputs. Do not redirect the user to an editor
menu to find cloning themselves. After sending the entry and reminder, stop
and wait for the user to report completion or send the draft created by the
dialog's Apply action.

When the user clicks Apply, the editor inserts both the cloned-voice attachment
and the localized equivalent of `Continue generating with this voice.` into
the prompt draft. Wait for the user to send that draft. The attached hidden
voice context contains the exact custom voice id; continue the existing request
with `submit_voice provider="fish-audio"` when final speech is actually
requested. Do not emit raw `<audio>` HTML or a second Retry / Apply widget.

### Agent-driven cloning from an available reference

Use this flow in either of these cases:

- the native ChatCut Agent already has a readable attachment or accessible
  ChatCut audio asset that the user supplied for this cloning request; or
- an external Codex / Claude host is collecting the reference as a conversation
  attachment because it cannot open the native editor dialog.

Never call the clone action until the user submits explicit permission. ChatCut
durably archives every submitted clone-source recording in its own user-file
storage before Fish Audio registration. This source is retained for future
provider migration even if the project asset is later deleted; never ask a
user to re-record or reselect an already accessible reference merely to create
the voice.

After the user chooses `clone_voice` or otherwise explicitly asks to clone from
an already available reference, resolve any reference already supplied for this
cloning request before asking for another upload. Do not use
`providerAvailable:false` to block intake; the provider integration may be
enabled after the choice was rendered, and the clone action is the
authoritative availability check. Use entitlement and slot fields to explain
an upgrade or a full quota before asking for unnecessary inputs when the
account cannot create another clone. If the clone action itself returns
`CUSTOM_VOICE_PROVIDER_NOT_CONFIGURED`, explain that cloning is temporarily
unavailable and keep the requested name and imported reference asset in context
so the flow can be retried later.

Before asking for a name, recording, upload, or consent, call `manage_voice
voiceType="custom" action="check-create-access"` and treat that fresh result as
the authoritative creation preflight. Do not use `list` alone as an entitlement
check: `check-create-access` deliberately emits the runtime's feature-gated
upgrade card for an ineligible free account.

- Free account with `freeTrialAvailable:false`: do not collect another
  reference. Explain that the Free custom-voice slot is occupied and that the
  user must delete the existing voice or upgrade to create another. In the
  native editor the product opens its pricing dialog; in the Agent conversation,
  the existing feature-gated upgrade card is the equivalent interaction. Do not
  invent a purchase URL or continue cloning behind it.
- Pro account with `activeVoiceCount >= voiceSlotLimit`: do not collect another
  reference. Say exactly how many voices are active and that the current plan
  limit has been reached, then offer two conversational actions: upgrade the
  plan through the runtime's upgrade card, or close/continue with an existing
  voice. Do not call `clone` until the user has upgraded or freed a slot.
- Otherwise continue with the intake below. A stale preflight never overrides
  the clone endpoint: if `clone` still returns `FEATURE_NOT_INCLUDED` or
  `CUSTOM_VOICE_QUOTA_EXCEEDED`, follow the same recovery path and do not retry
  automatically.

1. Explain that the reference audio will be securely processed to create a
   reusable cloned voice. Do not identify the underlying provider.
2. Require a name, explicit authorization/risk confirmation, preview text, and
   one valid reference audio. When the user already supplied the reference,
   reuse it and ask only for the other missing fields. The user must confirm
   that the speaker is the user or the user has permission, and that the voice
   will not be used for impersonation, fraud, or unlawful activity. Uploaded
   audio is language-detected by ChatCut, so do not ask the user to choose its
   language. The stable ChatCut formats are AAC, FLAC, M4A, MP3, OGG, WAV, and
   WebM. Require 10 seconds to 3 minutes; recommend 30-60 seconds of clean solo
   speech without music, reverb, or background noise.
3. Load `widget-forms` and request only the missing parts of this host-neutral
   intake contract:
   - `voice_name`: `short_text`, required.
   - `voice_reference`: `audio_reference`, exactly one required recording or
     attachment only when no usable reference has already been resolved.
   - `preview_text`: `short_text`, editable and no more than 100 Unicode
     characters. Prefill the localized default below, but let the user replace
     it with any text they want to hear in the cloned-voice preview.
   - `voice_consent`: `explicit_consent`, required and initially unselected.
     Ask the adapter to keep the missing fields in one intake when its host
     supports media fields. If the host collects conversation attachments
     separately, follow the adapter's attachment flow and do not treat attaching
     a file as consent. Do not re-ask a field the user has already supplied
     clearly in the current cloning request.

Authorization and the use commitment are a hard continuation gate. After the
intake returns, independently verify that the user affirmatively submitted the
exact localized authorization option above. Until that confirmation is
present, stop: do not import the attached reference, do not call
`manage_voice voiceType="custom" action="clone"` or `action="preview"`, and do
not claim that cloning has started. A supplied name, an attachment, a generic “yes”, a
previous unrelated approval, or widget state showing other completed fields is
not authorization. If authorization is missing or ambiguous, ask only for the
full authorization confirmation again and wait for the user's answer.

Normalize the submitted result before cloning:

```ts
{
  voiceName: "<submitted name>",
  sourceAssetId: "<ChatCut audio assetId>",
  previewText: "<submitted preview text or localized default>",
  confirmedConsent: true,
}
```

Localize all visible copy to the user's conversation language. For the
authorization checkbox, use the same product copy as the Create Voice
dialog rather than paraphrasing it:

- Chinese: `我确认拥有该音色或已获得克隆授权,并承诺不将其用于冒充他人、欺诈或其他违法用途。`
- English: `I confirm that I own this voice or have permission to clone it, and will not use it for impersonation, fraud, or unlawful purposes.`
- Spanish: `Confirmo que esta voz me pertenece o que tengo permiso para clonarla, y que no la utilizaré para suplantar identidades, cometer fraude ni otros fines ilícitos.`

Use ChatCut's detected language tag rather than asking the user to identify the
language manually.

The intake must resolve to a ChatCut audio `assetId`. If the host returns a
readable attachment instead, load `asset-import`, import it into the targeted
project, and use the returned audio `assetId`. Never pass a local path,
attachment URL, or raw bytes to `manage_voice`. If no readable audio is
available, stop and ask for it. Use the user's submitted `preview_text` after
trimming surrounding whitespace. If it is blank, use the localized default:

- Chinese: `这是你的克隆音色,希望你喜欢这个效果。`
- English: `This is your cloned voice. Hope you like it.`
- Spanish: `Tu voz clonada. Espero que te guste.`

Choose the default by the user's conversation language. The final preview text
must contain 1-100 Unicode characters. If the submitted value is longer, ask
the user to shorten it before cloning; do not silently truncate it or replace
it with the default. Preserve the user's wording rather than rewriting it from
conversation context.

Then call:

```ts
mcp__skill__manage_voice({
  voiceType: "custom",
  action: "clone",
  sourceAssetId: "<uploaded audio asset id>",
  name: "<submitted voice name>",
  confirmedConsent: true,
  previewText: "<submitted preview text or localized default>",
});
```

The tool waits for both cloning and preview generation. After it returns a
ready voice and `previewUrl`, ask the loaded `widget-forms` adapter to render
one `playable_preview` using the full `customVoiceId` as its stable value, the
submitted voice name as its display name, the submitted `previewText` as its
summary, and the exact tool-returned `previewUrl` as audio media. This is a
display-only preview, not another intake or selection form. Never put the URL
in prose, emit a raw `<audio>` element, or substitute an unsupported Markdown
audio/link syntax.

Immediately after the playable preview, ask for one single branch decision:
`Retry` / `重试` / `Reintentar`, or `Apply` / `应用` / `Aplicar`. This branch
decision uses `<choices/>`, not another form widget. Keep the full
`customVoiceId`, submitted name, submitted `previewText`, and `previewUrl`
mapped in context while waiting.

- `Retry` means collect a replacement recording or upload for this same voice;
  do not consume another slot and do not discard the previous ready voice until
  replacement succeeds. If the current tool surface cannot replace the source
  in place, explain that limitation instead of creating a second voice.
- `Apply` means select the cloned voice for the current AI draft/request. Call
  `manage_voice voiceType="custom" action="apply"` with the selected
  `customVoiceId`. This action performs the authoritative Pro/readiness check without generating
  audio or consuming credits. On success, retain the custom voice id as the
  selected voice and continue the conversation; only call `submit_voice` with
  `provider: "fish-audio"` when the user actually asks to generate final speech.
  If it returns `FEATURE_NOT_INCLUDED`, do not present the voice as applied and
  do not synthesize. Explain that the free clone can be previewed but applying
  it for generated narration requires Pro; the runtime's feature-gated upgrade
  card is the Agent equivalent of the editor paywall.

The free plan can attempt one clone and listen to its preview, but cannot use a
cloned voice for final TTS. Pro custom-voice slots equal
`floor(monthly plan credits / 100)`; a ready or processing voice occupies one
slot. Slot access does not include free generation: final cloned-voice TTS uses
the normal voice-generation credits, deducted only after generation succeeds.
The backend is authoritative for all three rules.

When the user asks to generate final speech with a cloned voice, call
`submit_voice` with the ChatCut custom voice id:

```ts
mcp__skill__submit_voice({
  provider: "fish-audio",
  voiceId: "<confirmed custom voice id>",
  text: "<final narration text>",
  name: "Custom voiceover",
});
```

If the tool returns `FEATURE_NOT_INCLUDED` / `feature_not_included`, tell the
user that one custom-voice preview slot is available on Free but final cloned
voice TTS requires Pro. Do not retry, switch voices, or submit the provider id
through another tool to bypass the gate. The runtime emits the standard pricing
upgrade card for this blocker; invite the user to upgrade and retry afterward.
For a paid account, phrase this as: it has created `activeVoiceCount` voices and
has reached the current subscription plan limit of `voiceSlotLimit`; offer an
upgrade action and a close/continue action. Do not route a quota error through
the free-account paywall copy.

## Sound Effects

For ordinary editing sound effects (SFX), do **not** generate first. Use the
built-in Sound Effects library before spending credits:

1. Call `browse_library` with `category:"sound-effects"` and a query such as
   `"whoosh"`, `"camera shutter"`, `"notification"`, `"censor beep"`, or
   `"record scratch"`.
2. Inspect the returned `library:sound:<id>`.
3. Place it with `edit_item`, using `fromFrame` as the sound's
   anchor/editorial moment frame:

```ts
mcp__core__browse_library({
  category: "sound-effects",
  query: "short whoosh transition",
});

mcp__core__edit_item({
  adds: [
    {
      type: "audio",
      assetId: "library:sound:whoosh-short",
      fromFrame: 120,
      trackId: "A1",
    },
  ],
});
```

Only generate sound effects from text descriptions with `submit_sound` when:

- The user explicitly asks for a generated/original/custom sound.
- The requested sound is too specific for the existing Sound Effects library.
- `browse_library({ category:"sound-effects", query })` returns no suitable
  match.

```ts
// Custom/generated sound effect after the library has no suitable match
mcp__skill__submit_sound({ prompt: "A dog barking in the distance" });

// With custom duration (0.5-22 seconds)
mcp__skill__submit_sound({
  prompt: "Thunder and heavy rain",
  durationSeconds: 15,
});

// High prompt adherence
mcp__skill__submit_sound({
  prompt: "Sci-fi laser gun firing",
  promptInfluence: 0.8,
});
```

**Tips for better results:**

- Be specific: "A dog barking loudly" vs just "dog"
- Include context: "Footsteps on wooden floor in an empty room"
- Specify style: "Cinematic whoosh" or "8-bit game sound"

## Parameters

### TTS

| Field        | Description                             | Notes           |
| ------------ | --------------------------------------- | --------------- |
| `provider`   | `doubao`, `elevenlabs`, or `fish-audio` | Required        |
| `text`       | Text to synthesize                      | Required        |
| `voiceId`    | Curated id, or ChatCut custom voice id  | Required        |
| `speedRatio` | Speech speed                            | Doubao only     |
| `modelId`    | ElevenLabs model id                     | ElevenLabs only |
| `stability`  | ElevenLabs stability                    | ElevenLabs only |
| `speed`      | ElevenLabs speech speed                 | ElevenLabs only |
| `name`       | Asset name                              | Optional        |

### Custom voices

| Action                | Required fields                                            | Result                                                |
| --------------------- | ---------------------------------------------------------- | ----------------------------------------------------- |
| `list`                | none                                                       | Entitlement, slot count, existing voices              |
| `check-create-access` | none                                                       | Authoritative create gate before collecting reference |
| `apply`               | `customVoiceId`                                            | Pro/readiness check; selects voice without generation |
| `clone`               | `sourceAssetId`, `name`, `confirmedConsent`, `previewText` | Reusable Fish Audio voice + preview                   |
| `preview`             | `customVoiceId`, `previewText`                             | Refreshed short audition                              |
| `rename`              | `customVoiceId`, `name`                                    | Updated product/provider display name                 |
| `delete`              | `customVoiceId`, `confirmedDelete:true`                    | Permanently deletes provider voice and frees the slot |

Do not use or expose a separate `manage_fish_audio_voice` tool. Fish model IDs,
public/unlisted publishing, arbitrary tags, cover images, and provider-level
catalog management are internal integration details. Delete only after an
explicit user confirmation; use `rename` for a simple name change.

### Sound Effects

| Field             | Description       | Notes          |
| ----------------- | ----------------- | -------------- |
| `prompt`          | Sound description | Required       |
| `durationSeconds` | Duration          | 0.5-22 seconds |
| `promptInfluence` | Prompt adherence  | 0-1            |
| `name`            | Asset name        | Optional       |

## Voices

Use `manage_voice action="list" voiceType="all"` for the current unified voice
candidate list, localized display names, previews, and submission ids.

### Voice presets are provider-specific — do NOT mix them

Official voice ids are provider-specific. Always keep the returned `provider`
and `voiceId` together; mixing values from different catalog tuples will fail.

If you need a specific voice and a particular language:

- For any narration language, query the catalog with that `language` and use a
  returned provider/voice pair. Do not guess a raw provider id.

## Hard rules — what you must NOT do

1. Never use a voice preset name from a different provider.
2. Never render official voice recommendations before the mandatory custom
   voice `list` gate when no concrete voice was named.
3. Never use explicit or inferred voice traits to exclude a ready custom voice
   with a playable preview from the combined audition surface.
4. Never omit the final clone action card from a voice recommendation or
   selection surface.
5. Never submit TTS when the voice is only described broadly and the user has
   not confirmed a concrete preset.
6. Never recommend or render a TTS voice option before calling `manage_voice
action="list" voiceType="all"` for the current request.
7. Never claim stable age, regional accent, pronunciation dictionary, or exact
   duration controls; the current tool does not expose those as reliable fields.
8. Never replace original recorded speech with TTS unless the user asks.
9. Never import or clone a reference voice until the user has affirmatively
   submitted the full ownership/permission and lawful-use commitment required
   by the cloning intake. Never infer this confirmation from an attachment,
   another completed field, a generic approval, or prior unrelated context.
10. Never bypass cloned-voice Pro checks by passing a provider voice id to
    `submit_voice` or another generation tool.
11. Never force the native clone dialog or ask the user to reselect audio when
    the user has already supplied a usable reference for the current cloning
    request. Preflight access, collect the missing authorization, name, and
    preview text, then clone from its ChatCut audio asset id.
12. Never treat unrelated speech already present in the project as a cloning
    reference or as permission to clone it.
13. Never render a cloned-voice preview as raw `<audio>` HTML, a Markdown link,
    or a bare URL. Use `widget-forms` `playable_preview` with the exact
    tool-returned `previewUrl`; keep Retry / Apply in the separate branch
    control required by the current host.
14. Never emit `<clone-voice/>` in an intermediate message or more than once in
    one Agent run. Complete required tools first; if the native dialog remains
    the correct route, follow its dialog-entry sequence and then stop.
15. Never expose Doubao, ElevenLabs, Fish Audio, provider model names,
    provider-specific ids, or provider URLs to the user. Use `official voice`
    and `cloned voice` product terminology, and sanitize provider details from
    visible errors and status messages.

Referenced files: 2

widget-forms6.58 KB

View saved version →

---
name: widget-forms
description: Ask for structured ChatCut input using the current plugin host's supported form surface.
---

# Widget Forms Host Adapter

The canonical agent's raw `<widget>`, `<choices/>`, and `<visual-option>` tags render only inside ChatCut. Never emit those tags from the published Codex or Claude Code plugin.

## Semantic contract mapping

This adapter owns how plugin hosts implement a caller's host-neutral form
contract. It does not decide which fields the workflow requires.

Map semantic field types as follows:

- `short_text`: one text input.
- `explicit_consent`: one explicit, initially unselected confirmation control
  using the caller's full localized copy. An attachment or another submitted
  field is never consent.
- `audio_reference`: plugin forms cannot record or upload media. Render the
  other fields, ask the user to attach the audio to the conversation, then load
  `asset-import`. Return the imported ChatCut audio `assetId` to the calling
  workflow; never return a local path, attachment URL, or raw bytes as an asset
  id.
- `playable_single_choice`: one choice surface with stable option values,
  localized labels, and playable media when the host supports it. Keep the
  value-to-label map in context when a host returns the visible label.
- `playable_preview`: one display-only, non-submitting audio surface using the
  exact caller-provided runtime media URL. Use only a host-supported safe media
  renderer; never emit raw `<audio>` HTML or expose the URL as ordinary prose.

## Codex

Call the ChatCut MCP tool `ask_followup_questions`. Put related fields into one form, write visible text in the user's language, and stop after the call until the submitted answer appears in chat. Use the current tool schema for field and option shapes.

Map `playable_single_choice` to one `single` field with `variant:"voice"`.
Use each stable value as the option id and keep its localized label,
description, and `audioUrl` tied to the same caller-provided item. An action
without playable media has no `audioUrl`.

Map an MG reference style choice to one `single` field with `variant:"visual"`.
Use the reference ID as the option id, its short title as the label, and the
selected detail's `heroPreview.url` as `preview`. Include one `__other__` option
labeled "Other" in the user's language; the tool supplies its text input. That branch accepts
either a new preference or a request for another batch; do not add a separate
Show more option. Return the selected reference ID to the MG skill; it owns
application. Do not bind `apply_preset` directly to a reference ID.

Example field, replacing the reference placeholders with the fetched shortlist
and localizing the labels:

```json
{
  "id": "mg_style",
  "label": "Choose a Motion Graphics style",
  "type": "single",
  "variant": "visual",
  "options": [
    { "id": "REFERENCE_ID", "label": "TITLE", "preview": "HERO_PREVIEW_URL" },
    { "id": "__other__", "label": "Other" }
  ]
}
```

The Other answer returns to discovery, without applying a style or
starting production. Keep the image-option IDs unchanged in the next step.

Map an AI-avatar identity choice to one `single` field with `variant:"visual"`.
Use the live `identityId` as the option id, the returned `name` as its label,
`previewVideoUrl` as `previewVideo`, and `previewImageUrl` as `preview`. The
image remains the poster and automatic fallback if video loading fails. Keep
`create_avatar` as a label-only option. Selecting it returns that value to the
workflow; it does not open ChatCut's native avatar dialog, so request the source
attachment separately and use `asset-import`.

Do not call `ask_followup_questions` solely to display a `playable_preview`,
because a passive preview is not a question. Use Codex's supported audio
attachment/media rendering when available. If this host session has no safe
runtime-URL audio renderer, say that the preview is available in the ChatCut
project assets and continue with the caller's separate branch decision; do not
print HTML, a bare URL, or a fake widget tag.

Map `explicit_consent` to a single-select field containing only one confirmation
option. Use the caller's full confirmation copy as its label and do not
preselect it. Mark every blocking field as required and verify every required
answer after submission before continuing.

## Claude Code

Follow the structured-input recipe in `chatcut-plugin-basics-claude`: use one `visualize.show_widget` Elicitation form, submit only through `.elicit-submit`, and wait for the user to send the filled prompt. Never call ChatCut's `ask_followup_questions` in Claude Code because that host does not render its MCP-App result.

Map `explicit_consent` to an unchecked checkbox using the caller's complete
localized copy. Do not add a file chooser for `audio_reference`; request the
conversation attachment outside the form and load `asset-import` after the
user sends it.

For `playable_single_choice`, resolve bundled media keys through
`${CLAUDE_PLUGIN_ROOT}/assets/widget-media/manifest.json`; use a label-only card
when a key is absent and never embed or process the original source URL. Keep
the stable value in the label-to-id map rather than exposing it as the DOM audio
key. Runtime media URLs are not in the bundled manifest: do not embed them in
Elicitation HTML. Present them through the host's normal safe link/audio
surface, then use a label-only confirmation control.

AI-avatar previews are runtime catalog media and are not bundled in the plugin
manifest. Keep each authoritative `identityId` mapped to its returned `name`,
use label-only selectable cards in Elicitation, and tell the user that live
video previews are available in ChatCut's AI Avatars library. Never embed or
print the provider preview URL. A `create_avatar` selection returns to the
workflow and then follows the normal conversation-attachment plus
`asset-import` path.

For MG reference choices, look up the exact reference ID in the same manifest
(`entries[referenceId]`, for example a `mg-reference:style:…` key). An image entry
provides `path`; resolve it against the manifest's immutable `baseUrl` and show
that image beside its short title in the Elicitation form. Keep reference IDs
mapped to the displayed labels. Do not substitute a similarly named preset.

The manifest may not include newly added references yet. For missing entries,
use label-only choices and say their image previews are available in ChatCut;
do not invent CDN paths, download or republish media during the conversation,
or embed the runtime S3 URL. Include a single label-only Other choice with a
text input. A preference or request for another batch returns to discovery,
not style application.
Package details

Publisher declarations from the archived package. These are separate from our research and the live service's terms.

Package license
GPL-3.0-only
Package author
ChatCut
Keywords
chatcut, video, video-editing, editor, captions, subtitles, transcription, timeline, local-video, codex, mcp

Declared capabilities

  • Read
  • Write

Package observed Sep 30, 2026.

Technical details
First seen
Sep 30, 2026 · 22:02 UTC
Last seen
Oct 1, 2026 · 12:00 UTC
Collection status
Collected

plugins_6a8c039995c08191932d494269209307

Download plugin data (JSON)