← ChatCutCONTENT HISTORY

Update to ChatCut

Snapshot Sep 30, 2026 · 23:14 UTC · version 1.10.14

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "voice",
  "description": "Text-to-Speech (TTS), voice cloning, voiceover, narration placement/sync, and custom sound effects (SFX) generator. Use when the user wants generated speech from text, wants to clone a consented voice from uploaded reference audio, wants to correct a misspoken word or short phrase in recorded speech, wants to add/replace/align narration or voiceover for an existing video/timeline, wants to keep existing voiceover synced after visual retiming edits, needs voice audition/selection, or explicitly wants a newly generated/custom sound effect that is not available in the Sound Effects library.",
  "included_files": [
    {
      "relative_path": "references/fix-spoken-mistakes.md",
      "size_in_bytes": 5560
    },
    {
      "relative_path": "references/video-sync.md",
      "size_in_bytes": 8034
    }
  ],
  "skill_md_contents": "---\nname: voice\ndescription: Text-to-Speech (TTS), voice cloning, voiceover, narration placement/sync, and custom sound effects (SFX) generator. Use when the user wants generated speech from text, wants to clone a consented voice from uploaded reference audio, wants to correct a misspoken word or short phrase in recorded speech, wants to add/replace/align narration or voiceover for an existing video/timeline, wants to keep existing voiceover synced after visual retiming edits, needs voice audition/selection, or explicitly wants a newly generated/custom sound effect that is not available in the Sound Effects library.\nuser-invocable: true\n---\n\n# Voice & Sound Effects Generator\n\nGenerate voiceovers (TTS) and sound effects. For TTS, choose a concrete\nprovider and voice before calling `submit_voice`.\n\n## When to Use\n\n- Generate voiceover/narration from text\n- Create text-to-speech audio for videos\n- Add, replace, or redo narration/voiceover for an existing video, timeline,\n  screen recording, slide animation, product demo, B-roll edit, MG explainer, or\n  other visual sequence\n- Keep existing narration/voiceover aligned after trimming, speeding up, slowing\n  down, moving, reordering, or replacing the visuals it describes\n- Offer and audition TTS voice choices when the user has not picked a concrete voice\n- Clone the user's own or explicitly authorized voice from reference audio\n- Correct a real misspoken word or short phrase with the authorized speaker's\n  cloned voice while preserving picture and downstream timing\n- Generate custom sound effects from text descriptions only after checking the Sound Effects library first\n\n## TTS (Text-to-Speech)\n\nIf the user wants to correct a word or short phrase that was actually spoken\nincorrectly in existing recorded speech, read\n[references/fix-spoken-mistakes.md](references/fix-spoken-mistakes.md) before\nediting or generating. That workflow preserves the picture and protects later\ntiming while replacing only the faulty sound. Do not use synthesized speech\nwhen the recording is already correct and only its transcript is wrong; fix\nthe transcript instead. Use the ordinary narration path for a full rewrite or\ncomplete new voiceover.\n\nIf the current request has an existing visual target and the user wants\nnarration, voiceover, dubbing, or replacement speech for that target, read\n[references/video-sync.md](references/video-sync.md) before drafting new\nnarration, using existing narration text to generate TTS, or placing audio. Do\nthis even when the user did not explicitly say \"sync\" or \"match the visuals\";\nthe existence of a visual target means narration timing and meaning may need to\nfollow on-screen content. Use the normal standalone TTS path only when there is\nno visual target or the user just wants an audio asset from text.\n\nAlso read [references/video-sync.md](references/video-sync.md) when the timeline\nalready has narration/voiceover and the user asks to change the visuals while\nkeeping that voiceover aligned. This is a sync maintenance task even if no new\nTTS is needed.\n\nUse `manage_voice` as the single catalog and management entry, and\n`submit_voice` for final TTS. `manage_voice action=\"list\" voiceType=\"all\"`\nreturns official and cloned candidates in one normalized list. Only\n`voiceType=\"custom\"` supports create-access, apply, clone, preview, rename, or\ndelete actions. The current contracts are:\n\n- `provider` is required. Use `doubao` for Chinese-optimized narration and\n  `elevenlabs` for English or multilingual narration, or `fish-audio` for a\n  ready ChatCut custom voice.\n- `voiceId` is required and provider-specific. Do not mix catalogs.\n- `submit_voice` creates an audio asset only. Timeline placement, replacement,\n  trimming, and alignment happen later with timeline tools.\n- For long narration, multiple `submit_voice` calls can be useful: split at\n  natural pauses, sentence groups, or script beat boundaries when the workflow\n  benefits from separately timed or placed voice clips, such as storyboard beats,\n  scene-level ad segments, or a user request for separate assets.\n- For Doubao, `speedRatio`, `loudnessRatio`, `pitch`, `emotion`,\n  `emotionScale`, `performancePrompt`, and `explicitDialect` are supported\n  knobs, but not every voice supports every expressive control. Use the\n  selected official entry returned by `manage_voice action=\"list\"` as the\n  authority.\n- For ElevenLabs, `modelId`, `speed`, and `stability` are the supported voice\n  knobs. For `eleven_v3`, inline audio tags are available for expressive\n  delivery such as emotion, tone, nonverbal cues, accent hints, pauses, or\n  local pacing.\n- For `fish-audio`, pass the ready voice's ChatCut `customVoiceId` as\n  `voiceId`. `submit_voice` performs the authoritative Pro/readiness check,\n  creates a generation job, and consumes standard TTS credits after success.\n  Never pass or request the provider's internal Fish model id. The current\n  Fish S2.1 model supports inline square-bracket cues in `text` for local\n  emotion, delivery, and paralinguistic control.\n\n### Keep voice providers private\n\nDoubao, ElevenLabs, Fish Audio, their model names, and other provider identity\nare internal implementation details. Never expose or attribute them in\nuser-facing replies, progress updates, voice recommendations, audition cards,\nclone instructions, success summaries, or errors. Use ChatCut product language\ninstead:\n\n- Call curated catalog voices `official voices` / `官方音色`.\n- Call saved custom voices `cloned voices` / `克隆音色`, `My Voices` /\n  `我的音色`, or use the voice's user-visible saved name.\n- Describe status and failures at the ChatCut feature level. If a tool or\n  provider error contains a provider name, model name, provider voice id, or\n  provider URL, preserve the actionable meaning but remove those details\n  before replying.\n\nProvider names and provider-specific ids remain valid only in internal tool\narguments, tool-result interpretation, and these implementation instructions.\nDo not copy them from tool output into visible UI metadata or prose.\n\nOfficial voice controls come from `manage_voice action=\"list\"`:\n\n- Call it for the target narration language and the user's visible locale.\n  Treat each returned `(provider, voiceId, name, summary, sampleUrl,\ncapabilities)` tuple as atomic.\n- Use only fields in that entry's `capabilities.supportedControls`; never infer\n  capability support from a remembered voice name or family.\n- Controls are not guarantees of a specific acting style. Use the returned\n  tags and sample to pick a naturally suitable voice, then use supported\n  controls for moderate delivery changes.\n- For entries that advertise `audioTags`, inline audio tags are available when\n  the user asks for expressive delivery such as emotion, tone, nonverbal cues,\n  accent hints, or local pacing. Official examples fit these useful TTS\n  categories:\n- Emotion/tone tags include `[happy]`, `[sad]`, `[angry]`, `[excited]`,\n  `[curious]`, `[sarcastic]`, `[crying]`, `[annoyed]`, `[appalled]`,\n  `[thoughtful]`, `[surprised]`, and `[mischievously]`; vocal delivery and\n  nonverbal cue tags such as `[whispers]`, `[laughs]`, `[sighs]`, `[exhales]`,\n  `[inhales deeply]`, `[clears throat]`, `[snorts]`, `[swallows]`,\n  `[wheezing]`, and `[coughs]`;\n  pacing/pause/local speed tags such as `[slowly]`, `[pause]`,\n  `[short pause]`, `[long pause]`, `[rushed]`, and `[drawn out]`; and\n  accent/special-performance tags such as\n  `[strong X accent]`, for example `[strong French accent]`, plus `[sings]`,\n  `[singing]`, `[woo]`, and `[pirate voice]`. Official examples are\n  non-exhaustive; similar auditory tags can be tried when the user explicitly\n  asks for that delivery and the tag describes how the voice should sound, not\n  a visual action. Write tags directly in `text`, close to the short phrase\n  they should affect. Treat tags as local guidance, not paragraph-wide controls.\n- For pauses and pacing, use punctuation, text structure,\n  shorter generated segments, or local audio tags such as `[short pause]` and\n  `[slowly]` when needed.\n\nFish Audio control support for cloned voices:\n\n- The current `fish-audio` route uses Fish S2.1. Put concise natural-language\n  cues in square brackets directly in `submit_voice.text`, for example\n  `[happy]`, `[calm]`, `[angry]`, `[excited]`, `[whisper]`, `[laugh]`,\n  `[sigh]`, `[gasp]`, `[pause]`, `[emphasis]`, `[inhale]`, or `[exhale]`.\n  S2.1 is not limited to a fixed tag list, so a specific auditory description\n  such as `[whispers sweetly]` or `[laughing nervously]` is also valid.\n- Place a cue immediately before the phrase or moment it should affect. Fish\n  S2.1 accepts cues anywhere in the text, so multiple short cues can create\n  local transitions, for example\n  `[calm] 先别着急。[excited] 好消息是,我们已经找到解决办法了!` or\n  `I thought it was over [gasp] but then the lights came back [relieved].`\n- Use cues sparingly and only when the requested delivery benefits from them.\n  Prefer one clear instruction at a transition over stacking conflicting\n  directions. Treat the result as model guidance rather than a deterministic\n  editing boundary, and split into separate `submit_voice` calls when exact\n  clip-level timing or independent retries matter.\n- Do not use Fish S1's legacy `(parenthesis)` emotion syntax on this route, and\n  do not add a separate `emotion` argument: the S2.1 cue belongs inside\n  `text`. Keep the default voice-cloning preview line untagged. Do not add cues\n  to a custom preview on the user's behalf, but preserve cues the user\n  intentionally includes within the 100-character preview limit.\n\n```ts\n// English / multilingual via ElevenLabs\nmcp__skill__submit_voice({\n  provider: \"elevenlabs\",\n  text: \"Hello world\",\n  voiceId: \"peter\",\n});\n\n// Chinese via Doubao\nmcp__skill__submit_voice({\n  provider: \"doubao\",\n  text: \"你好世界\",\n  voiceId: \"liuchang\",\n});\n\n// With speed adjustment (Doubao only)\nmcp__skill__submit_voice({\n  provider: \"doubao\",\n  text: \"这是一段稍快的中文旁白。\",\n  voiceId: \"liuchang\",\n  speedRatio: 1.5,\n});\n\n// With expressive Doubao controls\nmcp__skill__submit_voice({\n  provider: \"doubao\",\n  text: \"这次事故提醒我们,安全永远不能侥幸。\",\n  voiceId: \"liuchang\",\n  emotion: \"sad\",\n  emotionScale: 3,\n  performancePrompt: \"痛心但克制,语速稍慢,像新闻专题旁白\",\n  pitch: -1,\n  speedRatio: 0.92,\n});\n\n// With ElevenLabs delivery controls\nmcp__skill__submit_voice({\n  provider: \"elevenlabs\",\n  text: \"The launch changed how teams plan their daily work.\",\n  voiceId: \"peter\",\n  speed: 0.95,\n  stability: 0.4,\n});\n\n// With local Fish Audio S2.1 delivery cues on a ready cloned voice\nmcp__skill__submit_voice({\n  provider: \"fish-audio\",\n  voiceId: \"<confirmed custom voice id>\",\n  text: \"[calm] 先别着急。[excited] 好消息是,我们已经找到解决办法了!\",\n  name: \"Expressive custom voiceover\",\n});\n```\n\n## Voice Audition Before Generation\n\n### Emit the native clone entry only as a final action\n\n`<clone-voice/>` is an executable editor action, not prose, code, or an\ninternal process label. Never quote it, wrap it in backticks, describe it as\nthe next step, or emit it in a progress update, pre-tool explanation, plan, or\nother intermediate assistant message.\n\nFinish every required inspection and tool call first. If the native dialog is\nstill the correct route afterward, follow the **dialog-entry sequence** under\n**Native ChatCut editor** below. That sequence must be the final user-facing\nassistant message for the run: do not call another tool, add anything after\nits completion reminder, or emit the tag again. A single Agent run may render\nat most one standalone native clone entry.\n\n### Route cloning by host and available reference\n\nDecide the host path before loading `widget-forms` or asking for clone inputs:\n\n- In the native ChatCut editor, first check whether the user has already\n  supplied a readable audio attachment or explicitly identified an accessible\n  audio-bearing ChatCut asset for this cloning request, including a timeline\n  audio or video item. If so, do not make the user choose the same source again\n  and do not emit `<clone-voice/>`. Resolve or derive its audio asset, run the\n  creation preflight, collect only the missing name, preview text, and explicit\n  authorization, then use the Agent-driven cloning flow below.\n- In the native ChatCut editor, use the editor dialog only when a usable\n  reference has not already been supplied. Follow the native dialog path under\n  **Custom Voice Cloning** below.\n- In external Codex / Claude hosts, never emit `<clone-voice/>`; those hosts do\n  not render or dispatch the native editor action. Use the Agent-driven\n  attachment flow below.\n\nAn arbitrary voice already present in the project is not consent or a cloning\nreference. Use the direct path only when the user supplied or identified the\naudio for the current cloning request and later gives the full authorization\nrequired below.\n\nWhen the user explicitly chooses an audio-bearing timeline item, resolve it and\nits source range with `preview_timeline`. Reuse a 10-second-to-3-minute audio\nasset directly. For video, `pull_asset`, extract the chosen range as supported\naudio with ffmpeg, then load `asset-import` and `push_asset` the result. For any\nsource over 3 minutes, use a clean 30–60-second excerpt; under 10 seconds, ask\nfor a longer reference. Continue with the resulting audio asset id without\nemitting `<clone-voice/>`; the source choice is not authorization.\n\n### Choose the interaction from live state\n\nTreat the unified voice lookup as a mandatory gate. Whenever the user requests\nTTS without naming a concrete official or custom voice, and has not already\nexplicitly chosen voice cloning:\n\n1. If `manage_voice` is not already loaded, use `ToolSearch` to load it.\n2. Call `manage_voice action=\"list\" voiceType=\"all\"` with the target narration\n   language, conversation locale, and any useful official-voice filters.\n3. Use the returned normalized `voices` list as the only candidate directory.\n   It already puts cloned voices before matching official voices.\n\nDo this before recommending or rendering any voice-selection UI. Never infer\nthe user's saved or official voices from chat history or a static Skill file.\n\nAn explicit request to clone a voice is already a concrete source choice. Do\nnot call the combined list merely to choose between official and cloned voices\nin that case. First resolve whether the current request already supplies a\nusable reference; perform `manage_voice voiceType=\"custom\"\naction=\"check-create-access\"`, then follow the selected host route. The\nstandalone native entry may appear only after that work is complete, under the\nfinal-action discipline above.\n\nThe Agent may infer useful tone, mood, delivery, or use-case suggestions from\nthe narration text. Treat these as recommendation signals, not as the user's\nchoice of voice source. A reasonable inference must never exclude a ready,\nplayable custom voice from the combined audition surface below.\n\nClassify the request before rendering anything:\n\n- **Concrete custom voice:** when the user names a ready custom voice or chooses\n  the only named custom-voice pill, use that exact custom voice. When the user\n  chooses `Choose my cloned voice` and several are ready, show only those ready\n  custom voices as playable choices before continuing.\n- **Concrete official preset:** use or confirm that exact preset; do not reopen\n  the source-choice branch.\n- **Clone action selected:** only now enter the host-specific cloning route\n  above in the very next reply. Do not wait for the user to ask again, tell\n  them to open a menu or panel, or leave the choice as an unhandled pill. In\n  the native editor, finish any required inspection or preflight before the\n  dialog entry; in external Codex / Claude hosts, begin the attachment flow.\n- **At least one ready custom voice, but no concrete voice named:** render one\n  combined playable audition surface. Put ready custom voices that have a real\n  `previewUrl` first, add 2-4 suitable official voices with playable samples,\n  and finish with `Clone another voice`. Do not ask a text-only saved-versus-\n  official source question and do not use `<choices/>`.\n- **No ready custom voice and no voice preference:** infer a small, varied set\n  of 2-4 generally suitable official voices, render them as playable cards,\n  and finish with `Clone my voice`. Do not use text-only source choices.\n- **No ready custom voice and voice requirements available:** use explicit\n  requirements such as \"middle-aged male\", \"warm female\", or \"professional\"\n  first; otherwise reasonable traits inferred from the narration may guide the\n  recommendation. Show 2-4 matching official playable cards and one clone\n  action card in the same selection surface. Do not append a standalone\n  `<clone-voice/>` button to this reply.\n\nLocalize the branch labels and adjust them to live state. Use these meanings:\n\n- Chinese:\n  - Named custom: `使用「<name>」`\n  - Custom list: `选择我的克隆音色`\n  - Curated catalog: `选择官方音色`\n  - First clone: `克隆我的音色`\n  - Additional clone: `克隆新音色`\n- English:\n  - Named custom: `Use “<name>”`\n  - Custom list: `Choose my cloned voice`\n  - Curated catalog: `Choose an official voice`\n  - First clone: `Clone my voice`\n  - Additional clone: `Clone another voice`\n- Spanish:\n  - Named custom: `Usar «<name>»`\n  - Custom list: `Elegir mi voz clonada`\n  - Curated catalog: `Elegir una voz oficial`\n  - First clone: `Clonar mi voz`\n  - Additional clone: `Clonar otra voz`\n\nNative Chinese example when no usable reference has been supplied:\n\n```text\n<widget>\n  <form-visual id=\"voiceId\" label=\"请选择并试听音色\" media-kind=\"audio\" required=\"true\">\n    <visual-option value=\"<returned voiceId>\" name=\"<returned localized name>\" summary=\"<returned summary>\" media=\"<returned sampleUrl>\" media-kind=\"audio\"/>\n    <visual-option value=\"clone_voice\" name=\"克隆我的音色\"/>\n  </form-visual>\n</widget>\n```\n\nWhen no usable reference has been supplied, the dialog action belongs inside\nthe playable grid:\n\n```html\n<visual-option value=\"clone_voice\" name=\"克隆我的音色\" />\n```\n\nAlways expose one appropriate path to cloning during ordinary voice selection,\nbut never show both a clone action card and the standalone native\n`<clone-voice/>` entry in the same reply. Provider availability, plan, and\nquota govern what happens after the user chooses cloning, not whether the\nchoice is visible. Offering the choice does not require consent; explicit\npermission and a supported reference are required only before the clone tool\ncall. A scenario reference may intentionally use a narrower audition without\nthe clone action and explain its direct-from-project fallback in prose.\n\nBefore recommending, rendering, or submitting a TTS voice option, use the\ncurrent `manage_voice action=\"list\" voiceType=\"all\"` result. It is the live\ncatalog shared with ChatCut Web, Desktop, and published plugins. Do not create\nvoice options from memory, translated names, filenames, or broad descriptions.\n\nFirst determine two separate languages:\n\n- User conversation language: the language the user used to talk to you. Use\n  this for surrounding copy, option names, and summaries.\n- Target narration language: the language of the text being synthesized. Use\n  this only to choose provider and voice catalog.\n\nLoad `widget-forms` before collecting input or rendering a choice. That skill\nowns the current host's form, attachment, and media-card behavior. This skill\nowns the voice candidates, required fields, safety rules, and the asset ids\npassed to voice tools. Do not embed host-specific UI instructions here.\n\n\"help me generate ... voice over in Chinese\" is an English conversation asking\nfor Chinese narration, so the audition widget copy stays in English while the\nvoice candidates come from Doubao.\n\nFor an official audition:\n\n1. Call `manage_voice action=\"list\" voiceType=\"all\"` with the target narration\n   `language`, user conversation `locale`, and any explicit gender, age, tone,\n   or use-case requirements. If none were explicit, a concise inferred\n   tone/use-case query may guide the official shortlist. Present inferred traits\n   as a suggestion, not as a stated user preference. If nothing official\n   matches, broaden only optional filters and clearly describe the closest\n   supported choices; never drop returned playable custom voices because they\n   lack official catalog tags.\n2. Use returned ready custom voices with previews first, followed by 2-4\n   returned official voices.\n3. Load `widget-forms` and request one required `playable_single_choice`. Give\n   every official option the `voiceId`, localized `name`, `summary`, and\n   matching `sampleUrl` from the same returned tuple. Pass `sampleUrl`\n   verbatim; never prepend an inferred S3, CDN, editor, localhost, or production\n   base URL, and never reconstruct it from the filename pattern. This differs\n   from a custom voice's `previewUrl`, which must be used exactly as returned by\n   `manage_voice`.\n4. Always add one no-media action option with stable value `clone_voice` as the\n   final card on every ordinary recommendation/selection surface. Label it\n   `Clone my voice` when no ready custom voice exists, or `Clone another voice`\n   when one does; localize it using the meanings above. In native ChatCut this\n   is a compact `<visual-option>` without `media`, not a separate button. In\n   external hosts it is the equivalent label-only option. A scenario-specific\n   reference may explicitly replace this ordinary surface with a narrower one.\n5. Wait for the user to choose. If the answer maps to `clone_voice`, render the\n   host-specific cloning entry/workflow on the next turn; do not start cloning\n   from the audition reply itself.\n6. For an official voice, call `submit_voice` with the selected tuple's\n   `provider` and `voiceId` verbatim.\n7. For an existing custom voice, call `manage_voice voiceType=\"custom\"\naction=\"apply\"` with its full mapped `customVoiceId`. When final speech is\n   requested, pass that same id to `submit_voice` as `voiceId` with\n   `provider: \"fish-audio\"`.\n\nFor a custom-voice-only audition, include every relevant `ready` custom voice\nthat has a `previewUrl`. Use the full `customVoiceId` as its stable value, the\nstored name as its display label, `previewText` as its summary when present,\nand the exact tool-returned `previewUrl` as audio media. Do not copy, download,\nrewrite, validate by hostname, or invent that URL. If a ready voice lacks a\npreview, omit it from the audition surface rather than making an unplayable\ntext-only card.\n\nKeep every option's type, submission id, display label, preview, summary, and\ncapabilities tied to the same `manage_voice` tuple. The target narration\nlanguage filters official entries; the conversation language controls all\nvisible copy. Keep the label-to-tuple map in context so a host that returns\nvisible labels can still map the answer without another confirmation.\n\n## Custom Voice Cloning\n\nVoice cloning is a separate consented flow. Proactively offering it during\nvoice selection is required and is not the same as starting a clone. Never\nclone a third party's voice merely because a clip is present in the project.\n\n### Native ChatCut editor\n\nWhen the user has not already supplied a usable reference audio attachment or\nChatCut audio asset for this cloning request, delegate creation to the existing\neditor dialog and follow the **dialog-entry sequence** below. Do not first ask\nfor the voice name, language, reference upload, or consent in an Agent widget.\nThe dialog owns fresh entitlement and slot checks, recording/upload,\nauthorization, durable source storage, Fish Audio registration, preview, and\nretry.\n\nWhen the user has already supplied a usable reference for this cloning request,\ndo not send them back through the dialog and do not ask them to reselect the\nfile. Follow **Agent-driven cloning from an available reference** below. After\nthe creation preflight succeeds, collect only missing fields: explicit\nauthorization, voice name, and preview text. Then call\n`manage_voice voiceType=\"custom\" action=\"clone\"` with the resolved ChatCut\naudio asset id.\n\nFor the dialog path, the Agent owns presenting the entry. Immediately after the\nuser chooses the clone branch:\n\n1. Send one short, natural sentence that tells the user they can click the\n   button below to start cloning.\n2. Render exactly `<clone-voice/>` on its own line.\n3. Add one short, localized sentence asking the user to send a message when\n   cloning is complete so the Agent can continue the request.\n\nAdapt the instruction and reminder to the conversation and language. This\nthree-part sequence must be the final assistant message for the run after all\nrequired tools have finished; never place it in a pre-tool or intermediate\nmessage. Examples:\n\n- Chinese: `可以点击下方按钮开始克隆你的音色啦。\\n<clone-voice/>\\n克隆完成后告诉我一声,我再继续。`\n- English: `Click the button below to start cloning your voice.\\n<clone-voice/>\\nLet me know when cloning is complete, and I'll continue.`\n- Spanish: `Haz clic en el botón de abajo para empezar a clonar tu voz.\\n<clone-voice/>\\nAvísame cuando termine la clonación y continuaré.`\n\nThese are examples, not fixed copy. Do not say “follow these steps” or imply\nthat the Agent will collect clone inputs. Do not redirect the user to an editor\nmenu to find cloning themselves. After sending the entry and reminder, stop\nand wait for the user to report completion or send the draft created by the\ndialog's Apply action.\n\nWhen the user clicks Apply, the editor inserts both the cloned-voice attachment\nand the localized equivalent of `Continue generating with this voice.` into\nthe prompt draft. Wait for the user to send that draft. The attached hidden\nvoice context contains the exact custom voice id; continue the existing request\nwith `submit_voice provider=\"fish-audio\"` when final speech is actually\nrequested. Do not emit raw `<audio>` HTML or a second Retry / Apply widget.\n\n### Agent-driven cloning from an available reference\n\nUse this flow in either of these cases:\n\n- the native ChatCut Agent already has a readable attachment or accessible\n  ChatCut audio asset that the user supplied for this cloning request; or\n- an external Codex / Claude host is collecting the reference as a conversation\n  attachment because it cannot open the native editor dialog.\n\nNever call the clone action until the user submits explicit permission. ChatCut\ndurably archives every submitted clone-source recording in its own user-file\nstorage before Fish Audio registration. This source is retained for future\nprovider migration even if the project asset is later deleted; never ask a\nuser to re-record or reselect an already accessible reference merely to create\nthe voice.\n\nAfter the user chooses `clone_voice` or otherwise explicitly asks to clone from\nan already available reference, resolve any reference already supplied for this\ncloning request before asking for another upload. Do not use\n`providerAvailable:false` to block intake; the provider integration may be\nenabled after the choice was rendered, and the clone action is the\nauthoritative availability check. Use entitlement and slot fields to explain\nan upgrade or a full quota before asking for unnecessary inputs when the\naccount cannot create another clone. If the clone action itself returns\n`CUSTOM_VOICE_PROVIDER_NOT_CONFIGURED`, explain that cloning is temporarily\nunavailable and keep the requested name and imported reference asset in context\nso the flow can be retried later.\n\nBefore asking for a name, recording, upload, or consent, call `manage_voice\nvoiceType=\"custom\" action=\"check-create-access\"` and treat that fresh result as\nthe authoritative creation preflight. Do not use `list` alone as an entitlement\ncheck: `check-create-access` deliberately emits the runtime's feature-gated\nupgrade card for an ineligible free account.\n\n- Free account with `freeTrialAvailable:false`: do not collect another\n  reference. Explain that the Free custom-voice slot is occupied and that the\n  user must delete the existing voice or upgrade to create another. In the\n  native editor the product opens its pricing dialog; in the Agent conversation,\n  the existing feature-gated upgrade card is the equivalent interaction. Do not\n  invent a purchase URL or continue cloning behind it.\n- Pro account with `activeVoiceCount >= voiceSlotLimit`: do not collect another\n  reference. Say exactly how many voices are active and that the current plan\n  limit has been reached, then offer two conversational actions: upgrade the\n  plan through the runtime's upgrade card, or close/continue with an existing\n  voice. Do not call `clone` until the user has upgraded or freed a slot.\n- Otherwise continue with the intake below. A stale preflight never overrides\n  the clone endpoint: if `clone` still returns `FEATURE_NOT_INCLUDED` or\n  `CUSTOM_VOICE_QUOTA_EXCEEDED`, follow the same recovery path and do not retry\n  automatically.\n\n1. Explain that the reference audio will be securely processed to create a\n   reusable cloned voice. Do not identify the underlying provider.\n2. Require a name, explicit authorization/risk confirmation, preview text, and\n   one valid reference audio. When the user already supplied the reference,\n   reuse it and ask only for the other missing fields. The user must confirm\n   that the speaker is the user or the user has permission, and that the voice\n   will not be used for impersonation, fraud, or unlawful activity. Uploaded\n   audio is language-detected by ChatCut, so do not ask the user to choose its\n   language. The stable ChatCut formats are AAC, FLAC, M4A, MP3, OGG, WAV, and\n   WebM. Require 10 seconds to 3 minutes; recommend 30-60 seconds of clean solo\n   speech without music, reverb, or background noise.\n3. Load `widget-forms` and request only the missing parts of this host-neutral\n   intake contract:\n   - `voice_name`: `short_text`, required.\n   - `voice_reference`: `audio_reference`, exactly one required recording or\n     attachment only when no usable reference has already been resolved.\n   - `preview_text`: `short_text`, editable and no more than 100 Unicode\n     characters. Prefill the localized default below, but let the user replace\n     it with any text they want to hear in the cloned-voice preview.\n   - `voice_consent`: `explicit_consent`, required and initially unselected.\n     Ask the adapter to keep the missing fields in one intake when its host\n     supports media fields. If the host collects conversation attachments\n     separately, follow the adapter's attachment flow and do not treat attaching\n     a file as consent. Do not re-ask a field the user has already supplied\n     clearly in the current cloning request.\n\nAuthorization and the use commitment are a hard continuation gate. After the\nintake returns, independently verify that the user affirmatively submitted the\nexact localized authorization option above. Until that confirmation is\npresent, stop: do not import the attached reference, do not call\n`manage_voice voiceType=\"custom\" action=\"clone\"` or `action=\"preview\"`, and do\nnot claim that cloning has started. A supplied name, an attachment, a generic “yes”, a\nprevious unrelated approval, or widget state showing other completed fields is\nnot authorization. If authorization is missing or ambiguous, ask only for the\nfull authorization confirmation again and wait for the user's answer.\n\nNormalize the submitted result before cloning:\n\n```ts\n{\n  voiceName: \"<submitted name>\",\n  sourceAssetId: \"<ChatCut audio assetId>\",\n  previewText: \"<submitted preview text or localized default>\",\n  confirmedConsent: true,\n}\n```\n\nLocalize all visible copy to the user's conversation language. For the\nauthorization checkbox, use the same product copy as the Create Voice\ndialog rather than paraphrasing it:\n\n- Chinese: `我确认拥有该音色或已获得克隆授权,并承诺不将其用于冒充他人、欺诈或其他违法用途。`\n- English: `I confirm that I own this voice or have permission to clone it, and will not use it for impersonation, fraud, or unlawful purposes.`\n- Spanish: `Confirmo que esta voz me pertenece o que tengo permiso para clonarla, y que no la utilizaré para suplantar identidades, cometer fraude ni otros fines ilícitos.`\n\nUse ChatCut's detected language tag rather than asking the user to identify the\nlanguage manually.\n\nThe intake must resolve to a ChatCut audio `assetId`. If the host returns a\nreadable attachment instead, load `asset-import`, import it into the targeted\nproject, and use the returned audio `assetId`. Never pass a local path,\nattachment URL, or raw bytes to `manage_voice`. If no readable audio is\navailable, stop and ask for it. Use the user's submitted `preview_text` after\ntrimming surrounding whitespace. If it is blank, use the localized default:\n\n- Chinese: `这是你的克隆音色,希望你喜欢这个效果。`\n- English: `This is your cloned voice. Hope you like it.`\n- Spanish: `Tu voz clonada. Espero que te guste.`\n\nChoose the default by the user's conversation language. The final preview text\nmust contain 1-100 Unicode characters. If the submitted value is longer, ask\nthe user to shorten it before cloning; do not silently truncate it or replace\nit with the default. Preserve the user's wording rather than rewriting it from\nconversation context.\n\nThen call:\n\n```ts\nmcp__skill__manage_voice({\n  voiceType: \"custom\",\n  action: \"clone\",\n  sourceAssetId: \"<uploaded audio asset id>\",\n  name: \"<submitted voice name>\",\n  confirmedConsent: true,\n  previewText: \"<submitted preview text or localized default>\",\n});\n```\n\nThe tool waits for both cloning and preview generation. After it returns a\nready voice and `previewUrl`, ask the loaded `widget-forms` adapter to render\none `playable_preview` using the full `customVoiceId` as its stable value, the\nsubmitted voice name as its display name, the submitted `previewText` as its\nsummary, and the exact tool-returned `previewUrl` as audio media. This is a\ndisplay-only preview, not another intake or selection form. Never put the URL\nin prose, emit a raw `<audio>` element, or substitute an unsupported Markdown\naudio/link syntax.\n\nImmediately after the playable preview, ask for one single branch decision:\n`Retry` / `重试` / `Reintentar`, or `Apply` / `应用` / `Aplicar`. This branch\ndecision uses `<choices/>`, not another form widget. Keep the full\n`customVoiceId`, submitted name, submitted `previewText`, and `previewUrl`\nmapped in context while waiting.\n\n- `Retry` means collect a replacement recording or upload for this same voice;\n  do not consume another slot and do not discard the previous ready voice until\n  replacement succeeds. If the current tool surface cannot replace the source\n  in place, explain that limitation instead of creating a second voice.\n- `Apply` means select the cloned voice for the current AI draft/request. Call\n  `manage_voice voiceType=\"custom\" action=\"apply\"` with the selected\n  `customVoiceId`. This action performs the authoritative Pro/readiness check without generating\n  audio or consuming credits. On success, retain the custom voice id as the\n  selected voice and continue the conversation; only call `submit_voice` with\n  `provider: \"fish-audio\"` when the user actually asks to generate final speech.\n  If it returns `FEATURE_NOT_INCLUDED`, do not present the voice as applied and\n  do not synthesize. Explain that the free clone can be previewed but applying\n  it for generated narration requires Pro; the runtime's feature-gated upgrade\n  card is the Agent equivalent of the editor paywall.\n\nThe free plan can attempt one clone and listen to its preview, but cannot use a\ncloned voice for final TTS. Pro custom-voice slots equal\n`floor(monthly plan credits / 100)`; a ready or processing voice occupies one\nslot. Slot access does not include free generation: final cloned-voice TTS uses\nthe normal voice-generation credits, deducted only after generation succeeds.\nThe backend is authoritative for all three rules.\n\nWhen the user asks to generate final speech with a cloned voice, call\n`submit_voice` with the ChatCut custom voice id:\n\n```ts\nmcp__skill__submit_voice({\n  provider: \"fish-audio\",\n  voiceId: \"<confirmed custom voice id>\",\n  text: \"<final narration text>\",\n  name: \"Custom voiceover\",\n});\n```\n\nIf the tool returns `FEATURE_NOT_INCLUDED` / `feature_not_included`, tell the\nuser that one custom-voice preview slot is available on Free but final cloned\nvoice TTS requires Pro. Do not retry, switch voices, or submit the provider id\nthrough another tool to bypass the gate. The runtime emits the standard pricing\nupgrade card for this blocker; invite the user to upgrade and retry afterward.\nFor a paid account, phrase this as: it has created `activeVoiceCount` voices and\nhas reached the current subscription plan limit of `voiceSlotLimit`; offer an\nupgrade action and a close/continue action. Do not route a quota error through\nthe free-account paywall copy.\n\n## Sound Effects\n\nFor ordinary editing sound effects (SFX), do **not** generate first. Use the\nbuilt-in Sound Effects library before spending credits:\n\n1. Call `browse_library` with `category:\"sound-effects\"` and a query such as\n   `\"whoosh\"`, `\"camera shutter\"`, `\"notification\"`, `\"censor beep\"`, or\n   `\"record scratch\"`.\n2. Inspect the returned `library:sound:<id>`.\n3. Place it with `edit_item`, using `fromFrame` as the sound's\n   anchor/editorial moment frame:\n\n```ts\nmcp__core__browse_library({\n  category: \"sound-effects\",\n  query: \"short whoosh transition\",\n});\n\nmcp__core__edit_item({\n  adds: [\n    {\n      type: \"audio\",\n      assetId: \"library:sound:whoosh-short\",\n      fromFrame: 120,\n      trackId: \"A1\",\n    },\n  ],\n});\n```\n\nOnly generate sound effects from text descriptions with `submit_sound` when:\n\n- The user explicitly asks for a generated/original/custom sound.\n- The requested sound is too specific for the existing Sound Effects library.\n- `browse_library({ category:\"sound-effects\", query })` returns no suitable\n  match.\n\n```ts\n// Custom/generated sound effect after the library has no suitable match\nmcp__skill__submit_sound({ prompt: \"A dog barking in the distance\" });\n\n// With custom duration (0.5-22 seconds)\nmcp__skill__submit_sound({\n  prompt: \"Thunder and heavy rain\",\n  durationSeconds: 15,\n});\n\n// High prompt adherence\nmcp__skill__submit_sound({\n  prompt: \"Sci-fi laser gun firing\",\n  promptInfluence: 0.8,\n});\n```\n\n**Tips for better results:**\n\n- Be specific: \"A dog barking loudly\" vs just \"dog\"\n- Include context: \"Footsteps on wooden floor in an empty room\"\n- Specify style: \"Cinematic whoosh\" or \"8-bit game sound\"\n\n## Parameters\n\n### TTS\n\n| Field        | Description                             | Notes           |\n| ------------ | --------------------------------------- | --------------- |\n| `provider`   | `doubao`, `elevenlabs`, or `fish-audio` | Required        |\n| `text`       | Text to synthesize                      | Required        |\n| `voiceId`    | Curated id, or ChatCut custom voice id  | Required        |\n| `speedRatio` | Speech speed                            | Doubao only     |\n| `modelId`    | ElevenLabs model id                     | ElevenLabs only |\n| `stability`  | ElevenLabs stability                    | ElevenLabs only |\n| `speed`      | ElevenLabs speech speed                 | ElevenLabs only |\n| `name`       | Asset name                              | Optional        |\n\n### Custom voices\n\n| Action                | Required fields                                            | Result                                                |\n| --------------------- | ---------------------------------------------------------- | ----------------------------------------------------- |\n| `list`                | none                                                       | Entitlement, slot count, existing voices              |\n| `check-create-access` | none                                                       | Authoritative create gate before collecting reference |\n| `apply`               | `customVoiceId`                                            | Pro/readiness check; selects voice without generation |\n| `clone`               | `sourceAssetId`, `name`, `confirmedConsent`, `previewText` | Reusable Fish Audio voice + preview                   |\n| `preview`             | `customVoiceId`, `previewText`                             | Refreshed short audition                              |\n| `rename`              | `customVoiceId`, `name`                                    | Updated product/provider display name                 |\n| `delete`              | `customVoiceId`, `confirmedDelete:true`                    | Permanently deletes provider voice and frees the slot |\n\nDo not use or expose a separate `manage_fish_audio_voice` tool. Fish model IDs,\npublic/unlisted publishing, arbitrary tags, cover images, and provider-level\ncatalog management are internal integration details. Delete only after an\nexplicit user confirmation; use `rename` for a simple name change.\n\n### Sound Effects\n\n| Field             | Description       | Notes          |\n| ----------------- | ----------------- | -------------- |\n| `prompt`          | Sound description | Required       |\n| `durationSeconds` | Duration          | 0.5-22 seconds |\n| `promptInfluence` | Prompt adherence  | 0-1            |\n| `name`            | Asset name        | Optional       |\n\n## Voices\n\nUse `manage_voice action=\"list\" voiceType=\"all\"` for the current unified voice\ncandidate list, localized display names, previews, and submission ids.\n\n### Voice presets are provider-specific — do NOT mix them\n\nOfficial voice ids are provider-specific. Always keep the returned `provider`\nand `voiceId` together; mixing values from different catalog tuples will fail.\n\nIf you need a specific voice and a particular language:\n\n- For any narration language, query the catalog with that `language` and use a\n  returned provider/voice pair. Do not guess a raw provider id.\n\n## Hard rules — what you must NOT do\n\n1. Never use a voice preset name from a different provider.\n2. Never render official voice recommendations before the mandatory custom\n   voice `list` gate when no concrete voice was named.\n3. Never use explicit or inferred voice traits to exclude a ready custom voice\n   with a playable preview from the combined audition surface.\n4. Never omit the final clone action card from a voice recommendation or\n   selection surface.\n5. Never submit TTS when the voice is only described broadly and the user has\n   not confirmed a concrete preset.\n6. Never recommend or render a TTS voice option before calling `manage_voice\naction=\"list\" voiceType=\"all\"` for the current request.\n7. Never claim stable age, regional accent, pronunciation dictionary, or exact\n   duration controls; the current tool does not expose those as reliable fields.\n8. Never replace original recorded speech with TTS unless the user asks.\n9. Never import or clone a reference voice until the user has affirmatively\n   submitted the full ownership/permission and lawful-use commitment required\n   by the cloning intake. Never infer this confirmation from an attachment,\n   another completed field, a generic approval, or prior unrelated context.\n10. Never bypass cloned-voice Pro checks by passing a provider voice id to\n    `submit_voice` or another generation tool.\n11. Never force the native clone dialog or ask the user to reselect audio when\n    the user has already supplied a usable reference for the current cloning\n    request. Preflight access, collect the missing authorization, name, and\n    preview text, then clone from its ChatCut audio asset id.\n12. Never treat unrelated speech already present in the project as a cloning\n    reference or as permission to clone it.\n13. Never render a cloned-voice preview as raw `<audio>` HTML, a Markdown link,\n    or a bare URL. Use `widget-forms` `playable_preview` with the exact\n    tool-returned `previewUrl`; keep Retry / Apply in the separate branch\n    control required by the current host.\n14. Never emit `<clone-voice/>` in an intermediate message or more than once in\n    one Agent run. Complete required tools first; if the native dialog remains\n    the correct route, follow its dialog-entry sequence and then stop.\n15. Never expose Doubao, ElevenLabs, Fish Audio, provider model names,\n    provider-specific ids, or provider URLs to the user. Use `official voice`\n    and `cloned voice` product terminology, and sanitize provider details from\n    visible errors and status messages.\n"
}

SHA-256: 2299ac747a611c664d341ac17dba5add8794791ae66d5e9e93117612b339be99