Update to Creative Claw
Snapshot Oct 3, 2026 · 06:03 UTC · version 5.3.0
Collection source: downloaded plugin package. These snapshots do not have a confirmed matching collection source. Differences in file lists alone do not establish changes to the package.
Instructions updated for creativeclaw-generate-voiceover
Instruction wording changed from “Generate narration, dialogue, expressive speech, or a new synthetic voice with Creative Claw. Use for spoken audio, a saved Character voice, or voice selection. For cloning a personal voice from a recording, load creativeclaw-clone-voice.” to “"Generate narration, dialogue, or expressive speech with Creative Claw, or design a new custom synthetic voice from a description. Use for spoken audio, voice selection, a saved Character voice, or 'make me a new voice'. To copy a real p...”. 16 additional added or edited lines are in the evidence.
Observed in instructions or declared skills. Runtime behavior has not been tested.
Product description
expressive speech, or a new synthetic voice with Creative Claw. Use for spoken audio, a saved Character voice, or voice selection. For cloning a personal voice from a recording, load
or expressive speech with Creative Claw, or design a new custom synthetic voice from a description. Use for spoken audio, voice selection, a saved Character voice, or 'make me a new voice'. To copy a real person's voice from a recording,...
Skill instructions
Generate narration, dialogue, expressive speech, or a new synthetic voice with Creative Claw. Use for spoken audio, a saved Character voice, or voice selection. For cloning a personal voice from a recording, load creativeclaw-clone-voice...
"Generate narration, dialogue, or expressive speech with Creative Claw, or design a new custom synthetic voice from a description. Use for spoken audio, voice selection, a saved Character voice, or 'make me a new voice'. To copy a real p...
Supporting files
[{"relative_path":"agents/openai.yaml","size_in_bytes":513},{"relative_path":"references/job-recovery.md","size_in_bytes":3786},{"relative_path":"references/media-assembly.md","size_in_bytes":3876},{"relative_path":"references/platform-u...
[{"relative_path":"agents/openai.yaml","size_in_bytes":525},{"relative_path":"references/job-recovery.md","size_in_bytes":3786},{"relative_path":"references/media-assembly.md","size_in_bytes":8654},{"relative_path":"references/platform-u...
Compare saved observations
Download comparison JSONFull technical diff · 3 changed fields
changed /description
"Generate narration, dialogue, expressive speech, or a new synthetic voice with Creative Claw. Use for spoken audio, a saved Character voice, or voice selection. For cloning a personal voice from a recording, load creativeclaw-clone-voice."
"Generate narration, dialogue, or expressive speech with Creative Claw, or design a new custom synthetic voice from a description. Use for spoken audio, voice selection, a saved Character voice, or 'make me a new voice'. To copy a real person's voice from a recording, use creativeclaw-clone-voice."
changed /included_files
[
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 513
},
{
"relative_path": "references/job-recovery.md",
"size_in_bytes": 3786
},
{
"relative_path": "references/media-assembly.md",
"size_in_bytes": 3876
},
{
"relative_path": "references/platform-upload.md",
"size_in_bytes": 1416
},
{
"relative_path": "references/voices/alternatives.md",
"size_in_bytes": 1799
},
{
"relative_path": "references/voices/cartesia-sonic-guide.md",
"size_in_bytes": 4950
},
{
"relative_path": "references/voices/cartesia.md",
"size_in_bytes": 2200
},
{
"relative_path": "references/voices/chatterbox-guide.md",
"size_in_bytes": 3410
},
{
"relative_path": "references/voices/cloning.md",
"size_in_bytes": 4020
},
{
"relative_path": "references/voices/elevenlabs-v2-guide.md",
"size_in_bytes": 7975
},
{
"relative_path": "references/voices/elevenlabs-v2.md",
"size_in_bytes": 1784
},
{
"relative_path": "references/voices/elevenlabs-v3-guide.md",
"size_in_bytes": 9349
},
{
"relative_path": "references/voices/elevenlabs-v3.md",
"size_in_bytes": 1795
},
{
"relative_path": "references/voices/languages.md",
"size_in_bytes": 2776
},
{
"relative_path": "references/voices/minimax-speech-guide.md",
"size_in_bytes": 3453
},
{
"relative_path": "references/voices/xai-tts-guide.md",
"size_in_bytes": 11625
},
{
"relative_path": "references/workflow-basics.md",
"size_in_bytes": 7159
}
][
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 525
},
{
"relative_path": "references/job-recovery.md",
"size_in_bytes": 3786
},
{
"relative_path": "references/media-assembly.md",
"size_in_bytes": 8654
},
{
"relative_path": "references/platform-upload.md",
"size_in_bytes": 1416
},
{
"relative_path": "references/video/voice-in-video.md",
"size_in_bytes": 3328
},
{
"relative_path": "references/voices/alternatives.md",
"size_in_bytes": 1807
},
{
"relative_path": "references/voices/cartesia-sonic-guide.md",
"size_in_bytes": 5007
},
{
"relative_path": "references/voices/cartesia.md",
"size_in_bytes": 2239
},
{
"relative_path": "references/voices/chatterbox-guide.md",
"size_in_bytes": 3493
},
{
"relative_path": "references/voices/cloning.md",
"size_in_bytes": 4372
},
{
"relative_path": "references/voices/elevenlabs-v2-guide.md",
"size_in_bytes": 8088
},
{
"relative_path": "references/voices/elevenlabs-v2.md",
"size_in_bytes": 1792
},
{
"relative_path": "references/voices/elevenlabs-v3-guide.md",
"size_in_bytes": 9402
},
{
"relative_path": "references/voices/elevenlabs-v3.md",
"size_in_bytes": 1938
},
{
"relative_path": "references/voices/elevenlabs-v4-guide.md",
"size_in_bytes": 8693
},
{
"relative_path": "references/voices/elevenlabs-v4.md",
"size_in_bytes": 2069
},
{
"relative_path": "references/voices/languages.md",
"size_in_bytes": 2853
},
{
"relative_path": "references/voices/minimax-speech-guide.md",
"size_in_bytes": 3449
},
{
"relative_path": "references/voices/voice-design.md",
"size_in_bytes": 2243
},
{
"relative_path": "references/voices/xai-tts-guide.md",
"size_in_bytes": 11625
},
{
"relative_path": "references/workflow-basics.md",
"size_in_bytes": 7260
}
]changed /skill_md_contents
"---\nname: creativeclaw-generate-voiceover\ndescription: Generate narration, dialogue, expressive speech, or a new synthetic voice with Creative Claw. Use for spoken audio, a saved Character voice, or voice selection. For cloning a personal voice from a recording, load creativeclaw-clone-voice.\n---\n\n# Generate voiceover\n\nRead [shared execution guidance](references/workflow-basics.md) before tools. The speech tool is `generate_speech`.\n\n## Personal voice or stock voice\n\n- \"Clone my voice\", \"use my recording\", or an unsaved personal voice: load $creativeclaw-clone-voice when available. If this skill is installed alone, read [the complete cloning workflow](references/voices/cloning.md). Obtain explicit consent, import privately, clone, test and save a reusable Character. Do not send source audio directly to ElevenLabs or Cartesia speech generation as a substitute for cloning.\n- Existing clone: find the intended Character and pass `character_id`. Do not clone again.\n- Stock voice: fetch the selected model's current `get_model_params` voice catalog and use its exact `voice_id`. Never pass both selectors.\n- Designing a new synthetic voice is different from cloning a person. Use the available voice-design capability and current schema when that is what the user requests.\n\n## Choose and load one model guide\n\n| Need | Recommended starting point |\n| --- | --- |\n| General stock narration, expressive speech, broad language coverage | [ElevenLabs v3](references/voices/elevenlabs-v3.md) |\n| Steady narration from an existing clone | [ElevenLabs v2](references/voices/elevenlabs-v2.md) |\n| Fast natural stock or cloned speech | [Cartesia Sonic](references/voices/cartesia.md) |\n| A named alternative or a specific dialect/voice match | [MiniMax and xAI](references/voices/alternatives.md) |\n\nThese are task-based defaults, not a universal quality ranking. V2 is clone-only. If a user requests v2 without a clone, explain and offer v3 stock speech or the consented cloning workflow, without silently switching. Do not automatically route stock corporate or long-form narration to v2.\n\nFor non-English, mixed-language or less common languages, read [language routing](references/voices/languages.md) before selecting a voice. A model supporting a language does not guarantee every stock voice has a native accent.\n\n## Execute and deliver\n\n1. Establish text, target language/accent and delivery. Preserve supplied words unless rewriting is requested.\n2. Read the selected model's reference and current `get_model_params`. Reuse a current schema already loaded in this task. Use `list_models({ category: \"speech\" })` only when discovering alternatives.\n3. Select a language-appropriate stock voice or saved Character and only settings accepted by that model. Never copy prompting tags across models.\n4. Generate the requested take. For a new clone or uncertain pronunciation, offer a short audition before a long production, not an unrequested paid model sweep.\n5. Follow queued results with `check_job` as directed by [job recovery](references/job-recovery.md). Show the completed audio through the available native preview.\n6. Keep the Character ID and voice/model choice available for the requested continuation. Speech added to a video is an audio overlay, not automatic lip synchronization. A video's Character reference does not automatically select the saved clone for native video audio.\n""---\nname: creativeclaw-generate-voiceover\ndescription: \"Generate narration, dialogue, or expressive speech with Creative Claw, or design a new custom synthetic voice from a description. Use for spoken audio, voice selection, a saved Character voice, or 'make me a new voice'. To copy a real person's voice from a recording, use creativeclaw-clone-voice.\"\n---\n\n# Generate voiceover\n\nRead [shared execution guidance](references/workflow-basics.md) before tools. The speech tool is `generate_speech`.\n\n## Which voice\n\n- \"Clone my voice\", \"use my recording\", or an unsaved personal voice: load creativeclaw-clone-voice when available. If this skill is installed alone, read [the complete cloning workflow](references/voices/cloning.md). Obtain explicit consent, import privately, clone, test and save a reusable Character. Do not send source audio directly to ElevenLabs or Cartesia speech generation as a substitute for cloning.\n- Existing Character voice (cloned, designed or stock): find the Character with `list_characters` and pass `character_id`. Do not clone or design again.\n- Stock voice: fetch the selected model's current `get_model_params` voice catalog and use its exact `voice_id`. For more Cartesia or ElevenLabs choices, call `get_model_params` again with `include_voice_catalog: true`, or read the public Markdown catalog linked in `voiceCatalog.extendedCatalogMarkdownUrl`. Never pass both selectors.\n- Change who is speaking in an existing speech recording while keeping the performance: call `generate_speech` with `model: \"speech/cartesia-voice-changer\"`, `audio_url` = the workspace speech audio, and a Cartesia `voice_id` or a Character's `character_id`. Pass no `text`.\n- New voice from a description (no recording): call `design_voice({ prompt })` with age, accent, timbre, pace and attitude, never a real person's identity. It returns three auditions; after the user picks, `design_voice({ action: \"save\", preview_id, character_id | character_name })`. Speak with `generate_speech({ character_id })` using `speech/elevenlabs-v4`, or `speech/gemini-3.8-flash-tts` for `provider: \"google\"`. If the save hits the limit, offer `replace_voice_option_id`; never quote plan prices. For a stock voice use `manage_character({ id, voice_model, voice_id })`. Details: [voice design](references/voices/voice-design.md).\n\n## Choose and load one model guide\n\n| Need | Recommended starting point |\n| --- | --- |\n| Stock narration, designed ElevenLabs voices, expressive speech, broad language coverage, and dialogue | [ElevenLabs v4](references/voices/elevenlabs-v4.md) |\n| Google-designed voice | `speech/gemini-3.8-flash-tts` (read its `get_model_params`) |\n| A Cartesia clone | [Cartesia Sonic](references/voices/cartesia.md) |\n| An ElevenLabs clone, steady read | [ElevenLabs v2](references/voices/elevenlabs-v2.md) |\n| An ElevenLabs clone, expressive delivery or dialogue | [ElevenLabs v4](references/voices/elevenlabs-v4.md) |\n| An explicitly requested v3 workflow | [ElevenLabs v3](references/voices/elevenlabs-v3.md) |\n| A named alternative or a specific dialect/voice match | [MiniMax and xAI](references/voices/alternatives.md) |\n\nDefault to ElevenLabs v4 for stock and designed ElevenLabs voices. Speak a clone with its provider's model: Cartesia Sonic for a Cartesia clone; v2 for a steady read or v4 for expressive delivery or dialogue from an ElevenLabs clone. Pass the model explicitly rather than relying on the omitted-model v4 default, and keep it for the whole project. V3 is legacy; use it only when the user asks for it. V2 can use public stock IDs, but v4 is the stock choice. If a user explicitly requests v2 stock speech, use a compatible public `voice_id` instead of silently switching. Do not automatically route stock corporate or long-form narration to v2.\n\n## Multiple speakers in one run\n\nElevenLabs v4 supports up to 10 voices in one `generate_speech` request through `extras.dialogue`. Pass ordered turns with `speaker`, `text`, and either `voice_id` or a saved Character voice (cloned or designed) `character_id` for each turn. Google Flash and Flash-Lite TTS also support two-speaker dialogue. Read the selected model's current schema and [v4 dialogue examples](references/voices/elevenlabs-v4-guide.md) before submission. Omit top-level voice selectors and `text` for dialogue. If a cached OpenAI tool schema still requires `text`, pass `text: \"\"`.\n\nFor non-English, mixed-language or less common languages, read [language routing](references/voices/languages.md) before selecting a voice. A model supporting a language does not guarantee every stock voice has a native accent.\n\n## Execute and deliver\n\n1. Establish text, target language/accent and delivery. Preserve supplied words unless rewriting is requested.\n2. Read the selected model's reference and current `get_model_params`. Reuse a current schema already loaded in this task. Use `list_models({ category: \"speech\" })` only when discovering alternatives.\n3. Select a language-appropriate stock voice or saved Character and only settings accepted by that model. Never copy prompting tags across models.\n4. Generate the requested take. For a new clone or uncertain pronunciation, offer a short audition before a long production, not an unrequested paid model sweep.\n5. Follow queued results with `check_job` as directed by [job recovery](references/job-recovery.md). Show the completed audio through the available native preview.\n6. Keep the Character ID and voice/model choice available for the requested continuation. Speech added to a video is an audio overlay, not lip-sync, and a video's Character reference does not bring its voice into native video audio. For speech in video, read [voice in video](references/video/voice-in-video.md).\n"SKILL.md line diff
--- before +++ after @@ -1,29 +1,37 @@ --- name: creativeclaw-generate-voiceover -description: Generate narration, dialogue, expressive speech, or a new synthetic voice with Creative Claw. Use for spoken audio, a saved Character voice, or voice selection. For cloning a personal voice from a recording, load creativeclaw-clone-voice. +description: "Generate narration, dialogue, or expressive speech with Creative Claw, or design a new custom synthetic voice from a description. Use for spoken audio, voice selection, a saved Character voice, or 'make me a new voice'. To copy a real person's voice from a recording, use creativeclaw-clone-voice." --- # Generate voiceover Read [shared execution guidance](references/workflow-basics.md) before tools. The speech tool is `generate_speech`. -## Personal voice or stock voice +## Which voice -- "Clone my voice", "use my recording", or an unsaved personal voice: load $creativeclaw-clone-voice when available. If this skill is installed alone, read [the complete cloning workflow](references/voices/cloning.md). Obtain explicit consent, import privately, clone, test and save a reusable Character. Do not send source audio directly to ElevenLabs or Cartesia speech generation as a substitute for cloning. -- Existing clone: find the intended Character and pass `character_id`. Do not clone again. -- Stock voice: fetch the selected model's current `get_model_params` voice catalog and use its exact `voice_id`. Never pass both selectors. -- Designing a new synthetic voice is different from cloning a person. Use the available voice-design capability and current schema when that is what the user requests. +- "Clone my voice", "use my recording", or an unsaved personal voice: load creativeclaw-clone-voice when available. If this skill is installed alone, read [the complete cloning workflow](references/voices/cloning.md). Obtain explicit consent, import privately, clone, test and save a reusable Character. Do not send source audio directly to ElevenLabs or Cartesia speech generation as a substitute for cloning. +- Existing Character voice (cloned, designed or stock): find the Character with `list_characters` and pass `character_id`. Do not clone or design again. +- Stock voice: fetch the selected model's current `get_model_params` voice catalog and use its exact `voice_id`. For more Cartesia or ElevenLabs choices, call `get_model_params` again with `include_voice_catalog: true`, or read the public Markdown catalog linked in `voiceCatalog.extendedCatalogMarkdownUrl`. Never pass both selectors. +- Change who is speaking in an existing speech recording while keeping the performance: call `generate_speech` with `model: "speech/cartesia-voice-changer"`, `audio_url` = the workspace speech audio, and a Cartesia `voice_id` or a Character's `character_id`. Pass no `text`. +- New voice from a description (no recording): call `design_voice({ prompt })` with age, accent, timbre, pace and attitude, never a real person's identity. It returns three auditions; after the user picks, `design_voice({ action: "save", preview_id, character_id | character_name })`. Speak with `generate_speech({ character_id })` using `speech/elevenlabs-v4`, or `speech/gemini-3.8-flash-tts` for `provider: "google"`. If the save hits the limit, offer `replace_voice_option_id`; never quote plan prices. For a stock voice use `manage_character({ id, voice_model, voice_id })`. Details: [voice design](references/voices/voice-design.md). ## Choose and load one model guide | Need | Recommended starting point | | --- | --- | -| General stock narration, expressive speech, broad language coverage | [ElevenLabs v3](references/voices/elevenlabs-v3.md) | -| Steady narration from an existing clone | [ElevenLabs v2](references/voices/elevenlabs-v2.md) | -| Fast natural stock or cloned speech | [Cartesia Sonic](references/voices/cartesia.md) | +| Stock narration, designed ElevenLabs voices, expressive speech, broad language coverage, and dialogue | [ElevenLabs v4](references/voices/elevenlabs-v4.md) | +| Google-designed voice | `speech/gemini-3.8-flash-tts` (read its `get_model_params`) | +| A Cartesia clone | [Cartesia Sonic](references/voices/cartesia.md) | +| An ElevenLabs clone, steady read | [ElevenLabs v2](references/voices/elevenlabs-v2.md) | +| An ElevenLabs clone, expressive delivery or dialogue | [ElevenLabs v4](references/voices/elevenlabs-v4.md) | +| An explicitly requested v3 workflow | [ElevenLabs v3](references/voices/elevenlabs-v3.md) | | A named alternative or a specific dialect/voice match | [MiniMax and xAI](references/voices/alternatives.md) | -These are task-based defaults, not a universal quality ranking. V2 is clone-only. If a user requests v2 without a clone, explain and offer v3 stock speech or the consented cloning workflow, without silently switching. Do not automatically route stock corporate or long-form narration to v2. +Default to ElevenLabs v4 for stock and designed ElevenLabs voices. Speak a clone with its provider's model: Cartesia Sonic for a Cartesia clone; v2 for a steady read or v4 for expressive delivery or dialogue from an ElevenLabs clone. Pass the model explicitly rather than relying on the omitted-model v4 default, and keep it for the whole project. V3 is legacy; use it only when the user asks for it. V2 can use public stock IDs, but v4 is the stock choice. If a user explicitly requests v2 stock speech, use a compatible public `voice_id` instead of silently switching. Do not automatically route stock corporate or long-form narration to v2. + +## Multiple speakers in one run + +ElevenLabs v4 supports up to 10 voices in one `generate_speech` request through `extras.dialogue`. Pass ordered turns with `speaker`, `text`, and either `voice_id` or a saved Character voice (cloned or designed) `character_id` for each turn. Google Flash and Flash-Lite TTS also support two-speaker dialogue. Read the selected model's current schema and [v4 dialogue examples](references/voices/elevenlabs-v4-guide.md) before submission. Omit top-level voice selectors and `text` for dialogue. If a cached OpenAI tool schema still requires `text`, pass `text: ""`. For non-English, mixed-language or less common languages, read [language routing](references/voices/languages.md) before selecting a voice. A model supporting a language does not guarantee every stock voice has a native accent. @@ -34,4 +42,4 @@ 3. Select a language-appropriate stock voice or saved Character and only settings accepted by that model. Never copy prompting tags across models. 4. Generate the requested take. For a new clone or uncertain pronunciation, offer a short audition before a long production, not an unrequested paid model sweep. 5. Follow queued results with `check_job` as directed by [job recovery](references/job-recovery.md). Show the completed audio through the available native preview. -6. Keep the Character ID and voice/model choice available for the requested continuation. Speech added to a video is an audio overlay, not automatic lip synchronization. A video's Character reference does not automatically select the saved clone for native video audio. +6. Keep the Character ID and voice/model choice available for the requested continuation. Speech added to a video is an audio overlay, not lip-sync, and a video's Character reference does not bring its voice into native video audio. For speech in video, read [voice in video](references/video/voice-in-video.md).
Full snapshot data
{
"description": "Generate narration, dialogue, or expressive speech with Creative Claw, or design a new custom synthetic voice from a description. Use for spoken audio, voice selection, a saved Character voice, or 'make me a new voice'. To copy a real person's voice from a recording, use creativeclaw-clone-voice.",
"included_files": [
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 525
},
{
"relative_path": "references/job-recovery.md",
"size_in_bytes": 3786
},
{
"relative_path": "references/media-assembly.md",
"size_in_bytes": 8654
},
{
"relative_path": "references/platform-upload.md",
"size_in_bytes": 1416
},
{
"relative_path": "references/video/voice-in-video.md",
"size_in_bytes": 3328
},
{
"relative_path": "references/voices/alternatives.md",
"size_in_bytes": 1807
},
{
"relative_path": "references/voices/cartesia-sonic-guide.md",
"size_in_bytes": 5007
},
{
"relative_path": "references/voices/cartesia.md",
"size_in_bytes": 2239
},
{
"relative_path": "references/voices/chatterbox-guide.md",
"size_in_bytes": 3493
},
{
"relative_path": "references/voices/cloning.md",
"size_in_bytes": 4372
},
{
"relative_path": "references/voices/elevenlabs-v2-guide.md",
"size_in_bytes": 8088
},
{
"relative_path": "references/voices/elevenlabs-v2.md",
"size_in_bytes": 1792
},
{
"relative_path": "references/voices/elevenlabs-v3-guide.md",
"size_in_bytes": 9402
},
{
"relative_path": "references/voices/elevenlabs-v3.md",
"size_in_bytes": 1938
},
{
"relative_path": "references/voices/elevenlabs-v4-guide.md",
"size_in_bytes": 8693
},
{
"relative_path": "references/voices/elevenlabs-v4.md",
"size_in_bytes": 2069
},
{
"relative_path": "references/voices/languages.md",
"size_in_bytes": 2853
},
{
"relative_path": "references/voices/minimax-speech-guide.md",
"size_in_bytes": 3449
},
{
"relative_path": "references/voices/voice-design.md",
"size_in_bytes": 2243
},
{
"relative_path": "references/voices/xai-tts-guide.md",
"size_in_bytes": 11625
},
{
"relative_path": "references/workflow-basics.md",
"size_in_bytes": 7260
}
],
"name": "creativeclaw-generate-voiceover",
"skill_md_contents": "---\nname: creativeclaw-generate-voiceover\ndescription: \"Generate narration, dialogue, or expressive speech with Creative Claw, or design a new custom synthetic voice from a description. Use for spoken audio, voice selection, a saved Character voice, or 'make me a new voice'. To copy a real person's voice from a recording, use creativeclaw-clone-voice.\"\n---\n\n# Generate voiceover\n\nRead [shared execution guidance](references/workflow-basics.md) before tools. The speech tool is `generate_speech`.\n\n## Which voice\n\n- \"Clone my voice\", \"use my recording\", or an unsaved personal voice: load creativeclaw-clone-voice when available. If this skill is installed alone, read [the complete cloning workflow](references/voices/cloning.md). Obtain explicit consent, import privately, clone, test and save a reusable Character. Do not send source audio directly to ElevenLabs or Cartesia speech generation as a substitute for cloning.\n- Existing Character voice (cloned, designed or stock): find the Character with `list_characters` and pass `character_id`. Do not clone or design again.\n- Stock voice: fetch the selected model's current `get_model_params` voice catalog and use its exact `voice_id`. For more Cartesia or ElevenLabs choices, call `get_model_params` again with `include_voice_catalog: true`, or read the public Markdown catalog linked in `voiceCatalog.extendedCatalogMarkdownUrl`. Never pass both selectors.\n- Change who is speaking in an existing speech recording while keeping the performance: call `generate_speech` with `model: \"speech/cartesia-voice-changer\"`, `audio_url` = the workspace speech audio, and a Cartesia `voice_id` or a Character's `character_id`. Pass no `text`.\n- New voice from a description (no recording): call `design_voice({ prompt })` with age, accent, timbre, pace and attitude, never a real person's identity. It returns three auditions; after the user picks, `design_voice({ action: \"save\", preview_id, character_id | character_name })`. Speak with `generate_speech({ character_id })` using `speech/elevenlabs-v4`, or `speech/gemini-3.8-flash-tts` for `provider: \"google\"`. If the save hits the limit, offer `replace_voice_option_id`; never quote plan prices. For a stock voice use `manage_character({ id, voice_model, voice_id })`. Details: [voice design](references/voices/voice-design.md).\n\n## Choose and load one model guide\n\n| Need | Recommended starting point |\n| --- | --- |\n| Stock narration, designed ElevenLabs voices, expressive speech, broad language coverage, and dialogue | [ElevenLabs v4](references/voices/elevenlabs-v4.md) |\n| Google-designed voice | `speech/gemini-3.8-flash-tts` (read its `get_model_params`) |\n| A Cartesia clone | [Cartesia Sonic](references/voices/cartesia.md) |\n| An ElevenLabs clone, steady read | [ElevenLabs v2](references/voices/elevenlabs-v2.md) |\n| An ElevenLabs clone, expressive delivery or dialogue | [ElevenLabs v4](references/voices/elevenlabs-v4.md) |\n| An explicitly requested v3 workflow | [ElevenLabs v3](references/voices/elevenlabs-v3.md) |\n| A named alternative or a specific dialect/voice match | [MiniMax and xAI](references/voices/alternatives.md) |\n\nDefault to ElevenLabs v4 for stock and designed ElevenLabs voices. Speak a clone with its provider's model: Cartesia Sonic for a Cartesia clone; v2 for a steady read or v4 for expressive delivery or dialogue from an ElevenLabs clone. Pass the model explicitly rather than relying on the omitted-model v4 default, and keep it for the whole project. V3 is legacy; use it only when the user asks for it. V2 can use public stock IDs, but v4 is the stock choice. If a user explicitly requests v2 stock speech, use a compatible public `voice_id` instead of silently switching. Do not automatically route stock corporate or long-form narration to v2.\n\n## Multiple speakers in one run\n\nElevenLabs v4 supports up to 10 voices in one `generate_speech` request through `extras.dialogue`. Pass ordered turns with `speaker`, `text`, and either `voice_id` or a saved Character voice (cloned or designed) `character_id` for each turn. Google Flash and Flash-Lite TTS also support two-speaker dialogue. Read the selected model's current schema and [v4 dialogue examples](references/voices/elevenlabs-v4-guide.md) before submission. Omit top-level voice selectors and `text` for dialogue. If a cached OpenAI tool schema still requires `text`, pass `text: \"\"`.\n\nFor non-English, mixed-language or less common languages, read [language routing](references/voices/languages.md) before selecting a voice. A model supporting a language does not guarantee every stock voice has a native accent.\n\n## Execute and deliver\n\n1. Establish text, target language/accent and delivery. Preserve supplied words unless rewriting is requested.\n2. Read the selected model's reference and current `get_model_params`. Reuse a current schema already loaded in this task. Use `list_models({ category: \"speech\" })` only when discovering alternatives.\n3. Select a language-appropriate stock voice or saved Character and only settings accepted by that model. Never copy prompting tags across models.\n4. Generate the requested take. For a new clone or uncertain pronunciation, offer a short audition before a long production, not an unrequested paid model sweep.\n5. Follow queued results with `check_job` as directed by [job recovery](references/job-recovery.md). Show the completed audio through the available native preview.\n6. Keep the Character ID and voice/model choice available for the requested continuation. Speech added to a video is an audio overlay, not lip-sync, and a video's Character reference does not bring its voice into native video audio. For speech in video, read [voice in video](references/video/voice-in-video.md).\n"
}SHA-256 of public snapshot: 8f7c037ec968a54930d4e56565456a03e34f4631a065e403b022161c65f74232