← ChatCutCONTENT HISTORY

Update to ChatCut

Snapshot Sep 30, 2026 · 23:14 UTC · version 1.10.14

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "video-translation",
  "description": "Translate, dub, and localize speech in an existing video while optionally preserving speaker voices, generating translated captions, or synchronizing the speaker's lip movements. Use when the user asks for 视频译制、多语言配音、 把视频里的中文变成英文、让视频里的人说另一种语言、保留原音色、翻译声音、 口型同步、translated video, video dubbing, voice translation, or lip-synced localization, as well as traducción de video, doblaje de video, traducir un video, or sincronización labial. This Skill MUST be loaded before composing treatment choices for an ambiguous video-translation request in any language, including a turn that only asks the user to choose a treatment. Do not use when the user only wants subtitles translated, wants an SRT/VTT file, wants text or a script translated before generating a new avatar video, or wants ordinary TTS with a newly selected voice.",
  "included_files": [],
  "skill_md_contents": "---\nname: video-translation\ndescription: Translate, dub, and localize speech in an existing video while optionally preserving speaker voices, generating translated captions, or synchronizing the speaker's lip movements. Use when the user asks for 视频译制、多语言配音、 把视频里的中文变成英文、让视频里的人说另一种语言、保留原音色、翻译声音、 口型同步、translated video, video dubbing, voice translation, or lip-synced localization, as well as traducción de video, doblaje de video, traducir un video, or sincronización labial. This Skill MUST be loaded before composing treatment choices for an ambiguous video-translation request in any language, including a turn that only asks the user to choose a treatment. Do not use when the user only wants subtitles translated, wants an SRT/VTT file, wants text or a script translated before generating a new avatar video, or wants ordinary TTS with a newly selected voice.\nuser-invocable: true\n---\n\n# Video Translation\n\nCreate a new localized video asset from an existing project video. Treat\nspeech translation, dubbing, voice preservation, lip synchronization, captions,\nquality, authorization, generation, and verification as one workflow.\n\nLip-synced translation starts with source-transcript review in an editable\nChatCut Widget. The recognized source wording is never treated as final until\nthe user has edited or approved that form and submitted it. Audio-only\ntranslation keeps the existing direct-submit flow and skips this review.\n\n## When to Use\n\nUse this skill when the user wants an existing video to:\n\n- Make its speakers speak another language.\n- Produce a dubbed or localized version.\n- Preserve the original speakers' vocal identity across languages.\n- Synchronize visible mouth movements with translated speech.\n- Translate only the spoken audio while keeping the original picture.\n- Generate a translated video with optional translated captions.\n- Create a language-specific version for another market.\n\nTypical triggering requests include:\n\n- \"让视频里这个人说英语。\"\n- \"把中文口播做成英文版,口型也对上。\"\n- \"把这段采访译制成日语,保留每个人的声音。\"\n- \"做一个西班牙语配音版。\"\n- \"只把声音翻译成英语,画面不要变。\"\n- \"把这个中文数字人成片改成英文。\"\n\nRoute adjacent requests elsewhere:\n\n- Only translate, add, edit, or export captions: use caption translation.\n- Translate a script before generating a new avatar video: translate the text,\n  then use the Digital Human Skill.\n- Generate speech with a chosen replacement voice: use the Voice Skill.\n- Transcribe spoken content without changing it: use the Transcription Skill.\n- Edit pauses, mistakes, framing, or B-roll in talking-head footage: use the\n  Talking Head Guide.\n- Generate new footage rather than localize an existing video: use Video Gen.\n\nTreat these requests as ambiguous:\n\n- \"把这个视频翻译成英文。\"\n- \"帮我做一个英文版。\"\n- \"中文改成英语。\"\n- \"做一个海外版。\"\n- \"把这个视频国际化一下。\"\n\nFor an ambiguous request, ask one treatment question with text choices. Keep\nthe wording natural for the conversation, and include all applicable paths:\n\n- Lip-sync translation: continue through the editable source-transcript review\n  and submit with `audioOnly: false` or omit `audioOnly`.\n- Audio-only translation: keep the existing direct-submit path and submit with\n  `audioOnly: true`.\n- Subtitle-only translation: use caption translation and do not submit a video\n  translation job.\n\nKeep user-facing terminology in one language. These are localized concept names,\nnot fixed full option sentences; descriptions may stay natural for the context:\n\n| Conversation language | Lip-sync path                        | Audio-only path          | Subtitle-only path            |\n| --------------------- | ------------------------------------ | ------------------------ | ----------------------------- |\n| Chinese               | 口型同步翻译                         | 只翻译声音               | 只翻译字幕                    |\n| English               | Lip-synced translation               | Audio-only translation   | Subtitle-only translation     |\n| Spanish               | Traducción con sincronización labial | Traducción solo de audio | Traducción solo de subtítulos |\n\nThe target translation language does not control this copy; the user's\nconversation language does. Never append a second-language gloss such as\n`(lip-sync)` to a localized label.\n\nDo not omit lip-sync when the source has a visible speaker. Do not relabel\naudio-only translation as generic \"AI voice replacement\"; choosing a new\nsynthetic voice is a separate Voice Skill workflow. If the user already stated\none treatment, skip this question and follow that route directly.\n\nDo not submit a paid video-translation job until changing the spoken audio is\nexplicitly requested or confirmed.\n\n## Workflow\n\n1. **Resolve the source video.** It must be an imported project video asset;\n   use the exact asset id from the attachment, selection, or `browse_assets`.\n   Local or external media must be imported first.\n2. **Resolve the target language from the live catalog.** Call\n   `submit_video_translation` with `action: \"list_languages\"` and no other\n   arguments. This only reads the catalog; it requires no project, transcript\n   review, confirmation, or credits. Then copy the matching returned value\n   verbatim into `targetLanguage`, including any dialect, region,\n   or script qualifier. Do not invent a language name or send an ISO code.\n   Match the user's requested language and variant against the returned catalog;\n   confirm the language when the user only implied a market (\"海外版\"). If a\n   submission reports an unsupported language, use its suggested candidates or\n   refresh the catalog with `action: \"list_languages\"` instead of retrying\n   guessed names. The `action` argument is required on every call.\n3. **Pick the treatment:**\n   - Lip-sync: `mode: \"speed\"` — translated dubbing that preserves\n     the speakers' vocal identity, plus synchronized mouth movements. Continue\n     through the source-transcript review below.\n   - `audioOnly: true` — translate the audio only and keep the picture\n     untouched. Use for screen recordings, voice-over footage, or when the user\n     says the picture must not change. Prefer this for transparent WebM when\n     preserving the alpha channel matters. **Skip steps 4–6 and keep the existing\n     direct-submit flow; do not require or pass `reviewedSourceTranscript`.**\n4. **Prepare the source transcript for lip-sync only.** If word-level\n   transcription is not complete, call `trigger_transcript`, wait for it with\n   `track_progress`, and retry only when it is ready. Call `read_script`, then\n   read the matching `library/<filename>.md` source transcript. Do not use the\n   editable `timeline.md` cut as the translation source. **Do not substitute\n   `inspect_asset` transcript ranges for `read_script`; the range result may be\n   partial and is not the canonical complete source transcript.**\n5. **Render the editable source-text review for lip-sync only.** Load the\n   `widget-forms` Skill and reuse the same `<form-textarea>` confirmation pattern\n   used by the Digital Human Skill. Strip the library's `[sN]` addresses and\n   speaker-rendering rows, but keep **exactly one recognized transcript segment\n   per line**, in source order. Put that complete text in the textarea's\n   `default`; use a localized label that asks the user to check and correct the\n   recognized original text. Do not expose segment ids, word indices, file\n   syntax, or timestamps.\n\n   Hard preflight before emitting the Widget:\n   - The field id is exactly `reviewedSourceTranscript`.\n   - The prefill attribute is exactly `default`, never `default-value`,\n     `defaultValue`, `value`, or `placeholder`.\n   - `default` contains the actual complete non-empty recognized transcript,\n     not a placeholder or an omitted value. If the transcript text has not been\n     read successfully, do not render the Widget; read the canonical library\n     document first.\n   - Preserve one recognized source segment per line. Verify that the first and\n     last non-empty source segments are both present before sending the form.\n   - In the embedded raw-tag route, apply the `widget-forms` XML attribute\n     escaping rule to the complete transcript before placing it in `default`.\n     Never put unescaped recognized text inside the tag.\n\n   Embedded ChatCut example (localize visible copy):\n\n   ```text\n   Please review the recognized source text and correct any mistakes:\n\n   <widget>\n     <form-textarea id=\"reviewedSourceTranscript\" label=\"Please review and correct the source text\" rows=\"12\" required=\"true\" default=\"<recognized source text; one segment per line>\"/>\n   </widget>\n   ```\n\n   Follow `widget-forms` for the active host rather than emitting raw tags in a\n   host that does not support the embedded protocol. Do not add a separate\n   yes/no question or a handwritten submit button. **Stop here and wait. Never\n   submit a translation in the same turn that first presents the Widget.**\n\n6. **Use the submitted revision exactly.** The Widget submission is explicit\n   confirmation. Read the complete `reviewedSourceTranscript` textarea answer\n   from the user's next message and preserve its wording, punctuation, and\n   order exactly. Pass through the returned line breaks when present, but do\n   not reject or rewrite normal line-break edits made inside the textarea; the\n   tool realigns them to the source timing. Do not paraphrase the text and do\n   not ask the user to confirm the same text again. If the user replies outside\n   the Widget with further corrections, reopen the same editable Widget with\n   those corrections applied rather than reverting to a prose transcript.\n7. **Duration behavior.** By default the output may run slightly longer or\n   shorter than the source so the translated speech keeps a natural pace. Pass\n   `keepDuration: true` only when the user needs the exact original length\n   (for example to swap it into an existing timeline slot), and mention that\n   pacing may sound faster.\n8. **Submit** with `submit_video_translation` using `action: \"submit\"`. For lip-sync, pass\n   `reviewedSourceTranscript` as the exact complete textarea value returned in\n   step 6; the tool combines those reviewed lines with the original segment\n   timestamps and creates source subtitles. For `audioOnly: true`, omit\n   `reviewedSourceTranscript` and preserve the pre-existing direct-submit path.\n   The tool shows the user a paid confirmation before the job starts; do not\n   resubmit after a denial.\n9. **Track** with `track_progress(action=\"wait\")` until the translated video\n   asset lands in the library, then hand it back (place on the timeline only\n   when asked).\n\n## Constraints and cost\n\n- The whole source video is translated and billed by its full duration. To\n  localize only a section, trim/export that section into its own asset first,\n  then translate the shorter asset.\n- Optional translated captions: `enableCaption: true` burns subtitles into the\n  output video.\n- ChatCut always requests preservation of the source resolution and bitrate.\n  This is best-effort for lip-sync because the picture is re-rendered. A\n  transparent WebM may become opaque, and 4K or >30 fps media may be\n  re-encoded. The tool confirmation calls these cases out; never promise that\n  lip-sync will preserve the container, alpha channel, HDR, codec, bitrate, or\n  frame rate exactly.\n- When the picture must not be lip-synced, choose `audioOnly: true`; this avoids\n  facial re-rendering and is the safest treatment for transparent sources, but\n  do not promise byte-for-byte container preservation. After completion ChatCut\n  compares the source and output resolution, frame rate, and alpha pixel format\n  when media probing is available, and reports detected changes in generation\n  status.\n- Multiple speakers are supported; pass `speakerCount` when the user states it.\n- For lip-sync, the editable review starts with one source segment per line so\n  the submitted text can reuse the original word-level timing. Pass the complete\n  Widget value directly to the tool; do not merge, split, reorder, or silently\n  normalize it in the conversation layer. Do not invoke this review for\n  `audioOnly: true`.\n- Indicative cost: ≈ 8 credits per minute of source video.\n- Speech is dubbed with voices matched to the original speakers; it is a\n  translation of the recorded voice, so confirm the user has rights to the\n  footage and its speakers when the material is clearly someone else's.\n- Never reveal or discuss the underlying provider; present this as ChatCut's\n  AI video translation.\n"
}

SHA-256: c2481ffe682d2d4d117533a12dd85fcdc8b6cc0c4c9504efc635d03d5ef28289