← ChatCutCONTENT HISTORY

Update to ChatCut

Snapshot Sep 30, 2026 · 23:14 UTC · version 1.10.14

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "multicam-sync",
  "description": "Synchronize footage from a multi-camera / multi-recorder shoot — several cameras plus separate audio recorders covering one session, imported as loose clips — and optionally turn all or part of it into speaker-follow footage for a larger edit. Use when a user drops in multiple clips from the same recording and wants them aligned, asks for multicam / 多机位 / multi-angle sync, wants a \"cut to whoever is talking\" edit, or refers to camera A/B, angles, or separate lav or field recordings that need to line up with picture.",
  "included_files": [
    {
      "relative_path": "scripts/transcript-offset.mjs",
      "size_in_bytes": 7839
    }
  ],
  "skill_md_contents": "---\nname: multicam-sync\ndescription: Synchronize footage from a multi-camera / multi-recorder shoot — several cameras plus separate audio recorders covering one session, imported as loose clips — and optionally turn all or part of it into speaker-follow footage for a larger edit. Use when a user drops in multiple clips from the same recording and wants them aligned, asks for multicam / 多机位 / multi-angle sync, wants a \"cut to whoever is talking\" edit, or refers to camera A/B, angles, or separate lav or field recordings that need to line up with picture.\nuser-invocable: true\n---\n\n# Multicam Sync\n\nParts 1–3 always run: they produce the synced master, a verifiable fact that every\nlater edit is rebuilt from. Part 4 cuts a draft from it, and runs only on request.\nTranscripts do almost all of the work. AI locates; arithmetic decides frames.\n\nUse ordinary timelines, tracks, and items only. Never create a compound or nested\nmulticam clip, and never depend on a camera-switcher item. If a main edit already\nexists, leave it untouched while building the synced master and speaker-follow\ncutting board on separate timelines.\n\n## Part 1 — Discover the structure\n\nWork out what you actually have before aligning anything. Never ask the user how many\ncameras there are. The material already answers it.\n\n### The two rules that do the work\n\n**Same words at the same moment → cannot be sequential spans of one camera.** Overlapping\ntranscripts mean two devices recorded one event simultaneously, so they are different\nsources or simultaneous angles.\n\n**No shared words, and one ends where the next begins → one camera that stopped and\nrestarted.** These are sequential spans of a single capture source, not separate angles.\n\n### Steps\n\n1. Read every asset's transcript, build a pairwise overlap matrix — how much text\n   each pair shares, and where — and split by the rules above.\n2. Video assets are **angles**; audio-only assets are **recorders**.\n3. Corroborate with filenames, folder structure, camera-model metadata, durations —\n   tie-breakers only, never over transcript evidence. Real shoots ship three cameras\n   all named `C0001.MP4`.\n\n### Report before acting\n\nState the structure in the user's terms and wait if anything is ambiguous:\n\n> 2 camera angles (Pocket3, 3 spans · Pocket4, 2 spans), 2 audio recorders\n> (MIC1, MIC2). Session runs 53 minutes.\n\nWhen evidence is thin — a source with little speech, an overlap resting on a handful\nof matched lines — say so. Do not fill the gap with a guess.\n\n## Part 2 — Align and place\n\nOne number per clip: timeline time minus source time. Keep that number constant for\nevery piece cut from the clip. Do not time-stretch automatically: a changing offset\nmay be a bad match, a clock difference, or dropped frames, and each needs review.\n\n### Steps\n\n1. **Reference**: the source that covers the whole session with the richest transcript —\n   usually a dedicated audio recorder, not a camera.\n2. **Use the renderer when available**: on the separate master timeline, put the\n   untrimmed clips on ordinary source tracks, then call `multicam_sync` with the\n   same-take items and the reference. This is the preferred Web/Desktop path. It\n   uses existing source timing when decisive, otherwise audio correlation; it does\n   not create a multicam object.\n\n   Accept only `applied`, `already_synced`, or an understood `partial` result. Read\n   `alignmentEvidence` for method, match confidence, correlation, overlap, and\n   relative offset. Read `placementEvidence` for the actual post-sync\n   `timelineSourceOffsetSeconds`. Do not use a skipped or low-confidence item. If\n   the tool is unavailable or finds no confident alignment, leave the master\n   untouched and use the transcript fallback below.\n\n3. **Transcript fallback**: run\n   `scripts/transcript-offset.mjs <utterances.json>` from this skill. Do not\n   improvise the math. It first lets shared phrases vote on a coarse offset, then\n   takes the median of near-identical utterance-pair deltas. It also rejects too few\n   pairs, inconsistent pairs, and early/late drift. Only place files whose output is\n   `confident:true`; report every printed issue for the others. `--help` documents\n   the input and sign convention.\n4. **Continuity check** (free, run it): spans of one camera were recorded\n   back-to-back, so span N+1's offset minus span N's must equal span N's duration.\n   Each offset was measured independently, so agreement confirms both. Sub-frame\n   agreement is the norm; a multi-frame gap means dropped frames or a mismatch —\n   say which. **An overlap or gap between placed spans of one camera means an\n   offset is wrong. Recompute it — never trim a span to make it fit.**\n5. **Place transcript-fallback results**: shift everything so the earliest clip\n   starts at 0. One video track per camera, one audio track per recorder.\n   Convert seconds → frames once at the end, never accumulating, and **floor at\n   an asset's tail, never round up** — a\n   rounded-up final frame claims source the file doesn't have; renders tolerate it\n   silently, Script editing later refuses it.\n\n### Report\n\nPer clip: actual timeline-source offset, method and supporting evidence (renderer\nconfidence/correlation/overlap or transcript pair count/spread/drift), and the\ncontinuity-check result.\n\n## Part 3 — Identify and label\n\nTurn \"track 1 / track 2\" into who is actually on it. Everything here is evidence-first:\na wrong confident claim is a failure; an honest \"can't tell\" is not.\n\n### Who is in the session\n\nSpeakers usually name themselves or each other — self-introductions, banter. Pull real\nnames from the transcript. If none appear, A/B is fine; never invent names.\n\n### Which camera frames whom\n\nFrom the transcript, pick 3–4 moments where only one person is talking, spread across\nthe session. View the synced frame from **every** camera at the same wall-clock moment:\nthe talking face identifies the person; clothing and seating anchor identity across\nangles. Two-shots are identity anchors, not noise. Classify framing while you're\nthere: close-up / medium / two-shot; note reframing if the dominant framing changes.\n\nAngle labels are summaries, not promises: framing can drift, fail, then recover.\nInspect the exact synced interval before every cut; never choose from the label alone.\n\nHard check before concluding: at each self-introduction, the face whose lips are\nmoving is that name's owner. \"Both cameras frame the same person\" contradicts a\ntwo-camera, two-mic, two-voice structure — treat that conclusion as an error until\nframes at both introductions prove it.\n\n### Which mic belongs to whom\n\nA lav is dramatically louder for its wearer — typically 15 dB or more. Measure, don't\ninfer:\n\n- During one person's solo speech vs the other's, compare **the same track against\n  itself**. The in-track contrast cancels recorder gain. Above ~4 dB it decides;\n  below, say \"indistinct\" — that itself is a finding (ambient mic, shared mic).\n- Judge the **distribution**, not one moment: consistent → owned mic; 50/50 → not a\n  personal mic; flips mid-session → handheld passed around or seats changed.\n- **Never use utterance counts or transcript volume as evidence of mic ownership.**\n  Crosstalk transcribes fine; both speakers appear fully on both mics' transcripts.\n  The overall loudness difference between two mics is not evidence either — only\n  the in-track solo-vs-solo contrast is.\n- Fragmented diarization heals mechanically: cluster speaker-ids by their median\n  cross-track level difference. Never hand-reconcile speaker ids across assets —\n  they are per-asset serials.\n\n### Which source is the program audio\n\nDecide what the viewer will hear. Default: dedicated recorders beat camera embedded\naudio. The user's word beats everything — \"camera A has the good audio\" is a program\naudio assignment; that camera's audio track then behaves exactly like a recorder\n(its offset is already known from Part 2). Record the assignment in the report.\n\nWith no clean recorder, compare camera mics at the same solo-speech moments. Prefer\none continuous source that is good enough; switch for a speaker or passage only when\nanother is clearly better for a sustained stretch and the handoff is inaudible.\nJudge intelligibility, noise, clipping, and reverb — not labels or utterance counts.\n\n### Label\n\nRename tracks in place with the shortest evidenced label: `Cam · <subject or view>`\nand `Audio · <speaker or source>`. Append a shot size only when it distinguishes\notherwise similar angles: `WS` (wide shot), `MS` (medium shot), or `CU`\n(close-up), for example `Cam · Speaker A · MS` or `Cam · Two-shot · WS`.\nOtherwise omit it. Fall back to numbered labels when identity is uncertain.\n\nDeliver an evidence table alongside: claim | evidence (timestamp + what was seen or\nmeasured) | confidence. Decline to label what the evidence doesn't support —\nfragmented transcript speaker ids are usually in that category — and say so.\n\n### Deliver the master, then offer the cut\n\nThe master is the deliverable — report structure, offsets, identities, and evidence.\nNote that it stacks angles and is a reference, not something to watch: the top track\ncovers the others. If the user only asked to sync, stop here, but **offer** the\nspeaker-follow draft rather than leaving them with tracks and no next step.\n\n## Part 4 — Cut to the speaker (on request)\n\nTurn the synced master into a watchable draft: one video track that follows the\nconversation, program audio continuous underneath.\n\n### The master is never edited\n\nAll cutting happens on a **new timeline** (same fps/canvas). The synced master is the\nsource of truth every derived cut can be rebuilt from — if it changes, every offset\nbecomes unverifiable. When any instruction, taken literally, would break sync\n(e.g. \"start the audio at frame 0\" when its synced position is not 0), keep sync and\nsay why. Preserving sync outranks literal wording.\n\nTreat the speaker-follow timeline as a source or cutting board, not as an opaque clip.\nWhen the user wants part of it in an existing edit, materialize only those ordinary\npicture and audio ranges into the main timeline. To substitute another angle later,\ntake the same wall-clock interval from that camera using the master offset. The master\nis the timing record; no hidden multicam data structure is required.\n\n### The conversation drives the cut\n\nBuild one conversation document first: each person's speech taken from their\nassigned program source (Part 3), crosstalk dropped, interleaved by wall-clock via\nthe Part 2 offsets, with real names and timestamps. Then write the angle plan —\nstart–end, angle, reason — **before placing anything**.\n\nIf the user also asks to remove fillers, shorten answers, or restructure the\nconversation, use the talking-head workflow to decide which speech ranges remain.\nMulticam owns sync and angle choice; talking-head owns content. Finalize the content\nplan before materializing picture cuts.\n\n### Editorial objective\n\nKeep the viewer oriented, emotionally informed, and visually awake with the fewest\ncuts that add something.\n\n**Meaning chooses what to show. Rhythm chooses when to cut. Orientation chooses how\nwide to go.** Stay while the frame is still revealing; move when another frame gives\nmore; return to the room when the relationship needs refreshing.\n\nAngle rules (defaults, user's brief wins):\n\n- **Carrier of the moment**: the active speaker is the baseline, not the law. Show\n  the speaker when information originates there, the listener when the reaction is\n  the meaning, a pair or subgroup when the relationship carries the beat, and the\n  room when attention is divided.\n- **Shot scale**: use the smallest grouping that preserves the beat. `CU` isolates\n  thought or emotion; `MS` is the conversational default; a two-shot or group shot\n  shows a relationship; `WS` restores geography. Use a wider view at a new\n  question or topic, a participant change, overlap, shared laughter or silence,\n  physical movement, or after a long run of isolated singles. Hold it long enough\n  to read — usually 2–5s — then tighten when attention concentrates.\n- **Duration pressure**: a normal shot needs about 2–3s to arrive. Treat 4–10s as\n  a useful conversational range, not a metronome. After roughly 8–12s on an\n  unchanged single, actively look for a motivated alternative. By 20–25s the hold\n  should be deliberate. There is no maximum while performance, emotion, or visual\n  information is still developing.\n- **Speaker changes**: do not chase every sound. A <2s interjection usually stays\n  on the current shot. Let an unanticipated new speaker begin for about a second\n  before cutting; a direct question may motivate showing the respondent while they\n  prepare to answer. A short important line — introduction, direct address,\n  punchline — earns a shot; widen around it if necessary.\n- **Reactions**: use a visible, truthful reaction when it changes how the line lands\n  or refreshes a static hold. A reaction is usually 1.5–3s. Do not insert a generic\n  nod merely for variety, and never borrow a reaction from another wall-clock moment.\n- **Rhythm**: cut on a thought, breath, gesture, look, laugh, or relationship change,\n  not on a timer or arbitrary word boundary. Sentence boundaries are safe, not\n  mandatory.\n- **Seams**: one clean speech seam at a sentence boundary needs nothing. Cover a\n  cluster of visible jump cuts with one wall-clock-synced reaction or relationship\n  shot spanning the cluster. Never break source sync to manufacture coverage.\n- A camera restart under an unchanged angle is a same-angle join, not an editorial\n  cut. Count and report it separately.\n\n### Materialize flat\n\n- **One video track**, alternating angle segments, no gaps. Each segment's source\n  offset comes from the master placement math, never re-derived from transcripts.\n- For a full-length speaker-follow draft, keep program audio continuous and untrimmed\n  at its synced position. For a shortened or reordered cut, use identical kept\n  source-time ranges on every program mic so picture and audio cannot drift.\n- Mute every camera segment's embedded audio.\n- If audio outruns picture (recorders stopped later), keep the audio and leave the\n  tail dark. Never fabricate or freeze picture to cover it.\n\n### Keep program audio coherent\n\nKeep every isolated program mic open across each retained range. When program audio\ncomes from camera mics, use only the chosen source for that passage, crossfade at a\nquiet boundary, and do not stack them. Word timestamps carry ±30–60ms of ASR noise,\nso boundaries need natural handles; hard cuts on word\ntimestamps clip breaths and word onsets. Backchannel (\"嗯\", laughter) on an idle\nmic is part of the conversation. After content editing, run the talking-head\nworkflow's audio smoothing step. If mic isolation is weak, flag it for an audio pass\nrather than gating speakers independently.\n\n### Verify last, and prove it with numbers\n\nVerification is the **final action**, after every edit and every visual check.\nAnything that touches the timeline — including dragging in a browser to look at a\nframe — can move an item. **Inspection is not read-only.** If you interacted with\nthe editor UI at any point, re-run these checks afterwards; a pass from before that\ninteraction is void.\n\n- **Wall-clock invariant**, every segment and every program-audio item: timeline\n  time − source time must equal that source's master offset.\n- **Angle spot check**: at a few sampled shots — include the longest — the dominant\n  speaker in the window must match the angle's person.\n\nReport the **actual values**, not the word \"verified\": `MIC1 at 111, MIC2 at 111,\noffsets 3.700 / 3.700`. A claim you cannot print numbers for has not been checked.\n\nThe same rule governs recovery. If you disturb an item, **restore every field and\nre-read the row to prove it** — position, duration, and source offset each fail\nindependently, and fixing the one you noticed is not a restore.\n\nAlso report: cut count, shot length min/avg/max, the conversation document, the\nplan, and every place a rule conflicted with the material plus what you chose.\n"
}

SHA-256: b96943bba7d2343e3d3c9ffd3c80799b1ac34e8bc70a671c1e4164e672318777