← Files HiggsfieldARCHIVED FILE
skills/ai-host-video/references/anchor-prep.md
3.71 KB · Oct 5, 2026 · 12:03 UTC
# Measured speech and phrase anchors Plan meaning with exact phrases before generation. After download, derive timing from the actual audio and footage using the installed transcription route and ffprobe. A planned duration, script word count or video scene summary is not a word timestamp. ## Source measurements For each accepted host source, retain its plan id, actual provider id, source URL, local path, duration and word-timestamp transcript. Store measured results once and reuse them while the source bytes are unchanged. Keep a replacement's transcript separate; changing footage invalidates its trims and all downstream timing. Compare the transcript with the approved words. Normalize case/punctuation only for matching. Recognizer spelling, number expansion and false word splits need actual audio inspection; reuse the transcript and correct the phrase mapping instead of deleting it or rerunning ASR for a spelling mismatch. Do not rewrite the approved spoken text to match ASR. A real omitted spoken phrase is a generation defect handled by operations' replacement policy. If ASR is unavailable, resolve a reliable timing route before speech-synced motion; do not silently estimate anchors or claim measured alignment. Choose source_from/source_to around complete speech, keeping onset, final phonemes and intentional pauses. Trim only verified idle lead/tail; do not remove internal speech pauses or change playback rate. Verify 0 <= source_from < source_to <= probed source duration and that every required spoken word remains inside the trim. ## Timeline arithmetic For each source, kept_duration = source_to - source_from. Place hosts in script order. Each timeline_start is the sum of preceding kept host durations plus real PLAYBACK inserts before it. Supporting visuals over speech add no duration. Keep structural inserts between complete host jobs by default. If the user explicitly requires an insert inside a host job, represent it as two measured host spans with the insert between them. Never use a single unadjusted offset across the gap. For a word/phrase inside a retained source span: ```text timeline_phrase_start = timeline_start + source_phrase_start - source_from timeline_phrase_end = timeline_start + source_phrase_end - source_from ``` Record source and timeline ranges together with the matched words and occurrence number. Repeated phrases must select the intended occurrence. Validate anchors against their retained source span and final timeline. Recompute later placements when a trim, insert or selected source changes. The references use clip-bounds.json for trims, spine-map.json for placements, cuts.json for observed cuts and anchors.json for phrase ranges. These are convenient agent-written records, not files generated by a supplied resolver. A single compact measurement table is equivalent when it retains the same source/time facts. ## Apply anchors to motion A phrase identifies the idea; its first word is not automatically the picture handoff and its last word is not automatically the exit. Place entrance before a word impact, leave a readable fully-revealed hold, then exit or advance when the idea finishes. For a child starting at master time S and a word at master time A, its local animation time is A-S, accounting for every ancestor's start. If a planned phrase cannot be matched, inspect the audio and neighboring timed words, then re-anchor to the same measured spoken moment. Do not silently omit its visual or regenerate the host because recognition differs. If no reliable anchor can be obtained, explain the limitation and resolve a simpler non-word-synced treatment with the user when that changes the requested result. Preserve completed sources and review only ranges affected by the timing correction.
SHA-256: ece0a58bd5453b3374f754d43d801e3714debc6c9820b071f47676262ae863ab