← Files HiggsfieldARCHIVED FILE

skills/ai-host-video/references/anchor-prep.md

3.71 KB · Oct 3, 2026 · 06:02 UTC

↓ Download file

# Measured speech and phrase anchors

Plan meaning with exact phrases before generation. After download, derive timing
from the actual audio and footage using the installed transcription route and
ffprobe. A planned duration, script word count or video scene summary is not a
word timestamp.

## Source measurements

For each accepted host source, retain its plan id, actual provider id, source URL,
local path, duration and word-timestamp transcript. Store measured results once
and reuse them while the source bytes are unchanged. Keep a replacement's transcript
separate; changing footage invalidates its trims and all downstream timing.

Compare the transcript with the approved words. Normalize case/punctuation only
for matching. Recognizer spelling, number expansion and false word splits need
actual audio inspection; reuse the transcript and correct the phrase mapping instead
of deleting it or rerunning ASR for a spelling mismatch. Do not rewrite the approved spoken text to match ASR.
A real omitted spoken phrase is a generation defect handled by operations' replacement
policy. If ASR is unavailable, resolve a reliable timing route before speech-synced
motion; do not silently estimate anchors or claim measured alignment.

Choose source_from/source_to around complete speech, keeping onset, final phonemes
and intentional pauses. Trim only verified idle lead/tail; do not remove internal
speech pauses or change playback rate. Verify 0 <= source_from < source_to <=
probed source duration and that every required spoken word remains inside the trim.

## Timeline arithmetic

For each source, kept_duration = source_to - source_from. Place hosts in script
order. Each timeline_start is the sum of preceding kept host durations plus real
PLAYBACK inserts before it. Supporting visuals over speech add no duration.
Keep structural inserts between complete host jobs by default. If the user explicitly
requires an insert inside a host job, represent it as two measured host spans
with the insert between them. Never use a single unadjusted offset across the gap.

For a word/phrase inside a retained source span:

```text
timeline_phrase_start = timeline_start + source_phrase_start - source_from
timeline_phrase_end   = timeline_start + source_phrase_end   - source_from
```

Record source and timeline ranges together with the matched words and occurrence
number. Repeated phrases must select the intended occurrence. Validate anchors
against their retained source span and final timeline. Recompute later placements
when a trim, insert or selected source changes.

The references use clip-bounds.json for trims, spine-map.json for placements,
cuts.json for observed cuts and anchors.json for phrase ranges. These are convenient
agent-written records, not files generated by a supplied resolver. A single compact
measurement table is equivalent when it retains the same source/time facts.

## Apply anchors to motion

A phrase identifies the idea; its first word is not automatically the picture
handoff and its last word is not automatically the exit. Place entrance before a
word impact, leave a readable fully-revealed hold, then exit or advance when the
idea finishes. For a child starting at master time S and a word at master time A,
its local animation time is A-S, accounting for every ancestor's start.

If a planned phrase cannot be matched, inspect the audio and neighboring timed words,
then re-anchor to the same measured spoken moment. Do not silently omit its visual or
regenerate the host because recognition differs. If no reliable anchor can be obtained,
explain the limitation and resolve a simpler non-word-synced treatment with the user
when that changes the requested result. Preserve completed sources and review only
ranges affected by the timing correction.

SHA-256: ece0a58bd5453b3374f754d43d801e3714debc6c9820b071f47676262ae863ab