← Files Creatify Ad AgentARCHIVED FILE

skills/ad-agent/references/audio-sync.md

8.22 KB · Oct 4, 2026 · 00:02 UTC

↓ Download file

# Audio: music, sync, speech, and checking what you'll ship

## Background music (`composer_generate_music`)

```
composer_generate_music(project_id, prompt="<music prompt>", out="gen/music.mp3")
# when done, result.probe has the real duration
```

- **Lyria 3 Pro: the prompt is the ONLY parameter.** Write the duration IN WORDS in the prompt ("a 14-second upbeat ad jingle with a clean ending"), an explicit tempo ("at 120 BPM, steady four-on-the-floor kick"), and a drop/lift where the hero moment goes ("builds, then drops at about 9 seconds"). Add genre/mood/instrumentation and, unless the brief asks for a song, **"instrumental only, no vocals"** (vocal music under dialogue is a mess). Long prompts are fine; a one-liner like "upbeat music" yields generic mush.
- Track length is model-determined. DON'T loop it to fit: fade it in and out where it naturally fits (`fade_in`/`fade_out` on its `Stage.audio`). Dip the bed to ~0.15–0.3 under any speech, end it no later than the video does, and never cut it hard. Silence for the rest of the runtime is fine; an abrupt music cut is a defect. When the prompt says "no music" (e.g. raw UGC realism), use no music.
- Start the bed EARLY, in parallel with the first shots, and assemble around it. Then measure it with `composer_exec`: `python3 -m lib.beats gen/music.mp3 --out gen/beats.json` (see below).

## Audio sync: everything lands on something measured

Viewers feel a cut 2 frames off the beat, and a caption that leads or trails the voice, even when they can't say why. Nothing is timed by ear or by estimate:

- **Music: measure, then cut to it.** `python3 -m lib.beats gen/music.mp3 --out gen/beats.json` (via `composer_exec`) → `bpm`, `beats`, `downbeats`, `hits` (strong transients), `drop`, all in the track's own time. Lyria drifts from the BPM you asked for and rarely starts on a downbeat, which is why you measure. Declare it on the bed: `Stage.audio({src: 'gen/music.mp3', at: 0, …, beats: 'gen/beats.json', duck: true})`. A track placed at `at` with `in` shifts every time by `at - in`.
  - Cuts land on beats; section changes (hook → product, product → offer, → end card) land on downbeats; the hero moment (reveal, offer, logo slam) lands on `drop`; graphic accents land on `hits`. `Stage.beatAt(t, beats)` gives `{i, since}` for pulses that fire on every beat.
  - Plan shot lengths as beat counts: at 120 BPM a 4-beat shot is 2.0s, 6 beats is 3.0s. If the music's grid and the script disagree, move the cut to the nearest beat inside a speech gap (below). Never cut mid-word to hit a beat.
  - `lib.compose render` / `check` print `sync`: every cut's `to_beat_ms` / `to_downbeat_ms`. Target `on_beat_share: 1.0` (every cut within 45 ms of a beat, i.e. ±1 frame). Fix listed `off_beat` cuts by adjusting `dur`/`at`.
- **Speech: measure every clip that talks.** `composer_transcribe(project_id, media="gen/shotN.mp4", script="<its lines>", out="gen/shotN.words.json")` (ElevenLabs Scribe) → every spoken word with start/end, plus `problems` = script words that weren't said as written (wrong/missing/added words mean a regen; see the Loop). For a PROVIDED voiceover/dialogue file add `forced=true`: forced alignment times the exact script to ±10 ms.
  - **Declare every line:** `Stage.clip(el, {src, at, in, dur, keep_audio: true, says: "Wanna go see the grandkids?"})` (and `says` on a `Stage.audio` dialogue track). The ship gate transcribes exactly that window and fails on any word not said as written: a mispronunciation ("grandkrids"), a dropped word, an ad-lib inside the window. Fix: regenerate the take once, or move `in`/`dur` to a clean stretch. Never put pronunciation hints in an H3 prompt; it speaks them ("Grand N Kai kids"). **A word that fails twice is not a prompt problem.** Give the line exact audio and lip-sync it: `composer_voice_line(project_id, like=["<clean clips/recordings of that character>"], text=["<line>"], lead=0.4, out="gen/line.wav")`, then `composer_generate_clip(mode="a2v", image="<that shot's keyframe>", audio="gen/line.wav", …)`, and put `says` on the new clip. Never ship a talking window without `says`. `says` is the script's line verbatim. Never edit it to match what a take said (spelling variants like "all right"/"alright" or hyphenation are already treated as equal).
  - The sandbox has no transcription key, so a `lib.compose render` run through `composer_exec` can't check `says` windows (they come back as errors). While iterating, check speech with `composer_transcribe`; the ship gate runs the real check.
  - Captions are only ever `Stage.captions('#cap', <words array from that json>, {at: clip.at - clip.in, max: 3, upper: true})`. Chunks break on punctuation, pauses and `max` words, and each shows from its first word's start until the next chunk. They cannot drift from the voice, and they say what was actually said. Style the element in CSS (`transform: scale(calc(.9 + .1 * var(--p)))` pops each chunk in). Load the words JSON into the page (it is a project file next to `index.html`).
  - Cut between shots in the gaps between words (`end` of one word → `start` of the next), not inside them. Trim a talking shot so its first word starts within ~0.15s of the cut; dead air after a cut reads as a mistake.
  - Duck, don't hand-automate: `duck: true` on the music bed side-chains it under all dialogue and clip sound, so the bed dips while people talk and comes back in the gaps.
- **SFX land on their peak, not their file start.** Run `lib.beats` on a whoosh/impact to get its `hits[0]` (the transient) and place it at `at = event_time - hits[0]`.
- **Duration is a hard target.** The target duration ±0.5s (the ship gate checks it). If the script's measured speech plus the end card doesn't fit, tighten pauses and shot lengths. Don't let the edit run long.

## Judging audio: never ship audio you haven't checked

Frames alone can't judge it.

- `python3 -m lib.audio_check <video> [--reference assets/NNN_dialogue.wav]` (via `composer_exec`) is deterministic: audio present, loudness sane, and (with `--reference`) three facts about the PROVIDED track:
  - `best_corr`: ≥ 0.95 means preserved as-is; `< 0.5` means replaced, a FAIL you fix with a2v/muxing, not hope.
  - **`coverage` + `missing_last_seconds`**: the fraction of the reference audible in the output. `missing_last_seconds > 0.3` means the dialogue's TAIL was truncated. That is a hard FAIL, fixed by the a2v multi-segment protocol (generation.md), never by fading early.
  - Run it on any clip whose own soundtrack ships AND on the assembled video. For the assembled video, render it first in the sandbox (`python3 -m lib.compose render index.html -o out/video.mp4`, then `python3 -m lib.audio_check out/video.mp4 --reference …`), before you ship.
- **Speech: `says` + `composer_transcribe`, never "sounds fine".** Every talking window declares `says` (the gate transcribes it). While iterating, `composer_transcribe(media="gen/shotN.mp4", script="<its lines>")` shows what was actually said. Lip-sync comes from the mode, not from checking: a provided voiceover or a re-voiced line is a2v (synced by construction). For generated speech, look at a `strip` of the line: the mouth must move on the words and stop between them (a locked jaw over audible speech is a defect).
- **SCRIPTED dialogue** (the prompt dictates what is said, and no audio asset provides it): use r2v/i2v generative speech. r2v hallucinates words that SOUND like the target (fluent word salad: "had Pritchett like hyping lavender" for "had people hyping"), so every scripted shot MUST pass the `says` / transcript check. On mismatch, regen with a hardened prompt ("she says these exact words, clearly and completely, nothing added or omitted: <line>"). Shrink the line, since garble concentrates mid-sentence; an over-stuffed sentence can be cut into two clips and concatenated. Reserve r2v generative speech for scripts that are LOOSE (summarized, improvised-feeling) where exact wording isn't owed. Those only need coherence, not the word-diff.
- Whether "extra words" count as a FAIL depends on the brief. When the prompt quotes the exact line ("she says: X"), the model's ad-libbed additions are a FAIL. When speech content is up to the model, coherence+naturalness suffices.
- No dialogue in the brief → the clip must hear no words (`composer_transcribe` with no script); see product-fidelity.md rule 7.

SHA-256: cdc2f99a2ac9798facb4e040110a44403871ffe67a2db2cbdee7442affb2b88e