← Files PixVerseARCHIVED FILE

skills/pixverse-voiceover/references/production.md

5.2 KB · Oct 4, 2026 · 12:28 UTC

↓ Download file

# Narration Modes And Measured Delivery

Ordinary TTS is a direct take; do not impose the specialized retries below on every read.
Native UGC/host speech stays in its video workflow. Select the matching mode before audio
generation. Use real returned voice IDs and actual model capabilities, not source-platform
voice pairs, removed scripts, markup or rate flags copied into another engine.

## Fixed Video Windows

One numbered line maps to one fixed window. Lock voice ID, model, language and a single
delivery direction in the project record; reread that record for every submission.
Mood changes wording/performance, not voice identity. Preserve ordering and accepted takes.

This recipe targets 7.8–9.5s of speech in a 10s block, initially around 20–23 English
words (17–21 for animated children's delivery), no more than two sentences and little
punctuation. These are starting targets, not a claim that PixVerse TTS shares ElevenLabs'
calibration. Measure on the actual engine; localize unspaced languages by spoken duration.
If the engine supports timed delivery markup, express immediate onset and the window;
otherwise keep direction out of spoken text and use only verified controls.

Measure the voiced interval, internal silence and delivery rate, not total file duration.
For this 10s recipe, internal gaps ≥0.8s and speech faster than roughly 2.9 English words/s
require inspection/correction. Exclude real provider padding; inspect and repair an actual
tail click without blindly trimming a phoneme. Preserve the requested silence when it is
part of a performance instead of misclassifying it as failure.

Too long: shorten an editable clause; too short: add real information; pause-heavy: use
one flowing phrase. Never pad with filler, time-stretch, pitch-shift or cut speech to pass.
Locked copy cannot be rewritten without user authorization; use the agreed window/content
tradeoff. Limit to three attempts per failed line; a third duration attempt needs changed
copy. After a retry the 7.2–7.8s soft band can be disclosed; outside 7.2–9.5s remains a
failure for this recipe, scaled for shorter final windows. Do not promote the nearest
failed take. Accepted indices are immutable; regenerate only failing indices.

## Continuous Story Read

Keep a whole script or a few large paragraph-aligned chunks in one locked voice and
verbatim delivery direction. Read actual prompt-length limits; chunk below them without
splitting sentences. Source limits such as 2048 characters are not universal CLI values.
Join accepted chunks in order, losslessly only when formats match. Check seams, timbre,
garbling and unintended internal pauses; no per-line fixed-window gate in this mode.

The requested duration is a script-length target, not a reason to vary TTS speed. Generate
naturally, measure, and if editable copy misses the accepted range, estimate a revised
word budget as old_words × target_time / measured_time. One initial read plus at most two
changed-text duration corrections. Do not resubmit identical wording to chase duration;
same-text retries are for real provider/voice defects. After the limit, report the closest
clean read and exact miss rather than claiming a pass. Visuals follow its actual clock.

## Existing Video + Photo → On-Screen Narrator

This distinct mode preserves the base video's picture and full duration, with a moving
photo-derived narrator composited over it. A static portrait or audio-only file does not
satisfy the request. Preserve supplied script sentence order; otherwise transcribe the
base audio. Resolve real supported portrait animation, timing-preserving revoice if an
exact selected timbre is required, and clean compositor/matte capability before spending.
Reference conditioning is not equivalent to voice-change. If a required service is absent,
report that specific gap and prepare an explicit alternative; never fake its completion.

Source planning uses ceil(base_duration/10) blocks, one 10s talking take per block and
31–35 English words/full block, proportionally fewer in a partial last window. Confirm
that the selected engine can articulate that pace. A final partial block is trimmed only
after its words fit the real window; do not pad/loop base footage. A locked script that
cannot fit requires a content/range decision, not omitted sentences or a rushed take.

Create one identity-preserving portrait on uniform chroma green with clean hair edges,
no spill or beautification. Each take starts already speaking, fixed camera, small head/
shoulder gestures, exact new words once, no text or background changes. Use supported
voice conversion only when needed to meet the selected-voice contract; unrelated TTS
under moving lips is not a repair. Composite the moving presenter in a safe corner without
covering evidence. If cutout edges fail, use an honest designed badge instead of claiming
a clean cutout. Preserve base audio components that can be retained; explain a necessary
replacement when original speech/music are inseparable. Do not silently discard the mix.

Verify every block, opaque presenter, clean edges, complete script and final duration
within 0.1s of the base. Preserve all successful blocks. Deliver the actual narrated video
with voice, coverage, position/style and audio-treatment record. Captions are opt-in.

SHA-256: 6173cb5c02133396b4918068e3e6fb3322228c5999e7bb7e1c37604ca17fde97