← Files PixVerseARCHIVED FILE
skills/pixverse-explainer/references/branches.md
6.4 KB · Oct 2, 2026 · 00:28 UTC
# Specialized Faceless Branches ## Picture Story / Frame By Frame One continuous natural narration → actual word timestamps → dense numbered frame timeline. No video generation, per-phrase TTS, fixed 10s voice windows or slideshow substitution. Target 0.7–1.2s/image, maximum 1.5s: at least ceil(actual_voice_seconds/1.5) frames, typically 45–70/minute. Show the planned image count before spending. Adjust an editable script from measured speech to meet duration; never change audio speed. A moment is a burst of 2–3 near-identical images changing one detail. Roughly two-thirds of frames are literal edits of the immediately previous rendered frame as their ONLY reference; preserve crop/camera/cast/colors/background and change one named detail. New framing/action/place uses location→cast→props. Do not attach the roster to micro-edits, which would redraw the whole scene. Keep each variation dependency sequential. Mix wide/medium/close, show what speech names, and avoid consecutive close shots. The source's close-up micro-variation wording conflicts with its no-two-close-frames rule; resolve with one close reaction followed by a medium movement burst, keeping chains on wide/medium shots. Frames have the output aspect, no roster sheets. Measured contiguous segments have no gap/overlap and no hold >1.5s. Preserve continuous voice and deliver one film. ## Kids Narration Block 1 asks the central question and answers it in ordinary words in the first sentence; block 2 gives one visible example. Middle: what→where→does→without. Last: sourced surprise then four/five-word callback answer. Read the first sentence of every block alone: they must answer the opening question. One idea/sentence, at most two/block, familiar physical analogies, no stacked abstraction, needless decimals or dates. Explain action before terms. Every block changes the physical through-line. Narrator addresses named cast AND viewer; cast visibly waves/nods/looks/responds on the beat, mouths closed. Later participation questions may close a block, with the next block opening on their answer; this does not delay the opening central question's same-sentence answer. Do not leave a long pause inside either voice take. Warmth comes from words and energy, not stretched syllables. Plan 17–21 English words/full 10s. Four shots: context, character reaction, extreme detail and medium resolution, with order and opening coverage varied between blocks. One real action spans them. Add 2–4 playful diegetic cues tied to motion, no generated voice/music in narration clips. Kids looks normally include a wordless bed, even on another channel using that look, with mood matched to its content. Supplied music wins; use a supported music route otherwise. Start quietly (~0.05 linear gain), duck and listen; gain is not a measured loudness claim. Explicit no-music wins; disclose an unavailable optional bed while delivering the completed film. ## Kids Talking Characters Opt-in. Odd blocks have external narration/closed mouths; even blocks have native dialogue from at most two named speakers, 16–22 English words total/full10s. Dialogue is caused by the preceding narration and sets up the next event. Keep identity/wardrobe/style/delivery. Native timbre is not guaranteed identical across jobs; check if exact voice is required. Keep dialogue's full-length audio aligned to lips; do not center or replace it with TTS. Record block_kind and ordered speaker lines. Narration keeps its normal gate, all blocks keep four-shot/music rules. Song intent selects Song instead of combining these modes. ## Kids Song The primary audio must be a real sung song. Resolve genuine singing/music capability before spending; TTS/instrumental music is not equivalent. Ignore narrator voice metadata, skip narration takes and a second bed. Write visual lyrics in the requested language, verses plus a chorus repeated verbatim. Generate song FIRST, then measure sections and choreograph. Source targets 60/120s with ±3s tolerance and at most three attempts; explicit duration wins. Gentle: 100 BPM major-key ukulele/guitar/bells/light flute. Dance: 112 BPM major-key claps, plucked bass/marimba/bells. Keep steady 4/4 from start to end, no tempo changes or section gaps; dynamics grow through instrumentation. Use regular singable lines, often eight syllables/alternating stresses. Resolve the source's inconsistent eight-syllable/one-per- beat/four-bar instruction: eight beats means two 4/4 bars; longer phrases need held notes or rests. Specify an internally consistent meter instead of copying contradictory math. 60s: intro/verse, chorus, verse, fuller chorus. 120s: intro/verse 1, chorus 1, verse 2, chorus 2, verse 3, final chorus. Verses advance story/show named nouns; choruses repeat signature dance and location with growing cast/staging. Gesture/dance on beat, no articulated singing mouths. Four-shot blocks, diegetic cues only; song owns the final audio, SFX barely audible (~0.06 starting gain). No automatic singing subtitles. Requested lyric captions need verified singing alignment or actual manual timing, not an untested Whisper assumption. ## Long History (About 10 min+) Outline chapters, eras, through-line, vignettes and asset counts first. Source skeleton: two-block crisis opening, rewind around block 3, then 5–8 chronological chapters of 8–14 blocks where runtime fits. Each chapter has setup→pressure→turn and a handoff hook, one documented human vignette, a map beat and quantity comparison when figures are spoken. End with meaning/consequence and one brief optional CTA. Adapt counts to actual runtime explicitly. Verify 4–6 concretes/chapter and state historical uncertainty. One cast asset/era with fixed identity invariants plus actual aging/clothing changes. Every later incarnation references the first; for gradual 3+ eras also use the preceding era if needed. Match era variants to blocks. Locations need variants only when changed. Sum cast×eras+places+coverage+props before spending (often 25–40 assets/10 min, not a fixed quote). Rotate scene→map→quantity→vignette; never three identical device classes in succession. Action motifs occur at most once/chapter and about three/film, except through-line states which must advance. Don't use seals/stamps as generic filler for named real content. Produce chapter batches and preserve completed chapters; assemble one whole film. Mannequin reduces eras to at most small/adult variants with stable role colors, no facial/costume aging.
SHA-256: 8bcb1e333bbf8a5b04b36e06052ab0bfd61ea945be56b7a92f510f6ec6a56d9f