← Files PixVerseARCHIVED FILE

skills/pixverse-howto-video/references/production.md

4.89 KB · Oct 2, 2026 · 00:28 UTC

↓ Download file

# Four-Step Creator Tutorial

Read `../../../skills-shared/ugc-production.md`. The standard creator tutorial uses a
four-panel one-row 21:9 board per 4–15 second clip, a locked creator and product, a required
board refinement pass, exact persistent step headings and native spoken explanation.

## Real Procedure And Numbering

Derive mechanics from the manual, approved instructions or observed behavior. Record each
step's starting state, contact/action, expected visible result and relevant real caution.
Do not invent a mechanism because it looks plausible. A precision first-frame requirement
uses a supported image-to-video mode; reference conditioning alone is not a first-frame lock.

Plan 4N genuine steps for N clips. Board K carries global steps 4(K−1)+1 through 4K;
never restart at Step 1 on the next board. If the procedure is short, include meaningful
preparation or finishing; if long, merge adjacent micro-actions at an intelligible state
boundary. Never manufacture an unnecessary action to fill a slot. An explicit director's
step list maps 1:1; explain and use the necessary adapted panel/clip plan if it differs.

Typical slot roles are preparation with creator face/torso visible, core mechanism,
application, result. Slot 1 establishes the same identity and outfit, not anonymous hands.
Each subsequent frame must show the point of contact and the state required by the next.
Keep one main action per cut, supported weight and no more than two concurrent hand roles.

## Text Is Part Of The Board

Use the exact heading format `Step N — Heading`, with a 1–4 word Title Case heading in
English by default; the project's explicit language wins. Author all headings before
board creation. The board renders each heading once, and refinement and video prompts
must preserve it unchanged in its matching cut. This is required instructional text,
distinct from optional hook text or spoken-word subtitles added later.

Choose one category-appropriate font direction: restrained editorial serif for beauty,
fragrance or home; clean geometric sans for tech; stronger condensed sans for fitness or
outdoor; a warm readable face for food/coffee; a light editorial face for fashion. Legibility
wins over the style cue. Keep one font, size, color and position throughout all boards.
Place top-center with 5–7% top padding; longest heading spans about 60–70% of panel width.
Use white with a light shadow or dark text on a light area. Maximum two lines; if one label
wraps, use a consistent two-line system. If the top conflicts with a face, move every
heading to the same lower position. Add no extra symbols, caption duplicates or watermark.

Inspect the raw and refined text. If exact glyphs fail, repair the board's labels with an
editable compositor before motion and preserve the chosen appearance; do not silently
drop labels or deliver misspellings. Final overlay repair is available for model text drift,
but does not justify skipping the planned instructional board.

## Time, Camera And Speech

Use four timed cut sections and three hard-cut boundaries. At 12s allocate 3s each; at
10s use 2.5s each; at 15s use 3.5/4/4/3.5s. Other supported durations stay near-equal,
adjusted for the real step's complexity. Every cut needs an achievable action, exact
hand allocation, 4–10 useful sentences and five integrated micro-observations. Describe
the heading as unchanged and visible for that entire step; never generate Step 5 in cut 4.

A Set-Down or Pick-Up is opt-in for an explicit one-take brief, only at an adjacent legal
selfie/fixed boundary without a state jump. It replaces exactly one hard cut, consumes
its own time and carries continuous speech. Do not borrow the unboxing default Pick-Up.

Keep the final 0.5–1s of the last cut for a concise requested/default tutorial CTA, such
as following for the next step, while the last step's result and heading remain. It is
not a fifth operation or a new caption. Respect explicit no-CTA and exact-script requests.
Native speech explains what the viewer can actually see; no separate voice pasted over
visible lips. Keep meaningful room sound and real contact sounds, no music by default.

Generate a new product-free creator portrait only if no accepted identity exists. Board
references are product, creator, previous clean board; video references are clean board,
creator, product. Preserve sequence and caption styling through every refinement.

## Required Review And Prompt-Only Output

Check step order, actual mechanics, contact, consistent states, creator/product identity,
all headings at normal viewing size, unwanted text and 2–3 mid-word frames per clip.
Optional hook/subtitle layers use final audio and avoid the step-heading band. A mechanism
that cannot be rendered accurately needs verified real footage or a clear honest diagram.
Deliver prompt-only work with globally numbered steps, board and refinement prompts,
all four motion sections per board, exact spoken copy, reference roles and timing.

SHA-256: 6dbd713a8ebcd6ebc2597877866584b9b2811f47bd3a0c122253b40720f11e3d