← Files HiggsfieldARCHIVED FILE
skills/ugc-review-video/references/monologue-craft.md
13 KB · Oct 5, 2026 · 12:03 UTC
# Monologue craft — what the creator actually says
Load this at the monologue step, before splitting text into board segments. Word density first:
<=10s ~ 12-20 words, 11-12s ~ 20-28, 13-15s ~ 28-35. Sum across clips for the total, then split
into N board segments (the 8 slot-level beats inside a board are the board/clip prompt's job, not
yours). When a product is present, greetings and product introduction are allowed ONLY in board 1;
later boards continue mid-thought.
## Truth contract — mandatory for every run
The parent skill's safety and truth gate always applies. `approved_claims` is the complete claim
allowlist, including when it is empty. Never invent or imply purchase, ownership, use, results,
before/after outcomes, ratings, reviews, social proof, relationships, or lived experience.
- **Generated creator:** write a host/demonstrator script using observable product mechanics and exact
allowlisted claims only. Never give the fictional creator a customer history.
- **Authorized user as creator:** first-person experience is legal only when the user supplied the
exact first-person script and confirmed it is their own experience. Do not embellish it.
- **Productless mode:** fictional storytelling is allowed, but do not impersonate a real person or
invent claims about a real person, organization, product, or service.
- Quoted lines and pattern examples below are structural examples, never factual source material.
**Productless mode:** when both product fields are null, the creator's topic, routine, experience,
or story carries the whole arc. Never invent a product, brand, review claim, product entrance, or
sales CTA. Ignore every product-entry, product-pressure, product-mechanic, and product-claim rule in
this reference; choose a story shape that works without an object and end on the human resolution.
## Voice archetype (a default, not a question — state it for veto)
Every video gets ONE performance persona that colours HOW the creator talks. It is picked from the
tier of the already-resolved register and NEVER changes that register. Resolve the register FIRST
(NATURAL by default; HYPED or CALM only on an explicit signal), THEN pick a persona from its tier:
- **NATURAL tier (the default — use unless a signal says otherwise):** a warm, genuine, engaged
creator — girl-next-door / easy-going relatable / dry-witty. Lively and real, ONE honest
human-scale peak (a real grin, a delighted "oh"), NEVER squealing or screaming. Beauty, fashion,
and most briefs get THIS by default — not It-Girl.
- **HYPED tier (only on a hyped signal):** 2000s It-Girl (stretched vowels, squeals on the reveal),
Hype Beast (explosive bursts), Drama Queen.
- **CALM tier (only on a calm signal):** Deadpan Contrast (flat, one micro-crack at the peak), quiet
Gossip Whisperer, Southern Sweetheart (soft warm).
HARD RULE: under the NATURAL default the persona sentence carries NO energy words — never `squeal`,
`scream`, `stretched vowels`, `explosive`, `screaming`. Energy comes from the REGISTER; the persona
adds only attitude, melody, and cadence, never volume. Write the chosen persona as ONE sentence and
carry it verbatim into the character prompt, every board prompt, and every clip prompt — restated
each time, or it drifts. Archetype and accent stack independently.
## Accent and quirk (opt-in)
If the brief names an identifiable origin for the creator, or asks for "weird" / "viral" character
energy, offer it inside the single intake question (accent consent + one quirk). Never add an accent
the user did not approve; the default is neutral English with no quirk. If opted in, write the
persona/quirk sentence once (identity + attitude, e.g. "a young Korean it-girl — English is her
second language and you hear it in every sentence") and restate it verbatim everywhere downstream.
Text-only accent enforcement lands roughly one render in three: if the user can drop a 5-10s voice
sample, attach it to the clip generation as an audio reference ("accent and vocal delivery reference
only — do not copy words, only the accent, melody, and timbre"). Offer that once; never block on it.
## Story mode is the default
ONE scenario spans the whole video: situational stakes first, framed without claiming that a
generated creator really experienced them. If
a product is present, it enters as a SUPPORTING ACTOR at 40-60% of total runtime, mapped onto the arc roles —
stakes live in the HOOK board, the single "but then" twist opens the mid arc, and the CTA rides
INSIDE the closer's resolution, never as an outro beat. The story must survive with the product
deleted. Plain review structure only on explicit request.
**Story-shape menu — pick ONE spine (it decides WHAT happens; the hook pattern decides how it
OPENS). Brainstorm at least 3 candidates, then choose; combine two only if it deepens the gap
without clutter:**
- S1 Demonstration Discovery (problem appears → feature is demonstrated)
- S2 Mechanism Loop (open on an observable action, then show how it works)
- S3 Alternate Use (show a second visible use without claiming personal ownership)
- S4 Guided Comparison (compare only visible mechanics or exact allowlisted claims)
- S5 Behind-the-Concept (explain the creative or product setup without a testimonial)
- S6 Interrupted Storytime (a story about something else; the product interrupts, the story never
quite finishes)
- S7 GRWM With Stakes (getting ready FOR a charged thing; the stakes carry it)
- S8 New-Place Mini-Vlog
- S9 Day-in-the-Life With a Twist
- S10 Product Reveal (clearly staged reveal, no claimed first reaction)
- S11 POV Frame (talks TO the viewer as a character)
- S12 Process Win (real work, one small win — native for B2B and small business)
- S13 FAQ Reply (answer a product question with observable evidence)
- S14 Green-Screen Commentary (talk over on-screen evidence)
- S15 Feature Checklist (show a set of observable attributes)
- S16 Host Question (one direct question answered by the demonstrator)
- S17 Baseline-to-Demo (show an observable setup change without invented results)
- S18 Silent Flex
- S19 ASMR (see the silent/ASMR mode below)
Every shape carries exactly ONE "but then" twist (the mid-arc peak) — no twist is a flat anecdote,
two will not fit in 15s. Realize the shape through the arc-role boards, never as a new field.
**Review-style / demo mode** (only on an explicit "review it / plain demo" ask): convert the request
into a truthful demonstration, not a customer testimonial. Give the body ONE
compact skeleton — PAS (Problem → Agitate → Solve), BAB (Before → After → Bridge),
Hook-Story-Offer, Us-vs-Them (against a lazy pricey default), or Demonstration ("watch this" — the
product does something visible in ~5s). Two or three spoken lines; even here the FIRST line carries
one observable detail or exact allowlisted claim. Do not invent a before/after result.
## Hook-pattern menu — the HOOK board's opener
Pick ONE pattern per video, never blend. The pattern decides what KIND of moment the opener is; the
first-word constraint below still governs its words. H2/H5/H7 are text-led — the opener LINE
embodies them. H1/H3/H4 are visual-led — the line lands mid-event and the staging carries it.
- **H1 Impact Action** — frame one is something physical mid-peak (box mid-rip, product mid-catch,
mid-stumble); the first word lands DURING the action.
- **H2 Mid-Sentence Confession** — opens on word four of a sentence, as if the viewer walked in on
it; confessional volume.
- **H3 Pattern Interrupt** — a normal setting with one thing deeply wrong, delivered with total
normalcy; the wrongness is visual, the voice ignores it.
- **H4 Freeze-Reaction** — frame one: the face already in full reaction, locked on something; a
performed hold (<=0.7s), THEN the first line.
- **H5 Hostile Open** — the first line is a challenge or accusation aimed at the viewer; slight
lean-in, finger already pointing.
- **H6 Quirk-First** — the quirk IS the literal first event, before any context. Only when a quirk
was opted into.
- **H7 Mechanism-First** — an observable product action is already underway in frame one; the first
line names that visible action or repeats an exact allowlisted claim. Never imply a personal
outcome, transformation, or before/after result.
- **H8 Product Cold Open** (advanced, story mode only, needs a product image) — board 1 slot 1 is
clearly staged product-only footage (rougher light, slight compression, no face); ONE hard cut
into slot 2 where the creator is mid-reaction with the product already in hand and the first line
lands on the cut. The character enters one slot later than usual. Never frame the opening as
found/reposted third-party footage and never add social proof. The slot-1 footage is a produced
cold open, not the product's story entrance, so the 40-60% product entry still stands.
Write the chosen pattern into the HOOK board's prompt as ONE plain visual staging sentence — same
mechanism as the persona sentence, on board 1 only, never a pattern code.
## Anti-slop pass on every line
- The first words are never `Okay wait / Okay so / OMG / Hey guys / So basically / Stop scrolling /
You NEED this / Story time`.
- Banned anywhere: "literally", "obsessed", "game-changer", "holy grail", "changed my life", "hits
different", and corporate words (elevate / seamless / effortless).
- Demonstration-first openers beat unsupported enthusiasm ("Here is the part that twists open.").
- Every product claim must be an exact `approved_claims` string. With an empty allowlist, use only
observable mechanics or sensory descriptions visible from the supplied product; never invent a
number, duration, comparison, result, purchase count, or measurement for specificity.
- At most ONE peak reaction per clip in the monologue (the clip prompt may stage up to two — the arc
peak plus an optional product-motivated closer beat).
- Echoes and repeats count toward the word budget but carry no information — cut them first.
**First-word constraint (positional, every board's segment):** the literal FIRST WORD must be hook
content — never `OK / Okay / Okay so / Alright / Alright so / So / Yeah so / Right so / Um / Well /
Like / Wait / Wait what / Hold on`. Those read as recording warmup. They are fine mid-sentence
later; the constraint is the first word only.
**Performed dialogue** (emotion lives IN the words — the video model under-renders flat prose):
vowel-stretch on the peak word ("it's SO good"), 1-2 CAPS volume spikes per line max, one broken
sentence at the peak ("it's— okay wait. LOOK."), a whisper-to-spike swing, and a flat → spike →
settle arc across the clip. Under the NATURAL default the peak stays human-scale (a real gasp, a
breaking grin), never staged screaming. NEVER write engineered or dramatic pauses — they bloat the
line and break the render.
## Optional registers
- **Quiet process-led** (only on an explicit "quiet routine" / "aesthetic process video" ask): the
monologue drops to soft half-thoughts, <=20 words per 15s clip — the density floor is waived,
ceilings still bind. Sound carries the clip. The product enters as one honest routine step.
Skincare and beauty routines start bare-faced. Tell the boards and clips the register is quiet
process-led with ONE plain staging sentence.
- **Silent / ASMR** (shapes S18 / S19; best for products with satisfying physical feedback): spoken
budget drops to ZERO (S18) or <=15 whispered words (S19). The SFX line becomes the script — each
trigger named and timed; micro-movement density doubles; framing goes macro-close for S19. The
absence of a synthetic voice is itself the anti-slop move. The CTA lives only in the post package.
For S19 restate "no on-screen words of any kind" — the model loves to bake karaoke subtitles into
ASMR.
- **Music is opt-in.** When the user asks for music, pass the request into the clip prompts; they
carry the music line. Never invent music uninvited.
## Series mode (opt-in — a content calendar, not one long video)
Triggers only on an explicit ask for a SERIES / "a week of content" / "N posts from one creator".
Different from a >15s multi-board video (that is ONE stitched video): a series is N SEPARATE posts,
each its own full run, sharing one creator and a week arc. It costs N times the renders — always
output the arc TABLE first, get a nod, then generate post by post.
- Carry forward VERBATIM across all posts: the persona sentence, voice archetype, accent delivery,
quirk, and wardrobe anchors. Reuse the SAME `character_media_id` — no regeneration.
- 7-post skeleton (adapt the count): 1 = S10 Product Reveal · 2 = S1/S8 (creator earns attention,
product barely mentioned) · 3 = S7 GRWM With Stakes (first soft sell) · 4 = S4 Guided Comparison
· 5 = S3 Alternate Use (rewatch fuel) · 6 = S2 Mechanism Loop (strongest demonstrable CTA)
· 7 = S5 Behind-the-Concept / quirk payoff (community post).
- Product-pressure curve: light (1-2) → selling (3-6) → community (7). A series that sells in every
post burns the account.
- Quirk escalation: identical in posts 1-4, NOTICED by the creator in post 5, paid off or subverted
in post 7 — that anticipation is the follow mechanic.
SHA-256: 53e420441a9378f6dd56c1668da916c95f3362ae5f44f054f2ccd714a9cac42f