← Files PixVerseARCHIVED FILE
skills/pixverse-cover-art/references/thumbnail-recipes.md
6.17 KB · Oct 5, 2026 · 18:29 UTC
# Thumbnail Concept And Rendering
## Select A Truthful Information Gap
Read the actual video promise and brainstorm at least five concepts internally across
these 16 frameworks before selecting the strongest. Combining frameworks is useful only
if the result reads within one second at about 120px wide.
Before/after; generic social-message UI; three-step progression; a compelling real video
frame; posed portrait; posed action; a real significant day; diagram/graphical relation;
landscape hero; map/aerial; product hero; exact-text callout; repeated objects; size contrast;
generic news-clip treatment; amplified reality around one actual story element.
Use real source pixels for a compelling screenshot. Generic social/news treatments must
not fabricate messages, reviews, news events or a broadcaster endorsement. Amplification
still has to represent the video's truth. A graphical concept replaces photoreal language.
Do not choose a split merely because a title contains “versus”; use an explicit split,
before/after/side-by-side request or an accepted split reference.
## Input Roles And Counts
User choices override analyzed reference fields, which override defaults. A style-reference
thumbnail is analyzed for brief, pose/action, key elements, location, composition,
background, split/panel count, people count and expression; it is never an identity source
or generation attachment. Respect already stated match/inspiration intent. Match preserves
the selected composition/style; unique take uses only loose direction. Do not add strangers
because a style reference contains more people than supplied character references.
For 0–3 supplied people/characters, bind each exact identity in stable order. A people-led
concept without a person reference requires a resolved choice of supplied identity,
generated character or people-free concept. Delegated people-free design can proceed;
never silently substitute a new face for an attached one. Faces precede the optional logo.
State actual reference order when multiple inputs exist. Keep logo and style roles distinct.
Respect exact variant count. A set is distinct concepts/emotions/camera takes, one prompt
per final image, not one contact sheet. An emotion×take matrix is deliberate and capped
at 16 in this recipe; it must not multiply an explicit count. Emotion options:
shock, hype, rage, awe, laugh, fear, smug, charisma, confusion, determination, disgust.
Shock is the unresolved person-expression default, but explicit calm/natural direction
and truthful content win. Preserve the chosen emotion within a take set.
## Eleven Prompt Blocks
Use Sunburst 2K/high under the image quality policy, default 16:9 or requested 9:16/4:5.
Keep any explicit user model/format choice. Write resolved English direction and exact
literal on-image text in its original language. The main prompt order is:
1. Frame: bold high-impact unified thumbnail, one coherent image; explicit split only
when selected. For vertical, keep faces in the upper two-thirds. Non-photoreal media
keep their exact material/design lock instead of photoreal finish.
2. Scene: truthful subject and event from the brief.
3. Text policy: no text/readable UI labels/watermark unless text is explicitly requested.
4. Subjects: one lock per referenced person, preserving facial structure, eyes/nose/lips,
skin/hair and design, without beautifying/averaging; distinct pose and expression.
5. Key elements: only signature objects that explain the gap.
6. Logo when supplied: exact outline, letterforms, proportions and color, clear of faces.
7. Location when relevant: real spatial context, time/weather and atmosphere.
8. Composition: dominant foreground hero, usually 40–60% of frame, power-third placement,
clear scale hierarchy and separated planes; do not clutter the competing area.
9. Background: vivid subject-related field/environment with cohesive contrast and edge
falloff unless explicitly muted; no accidental divide.
10. For photoreal people: sculpting key, soft shadow fill, defined back/hair rim. Colored
rim only when chosen; do not tint the entire face by habit.
11. Cohesive bright poster finish, crisp highlights/deep blacks and readable faces, adapted
to the user's medium. Illustrated handoffs omit photographic rig/gloss.
Do not attach a style-only thumbnail to the model. A split's panels need independent
subject/state descriptions, an exact divider and no identity bleed. A requested dimensional
logo look uses a separate generated image reference: show plausible depth while preserving
the supplied logo's outline, letters and colors. Reuse that accepted image in relevant
finals. Do not add this stage to a flat-logo request.
## Text And Surgical Revisions
No headline by default. Explicit overlay uses a controlled text layer; explicit baked text
is integrated into the generation and checked character-for-character. An automatically
chosen text-friendly framework does not authorize added words. A calling episode's locked
headline/text-bake choice is already answered. If baked glyphs fail, create/reuse a clean
equivalent and add exact text deterministically, recording any new paid image attempt.
For an overlay, use 2–4 words, keep off faces, cap height roughly 12–18%H, comfortable
margins, tight tracking and ~0.9 line height. Bold condensed type needs its 8–14% cap-size
stroke painted UNDER fill, a dark offset shadow and adequate contrast. Five choices:
Beast (white/black stroke), Fire (yellow→orange→red), Neon Lime (acid lime/glow), Clean
Glass (heavy sans on a restrained translucent plate), Marker (black on lime line boxes).
Use FFmpeg with actual font files/glyphs for controlled text overlays and export the
final PNG at native resolution. Inspect the actual exported pixels.
For expression-only, background-only, background-color-only or rim-color-only revisions,
use the picked final as the single image source and protect all other properties. The
edit uses Sunburst 2K/high with the accepted image as its reference; keep the same
aspect and protect all unchanged content. Do not downgrade on failure.
Use accepted edits as the next source. Inspect identity, exact text, logo and 120px clarity;
preserve successful variants. Do not claim click-through improvement without an experiment.
SHA-256: 6a2a815b49e8918a3ea833beb92509bd4cd548c64ccd5262525809e5ee2e76cc