← Files Creatify Ad AgentARCHIVED FILE
skills/ad-agent/references/product-fidelity.md
8.1 KB · Oct 2, 2026 · 18:02 UTC
# Product fidelity protocol: when the references are a product or an app screen
Product fidelity (shape, parts, label and on-screen text) is the thing Seedance loses on, and it's the thing you must win on. Footage is `boreal-h3` only (`composer_generate_clip`): **every delivered video contains H3-rendered footage.** A still image animated in assembly (Ken Burns on the reference, a held screenshot) is never the whole deliverable. **Proof stays real pixels:** a reference the viewer must trust as real (a before/after, a finished job, the actual dish, the real venue or customer) is shown as itself with a move in the page, never re-rendered by H3. H3 re-draws it, and a re-drawn result is a false claim; H3 carries the shots around it. These rules decide how to use H3:
1. **Frame 0 carries the exact product: i2v from a keyframe, not free r2v.** r2v reimagines the product from the reference; i2v starts from pixels you control.
- The reference already IS the shot (packshot, product on a counter, app screenshot) and the brief only adds motion → the reference is the keyframe. Bring it to the output aspect by padding with its own background color (flat backgrounds: the pad recipe in media-recipes.md) or by outpainting with `composer_edit_image`.
- The brief puts the product into a new scene → build the keyframe with `composer_edit_image`, product image first.
- **Frame the product large.** Label text that is small in the keyframe (a wordmark a few dozen pixels wide in a wide shot) comes out garbled in motion ("OLIOM" → "ouom"). Put the product big in the frame (lower third of a medium shot, or a close-up) whenever its text or parts matter. Add a close-up cut if the brief wants a wide.
- The keyframe must satisfy the brief's people too: state their demographics, age and wardrobe from the prompt in the edit instruction ("Western customer" means you write it; an unstated person comes back arbitrary).
2. **Image models rewrite small text. Paste the real pixels back.** After an outpaint or any edit that keeps the product where it was, composite the original reference back over the product area with a feathered mask (a small python/ffmpeg script via `composer_exec`). Before rendering, `composer_view` a zoomed crop of the keyframe's label/logo/UI text (crop recipe in media-recipes.md) and compare it to the reference word by word. If a letter changed, fix the keyframe first. A garbled keyframe can only get worse in motion.
3. **Camera moves are anchored at both ends, so they stay real AND faithful.** A locked-off camera with a digital zoom passes fidelity but reads as a still photo, and loses the viewer (see dynamism.md). For a push-in, dolly or reveal, make the END frame yourself from real pixels: a crop of the keyframe around the product/logo, upscaled to the canvas (the anchored push-in recipe in media-recipes.md). Then render i2v with `image` = keyframe and `last_image` = that crop. H3 then renders a true camera move with parallax that starts and lands on exact product pixels. Same trick for an arc or a tilt: the end frame is a product view you trust (a second reference angle, or a `composer_edit_image` of the product that passed your crop check). Crop the middle of the clip at full resolution against the reference, and re-roll with a new `seed` if a letter drifted mid-move. Digital moves in assembly are for extra punch on top of H3 motion, never a substitute for it.
4. **Handling the product's parts (lids, closures, openings) → short anchored segments.** One long take morphs the part (a removable glass lid turns hinged). Keyframe each product state with `composer_edit_image`; each new keyframe gets the product image AND the previous keyframe in `images`, so the person and scene carry over. Then render ≤5s i2v segments with `image` = this state and `last_image` = the next state, and concat them in the page. Every seam is then pinned to an exact product frame.
5. **On-screen UI text is never generated.** H3 cannot type or keep UI labels intact, and on a static screen it invents decorations (glowing rings, stars, extra icons, a living room behind a dark UI). Render the screen as H3 i2v footage from the screenshot keyframe with a static prompt ("flat, perfectly static screen capture of this exact page; nothing moves or changes; no people, no room, no reflections, no added graphics, no cursor, no typing"). Compare full-resolution crops of its middle and last frames with the screenshot, and re-roll the seed until nothing was added or changed (≤3 takes; keep the cleanest). Then build the interaction (typed text, caret, cursor, button press) as overlays at pixel positions you measure on the screenshot (sample the box fill color to cover placeholder text). Typed text must read at a glance to a viewer AND to a judge sampling frames at low resolution: 2px larger than the UI's own label size and weight 600, never below ~20px on a 1080p frame. Thin letters get misread ("thermos" read as "themos").
- **A phone in someone's hand showing the app → `lib.screen_replace`, never a rectangle over a bounding box.** Generate the shot i2v from a keyframe that already has the real screenshot on the phone (H3 keeps the layout; it will garble the text), then put the real screen back, tracked to the phone's four corners, via `composer_exec`:
```bash
python3 -m lib.screen_replace gen/shot.mp4 --screen assets/<app.png> [--anchor <the screenshot the keyframe used>] \
[--screen2 assets/<next.png> --switch-at <s>] [--screen-video assets/<recording.mp4> --screen-in <s>] \
--out gen/shot_screen.mp4 --sheet out/check/shot_screen.jpg
```
It matches the anchor screenshot to every frame (so the screen follows tilt, perspective and push-ins), warps the real screen in with rounded corners, keeps a thumb that crosses the screen edge on top, and flashes a camera shutter at `--switch-at`. `composer_view` the `--sheet` (the tracked quad is drawn in yellow): the quad must sit on the screen edge inside the bezel in every tile. `lost` > 0, or a quad off the screen, means the anchor is wrong (pass the screenshot the keyframe actually used) or the phone turns too far; re-roll with a steadier hold. Use the `_screen.mp4` as the clip, with no overlay `<img>` on the phone in the page. It may run past 50 s; poll the returned task.
- H3 still nudges small UI text on a static render (a label smudged, a word struck through, a spinner added). After the take passes, still crop every text line of the screenshot (labels, buttons, placeholder, header) and compare it with the same crop of the H3 clip's middle and last frames. For every region that changed, lay the screenshot's own pixels over it in the composition for the whole clip (a positioned `<img>` crop in `index.html`); the rest of the frame stays H3 footage. Audio: quiet room tone, soft keyboard clicks, or none.
6. **Judge fidelity yourself before shipping, at full resolution, against the reference.** Extract 3 frames spread over every product/app shot, crop the product/logo/label region at full resolution, and `composer_view` each next to the same crop of the reference (`composer_exec`: `python3 -m lib.compose sheet index.html --at <t> --crop x,y,w,h --width 900 -o out/check/label.jpg`). Any letter or color that differs, any added/missing part, any morph is a defect (the last seconds too: a hero shot's colors drift as its light moves). Pull a lever above (bigger product, pasted pixels, locked camera + digital move, anchored segments, overlay, `lib.screen_replace`) and re-render the failing part. Do not ship a known fidelity defect.
7. **Audio is part of the deliverable.** H3 generates a soundtrack, and without dialogue in the brief it still invents mumbled "voiceover" gibberish. When the brief has no dialogue, say in the prompt "no speech, no voices, no voiceover — only <ambience>", and check the clip with `composer_transcribe(project_id, media="gen/shotN.mp4")`: it must hear no words. Garbled speech → regen with that line, or replace the track with quiet ambience in assembly. Never reverse or ping-pong audio. If you loop or boomerang footage to fill the duration, the audio stays forward (loop a clean ambience bed or crossfade), because a reversed soundtrack is an instant audio FAIL.
SHA-256: 30cd25e5b41aebbc0b1f3da009e364b4c15dd0be311fee3953883e761a01853d