← Files FrameoARCHIVED FILE
skills/ugc-try-on-video/SKILL.md
9.3 KB · Oct 10, 2026 · 06:03 UTC
---
name: ugc-try-on-video
description: "Wearing and fit-check for clothing, jewellery, glasses or shoes: the creator in the item, a turn, a close-up, a verdict."
---
Needs: product photos; a creator photo (the person who wears it) or a description; size or fit notes
Credits: about 500–650 for 30 seconds, plus a retake allowance of 15% of the image and clip total
Time: 15–25 minutes, plus about a minute per clip for the review
## When
Wearable products. The creator's photo matters more here than anywhere else: the item has
to sit on a consistent person. Not for products that are held or used (`ugc-review-video`).
## Ask first
1. **The product** — photos (`show_upload` for files on the user's device, `import_media_url` for a web link; images attached to the chat do not reach Frameo, and `create_upload_url` is for clients that send the file themselves), the name, and the one thing the video
must say about it.
2. **The creator** — design one (age, vibe, setting), use the user's photo, or reuse a creator
already in the project (`list_characters`: its first image is the step-1 still and its
`voice_id` the voice in step 3, with nothing to generate or choose).
3. **Frame and length** — 9:16, 15 / 30 / 45 s.
Say the rough cost in the same turn as the questions; the user's yes comes with the plan (Steps, **The plan first.**).
## Steps
**Before any paid step.** The project: `list_projects` (or `create_project`) gives the `project_id`
and, when the project has several modules, the `module_id`; every call below that takes a project —
`estimate_cost` included — gets that same pair; without it those tools answer `project_needed`. The
quote: one `estimate_cost(items=[…])` prices a stage in one call, with the same project, model,
size and number of `image_urls` or `reference_image_urls` (`reference_count`) as each generate call, and returns a `quote_id` per item plus the total. Each generate call
then passes its own item's `quote_id` and `confirmed_by_user=true`. A quote is single-use and lasts
15 minutes, so a long plan is priced stage by stage, right before each stage runs; a stage that
comes to more than the user approved is asked about again first. Generate calls return
`generation_ids`; `wait_task` returns the links, and `show_generations` shows each stage's running
and finished work in one card where the chat app displays Frameo cards: all the stage's ids at
once, before its first `wait_task`, and the finished result with `final: true`.
**The plan first.** Before the first paid call, the plan goes to the user in the chat as plain
text: what will be made, in order, one line per generation (for a script, the shot list; for a
set, each shot), with each line's credits and the total. The credits come from
`estimate_cost(items=[…])`, up to 10 items a call, so a long plan takes several calls; those
quotes may expire unused, since each stage is quoted again right before it runs. Nothing is
generated until the user says yes. The user can drop or change lines; a changed line is priced
again.
**Canvas rows.** Pass `shot_number` on every `generate_image` and `generate_video` of a shot (1, 2, 3… in story order; the same number for a retake and for that shot's video), so each shot gets its own row on the Frameo canvas. Cast, prop and location references take `placement_kind` (`character`, `prop` or `location`) and the subject's name as `placement_group` instead: they sit on their own board, and the shots built from them do not pile into their row.
**Review (on by default).** Every still and clip is checked before it is built on, and the
finished cut before it is handed over, following `review-shots`: its checklist and steps come
from `get_skill("review-shots")`, loaded before step 1. The checks run on the step-1 still
before step 2 builds on it, after step 2 (`review-shots` steps 0, 1 and 4), after step 4 (steps 2 and 4) and after the cut (step 5); the
product's label, shape and colour are checked against the product photos in every shot, and a
failed shot is retaken once. The plan carries the retake allowance as its own line for the user
to approve with everything else. The paid check of clips with speech (`review-shots` step 3)
runs only when the user asks for it, since the talking beats are lipsynced to one fixed voice,
and its rough cost (a few credits per minute of clip) is said and agreed first; without it, the
note on each talking clip says its sound was not checked.
The user can turn the review off.
**0. Script.** Write a script of the length chosen in Ask first in the creator's voice: hook in the first line, the one
claim, a close. Show it; the user edits before anything is generated.
**1. Creator still (~18).** `generate_image`: the creator full-length in front of a mirror or a
plain wall, wearing the item, phone-video framing. `image_urls=[creator photo, product
photos…]` in that order: the person is the base, the item the reference. With no creator
photo, the first product photo is the base and the prompt describes the person.
**2. Product-in-scene stills (~18 each).** `generate_image` for each beat that shows the item on
the creator (front, a half turn, a detail) with `image_urls=[step-1 creator still, product
photos…]` — the same person as the base every time, the item as the reference.
**3. Voice (~1 per 50 characters).** First the beat list, no tool call: estimate each beat
at 2.8 words per second (2.5 for Hindi) plus one second of headroom, say it is
approximate, and split any beat that is too long — the fit beats at the video model's max
(`list_models(kind="video")`; 30 s on the default) because the mix at the cut cannot stretch
the picture, the talking beats at the smaller of that max and 15 s, the lipsync's cap —
until every piece is within it, so every clip's `duration` is its length on the timeline. Then
`search_voices` for the creator's age and language, 3 candidates, then `generate_speech` per final beat (each
paid call, here and below, has its own quote); `wait_task` for the links. The creator the user keeps is saved with `save_character` (a name,
`image_urls=[creator still]` and the chosen `voice_id`), so the next video in this project
reuses the same face and voice. Each speech
result's `duration_seconds` is its measured spoken length; from here on the estimate means that
length when the result has one, and the words-per-second figure only when it has none. That estimate is the clip's `duration` in step 4 and the `duration` / `audio_duration`
of its lipsync.
**4. Clips (~65 per 5 s + ~20 per second of lipsync).** Verdict beats: `generate_video` from the creator still, then `generate_lipsync` with the step-3 speech. Fit beats: `generate_video` from the try-on stills with 'slow half turn' / 'steps closer' in the prompt, silent. Every `generate_video` takes `duration` = that beat's speech estimate rounded up to whole
seconds and kept inside the model's range (`list_models(kind="video")` gives it; 4–30 s on
the default), 5 s for a beat with no speech; every `generate_lipsync` takes that same whole
`duration` (never above 15 s — the lipsync's own cap, whatever the video model allows),
`audio_duration` = the estimate, and the clip's `aspect_ratio` and `resolution`. `wait_task` each. Test the first beat before
the rest.
On Seedance models (the default) the product photos go along as `reference_image_urls` on every
clip that shows the product, and the creator still on every product clip the creator is in,
which holds the label and the face through the motion. The frame is the first reference image,
so the prompt names the others from the second image on, and it counts toward the model's cap;
the quote carries `first_frame=true` and `reference_count` = the number of
`reference_image_urls`, not counting the frame. Each
later piece of a split beat opens on the previous piece's last frame (`extract_frames` on the
clip that goes into the cut, the lipsync result on a talking beat) once that piece's look has
been checked, instead of the beat's still, so the join reads as one take. Each prompt names where
the clip ends and nothing the beat does not show: on a lipsynced beat the creator facing the
lens; on a silent beat the product in frame with the creator's eyes on it, never on the lens,
since the voiceover plays over lips that do not move.
**5. Cut (free).** `run_ffmpeg` takes at most 10 inputs and 4 outputs, so the cut is passes, `wait_task`
after each (its outputs are the next pass's inputs): per beat — lipsynced beats bring only their clip — mix its step-3 speech over its clip (`amix`, speech on top) into `beatN.mp4` at exactly
the clip's `duration` (`tpad`/`apad` over-pad, `-t <duration>` sets the length; captions are
timed from those durations), up to four beats per call (one output
each) and fewer when their files would pass ten inputs
(a clip plus its speech is two inputs and four such beats fill a call); then concatenate all the beats in
order in groups of at most 10; then one last call joins the groups and burns captions
from an `.ass` sidecar in a casual style. Output `final.mp4`, review the finished cut (`review-shots` step 5), and return the link
and `open_in_frameo`.
## Done
The UGC clip, the script, and every still and clip in the Frameo project, with the review note
per shot: passed, retaken and why, or kept with a known flaw.
`open_in_frameo` on the result opens the project in the Frameo app, where every take sits on
the canvas, ready for retakes and edits.
## Files
None.
SHA-256: c99d50aede1c6acb0d10b3bdcbfd133da420dae88e3c604571dc9fb3c3cb583b