← Files FrameoARCHIVED FILE

skills/review-shots/SKILL.md

10.9 KB · Oct 10, 2026 · 06:03 UTC

↓ Download file

---
name: review-shots
description: "Check every image and clip against the plan before building on it — the right people, the asked-for action, clean hands and text, lips in sync — and retake a failed shot once."
---

Needs: the shots' generation_ids, the shot list (or the prompts), and the cast, product or location references the shots were made from
Credits: the checks of images and of the look of clips are free; a few credits per minute for clips with speech; every retake is paid like the shot it replaces
Time: a few seconds per image, about a minute per clip checked

## When

Inside `script-to-video`, `faceless-story-video`, the five `ugc-*` video skills and
`product-photoshoot`, where it is on by default, and inside any
other skill when the user asks for it ("check them before you show me"). On its own for shots
made earlier, from their `generation_ids` (the calls that made them returned those);
`read_project` lists a project's media but no generation ids, so an image without its id cannot
be previewed and is reported as not reviewed. A clip with only its link can still be checked
from its frames (step 2) when its length is known.

## Ask first

Nothing more when a workflow brings it in: that workflow's plan already carries the review
line and the retake allowance. On its own: which shots, and the retake allowance as a credit
figure the user names — `estimate_cost` for one retake of each shot shows what a figure
covers; there is no default, since what the shots first cost may not be known.

## Steps

**Before any paid step.** The project: `list_projects` (or `create_project`) gives the
`project_id` and, when the project has several modules, the `module_id`; every call below that
takes a project gets that same pair. A retake is a paid generation like any other: its own
`estimate_cost`, its `quote_id`, `confirmed_by_user=true`, then `wait_task`. Retakes started
together are shown in one `show_generations` card, with all their ids, before that `wait_task`.

**The retake allowance.** Retakes have their own budget, apart from the figures each stage of
the plan was approved for: by default 15% of the plan's image and clip total, shown as its own
line in the plan and approved with it. Every paid call a retake makes counts against it — the
new image, the new clip, on route B both the speech and the lipsync, and the paid recheck of a
retaken clip with speech. When the next retake would push retake spending past the allowance,
the user is asked first.

**0. The checklist, per shot.** Written from the shot list before the first check:
- **Who** — every character the beat names and no one else, each with the face, hair and
  outfit of their cast portrait (when the workflow made portraits; otherwise against the
  reference image or the description); the product with its real label and colour.
- **What** — the beat's action, setting and framing.
- **Clean** — whole hands and limbs (no extra fingers, no merged bodies), no garbled lettering
  or stray watermark, nothing that matters cut off at the frame edge.
- **Coherence** (clips) — physically and visually coherent, within the clip and across its
  cuts: identity (faces, bodies and outfits stay the same), props (they persist and keep their
  state), setting and light (the place, light direction and time of day hold), space (people
  keep their positions and screen direction), cause and effect (an action has a visible cause
  and result), physics (every movement could really happen; nothing appears, vanishes or
  doubles) and artefacts (no warping, melting or flicker).
- **Continuity** — the same outfit, props and time of day as the neighbouring shots, and each
  cut into the next shot reads as coherent (step 2).
- **Sound** (clips with speech only) — the line is heard, in the right language and voice, and
  the lips move with it.

Every shot on the list ends with one note: passed, failed (and why), or **not reviewed**. A
shot is never passed without its own preview or frames having been looked at.

**1. Images (free).** One call returns at most four previews (at most 384 px), so images are
checked in groups of four or fewer: `wait_task` on those shots' `generation_ids` (or
`get_task` once they have finished). Match each preview to its shot through `previews`, which
lists the links shown in order; a shot whose link is not among them is fetched again on its
own, and if it still has no preview it is noted **not reviewed**. A preview catches the wrong
person, a missing character, broken hands or bodies and garbled large text; fine lettering at
that size is the user's check, and the hand-over says so.

**2. The look of clips (free).** Per shot, `extract_frames(video_url, duration, at_seconds=[…])`
at four moments — half a second in, a third of the way, two thirds of the way, and half a
second before the end, the most one call previews — returns the frames (each with its `url`)
and previews of them; `previews` lists the frame links shown, in order. A clip is fully
reviewed only when every moment asked for has a frame whose link is in `previews`. Fewer frames than moments, or a `note` in the result, means a moment came back
empty — and then no frame says which moment it is — so each moment is asked for again, alone,
a little earlier (a quarter-second before); a frame without a preview is asked for again the
same way. Then the same checklist, without the sound. `duration` is the clip's length from
`wait_task` (`duration_seconds`). A clip with a moment still unseen is noted as partly
reviewed, and one with none seen is noted **not reviewed**. It takes up to a minute per clip
and charges nothing.

**The cuts.** Once two neighbouring clips are both checked, the cut between them is judged
from the earlier clip's last frame and the later clip's first frame, both already extracted.
When a cut jumps in time, angle or distance, the jump has to read as one coherent film: the
same characters on the same sides of the frame, the same light and time of day, props in the
state the earlier shot left them, and a change of time or place a viewer can follow. A cut that
jars fails the later shot, noted with what jumps, and that shot is retaken (step 4).

**3. Clips with speech (a few credits per minute).** Per clip with dialogue or lip-sync,
`analyze_media(media_url, question=…, confirmed_by_user=true)` — the one paid call with no
quote: it is charged by clip length, a few credits per minute; the user agreed to it with the
plan, and the finished run (`wait_task`) reports `credits_charged_estimate`. The question asks
for description, not a verdict — the tool sees only the clip, never the cast portraits:
"Transcribe what is said and by whom. Do the speaker's lips move in time with the words? Say
yes or no, and where it drifts. Describe each person's face, hair and clothes." The assistant
compares the answer with the shot's line and cast. Analysis, ffmpeg and motion-graphics runs
share one lane per module: while one is running, another is refused with `chat_busy`
(retryable), so the clips are analysed one after another, not between generations. Each
answer is also recorded in the module in Frameo. Every `analyze_media`, `run_ffmpeg` and
`render_motion_graphics` call counts against 30 runs per clock hour for the user, so a long
workflow analyses and cuts part by part: a part's clips are analysed right before that part is
cut, in the same hour, and its analyses, rechecks and cut calls together stay within 30
(`script-to-video` sizes its parts that way).

**4. Retake a failed shot, once.** Same `shot_number`, so the retake lands in the shot's own
canvas row. What changes depends on the failure and on what failed:
- **An image** with the wrong or drifting face — `generate_image` again with the cast portrait
  first in `image_urls` and the character named with their look in the prompt.
- **A clip** with the wrong or drifting face — the fix starts from its image: remake the beat's
  image as above, check it, then animate the new image (`first_frame_url`) as the clip was
  first made. On Seedance models the cast portraits go along as `reference_image_urls` (the
  frame counts as one of the model's reference images); every other model refuses a first frame
  together with reference images, so there the image alone carries the face.
- A clip with no image of its own (a `continues` beat) — retaken from the previous clip's last
  frame, as it was first made. When a retaken clip has a `continues` beat after it, that beat is
  remade from the new last frame too, from the same retake allowance.
- A cut that jars — the later clip remade to follow on from the earlier one: from the earlier
  clip's last frame (`first_frame_url`) when the action carries straight on, otherwise from its
  image remade with the positions, light and props the earlier shot ends on.
- Broken hands — a framing that keeps hands out of shot or a simpler gesture.
- Garbled text — no lettering in the prompt; add it at the cut with `run_ffmpeg`.
- Wrong action — the action first in the prompt, in plain words.
- Lips out of sync on a route-A clip — route B for that beat: `generate_speech`, then
  `generate_lipsync` at the clip's `aspect_ratio` and `resolution`, never above 15 s, each with
  its own quote.
  The lipsync's `audio_duration` is the speech result's `duration_seconds` (without one, 2.8
  words per second, 2.5 in Hindi, plus one second — approximate), and the clip keeps the
  `duration` it was made at: whole seconds inside the model's range. Route B can cost more
  than the clip it replaces; the quotes say how much before it runs.

Each retake has its own quote: `estimate_cost` for each paid call of it (with `reference_count` =
the number of `image_urls` or `reference_image_urls` it passes), its `quote_id`,
`confirmed_by_user=true` while it fits the allowance, `wait_task`, and the same checklist
again. A shot that still fails — a route-B clip included — keeps the better take and is
reported; it is not retaken a third time.

**5. The finished cut (free).** When the workflow ends in a cut, `extract_frames` on the
finished video at the middle of each shot, timed from the shot list's `length` column (or each
clip's `duration` where the workflow keeps no such column), and half a second before the end,
four moments a call (the most one call previews) with the end moment among them. Each frame has to show its shot's
action: a trim that cut the action away fails the cut, and that shot is cut again with a
`length` that keeps it. The last frame has to show the story's last moment, not an empty frame
or a camera that drifted off it; a closing shot that drifts is cut again shorter. A re-cut
changes only the cut and costs no credits, but it redoes that shot's mix pass, its group join and
the final join: at least three runs from the hour's 30.

## Done

A review note per shot, handed over with the result: passed, retaken (what was wrong and what
changed), kept with a known flaw the user may want to fix in the Frameo app, or not reviewed
(and why), and one for the finished cut. Retakes sit in the same canvas row as the shot they replaced.

## Files

None.

SHA-256: 3ddcfa041167851abf3516bad622dffe80141d524e1a645eec30da529b36ac19