{"id":28141,"plugin_id":"plugin_asdk_app_6ab639b28e2081919b704b4ada7e478f","kind":"skill","collection_source":"plugin_package","comparison_source":null,"observed_at":"2026-10-08T18:02:44.653Z","digest":"9189b8294936dc8edc3934239b8f01acfd3a35ef3e6a6cc0eaa521a287c2b3e7","against":null,"payload":{"description":"Check every image and clip against the plan before building on it — the right people, the asked-for action, clean hands and text, lips in sync — and retake a failed shot once.","included_files":[{"relative_path":"agents/openai.yaml","size_in_bytes":375},{"relative_path":"assets/icon.svg","size_in_bytes":398}],"name":"review-shots","skill_md_contents":"---\nname: review-shots\ndescription: \"Check every image and clip against the plan before building on it — the right people, the asked-for action, clean hands and text, lips in sync — and retake a failed shot once.\"\n---\n\nNeeds: the shots' generation_ids, the shot list (or the prompts), and the cast, product or location references the shots were made from\nCredits: the checks of images and of the look of clips are free; a few credits per minute for clips with speech; every retake is paid like the shot it replaces\nTime: a few seconds per image, about a minute per clip checked\n\n## When\n\nInside `script-to-video`, `faceless-story-video`, the five `ugc-*` video skills and\n`product-photoshoot`, where it is on by default, and inside any\nother skill when the user asks for it (\"check them before you show me\"). On its own for shots\nmade earlier, from their `generation_ids` (the calls that made them returned those);\n`read_project` lists a project's media but no generation ids, so an image without its id cannot\nbe previewed and is reported as not reviewed. A clip with only its link can still be checked\nfrom its frames (step 2) when its length is known.\n\n## Ask first\n\nNothing more when a workflow brings it in: that workflow's plan already carries the review\nline and the retake allowance. On its own: which shots, and the retake allowance as a credit\nfigure the user names — `estimate_cost` for one retake of each shot shows what a figure\ncovers; there is no default, since what the shots first cost may not be known.\n\n## Steps\n\n**Before any paid step.** The project: `list_projects` (or `create_project`) gives the\n`project_id` and, when the project has several modules, the `module_id`; every call below that\ntakes a project gets that same pair. A retake is a paid generation like any other: its own\n`estimate_cost`, its `quote_id`, `confirmed_by_user=true`, then `wait_task`. Retakes started\ntogether are shown in one `show_generations` card, with all their ids, before that `wait_task`.\n\n**The retake allowance.** Retakes have their own budget, apart from the figures each stage of\nthe plan was approved for: by default 15% of the plan's image and clip total, shown as its own\nline in the plan and approved with it. Every paid call a retake makes counts against it — the\nnew image, the new clip, on route B both the speech and the lipsync, and the paid recheck of a\nretaken clip with speech. When the next retake would push retake spending past the allowance,\nthe user is asked first.\n\n**0. The checklist, per shot.** Written from the shot list before the first check:\n- **Who** — every character the beat names and no one else, each with the face, hair and\n  outfit of their cast portrait (when the workflow made portraits; otherwise against the\n  reference image or the description); the product with its real label and colour.\n- **What** — the beat's action, setting and framing.\n- **Clean** — whole hands and limbs (no extra fingers, no merged bodies), no garbled lettering\n  or stray watermark, nothing that matters cut off at the frame edge.\n- **Coherence** (clips) — physically and visually coherent, within the clip and across its\n  cuts: identity (faces, bodies and outfits stay the same), props (they persist and keep their\n  state), setting and light (the place, light direction and time of day hold), space (people\n  keep their positions and screen direction), cause and effect (an action has a visible cause\n  and result), physics (every movement could really happen; nothing appears, vanishes or\n  doubles) and artefacts (no warping, melting or flicker).\n- **Continuity** — the same outfit, props and time of day as the neighbouring shots, and each\n  cut into the next shot reads as coherent (step 2).\n- **Sound** (clips with speech only) — the line is heard, in the right language and voice, and\n  the lips move with it.\n\nEvery shot on the list ends with one note: passed, failed (and why), or **not reviewed**. A\nshot is never passed without its own preview or frames having been looked at.\n\n**1. Images (free).** One call returns at most four previews (at most 384 px), so images are\nchecked in groups of four or fewer: `wait_task` on those shots' `generation_ids` (or\n`get_task` once they have finished). Match each preview to its shot through `previews`, which\nlists the links shown in order; a shot whose link is not among them is fetched again on its\nown, and if it still has no preview it is noted **not reviewed**. A preview catches the wrong\nperson, a missing character, broken hands or bodies and garbled large text; fine lettering at\nthat size is the user's check, and the hand-over says so.\n\n**2. The look of clips (free).** Per shot, `extract_frames(video_url, duration, at_seconds=[…])`\nat four moments — half a second in, a third of the way, two thirds of the way, and half a\nsecond before the end, the most one call previews — returns the frames (each with its `url`)\nand previews of them; `previews` lists the frame links shown, in order. A clip is fully\nreviewed only when every moment asked for has a frame whose link is in `previews`. Fewer frames than moments, or a `note` in the result, means a moment came back\nempty — and then no frame says which moment it is — so each moment is asked for again, alone,\na little earlier (a quarter-second before); a frame without a preview is asked for again the\nsame way. Then the same checklist, without the sound. `duration` is the clip's length from\n`wait_task` (`duration_seconds`). A clip with a moment still unseen is noted as partly\nreviewed, and one with none seen is noted **not reviewed**. It takes up to a minute per clip\nand charges nothing.\n\n**The cuts.** Once two neighbouring clips are both checked, the cut between them is judged\nfrom the earlier clip's last frame and the later clip's first frame, both already extracted.\nWhen a cut jumps in time, angle or distance, the jump has to read as one coherent film: the\nsame characters on the same sides of the frame, the same light and time of day, props in the\nstate the earlier shot left them, and a change of time or place a viewer can follow. A cut that\njars fails the later shot, noted with what jumps, and that shot is retaken (step 4).\n\n**3. Clips with speech (a few credits per minute).** Per clip with dialogue or lip-sync,\n`analyze_media(media_url, question=…, confirmed_by_user=true)` — the one paid call with no\nquote: it is charged by clip length, a few credits per minute; the user agreed to it with the\nplan, and the finished run (`wait_task`) reports `credits_charged_estimate`. The question asks\nfor description, not a verdict — the tool sees only the clip, never the cast portraits:\n\"Transcribe what is said and by whom. Do the speaker's lips move in time with the words? Say\nyes or no, and where it drifts. Describe each person's face, hair and clothes.\" The assistant\ncompares the answer with the shot's line and cast. Analysis, ffmpeg and motion-graphics runs\nshare one lane per module: while one is running, another is refused with `chat_busy`\n(retryable), so the clips are analysed one after another, not between generations. Each\nanswer is also recorded in the module in Frameo. Every `analyze_media`, `run_ffmpeg` and\n`render_motion_graphics` call counts against 30 runs per clock hour for the user, so a long\nworkflow analyses and cuts part by part: a part's clips are analysed right before that part is\ncut, in the same hour, and its analyses, rechecks and cut calls together stay within 30\n(`script-to-video` sizes its parts that way).\n\n**4. Retake a failed shot, once.** Same `shot_number`, so the retake lands in the shot's own\ncanvas row. What changes depends on the failure and on what failed:\n- **An image** with the wrong or drifting face — `generate_image` again with the cast portrait\n  first in `image_urls` and the character named with their look in the prompt.\n- **A clip** with the wrong or drifting face — the fix starts from its image: remake the beat's\n  image as above, check it, then animate the new image (`first_frame_url`) as the clip was\n  first made. On Seedance models the cast portraits go along as `reference_image_urls` (the\n  frame counts as one of the model's reference images); every other model refuses a first frame\n  together with reference images, so there the image alone carries the face.\n- A clip with no image of its own (a `continues` beat) — retaken from the previous clip's last\n  frame, as it was first made. When a retaken clip has a `continues` beat after it, that beat is\n  remade from the new last frame too, from the same retake allowance.\n- A cut that jars — the later clip remade to follow on from the earlier one: from the earlier\n  clip's last frame (`first_frame_url`) when the action carries straight on, otherwise from its\n  image remade with the positions, light and props the earlier shot ends on.\n- Broken hands — a framing that keeps hands out of shot or a simpler gesture.\n- Garbled text — no lettering in the prompt; add it at the cut with `run_ffmpeg`.\n- Wrong action — the action first in the prompt, in plain words.\n- Lips out of sync on a route-A clip — route B for that beat: `generate_speech`, then\n  `generate_lipsync` at the clip's `aspect_ratio` and `resolution`, never above 15 s, each with\n  its own quote.\n  The lipsync's `audio_duration` is the speech result's `duration_seconds` (without one, 2.8\n  words per second, 2.5 in Hindi, plus one second — approximate), and the clip keeps the\n  `duration` it was made at: whole seconds inside the model's range. Route B can cost more\n  than the clip it replaces; the quotes say how much before it runs.\n\nEach retake has its own quote: `estimate_cost` for each paid call of it (with `reference_count` =\nthe number of `image_urls` or `reference_image_urls` it passes), its `quote_id`,\n`confirmed_by_user=true` while it fits the allowance, `wait_task`, and the same checklist\nagain. A shot that still fails — a route-B clip included — keeps the better take and is\nreported; it is not retaken a third time.\n\n**5. The finished cut (free).** When the workflow ends in a cut, `extract_frames` on the\nfinished video at the middle of each shot, timed from the shot list's `length` column (or each\nclip's `duration` where the workflow keeps no such column), and half a second before the end,\nfour moments a call (the most one call previews) with the end moment among them. Each frame has to show its shot's\naction: a trim that cut the action away fails the cut, and that shot is cut again with a\n`length` that keeps it. The last frame has to show the story's last moment, not an empty frame\nor a camera that drifted off it; a closing shot that drifts is cut again shorter. A re-cut\nchanges only the cut and costs no credits, but it redoes that shot's mix pass, its group join and\nthe final join: at least three runs from the hour's 30.\n\n## Done\n\nA review note per shot, handed over with the result: passed, retaken (what was wrong and what\nchanged), kept with a known flaw the user may want to fix in the Frameo app, or not reviewed\n(and why), and one for the finished cut. Retakes sit in the same canvas row as the shot they replaced.\n\n## Files\n\nNone.\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}