← FrameoCONTENT HISTORY

Update to Frameo

Snapshot Oct 8, 2026 · 18:02 UTC · version 1.2.1

Collection source: downloaded plugin package.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "description": "Turn a script into a finished video, such as a micro-drama — a consistent cast, one shot per beat, characters speaking their lines, cut together with subtitles.",
  "included_files": [
    {
      "relative_path": "agents/openai.yaml",
      "size_in_bytes": 376
    },
    {
      "relative_path": "assets/icon.svg",
      "size_in_bytes": 487
    },
    {
      "relative_path": "shot-list.md",
      "size_in_bytes": 2232
    },
    {
      "relative_path": "subtitles.md",
      "size_in_bytes": 753
    }
  ],
  "name": "script-to-video",
  "skill_md_contents": "---\nname: script-to-video\ndescription: \"Turn a script into a finished video, such as a micro-drama — a consistent cast, one shot per beat, characters speaking their lines, cut together with subtitles.\"\n---\n\nNeeds: a script (scenes, characters, dialogue or narration); optionally photos of the people or a style reference\nCredits: about 2,000–3,000 for a 12–16 shot piece whose speakers have several lines each (speech + lipsync is the larger half; see Steps for the breakdown)\nTime: 20–40 minutes, mostly waiting on generations\n\n## When\n\nUse for any script with named characters who speak or act on screen: a micro-drama or\nserial episode, a short film, a dialogue-driven ad, a kids' story, a sketch, a training\nscenario. Any length the user wants, in clips — a long script is delivered in parts (step 6 sets\nthe size); 9:16 by default, 16:9 on request.\n\nNot for: narration over stills with no on-screen speaker (that is `faceless-story-video`),\none person talking to camera (`talking-presenter`), or editing footage the user already has\n(`subtitles-and-cut`).\n\n## Ask first\n\nThree questions, one turn, before anything is generated:\n\n1. **Aspect** — 9:16 for Reels/Shorts (default) or 16:9.\n2. **Look** — photoreal, stylised, or anime; any reference image or a Frameo project whose\n   look to match.\n3. **Cast** — photos of the people (`show_upload` for files on the user's device, `import_media_url` for a web link; images attached to the chat do not reach Frameo, and `create_upload_url` is for clients that send the file themselves) or design the cast\n   from the script's descriptions.\n\nThen say the rough cost from the script (Steps, below) in the same turn as the questions; the user's yes comes with the plan (Steps, **The plan first.**).\n\n## Steps\n\n**Before any paid step.** The project: `list_projects` (or `create_project`) gives the `project_id`\nand, when the project has several modules, the `module_id`; every call below that takes a project —\n`estimate_cost` included — gets that same pair; without it those tools answer `project_needed`. The\nquote: one `estimate_cost(items=[…])` prices a stage in one call, with the same project, model,\nsize and number of `image_urls` or `reference_image_urls` (`reference_count`) as each generate call, and returns a `quote_id` per item plus the total. Each generate call\nthen passes its own item's `quote_id` and `confirmed_by_user=true`. A quote is single-use and lasts\n15 minutes, so a long plan is priced stage by stage, right before each stage runs; a stage that\ncomes to more than the user approved is asked about again first. Generate calls return\n`generation_ids`; `wait_task` returns the links, and `show_generations` shows each stage's running\nand finished work in one card where the chat app displays Frameo cards: all the stage's ids at\nonce, before its first `wait_task`, and the finished result with `final: true`.\n\n**The plan first.** Before the first paid call, the plan goes to the user in the chat as plain\ntext: what will be made, in order, one line per generation (for a script, the shot list; for a\nset, each shot), with each line's credits and the total. The credits come from\n`estimate_cost(items=[…])`, up to 10 items a call, so a long plan takes several calls; those\nquotes may expire unused, since each stage is quoted again right before it runs. Nothing is\ngenerated until the user says yes. The user can drop or change lines; a changed line is priced\nagain.\n\n**Canvas rows.** Pass `shot_number` on every `generate_image` and `generate_video` of a shot (1, 2, 3… in story order; the same number for a retake and for that shot's video), so each shot gets its own row on the Frameo canvas. Cast, prop and location references take `placement_kind` (`character`, `prop` or `location`) and the subject's name as `placement_group` instead: they sit on their own board, and the shots built from them do not pile into their row.\n\n**Review (on by default).** Every image and clip is checked before it is built on, and the\nfinished cut before it is handed over, following `review-shots`: its checklist and steps come\nfrom `get_skill(\"review-shots\")`, loaded before step 2. The checks run after step 2, after\nstep 4 and after step 6, and a failed shot is retaken once. The plan\ncarries its two lines — the check of clips with speech and the retake allowance — for the user\nto approve with everything else; the user can turn the review off.\n\n**0. The list.** The script is the source of truth: read it once and write down the cast\n(everyone who appears in more than one beat, speaking or not, such as a waiter or a guest, plus\neach setting), the props (every object seen in more than one beat or that the action turns on:\na phone, a suitcase, a napkin) and the beats (one beat per continuous action or speech: an\naction that runs on without a cut, such as a trick from set-up to reveal, stays one beat and\none clip however many lines it spans: its `seconds` is the action's own running time, at most\nthe model's longest clip and 15 s when a route-B speaker talks in it, and only one character\nspeaks in it on route B, since a lipsync takes one visible speaker; a two-scene script is\nusually 12–16), using `shot-list.md`. A beat that carries straight on from the one before,\nwith no jump in time, place or angle, is marked `continues`: it opens on the previous clip's\nlast frame (step 4) and gets no image of its own in step 2. That list is the plan the user\napproves (**The plan first.**); every step below reads from it.\n\n**1. Cast (free-ish, ~18 credits per character).** `list_characters` first (free): a\ncharacter already in the project's cast is reused — its first image is the portrait, and its\n`voice_id` is the character's voice in step 3, with no voice to choose — and costs nothing\nhere. For each character not in the cast, `generate_image` a clean portrait in the chosen look\n(with `image_urls` when the user gave a photo), then `wait_task` and note its link in the shot\nlist. Do the setting the same way, and each prop as a clean image of the object alone\n(`placement_kind=\"prop\"`, its name as `placement_group`).\nShow the cast to the user; fix anyone who looks wrong before moving on. The portraits are\nwhat the review checks every later shot against. Each new character the user keeps is saved\nwith `save_character` (name, portrait link) so the next video in this project reuses it; a\nvoice chosen in step 3 is saved the same way (`save_character` with the character's name,\n`image_urls=[portrait link]` and `voice_id`), and the Frameo app's agent then uses that voice\ntoo. A character `list_characters` shows with `in_library: false` has only a voice so far: its\nportrait goes to `save_character` like a new one's, and its voice is kept.\n\n**2. One image per beat (~18 credits each).** For every beat not marked `continues`,\n`generate_image` with `image_urls` in this order, up to 14: the portraits of the characters in\nit, the props it shows, the latest beat image in the same scene (it carries the extras, props\nand positions across), and the setting last. The prompt is the action line in\nthe chosen look, naming each reference in words and nothing the beat does not name;\n`aspect_ratio` from Ask first. Test: do\nthe first beat alone, `wait_task`, show it; then the rest. Then review every beat image\n(`review-shots` steps 0, 1 and 4) before any clip is made from one.\n\n**3. Test one line of dialogue, two ways (~65–130 credits).** Pick the first spoken line on a\nbeat with its own image (not a `continues` beat).\nRoute A — *native audio*: `generate_video` from that beat's image (`first_frame_url`), the\nline and its delivery in the prompt, `generate_audio=true` (the quote must carry `generate_audio=true` as well — it is part of\nthe quote shape), on a model `list_models` marks as able to make sound. Route B — *speech + lipsync*: `search_voices` for the character in the\nscript's language, `generate_speech(text, voice_id)`, then\n`generate_lipsync(video_url, audio_url, duration, audio_duration, aspect_ratio, resolution)`\nover the beat's animated shot, at the size the clip was made at. The shot list's `speech s` column (the speech result's\n`duration_seconds`, else 2.5 words per second in Hindi, 2.8 in English, plus 1 s headroom) is the lipsync's\n`audio_duration`, and its `seconds` column (that estimate rounded up to whole seconds and\nkept inside the model's range, and never above 15 s on a route-B beat — the lipsync's own\ncap, whatever the video model allows) is the clip's `duration`. The two differ only when the range\nclamps a long line, and then the lipsync needs the real speech length. The text says both\nare approximate. Show both; the user\npicks the voice and confirms the routes. The rule of thumb: native audio invents the voice afresh in every clip from the words\nin the prompt, so it suits a character with a single line (cheaper, one step); a character who\nspeaks in more than one clip keeps one voice only on route B, with the same `voice_id` in every\n`generate_speech`, and so does a language that is weak in the video model. A character\n`list_characters` returned with a `voice_id` already has its voice, so its lines go route B with\nit. The choice can be per character, and each route-B voice is saved with `save_character`\n(step 1). A script with\nnarration instead of dialogue skips the lipsync: `generate_speech` for the narrator and mix\nit in at the cut.\n\n**4. All dialogue beats (the bulk: ~65 credits per 5 s clip, plus ~20 credits per second\nof lipsync on route B).** Tell the user the batch total — `estimate_cost` is free, so quote each beat at its own\n`seconds` from the shot list (plus its lipsync seconds on route B) and add them up — and\nget one yes for that figure. Then, for each beat in script order:\n`estimate_cost` for that clip, `generate_video` from its image (`first_frame_url`; a\n`continues` beat opens on the previous beat's last frame, from `extract_frames` on the file that\ngoes into the cut — the lipsynced clip on route B — at that beat's `length`)\nwith that `quote_id` (route A: line in the prompt with audio on, quoted with\n`generate_audio=true`; route B: silent, then `generate_speech` and `generate_lipsync`, each\nwith its own quote), `wait_task`. On Seedance models (the default) the portraits of the beat's\ncharacters and the props it shows go along as `reference_image_urls`, which holds faces and\nobjects through the motion. The frame is the first reference image, so the prompt names the portraits\nfrom the second image on, and it counts toward the model's cap; the quote carries\n`first_frame=true` and `reference_count` = the number of `reference_image_urls`, not counting the\nframe. A clip with a `continues` beat after it has its look checked (`review-shots` step 2,\nfree) before that beat is made from its last frame. The prompt names where the\npeople and the camera end up, since a clip with no end named can drift to an empty frame. The\nbatch approval is the\nconsent for each clip's `confirmed_by_user=true` — as long as that clip's quote, and the\nrunning total, stay within the approved figure; when either goes over, stop and ask again\nbefore the call. Retakes from the review are not part of this figure: they spend the plan's\nseparate retake allowance (`review-shots`).\n\nThen review the clips (`review-shots` steps 2–4) before they are cut: a bad take found after\nthe cut costs the cut again. The look of every clip (steps 2 and 4) is checked now, since it is\nfree; clips with speech are analysed part by part in step 6, right before their part is cut,\nbecause analyses share the cut's hourly run budget.\n\n**5. Sound (free from the library; ~1 credit per 50 characters of prompt when generated).** The script's\ncues (a crowd murmur, a door, a gasp) come from the free library first: `search_sound_effects`, then\n`add_sound_to_library`; `generate_sound_effect` makes only the cues the library lacks. `generate_music` for one bed if the tone wants it — at most 120 s (the tool's cap), looped\nat the join (`-stream_loop -1` on the bed input, `-shortest`) when the cut is longer.\n\n**6. Cut (free).** `run_ffmpeg` takes at most 10 inputs and 4 outputs, so the cut is passes,\n`wait_task` after each (its outputs are the next pass's inputs): per beat, mix that beat's\nspeech and cues over its clip into `beatN.mp4` at exactly the shot list's `length` —\n`tpad=stop_mode=clone:stop_duration=<length>` on the video plus `apad` on the mixed audio\nover-pad a short file (`stop_duration` is padding added, not a target) and `-t <length>`\nthen sets the final length; `-t` alone never extends anything —\nup to four beats per call (one output each) and fewer when their files would pass ten\ninputs: a clip plus every sound file counts, so a narrated beat with one cue is three\ninputs and three such beats fill a call, and a beat with more than eight cue files gets its\ncues pre-mixed into one stem first. Route-A and lipsynced beats already carry their sound\nand go through the same pass with only their clip, for the length. Then concatenate the\nbeats in order in groups of at most 10 into `partN.mp4`, and repeat the grouping until the\nfinal call — the parts plus the bed — fits; then that last call joins the parts, mixes the\nbed under the dialogue and burns subtitles from a `.ass` sidecar built from the script's\nlines (one cue per beat, timed from the `length` column of the shot list;\n`sidecars=[{\"name\": \"subs.ass\", \"content\": ...}]` and the `ass` filter; the skeleton is in\n`subtitles.md`; `timeout_seconds=600` on the group and final calls — a long join can pass the\n300 s default). Output `final.mp4`. Sixteen beats is four mix calls, two group concats and\none final join — seven calls, a few more when cues shrink the mix batches. The budget is 30 runs per clock hour and every `run_ffmpeg` call counts, so a\npart is at most 36 beats at one cue each — fewer with more sound — and the\nnext part waits for the next hour. With the review on, each analysed clip and each recheck is a\nrun too, taken from the same hour: a part's analyses, its expected rechecks and its cut calls\n(the mix passes for its cue count, the concats and the join) add up to at most 30 — at one cue\nper beat, a part whose beats all speak is at most 16 beats (16 analyses and 7 cut calls, room\nfor 7 rechecks and re-cuts together), and with more cues per beat fewer. Each part's clips are analysed right before\nthat part is cut, then the next part waits for the next hour.\n\n**7. Hand over.** `wait_task` on the cut, then the review of the finished cut (`review-shots`\nstep 5); return the video link and `open_in_frameo`.\n\n## Done\n\nThe user gets: the finished video, and a Frameo project with the cast portraits and\nevery shot on the canvas in script order, the dialogue clips, and the chat that records\neach step — editable in the app from there. With it, the review note per shot: passed,\nretaken and why, or kept with a known flaw. Say what was skipped (a beat that failed, a\nvoice the user may want to swap) rather than hiding it.\n\n## Files\n\n- `shot-list.md` — the cast and beat list template used in step 0.\n- `subtitles.md` — the `.ass` sidecar skeleton and the timing rule for step 6.\n\n## Worked example (a two-scene period drama, Hindi dialogue)\n\nCast: Chandrika, Arjun, the King of Avanti, Malti, Sudhan, an officer, a servant, a\nminister; setting: the royal council hall. Two scenes, 15 beats: Chandrika's proposal over\nthe map (3), the King's approval and Arjun's reaction (3), Arjun turning on Malti (4), the\naccusations (3), the banishment and Chandrika's silent close (2). Route for dialogue:\nChandrika, Arjun and the King speak more than once, so their lines go route B; step 3 tests\nChandrika's opening line both ways to pick her voice. Rough cost at 15 clips with a\n5 s Seedance clip each and 12 spoken beats on route B: cast 8 × ~18 + beats 15 × ~18 + clips\n15 × ~65 + lipsync 12 × 5 s × ~20 + sound ~5 = about 2,590 credits; a line on route A saves its\n~100 credits of lipsync.\n"
}

SHA-256 of public snapshot: 05ae6789fa1e7d02dec53a6b349b6e58bfa070fa37ab728af15ad421738734e0