← Files FrameoARCHIVED FILE

skills/script-to-video/SKILL.md

15.6 KB · Oct 10, 2026 · 06:03 UTC

↓ Download file

---
name: script-to-video
description: "Turn a script into a finished video, such as a micro-drama — a consistent cast, one shot per beat, characters speaking their lines, cut together with subtitles."
---

Needs: a script (scenes, characters, dialogue or narration); optionally photos of the people or a style reference
Credits: about 2,000–3,000 for a 12–16 shot piece whose speakers have several lines each (speech + lipsync is the larger half; see Steps for the breakdown)
Time: 20–40 minutes, mostly waiting on generations

## When

Use for any script with named characters who speak or act on screen: a micro-drama or
serial episode, a short film, a dialogue-driven ad, a kids' story, a sketch, a training
scenario. Any length the user wants, in clips — a long script is delivered in parts (step 6 sets
the size); 9:16 by default, 16:9 on request.

Not for: narration over stills with no on-screen speaker (that is `faceless-story-video`),
one person talking to camera (`talking-presenter`), or editing footage the user already has
(`subtitles-and-cut`).

## Ask first

Three questions, one turn, before anything is generated:

1. **Aspect** — 9:16 for Reels/Shorts (default) or 16:9.
2. **Look** — photoreal, stylised, or anime; any reference image or a Frameo project whose
   look to match.
3. **Cast** — photos of the people (`show_upload` for files on the user's device, `import_media_url` for a web link; images attached to the chat do not reach Frameo, and `create_upload_url` is for clients that send the file themselves) or design the cast
   from the script's descriptions.

Then say the rough cost from the script (Steps, below) in the same turn as the questions; the user's yes comes with the plan (Steps, **The plan first.**).

## Steps

**Before any paid step.** The project: `list_projects` (or `create_project`) gives the `project_id`
and, when the project has several modules, the `module_id`; every call below that takes a project —
`estimate_cost` included — gets that same pair; without it those tools answer `project_needed`. The
quote: one `estimate_cost(items=[…])` prices a stage in one call, with the same project, model,
size and number of `image_urls` or `reference_image_urls` (`reference_count`) as each generate call, and returns a `quote_id` per item plus the total. Each generate call
then passes its own item's `quote_id` and `confirmed_by_user=true`. A quote is single-use and lasts
15 minutes, so a long plan is priced stage by stage, right before each stage runs; a stage that
comes to more than the user approved is asked about again first. Generate calls return
`generation_ids`; `wait_task` returns the links, and `show_generations` shows each stage's running
and finished work in one card where the chat app displays Frameo cards: all the stage's ids at
once, before its first `wait_task`, and the finished result with `final: true`.

**The plan first.** Before the first paid call, the plan goes to the user in the chat as plain
text: what will be made, in order, one line per generation (for a script, the shot list; for a
set, each shot), with each line's credits and the total. The credits come from
`estimate_cost(items=[…])`, up to 10 items a call, so a long plan takes several calls; those
quotes may expire unused, since each stage is quoted again right before it runs. Nothing is
generated until the user says yes. The user can drop or change lines; a changed line is priced
again.

**Canvas rows.** Pass `shot_number` on every `generate_image` and `generate_video` of a shot (1, 2, 3… in story order; the same number for a retake and for that shot's video), so each shot gets its own row on the Frameo canvas. Cast, prop and location references take `placement_kind` (`character`, `prop` or `location`) and the subject's name as `placement_group` instead: they sit on their own board, and the shots built from them do not pile into their row.

**Review (on by default).** Every image and clip is checked before it is built on, and the
finished cut before it is handed over, following `review-shots`: its checklist and steps come
from `get_skill("review-shots")`, loaded before step 2. The checks run after step 2, after
step 4 and after step 6, and a failed shot is retaken once. The plan
carries its two lines — the check of clips with speech and the retake allowance — for the user
to approve with everything else; the user can turn the review off.

**0. The list.** The script is the source of truth: read it once and write down the cast
(everyone who appears in more than one beat, speaking or not, such as a waiter or a guest, plus
each setting), the props (every object seen in more than one beat or that the action turns on:
a phone, a suitcase, a napkin) and the beats (one beat per continuous action or speech: an
action that runs on without a cut, such as a trick from set-up to reveal, stays one beat and
one clip however many lines it spans: its `seconds` is the action's own running time, at most
the model's longest clip and 15 s when a route-B speaker talks in it, and only one character
speaks in it on route B, since a lipsync takes one visible speaker; a two-scene script is
usually 12–16), using `shot-list.md`. A beat that carries straight on from the one before,
with no jump in time, place or angle, is marked `continues`: it opens on the previous clip's
last frame (step 4) and gets no image of its own in step 2. That list is the plan the user
approves (**The plan first.**); every step below reads from it.

**1. Cast (free-ish, ~18 credits per character).** `list_characters` first (free): a
character already in the project's cast is reused — its first image is the portrait, and its
`voice_id` is the character's voice in step 3, with no voice to choose — and costs nothing
here. For each character not in the cast, `generate_image` a clean portrait in the chosen look
(with `image_urls` when the user gave a photo), then `wait_task` and note its link in the shot
list. Do the setting the same way, and each prop as a clean image of the object alone
(`placement_kind="prop"`, its name as `placement_group`).
Show the cast to the user; fix anyone who looks wrong before moving on. The portraits are
what the review checks every later shot against. Each new character the user keeps is saved
with `save_character` (name, portrait link) so the next video in this project reuses it; a
voice chosen in step 3 is saved the same way (`save_character` with the character's name,
`image_urls=[portrait link]` and `voice_id`), and the Frameo app's agent then uses that voice
too. A character `list_characters` shows with `in_library: false` has only a voice so far: its
portrait goes to `save_character` like a new one's, and its voice is kept.

**2. One image per beat (~18 credits each).** For every beat not marked `continues`,
`generate_image` with `image_urls` in this order, up to 14: the portraits of the characters in
it, the props it shows, the latest beat image in the same scene (it carries the extras, props
and positions across), and the setting last. The prompt is the action line in
the chosen look, naming each reference in words and nothing the beat does not name;
`aspect_ratio` from Ask first. Test: do
the first beat alone, `wait_task`, show it; then the rest. Then review every beat image
(`review-shots` steps 0, 1 and 4) before any clip is made from one.

**3. Test one line of dialogue, two ways (~65–130 credits).** Pick the first spoken line on a
beat with its own image (not a `continues` beat).
Route A — *native audio*: `generate_video` from that beat's image (`first_frame_url`), the
line and its delivery in the prompt, `generate_audio=true` (the quote must carry `generate_audio=true` as well — it is part of
the quote shape), on a model `list_models` marks as able to make sound. Route B — *speech + lipsync*: `search_voices` for the character in the
script's language, `generate_speech(text, voice_id)`, then
`generate_lipsync(video_url, audio_url, duration, audio_duration, aspect_ratio, resolution)`
over the beat's animated shot, at the size the clip was made at. The shot list's `speech s` column (the speech result's
`duration_seconds`, else 2.5 words per second in Hindi, 2.8 in English, plus 1 s headroom) is the lipsync's
`audio_duration`, and its `seconds` column (that estimate rounded up to whole seconds and
kept inside the model's range, and never above 15 s on a route-B beat — the lipsync's own
cap, whatever the video model allows) is the clip's `duration`. The two differ only when the range
clamps a long line, and then the lipsync needs the real speech length. The text says both
are approximate. Show both; the user
picks the voice and confirms the routes. The rule of thumb: native audio invents the voice afresh in every clip from the words
in the prompt, so it suits a character with a single line (cheaper, one step); a character who
speaks in more than one clip keeps one voice only on route B, with the same `voice_id` in every
`generate_speech`, and so does a language that is weak in the video model. A character
`list_characters` returned with a `voice_id` already has its voice, so its lines go route B with
it. The choice can be per character, and each route-B voice is saved with `save_character`
(step 1). A script with
narration instead of dialogue skips the lipsync: `generate_speech` for the narrator and mix
it in at the cut.

**4. All dialogue beats (the bulk: ~65 credits per 5 s clip, plus ~20 credits per second
of lipsync on route B).** Tell the user the batch total — `estimate_cost` is free, so quote each beat at its own
`seconds` from the shot list (plus its lipsync seconds on route B) and add them up — and
get one yes for that figure. Then, for each beat in script order:
`estimate_cost` for that clip, `generate_video` from its image (`first_frame_url`; a
`continues` beat opens on the previous beat's last frame, from `extract_frames` on the file that
goes into the cut — the lipsynced clip on route B — at that beat's `length`)
with that `quote_id` (route A: line in the prompt with audio on, quoted with
`generate_audio=true`; route B: silent, then `generate_speech` and `generate_lipsync`, each
with its own quote), `wait_task`. On Seedance models (the default) the portraits of the beat's
characters and the props it shows go along as `reference_image_urls`, which holds faces and
objects through the motion. The frame is the first reference image, so the prompt names the portraits
from the second image on, and it counts toward the model's cap; the quote carries
`first_frame=true` and `reference_count` = the number of `reference_image_urls`, not counting the
frame. A clip with a `continues` beat after it has its look checked (`review-shots` step 2,
free) before that beat is made from its last frame. The prompt names where the
people and the camera end up, since a clip with no end named can drift to an empty frame. The
batch approval is the
consent for each clip's `confirmed_by_user=true` — as long as that clip's quote, and the
running total, stay within the approved figure; when either goes over, stop and ask again
before the call. Retakes from the review are not part of this figure: they spend the plan's
separate retake allowance (`review-shots`).

Then review the clips (`review-shots` steps 2–4) before they are cut: a bad take found after
the cut costs the cut again. The look of every clip (steps 2 and 4) is checked now, since it is
free; clips with speech are analysed part by part in step 6, right before their part is cut,
because analyses share the cut's hourly run budget.

**5. Sound (free from the library; ~1 credit per 50 characters of prompt when generated).** The script's
cues (a crowd murmur, a door, a gasp) come from the free library first: `search_sound_effects`, then
`add_sound_to_library`; `generate_sound_effect` makes only the cues the library lacks. `generate_music` for one bed if the tone wants it — at most 120 s (the tool's cap), looped
at the join (`-stream_loop -1` on the bed input, `-shortest`) when the cut is longer.

**6. Cut (free).** `run_ffmpeg` takes at most 10 inputs and 4 outputs, so the cut is passes,
`wait_task` after each (its outputs are the next pass's inputs): per beat, mix that beat's
speech and cues over its clip into `beatN.mp4` at exactly the shot list's `length` —
`tpad=stop_mode=clone:stop_duration=<length>` on the video plus `apad` on the mixed audio
over-pad a short file (`stop_duration` is padding added, not a target) and `-t <length>`
then sets the final length; `-t` alone never extends anything —
up to four beats per call (one output each) and fewer when their files would pass ten
inputs: a clip plus every sound file counts, so a narrated beat with one cue is three
inputs and three such beats fill a call, and a beat with more than eight cue files gets its
cues pre-mixed into one stem first. Route-A and lipsynced beats already carry their sound
and go through the same pass with only their clip, for the length. Then concatenate the
beats in order in groups of at most 10 into `partN.mp4`, and repeat the grouping until the
final call — the parts plus the bed — fits; then that last call joins the parts, mixes the
bed under the dialogue and burns subtitles from a `.ass` sidecar built from the script's
lines (one cue per beat, timed from the `length` column of the shot list;
`sidecars=[{"name": "subs.ass", "content": ...}]` and the `ass` filter; the skeleton is in
`subtitles.md`; `timeout_seconds=600` on the group and final calls — a long join can pass the
300 s default). Output `final.mp4`. Sixteen beats is four mix calls, two group concats and
one final join — seven calls, a few more when cues shrink the mix batches. The budget is 30 runs per clock hour and every `run_ffmpeg` call counts, so a
part is at most 36 beats at one cue each — fewer with more sound — and the
next part waits for the next hour. With the review on, each analysed clip and each recheck is a
run too, taken from the same hour: a part's analyses, its expected rechecks and its cut calls
(the mix passes for its cue count, the concats and the join) add up to at most 30 — at one cue
per beat, a part whose beats all speak is at most 16 beats (16 analyses and 7 cut calls, room
for 7 rechecks and re-cuts together), and with more cues per beat fewer. Each part's clips are analysed right before
that part is cut, then the next part waits for the next hour.

**7. Hand over.** `wait_task` on the cut, then the review of the finished cut (`review-shots`
step 5); return the video link and `open_in_frameo`.

## Done

The user gets: the finished video, and a Frameo project with the cast portraits and
every shot on the canvas in script order, the dialogue clips, and the chat that records
each step — editable in the app from there. With it, the review note per shot: passed,
retaken and why, or kept with a known flaw. Say what was skipped (a beat that failed, a
voice the user may want to swap) rather than hiding it.

## Files

- `shot-list.md` — the cast and beat list template used in step 0.
- `subtitles.md` — the `.ass` sidecar skeleton and the timing rule for step 6.

## Worked example (a two-scene period drama, Hindi dialogue)

Cast: Chandrika, Arjun, the King of Avanti, Malti, Sudhan, an officer, a servant, a
minister; setting: the royal council hall. Two scenes, 15 beats: Chandrika's proposal over
the map (3), the King's approval and Arjun's reaction (3), Arjun turning on Malti (4), the
accusations (3), the banishment and Chandrika's silent close (2). Route for dialogue:
Chandrika, Arjun and the King speak more than once, so their lines go route B; step 3 tests
Chandrika's opening line both ways to pick her voice. Rough cost at 15 clips with a
5 s Seedance clip each and 12 spoken beats on route B: cast 8 × ~18 + beats 15 × ~18 + clips
15 × ~65 + lipsync 12 × 5 s × ~20 + sound ~5 = about 2,590 credits; a line on route A saves its
~100 credits of lipsync.

SHA-256: 48d59c5be5ff9c726e7b793627172cae2aef84f9af6fde913656cd879da86d98