← Files FrameoARCHIVED FILE
skills/faceless-story-video/SKILL.md
6.85 KB · Oct 10, 2026 · 06:03 UTC
---
name: faceless-story-video
description: "A narrated explainer or story with no on-screen speaker: one image per scene, gentle motion, a narrator voice, music, captions burned in."
---
Needs: a script or a topic outline; the narrator's language
Credits: about 800–1,300 for 60–90 seconds, about 200–250 per 15 seconds (images ~18 each, clips ~65 per 5 s, speech ~1 per 50 characters)
Time: 15–30 minutes
## When
Use for explainers, history and science shorts, kids' stories, listicles, "did you know"
channels — anything told by a narrator over scenes. Not for characters speaking on
screen (`script-to-video`) or a presenter (`talking-presenter`).
## Ask first
1. **Frame and length** — 9:16 or 16:9; target seconds.
2. **Look** — photoreal, illustrated, paper-cut, 3D; one reference image if they have one.
3. **Narrator** — 3 candidates from `search_voices` in the script's language; they pick.
Say the rough cost from the scene count in the same turn as the questions; the user's yes comes with the plan (Steps, **The plan first.**).
## Steps
**Before any paid step.** The project: `list_projects` (or `create_project`) gives the `project_id`
and, when the project has several modules, the `module_id`; every call below that takes a project —
`estimate_cost` included — gets that same pair; without it those tools answer `project_needed`. The
quote: one `estimate_cost(items=[…])` prices a stage in one call, with the same project, model,
size and number of `image_urls` or `reference_image_urls` (`reference_count`) as each generate call, and returns a `quote_id` per item plus the total. Each generate call
then passes its own item's `quote_id` and `confirmed_by_user=true`. A quote is single-use and lasts
15 minutes, so a long plan is priced stage by stage, right before each stage runs; a stage that
comes to more than the user approved is asked about again first. Generate calls return
`generation_ids`; `wait_task` returns the links, and `show_generations` shows each stage's running
and finished work in one card where the chat app displays Frameo cards: all the stage's ids at
once, before its first `wait_task`, and the finished result with `final: true`.
**The plan first.** Before the first paid call, the plan goes to the user in the chat as plain
text: what will be made, in order, one line per generation (for a script, the shot list; for a
set, each shot), with each line's credits and the total. The credits come from
`estimate_cost(items=[…])`, up to 10 items a call, so a long plan takes several calls; those
quotes may expire unused, since each stage is quoted again right before it runs. Nothing is
generated until the user says yes. The user can drop or change lines; a changed line is priced
again.
**Canvas rows.** Pass `shot_number` on every `generate_image` and `generate_video` of a shot (1, 2, 3… in story order; the same number for a retake and for that shot's video), so each shot gets its own row on the Frameo canvas. Cast, prop and location references take `placement_kind` (`character`, `prop` or `location`) and the subject's name as `placement_group` instead: they sit on their own board, and the shots built from them do not pile into their row.
**Review (on by default).** Every scene image and clip is checked before it is built on,
following `review-shots` (its steps come from `get_skill("review-shots")`, loaded before the
first scene image), and a failed scene is retaken once. The checks are free here — no
scene has speech on screen, so no clip needs the paid media analysis — but a retake is paid like the
scene it replaces, from the retake allowance the plan carries for the user to approve. With no
cast, "who" is checked against the reference image or the scene's visual line. The user can
turn the review off.
**0. Scenes.** Split the script into scenes of one idea each, 5–8 s of narration per scene.
Write a one-line visual for every scene. Keep the list.
**1. Narration (~1 per 50 characters).** `generate_speech` per scene (one file each keeps
the cut simple); `wait_task` for the links. Each result's `duration_seconds` is
the scene's measured spoken length; a result without one is estimated at 2.8 words per second
(2.5 for Hindi) plus one second of headroom, and that timing is approximate. The scene's estimate
is that length rounded up to whole seconds and kept inside the video model's range
(`list_models(kind="video")`; 4–30 s on the default).
**2. Images (~18 each).** `generate_image` per scene, each with its own quote, from its
visual line in the chosen look, `aspect_ratio` from Ask first, the reference image as `image_urls` when given. A subject that
recurs across scenes (a character, a mascot, a place) gets its own image first
(`placement_kind` and its name as `placement_group`), passed in `image_urls` on every scene that
shows it, so it stays the same from scene to scene.
`wait_task`; show the first two before the rest. Review every image (`review-shots` steps 0,
1 and 4) before it is animated.
**3. Motion (~65 per 5 s).** `generate_video` from each image (`first_frame_url`) with a
slow camera move in the prompt ("slow push in", "gentle drift") and the scene otherwise
unchanged: no new objects or people, ending on the same scene (the same subjects and setting), `duration` = the scene's
whole-second estimate. Silent (`generate_audio=false`) — the narrator carries the sound.
`wait_task` each, then review the clips (`review-shots` steps 2 and 4). Skip this step for a
lower-cost cut and let ffmpeg pan over the stills instead.
**4. Music (~1 per 50 characters of prompt).** `generate_music` for one bed matching the tone, `duration` =
the total, at most 120 s (the tool's cap); for a longer story make one 120 s bed and loop
it at the join (`-stream_loop -1` on the bed input, `-shortest`).
**5. Cut (free).** `run_ffmpeg` takes at most 10 inputs and 4 outputs, so the cut is passes, `wait_task`
after each (its outputs are the next pass's inputs): per scene, mix its narration over its clip (or pan over its
still) into `sceneN.mp4` at exactly the clip's `duration` (`tpad`/`apad` over-pad,
`-t <duration>` sets the length), up to four scenes per call (one output each) and fewer when their files would pass ten inputs
(a clip plus its narration is two inputs and four such scenes fill a call); then concatenate the scenes in order in groups of at
most 10; then one last call joins the groups, mixes the music under at low volume and
burns captions from an `.ass` sidecar (one cue per scene, timed from those clip
durations). Output `final.mp4`, review the finished cut (`review-shots` step 5), and return the
link and `open_in_frameo`.
## Done
The video and the review note per scene (passed, retaken and why, or kept with a known
flaw) and for the finished cut, plus every scene image and clip in the Frameo project.
`open_in_frameo` on the result opens the project in the Frameo app, where every take sits on
the canvas, ready for retakes and edits.
## Files
None.
SHA-256: d5fbb28629873aba9d5892110c91d510282874f99a79787d920ee98cf0ba71a0