← Files FrameoARCHIVED FILE

skills/talking-presenter/SKILL.md

6.77 KB · Oct 10, 2026 · 06:03 UTC

↓ Download file

---
name: talking-presenter
description: "One person on camera delivers a script: a face (photo or designed), a fixed voice, lipsynced delivery, optional title card, cut to length."
---

Needs: a script; a photo of the presenter or a description; optionally a brand colour or logo
Credits: about 1,000–2,100 for a 30–60 second piece (lipsync is the main cost: ~20 credits per second of clip)
Time: 10–20 minutes

## When

Use for one person speaking to camera: an explainer, a founder message, a course intro, an
announcement, a customer-support answer. Any language the voice catalog covers.

Not for: several characters in scenes (`script-to-video`), or narration over other footage
(`faceless-story-video`).

## Ask first

1. **Who** — a photo (`show_upload` for files on the user's device, `import_media_url` for a web link; images attached to the chat do not reach Frameo, and `create_upload_url` is for clients that send the file themselves) or a description to design from.
2. **Voice** — the user's preference in words (warm, brisk, older, Hindi, British…); pick 3
   with `search_voices` and show their names.
3. **Frame** — 9:16 or 16:9, and whether a title card opens the piece.

Say the rough cost (the script's length decides it) in the same turn as the questions; the user's yes comes with the plan (Steps, **The plan first.**).

## Steps

**Before any paid step.** The project: `list_projects` (or `create_project`) gives the `project_id`
and, when the project has several modules, the `module_id`; every call below that takes a project —
`estimate_cost` included — gets that same pair; without it those tools answer `project_needed`. The
quote: one `estimate_cost(items=[…])` prices a stage in one call, with the same project, model,
size and number of `image_urls` or `reference_image_urls` (`reference_count`) as each generate call, and returns a `quote_id` per item plus the total. Each generate call
then passes its own item's `quote_id` and `confirmed_by_user=true`. A quote is single-use and lasts
15 minutes, so a long plan is priced stage by stage, right before each stage runs; a stage that
comes to more than the user approved is asked about again first. Generate calls return
`generation_ids`; `wait_task` returns the links, and `show_generations` shows each stage's running
and finished work in one card where the chat app displays Frameo cards: all the stage's ids at
once, before its first `wait_task`, and the finished result with `final: true`.

**The plan first.** Before the first paid call, the plan goes to the user in the chat as plain
text: what will be made, in order, one line per generation (for a script, the shot list; for a
set, each shot), with each line's credits and the total. The credits come from
`estimate_cost(items=[…])`, up to 10 items a call, so a long plan takes several calls; those
quotes may expire unused, since each stage is quoted again right before it runs. Nothing is
generated until the user says yes. The user can drop or change lines; a changed line is priced
again.

**Canvas rows.** Pass `shot_number` on every `generate_image` and `generate_video` of a shot (1, 2, 3… in story order; the same number for a retake and for that shot's video), so each shot gets its own row on the Frameo canvas. Cast, prop and location references take `placement_kind` (`character`, `prop` or `location`) and the subject's name as `placement_group` instead: they sit on their own board, and the shots built from them do not pile into their row.

**1. The presenter still (~18 credits).** `generate_image`: a mid-shot of the presenter facing
camera, neutral background unless the user wants one, `image_urls=[photo]` when there is a
photo. `wait_task`, show it; adjust once if asked.

**2. The voice (~1 credit per 50 characters).** `generate_speech(text=script, voice_id)` in one
call when the whole script fits one clip (15 s, the lipsync's cap, or the video model's max from
`list_models(kind="video")` when that is lower); otherwise split at sentence breaks — and at clause breaks (commas, dashes) when one
sentence alone is still too long — so each file fits one clip. `wait_task` gives the link; play it back to the user before spending on video. The result's
`duration_seconds`, the file's measured spoken length, is its raw estimate, kept as is for the
lipsync; a result without one is estimated at 2.8 words per second (2.5 for Hindi) plus one
second of headroom, and that estimate is approximate. When the user picks a presenter
`list_characters` already has in the project, its first image is the still and its `voice_id` the
voice. A new
presenter is saved with `save_character` under a name not already in the cast (the same name
replaces that character's voice), with `image_urls=[presenter still]` and the `voice_id`, so later
videos in this project reuse the same face and voice.

**3. The clip (~65 credits per 5 s + ~20 per second of lipsync).** For each speech file:
`generate_video` from the still (`first_frame_url`) with a prompt like "presenter speaking
to camera, subtle natural head movement, steady framing", `duration` = that file's raw
estimate rounded up to whole seconds, kept inside the video model's range
(`list_models(kind="video")`; 4–30 s on the default) and at most 15 s — step 2 sized the files to
that, and a file that measures a little longer still goes in whole as `audio_duration`,
`wait_task`, then `generate_lipsync(video_url, audio_url, duration, audio_duration=the raw
estimate, aspect_ratio, resolution)` — the raw one, so a clamped clip is still extended to
the speech; the size the clip was made at; and `duration` never above 15 s, the lipsync's
own cap, whatever the video model allows — and
`wait_task`. Test the first clip
before the rest.

**4. Title card (free).** If wanted: `render_motion_graphics` with `template="title_card"`, the title
(and a `subtitle`), the piece's `size` and a `duration` of 2 to 3 s; `wait_task` returns the card's clip,
joined in front in the cut. It counts against the 30 runs per clock hour.

**5. Cut (free).** With a title card, one pass of its own first: the card comes at full HD in its
`size`, so it is scaled to the clips' size (`scale`, then `setsar=1`) and given a bounded silent
track (`-f lavfi -t <card seconds> -i anullsrc=r=<the clips' sample rate>:cl=stereo`); an
unbounded `anullsrc` never ends and stalls the join. Then `run_ffmpeg` takes at most 10 inputs and 4 outputs:
concatenate the lipsynced clips in
order, in groups of at most 10 — the card counts as an input, so nine clips per group when
one is joined in front — and repeat until the final call fits (`wait_task` after each pass;
its outputs are the next pass's inputs). Output `final.mp4`. `wait_task`,
return the link and `open_in_frameo`.

## Done

The finished clip, the presenter still, and every clip in the Frameo project (chat and
canvas) for edits in the app.

## Files

None.

SHA-256: 1f9dfca5f9413551ad6be89961264af7823ed24dece57329f2144c91c99cdc78