← Files FrameoARCHIVED FILE
skills/talking-presenter/SKILL.md
6.77 KB · Oct 10, 2026 · 06:03 UTC
--- name: talking-presenter description: "One person on camera delivers a script: a face (photo or designed), a fixed voice, lipsynced delivery, optional title card, cut to length." --- Needs: a script; a photo of the presenter or a description; optionally a brand colour or logo Credits: about 1,000–2,100 for a 30–60 second piece (lipsync is the main cost: ~20 credits per second of clip) Time: 10–20 minutes ## When Use for one person speaking to camera: an explainer, a founder message, a course intro, an announcement, a customer-support answer. Any language the voice catalog covers. Not for: several characters in scenes (`script-to-video`), or narration over other footage (`faceless-story-video`). ## Ask first 1. **Who** — a photo (`show_upload` for files on the user's device, `import_media_url` for a web link; images attached to the chat do not reach Frameo, and `create_upload_url` is for clients that send the file themselves) or a description to design from. 2. **Voice** — the user's preference in words (warm, brisk, older, Hindi, British…); pick 3 with `search_voices` and show their names. 3. **Frame** — 9:16 or 16:9, and whether a title card opens the piece. Say the rough cost (the script's length decides it) in the same turn as the questions; the user's yes comes with the plan (Steps, **The plan first.**). ## Steps **Before any paid step.** The project: `list_projects` (or `create_project`) gives the `project_id` and, when the project has several modules, the `module_id`; every call below that takes a project — `estimate_cost` included — gets that same pair; without it those tools answer `project_needed`. The quote: one `estimate_cost(items=[…])` prices a stage in one call, with the same project, model, size and number of `image_urls` or `reference_image_urls` (`reference_count`) as each generate call, and returns a `quote_id` per item plus the total. Each generate call then passes its own item's `quote_id` and `confirmed_by_user=true`. A quote is single-use and lasts 15 minutes, so a long plan is priced stage by stage, right before each stage runs; a stage that comes to more than the user approved is asked about again first. Generate calls return `generation_ids`; `wait_task` returns the links, and `show_generations` shows each stage's running and finished work in one card where the chat app displays Frameo cards: all the stage's ids at once, before its first `wait_task`, and the finished result with `final: true`. **The plan first.** Before the first paid call, the plan goes to the user in the chat as plain text: what will be made, in order, one line per generation (for a script, the shot list; for a set, each shot), with each line's credits and the total. The credits come from `estimate_cost(items=[…])`, up to 10 items a call, so a long plan takes several calls; those quotes may expire unused, since each stage is quoted again right before it runs. Nothing is generated until the user says yes. The user can drop or change lines; a changed line is priced again. **Canvas rows.** Pass `shot_number` on every `generate_image` and `generate_video` of a shot (1, 2, 3… in story order; the same number for a retake and for that shot's video), so each shot gets its own row on the Frameo canvas. Cast, prop and location references take `placement_kind` (`character`, `prop` or `location`) and the subject's name as `placement_group` instead: they sit on their own board, and the shots built from them do not pile into their row. **1. The presenter still (~18 credits).** `generate_image`: a mid-shot of the presenter facing camera, neutral background unless the user wants one, `image_urls=[photo]` when there is a photo. `wait_task`, show it; adjust once if asked. **2. The voice (~1 credit per 50 characters).** `generate_speech(text=script, voice_id)` in one call when the whole script fits one clip (15 s, the lipsync's cap, or the video model's max from `list_models(kind="video")` when that is lower); otherwise split at sentence breaks — and at clause breaks (commas, dashes) when one sentence alone is still too long — so each file fits one clip. `wait_task` gives the link; play it back to the user before spending on video. The result's `duration_seconds`, the file's measured spoken length, is its raw estimate, kept as is for the lipsync; a result without one is estimated at 2.8 words per second (2.5 for Hindi) plus one second of headroom, and that estimate is approximate. When the user picks a presenter `list_characters` already has in the project, its first image is the still and its `voice_id` the voice. A new presenter is saved with `save_character` under a name not already in the cast (the same name replaces that character's voice), with `image_urls=[presenter still]` and the `voice_id`, so later videos in this project reuse the same face and voice. **3. The clip (~65 credits per 5 s + ~20 per second of lipsync).** For each speech file: `generate_video` from the still (`first_frame_url`) with a prompt like "presenter speaking to camera, subtle natural head movement, steady framing", `duration` = that file's raw estimate rounded up to whole seconds, kept inside the video model's range (`list_models(kind="video")`; 4–30 s on the default) and at most 15 s — step 2 sized the files to that, and a file that measures a little longer still goes in whole as `audio_duration`, `wait_task`, then `generate_lipsync(video_url, audio_url, duration, audio_duration=the raw estimate, aspect_ratio, resolution)` — the raw one, so a clamped clip is still extended to the speech; the size the clip was made at; and `duration` never above 15 s, the lipsync's own cap, whatever the video model allows — and `wait_task`. Test the first clip before the rest. **4. Title card (free).** If wanted: `render_motion_graphics` with `template="title_card"`, the title (and a `subtitle`), the piece's `size` and a `duration` of 2 to 3 s; `wait_task` returns the card's clip, joined in front in the cut. It counts against the 30 runs per clock hour. **5. Cut (free).** With a title card, one pass of its own first: the card comes at full HD in its `size`, so it is scaled to the clips' size (`scale`, then `setsar=1`) and given a bounded silent track (`-f lavfi -t <card seconds> -i anullsrc=r=<the clips' sample rate>:cl=stereo`); an unbounded `anullsrc` never ends and stalls the join. Then `run_ffmpeg` takes at most 10 inputs and 4 outputs: concatenate the lipsynced clips in order, in groups of at most 10 — the card counts as an input, so nine clips per group when one is joined in front — and repeat until the final call fits (`wait_task` after each pass; its outputs are the next pass's inputs). Output `final.mp4`. `wait_task`, return the link and `open_in_frameo`. ## Done The finished clip, the presenter still, and every clip in the Frameo project (chat and canvas) for edits in the app. ## Files None.
SHA-256: 1f9dfca5f9413551ad6be89961264af7823ed24dece57329f2144c91c99cdc78