← Plugin catalog
Creativity

Frameo

Dashverse v1.2.1

Publisher description

From the marketplace listing

Frameo turns your ideas into videos. Describe a scene, share a photo or paste a story, and Frameo brings it to life: cinematic shots, characters who look the same from scene to scene, voices, music and sound effects, all cut together into a video that's ready to share. With Frameo, you can: • Put yourself in any scene: share a selfie and step into a film set, a far-off city or a cosy café. • Bring a photo to life, from a family snapshot to your pet's best portrait. • Turn a story, a script or a single line into a short film, an animated tale or a movie trailer. • Make reels, shorts and thumbnails for your social feeds, in vertical or widescreen. • Give your characters voices, then add music, sound effects, titles and subtitles. • Shoot your own products and designs, from T-shirts to candles, on models and in styled scenes. • Create with a wide range of image, video and audio models, all in one place. Everything you make is saved to your Frameo account, so you can keep editing it in the Frameo app. Connect your Frameo account to get started.

Language: English · Automatically detected from descriptions.

Screenshots

Provided by the publisher. Illustrations of the product, not our hands-on testing.

Show all 4 screenshots

Publisher keywords

Search terms declared by the publisher.

Matches for “share”

Exact text from the indicated source. A mention alone does not establish support for your task.

Publisher full description

Frameo turns your ideas into videos. Describe a scene, share a photo or paste a story, and Frameo brings it to life: cinematic shots, characters who look the same from scene to scene, voices, music and sound effects, all cut together into a video that's ready to share. With Frameo, you can: • Put yourself in any scene: share a selfie and step into a film set, a far-off city or a cosy café. • Bring a photo to life, from a family snapshot to your pet's best portrait. • Turn a story, a script or a single line into a short film, an animated tale or a movie trailer. • Make reels, shorts and thumbnails for your social feeds, in vertical or widescreen. • Give your characters voices, then add music, sound effects, titles and subtitles. • Shoot your own products and designs, from T-shirts to candles, on models and in styled scenes. • Create with a wide range of image, video and audio models, all in one place. Everything you make is saved to your Frameo account, so you can keep editing it in the Frameo app. Connect your Frameo account to get started.

Changes

Frameo

Oct 8, 2026 · 6 saved observations

Capabilities & instructions

Declared skills changed from “[]” to “[{"description":"A consistent character across views and expressions, on the canvas so every later image and video can keep the same face.","interface":{"brand_color":"#4C6CFF","default_prompt":null,"display_name":"Character sheet","icon...”.

Metadata evidence →Listing evidence →
Positioning & listing

Product description changed from “Frameo turns a prompt into a finished short video inside ChatGPT. It generates images, video clips, voiceovers, music, and sound effects in the user's Frameo workspace, keeps characters and style consistent across shots, and assembles th...” to “Create images, video, voiceover, music and sound effects in your Frameo projects, and assemble them into finished video.”.

Metadata evidence →Listing evidence →
Technical updates

Package contents changed in 22 files: assets/screenshot-fight-scene.jpg, assets/screenshot-fight-scene.png, assets/screenshot-short-film.jpg, …. Open the file diff to inspect the edits.

Files evidence →
1 more changes that day

Package contents changed in 52 files: .app.json, .codex-plugin/plugin.json, assets/icon.png, …. Open the file diff to inspect the edits.

Files evidence →

Files & skills

File archives

Plugin package51 files · 2.01 MBBrowse files →
Skill instructions
character-sheet4.09 KB

View saved version →

---
name: character-sheet
description: "A consistent character across views and expressions, on the canvas so every later image and video can keep the same face."
---

Needs: a description or a photo; the style (photoreal, anime, 3D, editorial)
Credits: about 90–130 (5–7 images at ~18 each)
Time: 10 minutes

## When

Before any multi-shot work with a recurring person or mascot: a series, a brand
character, a game or story cast. `script-to-video` calls this for its cast.

## Ask first

1. **Who** — a photo (`show_upload` for files on the user's device, `import_media_url` for a web link; images attached to the chat do not reach Frameo, and `create_upload_url` is for clients that send the file themselves) or a description: age, build, hair, clothes, one
   distinguishing detail.
2. **Style** — photoreal, editorial, anime, 3D, game concept.
3. **Views** — the default is front, three-quarter, profile, full body, two expressions.

## Steps

**Before any paid step.** The project: `list_projects` (or `create_project`) gives the `project_id`
and, when the project has several modules, the `module_id`; every call below that takes a project —
`estimate_cost` included — gets that same pair; without it those tools answer `project_needed`. The
quote: one `estimate_cost(items=[…])` prices a stage in one call, with the same project, model,
size and number of `image_urls` or `reference_image_urls` (`reference_count`) as each generate call, and returns a `quote_id` per item plus the total. Each generate call
then passes its own item's `quote_id` and `confirmed_by_user=true`. A quote is single-use and lasts
15 minutes, so a long plan is priced stage by stage, right before each stage runs; a stage that
comes to more than the user approved is asked about again first. Generate calls return
`generation_ids`; `wait_task` returns the links, and `show_generations` shows each stage's running
and finished work in one card where the chat app displays Frameo cards: all the stage's ids at
once, before its first `wait_task`, and the finished result with `final: true`.

**The plan first.** Before the first paid call, the plan goes to the user in the chat as plain
text: what will be made, in order, one line per generation (for a script, the shot list; for a
set, each shot), with each line's credits and the total. The credits come from
`estimate_cost(items=[…])`, up to 10 items a call, so a long plan takes several calls; those
quotes may expire unused, since each stage is quoted again right before it runs. Nothing is
generated until the user says yes. The user can drop or change lines; a changed line is priced
again.

**Canvas rows.** Every `generate_image` here takes `placement_kind="character"` and `placement_group=<the character's name>`, so the sheet is one row on the character board, and later shots that use it as a reference get rows of their own.

**1. The anchor (~18).** `generate_image`: front view, neutral expression, plain background,
in the style; `image_urls=[photo]` when given. `wait_task`, show it; this is the face
everything else follows — iterate here, not later.

**2. The views (~18 each).** `generate_image` per view, each with its own quote, with
`image_urls=[anchor]` first and the view in the prompt ("three-quarter view, same person,
same clothes"). Keep the background plain. `wait_task` for the links.

**3. The references.** The anchor and views stay on the project canvas; their links, anchor
first, are what later generations take as `image_urls` / `reference_image_urls`.

**4. Save it to the cast (free).** `save_character` with the character's name and the anchor
first, then the views in `image_urls`, six images at most (the most a saved character keeps;
with more views, the six that show the face and outfit best): later work in the project finds it with
`list_characters`, and it shows in the Frameo app's character panel.

## Done

The anchor and its views on the canvas in one row, with their links for later shots, and the
character in the project's cast.
`open_in_frameo` on the result opens the project in the Frameo app, where every take sits on
the canvas, ready for retakes and edits.

## Files

None.

Referenced files: 2

faceless-story-video6.85 KB

View saved version →

---
name: faceless-story-video
description: "A narrated explainer or story with no on-screen speaker: one image per scene, gentle motion, a narrator voice, music, captions burned in."
---

Needs: a script or a topic outline; the narrator's language
Credits: about 800–1,300 for 60–90 seconds, about 200–250 per 15 seconds (images ~18 each, clips ~65 per 5 s, speech ~1 per 50 characters)
Time: 15–30 minutes

## When

Use for explainers, history and science shorts, kids' stories, listicles, "did you know"
channels — anything told by a narrator over scenes. Not for characters speaking on
screen (`script-to-video`) or a presenter (`talking-presenter`).

## Ask first

1. **Frame and length** — 9:16 or 16:9; target seconds.
2. **Look** — photoreal, illustrated, paper-cut, 3D; one reference image if they have one.
3. **Narrator** — 3 candidates from `search_voices` in the script's language; they pick.

Say the rough cost from the scene count in the same turn as the questions; the user's yes comes with the plan (Steps, **The plan first.**).

## Steps

**Before any paid step.** The project: `list_projects` (or `create_project`) gives the `project_id`
and, when the project has several modules, the `module_id`; every call below that takes a project —
`estimate_cost` included — gets that same pair; without it those tools answer `project_needed`. The
quote: one `estimate_cost(items=[…])` prices a stage in one call, with the same project, model,
size and number of `image_urls` or `reference_image_urls` (`reference_count`) as each generate call, and returns a `quote_id` per item plus the total. Each generate call
then passes its own item's `quote_id` and `confirmed_by_user=true`. A quote is single-use and lasts
15 minutes, so a long plan is priced stage by stage, right before each stage runs; a stage that
comes to more than the user approved is asked about again first. Generate calls return
`generation_ids`; `wait_task` returns the links, and `show_generations` shows each stage's running
and finished work in one card where the chat app displays Frameo cards: all the stage's ids at
once, before its first `wait_task`, and the finished result with `final: true`.

**The plan first.** Before the first paid call, the plan goes to the user in the chat as plain
text: what will be made, in order, one line per generation (for a script, the shot list; for a
set, each shot), with each line's credits and the total. The credits come from
`estimate_cost(items=[…])`, up to 10 items a call, so a long plan takes several calls; those
quotes may expire unused, since each stage is quoted again right before it runs. Nothing is
generated until the user says yes. The user can drop or change lines; a changed line is priced
again.

**Canvas rows.** Pass `shot_number` on every `generate_image` and `generate_video` of a shot (1, 2, 3… in story order; the same number for a retake and for that shot's video), so each shot gets its own row on the Frameo canvas. Cast, prop and location references take `placement_kind` (`character`, `prop` or `location`) and the subject's name as `placement_group` instead: they sit on their own board, and the shots built from them do not pile into their row.

**Review (on by default).** Every scene image and clip is checked before it is built on,
following `review-shots` (its steps come from `get_skill("review-shots")`, loaded before the
first scene image), and a failed scene is retaken once. The checks are free here — no
scene has speech on screen, so no clip needs the paid media analysis — but a retake is paid like the
scene it replaces, from the retake allowance the plan carries for the user to approve. With no
cast, "who" is checked against the reference image or the scene's visual line. The user can
turn the review off.

**0. Scenes.** Split the script into scenes of one idea each, 5–8 s of narration per scene.
Write a one-line visual for every scene. Keep the list.

**1. Narration (~1 per 50 characters).** `generate_speech` per scene (one file each keeps
the cut simple); `wait_task` for the links. Each result's `duration_seconds` is
the scene's measured spoken length; a result without one is estimated at 2.8 words per second
(2.5 for Hindi) plus one second of headroom, and that timing is approximate. The scene's estimate
is that length rounded up to whole seconds and kept inside the video model's range
(`list_models(kind="video")`; 4–30 s on the default).

**2. Images (~18 each).** `generate_image` per scene, each with its own quote, from its
visual line in the chosen look, `aspect_ratio` from Ask first, the reference image as `image_urls` when given. A subject that
recurs across scenes (a character, a mascot, a place) gets its own image first
(`placement_kind` and its name as `placement_group`), passed in `image_urls` on every scene that
shows it, so it stays the same from scene to scene.
`wait_task`; show the first two before the rest. Review every image (`review-shots` steps 0,
1 and 4) before it is animated.

**3. Motion (~65 per 5 s).** `generate_video` from each image (`first_frame_url`) with a
slow camera move in the prompt ("slow push in", "gentle drift") and the scene otherwise
unchanged: no new objects or people, ending on the same scene (the same subjects and setting), `duration` = the scene's
whole-second estimate. Silent (`generate_audio=false`) — the narrator carries the sound.
`wait_task` each, then review the clips (`review-shots` steps 2 and 4). Skip this step for a
lower-cost cut and let ffmpeg pan over the stills instead.

**4. Music (~1 per 50 characters of prompt).** `generate_music` for one bed matching the tone, `duration` =
the total, at most 120 s (the tool's cap); for a longer story make one 120 s bed and loop
it at the join (`-stream_loop -1` on the bed input, `-shortest`).

**5. Cut (free).** `run_ffmpeg` takes at most 10 inputs and 4 outputs, so the cut is passes, `wait_task`
after each (its outputs are the next pass's inputs): per scene, mix its narration over its clip (or pan over its
still) into `sceneN.mp4` at exactly the clip's `duration` (`tpad`/`apad` over-pad,
`-t <duration>` sets the length), up to four scenes per call (one output each) and fewer when their files would pass ten inputs
(a clip plus its narration is two inputs and four such scenes fill a call); then concatenate the scenes in order in groups of at
most 10; then one last call joins the groups, mixes the music under at low volume and
burns captions from an `.ass` sidecar (one cue per scene, timed from those clip
durations). Output `final.mp4`, review the finished cut (`review-shots` step 5), and return the
link and `open_in_frameo`.

## Done

The video and the review note per scene (passed, retaken and why, or kept with a known
flaw) and for the finished cut, plus every scene image and clip in the Frameo project.
`open_in_frameo` on the result opens the project in the Frameo app, where every take sits on
the canvas, ready for retakes and edits.

## Files

None.

Referenced files: 2

product-photoshoot4.48 KB

View saved version →

---
name: product-photoshoot
description: "Packshots on several backgrounds or scenes, lifestyle shots, hero banners and carousel sets from a few product photos."
---

Needs: 2–5 product photos on any background; the brand's look in a sentence; which set they need
Credits: about 180 for a 10-image set (images ~18 each), plus a retake allowance of 15%
Time: 10–20 minutes

## When

E-commerce and ads: a clean packshot set, lifestyle scenes, a hero banner, a carousel, a
restyle of an old shot. Not for video (`ugc-product-video`) or a moving spin
(`product-spin`).

## Ask first

1. **Photos** — 2–5, through `show_upload` from the user's device or `import_media_url` from web links (images attached to the chat do not reach Frameo, and `create_upload_url` is for clients that send the file themselves); the more angles the better.
2. **Set** — packshots / lifestyle / hero / carousel / restyle, and how many.
3. **Look** — the brand in a sentence, a reference image if they have one, aspect ratio.

Say the rough cost in the same turn as the questions; the user's yes comes with the plan (Steps, **The plan first.**).

## Steps

**Before any paid step.** The project: `list_projects` (or `create_project`) gives the `project_id`
and, when the project has several modules, the `module_id`; every call below that takes a project —
`estimate_cost` included — gets that same pair; without it those tools answer `project_needed`. The
quote: one `estimate_cost(items=[…])` prices a stage in one call, with the same project, model,
size and number of `image_urls` or `reference_image_urls` (`reference_count`) as each generate call, and returns a `quote_id` per item plus the total. Each generate call
then passes its own item's `quote_id` and `confirmed_by_user=true`. A quote is single-use and lasts
15 minutes, so a long plan is priced stage by stage, right before each stage runs; a stage that
comes to more than the user approved is asked about again first. Generate calls return
`generation_ids`; `wait_task` returns the links, and `show_generations` shows each stage's running
and finished work in one card where the chat app displays Frameo cards: all the stage's ids at
once, before its first `wait_task`, and the finished result with `final: true`.

**The plan first.** Before the first paid call, the plan goes to the user in the chat as plain
text: what will be made, in order, one line per generation (for a script, the shot list; for a
set, each shot), with each line's credits and the total. The credits come from
`estimate_cost(items=[…])`, up to 10 items a call, so a long plan takes several calls; those
quotes may expire unused, since each stage is quoted again right before it runs. Nothing is
generated until the user says yes. The user can drop or change lines; a changed line is priced
again.

**Review (on by default).** Every image is checked before the set is handed over, following
`review-shots`: its checklist and steps come from `get_skill("review-shots")`, loaded before
step 1, and the checks here are its steps 0, 1 and 4. Each preview is compared with the product
photos — the label, its lettering, the shape and the colour — and with the shot's pattern,
including whole hands in an in-hand shot; a failed shot is retaken once. The checks are free; a
retake is paid like the shot it replaces, from the retake allowance the plan carries as its own
line for the user to approve. The user can turn the review off.

**1. One test (~18).** `generate_image` with `image_urls=[product photos…]` (the cleanest first
— it is the base) and a prompt for the first shot in the set; `image_model` from
`list_models(kind="image")` if the user names one. `wait_task`, show it; fix the product's
fidelity before the batch (more photos, a tighter prompt).

**2. The set (~18 each).** One `generate_image` per shot, each with its own quote, same
`image_urls`, prompts from the recipe patterns (`hero-shot`, `flatlay`, `in-hand`, `on-desk`, `minimalist-white`,
`luxury`, `kitchen-scene`, `angles-set`) — `get_skill` on any of them for the pattern. Then review the set.

**3. Size.** There is no image upscale; for print or large use, ask for `image_resolution`
at generation (and `image_model` from `list_models(kind="image")`).

**4. Deliver.** The links, the review note per shot, and `open_in_frameo` — every image is on
the canvas in one row.

## Done

The image set in the Frameo project.
`open_in_frameo` on the result opens the project in the Frameo app, where every take sits on
the canvas, ready for retakes and edits.

## Files

None.

Referenced files: 2

review-shots10.9 KB

View saved version →

---
name: review-shots
description: "Check every image and clip against the plan before building on it — the right people, the asked-for action, clean hands and text, lips in sync — and retake a failed shot once."
---

Needs: the shots' generation_ids, the shot list (or the prompts), and the cast, product or location references the shots were made from
Credits: the checks of images and of the look of clips are free; a few credits per minute for clips with speech; every retake is paid like the shot it replaces
Time: a few seconds per image, about a minute per clip checked

## When

Inside `script-to-video`, `faceless-story-video`, the five `ugc-*` video skills and
`product-photoshoot`, where it is on by default, and inside any
other skill when the user asks for it ("check them before you show me"). On its own for shots
made earlier, from their `generation_ids` (the calls that made them returned those);
`read_project` lists a project's media but no generation ids, so an image without its id cannot
be previewed and is reported as not reviewed. A clip with only its link can still be checked
from its frames (step 2) when its length is known.

## Ask first

Nothing more when a workflow brings it in: that workflow's plan already carries the review
line and the retake allowance. On its own: which shots, and the retake allowance as a credit
figure the user names — `estimate_cost` for one retake of each shot shows what a figure
covers; there is no default, since what the shots first cost may not be known.

## Steps

**Before any paid step.** The project: `list_projects` (or `create_project`) gives the
`project_id` and, when the project has several modules, the `module_id`; every call below that
takes a project gets that same pair. A retake is a paid generation like any other: its own
`estimate_cost`, its `quote_id`, `confirmed_by_user=true`, then `wait_task`. Retakes started
together are shown in one `show_generations` card, with all their ids, before that `wait_task`.

**The retake allowance.** Retakes have their own budget, apart from the figures each stage of
the plan was approved for: by default 15% of the plan's image and clip total, shown as its own
line in the plan and approved with it. Every paid call a retake makes counts against it — the
new image, the new clip, on route B both the speech and the lipsync, and the paid recheck of a
retaken clip with speech. When the next retake would push retake spending past the allowance,
the user is asked first.

**0. The checklist, per shot.** Written from the shot list before the first check:
- **Who** — every character the beat names and no one else, each with the face, hair and
  outfit of their cast portrait (when the workflow made portraits; otherwise against the
  reference image or the description); the product with its real label and colour.
- **What** — the beat's action, setting and framing.
- **Clean** — whole hands and limbs (no extra fingers, no merged bodies), no garbled lettering
  or stray watermark, nothing that matters cut off at the frame edge.
- **Coherence** (clips) — physically and visually coherent, within the clip and across its
  cuts: identity (faces, bodies and outfits stay the same), props (they persist and keep their
  state), setting and light (the place, light direction and time of day hold), space (people
  keep their positions and screen direction), cause and effect (an action has a visible cause
  and result), physics (every movement could really happen; nothing appears, vanishes or
  doubles) and artefacts (no warping, melting or flicker).
- **Continuity** — the same outfit, props and time of day as the neighbouring shots, and each
  cut into the next shot reads as coherent (step 2).
- **Sound** (clips with speech only) — the line is heard, in the right language and voice, and
  the lips move with it.

Every shot on the list ends with one note: passed, failed (and why), or **not reviewed**. A
shot is never passed without its own preview or frames having been looked at.

**1. Images (free).** One call returns at most four previews (at most 384 px), so images are
checked in groups of four or fewer: `wait_task` on those shots' `generation_ids` (or
`get_task` once they have finished). Match each preview to its shot through `previews`, which
lists the links shown in order; a shot whose link is not among them is fetched again on its
own, and if it still has no preview it is noted **not reviewed**. A preview catches the wrong
person, a missing character, broken hands or bodies and garbled large text; fine lettering at
that size is the user's check, and the hand-over says so.

**2. The look of clips (free).** Per shot, `extract_frames(video_url, duration, at_seconds=[…])`
at four moments — half a second in, a third of the way, two thirds of the way, and half a
second before the end, the most one call previews — returns the frames (each with its `url`)
and previews of them; `previews` lists the frame links shown, in order. A clip is fully
reviewed only when every moment asked for has a frame whose link is in `previews`. Fewer frames than moments, or a `note` in the result, means a moment came back
empty — and then no frame says which moment it is — so each moment is asked for again, alone,
a little earlier (a quarter-second before); a frame without a preview is asked for again the
same way. Then the same checklist, without the sound. `duration` is the clip's length from
`wait_task` (`duration_seconds`). A clip with a moment still unseen is noted as partly
reviewed, and one with none seen is noted **not reviewed**. It takes up to a minute per clip
and charges nothing.

**The cuts.** Once two neighbouring clips are both checked, the cut between them is judged
from the earlier clip's last frame and the later clip's first frame, both already extracted.
When a cut jumps in time, angle or distance, the jump has to read as one coherent film: the
same characters on the same sides of the frame, the same light and time of day, props in the
state the earlier shot left them, and a change of time or place a viewer can follow. A cut that
jars fails the later shot, noted with what jumps, and that shot is retaken (step 4).

**3. Clips with speech (a few credits per minute).** Per clip with dialogue or lip-sync,
`analyze_media(media_url, question=…, confirmed_by_user=true)` — the one paid call with no
quote: it is charged by clip length, a few credits per minute; the user agreed to it with the
plan, and the finished run (`wait_task`) reports `credits_charged_estimate`. The question asks
for description, not a verdict — the tool sees only the clip, never the cast portraits:
"Transcribe what is said and by whom. Do the speaker's lips move in time with the words? Say
yes or no, and where it drifts. Describe each person's face, hair and clothes." The assistant
compares the answer with the shot's line and cast. Analysis, ffmpeg and motion-graphics runs
share one lane per module: while one is running, another is refused with `chat_busy`
(retryable), so the clips are analysed one after another, not between generations. Each
answer is also recorded in the module in Frameo. Every `analyze_media`, `run_ffmpeg` and
`render_motion_graphics` call counts against 30 runs per clock hour for the user, so a long
workflow analyses and cuts part by part: a part's clips are analysed right before that part is
cut, in the same hour, and its analyses, rechecks and cut calls together stay within 30
(`script-to-video` sizes its parts that way).

**4. Retake a failed shot, once.** Same `shot_number`, so the retake lands in the shot's own
canvas row. What changes depends on the failure and on what failed:
- **An image** with the wrong or drifting face — `generate_image` again with the cast portrait
  first in `image_urls` and the character named with their look in the prompt.
- **A clip** with the wrong or drifting face — the fix starts from its image: remake the beat's
  image as above, check it, then animate the new image (`first_frame_url`) as the clip was
  first made. On Seedance models the cast portraits go along as `reference_image_urls` (the
  frame counts as one of the model's reference images); every other model refuses a first frame
  together with reference images, so there the image alone carries the face.
- A clip with no image of its own (a `continues` beat) — retaken from the previous clip's last
  frame, as it was first made. When a retaken clip has a `continues` beat after it, that beat is
  remade from the new last frame too, from the same retake allowance.
- A cut that jars — the later clip remade to follow on from the earlier one: from the earlier
  clip's last frame (`first_frame_url`) when the action carries straight on, otherwise from its
  image remade with the positions, light and props the earlier shot ends on.
- Broken hands — a framing that keeps hands out of shot or a simpler gesture.
- Garbled text — no lettering in the prompt; add it at the cut with `run_ffmpeg`.
- Wrong action — the action first in the prompt, in plain words.
- Lips out of sync on a route-A clip — route B for that beat: `generate_speech`, then
  `generate_lipsync` at the clip's `aspect_ratio` and `resolution`, never above 15 s, each with
  its own quote.
  The lipsync's `audio_duration` is the speech result's `duration_seconds` (without one, 2.8
  words per second, 2.5 in Hindi, plus one second — approximate), and the clip keeps the
  `duration` it was made at: whole seconds inside the model's range. Route B can cost more
  than the clip it replaces; the quotes say how much before it runs.

Each retake has its own quote: `estimate_cost` for each paid call of it (with `reference_count` =
the number of `image_urls` or `reference_image_urls` it passes), its `quote_id`,
`confirmed_by_user=true` while it fits the allowance, `wait_task`, and the same checklist
again. A shot that still fails — a route-B clip included — keeps the better take and is
reported; it is not retaken a third time.

**5. The finished cut (free).** When the workflow ends in a cut, `extract_frames` on the
finished video at the middle of each shot, timed from the shot list's `length` column (or each
clip's `duration` where the workflow keeps no such column), and half a second before the end,
four moments a call (the most one call previews) with the end moment among them. Each frame has to show its shot's
action: a trim that cut the action away fails the cut, and that shot is cut again with a
`length` that keeps it. The last frame has to show the story's last moment, not an empty frame
or a camera that drifted off it; a closing shot that drifts is cut again shorter. A re-cut
changes only the cut and costs no credits, but it redoes that shot's mix pass, its group join and
the final join: at least three runs from the hour's 30.

## Done

A review note per shot, handed over with the result: passed, retaken (what was wrong and what
changed), kept with a known flaw the user may want to fix in the Frameo app, or not reviewed
(and why), and one for the finished cut. Retakes sit in the same canvas row as the shot they replaced.

## Files

None.

Referenced files: 2

script-to-video15.6 KB

View saved version →

---
name: script-to-video
description: "Turn a script into a finished video, such as a micro-drama — a consistent cast, one shot per beat, characters speaking their lines, cut together with subtitles."
---

Needs: a script (scenes, characters, dialogue or narration); optionally photos of the people or a style reference
Credits: about 2,000–3,000 for a 12–16 shot piece whose speakers have several lines each (speech + lipsync is the larger half; see Steps for the breakdown)
Time: 20–40 minutes, mostly waiting on generations

## When

Use for any script with named characters who speak or act on screen: a micro-drama or
serial episode, a short film, a dialogue-driven ad, a kids' story, a sketch, a training
scenario. Any length the user wants, in clips — a long script is delivered in parts (step 6 sets
the size); 9:16 by default, 16:9 on request.

Not for: narration over stills with no on-screen speaker (that is `faceless-story-video`),
one person talking to camera (`talking-presenter`), or editing footage the user already has
(`subtitles-and-cut`).

## Ask first

Three questions, one turn, before anything is generated:

1. **Aspect** — 9:16 for Reels/Shorts (default) or 16:9.
2. **Look** — photoreal, stylised, or anime; any reference image or a Frameo project whose
   look to match.
3. **Cast** — photos of the people (`show_upload` for files on the user's device, `import_media_url` for a web link; images attached to the chat do not reach Frameo, and `create_upload_url` is for clients that send the file themselves) or design the cast
   from the script's descriptions.

Then say the rough cost from the script (Steps, below) in the same turn as the questions; the user's yes comes with the plan (Steps, **The plan first.**).

## Steps

**Before any paid step.** The project: `list_projects` (or `create_project`) gives the `project_id`
and, when the project has several modules, the `module_id`; every call below that takes a project —
`estimate_cost` included — gets that same pair; without it those tools answer `project_needed`. The
quote: one `estimate_cost(items=[…])` prices a stage in one call, with the same project, model,
size and number of `image_urls` or `reference_image_urls` (`reference_count`) as each generate call, and returns a `quote_id` per item plus the total. Each generate call
then passes its own item's `quote_id` and `confirmed_by_user=true`. A quote is single-use and lasts
15 minutes, so a long plan is priced stage by stage, right before each stage runs; a stage that
comes to more than the user approved is asked about again first. Generate calls return
`generation_ids`; `wait_task` returns the links, and `show_generations` shows each stage's running
and finished work in one card where the chat app displays Frameo cards: all the stage's ids at
once, before its first `wait_task`, and the finished result with `final: true`.

**The plan first.** Before the first paid call, the plan goes to the user in the chat as plain
text: what will be made, in order, one line per generation (for a script, the shot list; for a
set, each shot), with each line's credits and the total. The credits come from
`estimate_cost(items=[…])`, up to 10 items a call, so a long plan takes several calls; those
quotes may expire unused, since each stage is quoted again right before it runs. Nothing is
generated until the user says yes. The user can drop or change lines; a changed line is priced
again.

**Canvas rows.** Pass `shot_number` on every `generate_image` and `generate_video` of a shot (1, 2, 3… in story order; the same number for a retake and for that shot's video), so each shot gets its own row on the Frameo canvas. Cast, prop and location references take `placement_kind` (`character`, `prop` or `location`) and the subject's name as `placement_group` instead: they sit on their own board, and the shots built from them do not pile into their row.

**Review (on by default).** Every image and clip is checked before it is built on, and the
finished cut before it is handed over, following `review-shots`: its checklist and steps come
from `get_skill("review-shots")`, loaded before step 2. The checks run after step 2, after
step 4 and after step 6, and a failed shot is retaken once. The plan
carries its two lines — the check of clips with speech and the retake allowance — for the user
to approve with everything else; the user can turn the review off.

**0. The list.** The script is the source of truth: read it once and write down the cast
(everyone who appears in more than one beat, speaking or not, such as a waiter or a guest, plus
each setting), the props (every object seen in more than one beat or that the action turns on:
a phone, a suitcase, a napkin) and the beats (one beat per continuous action or speech: an
action that runs on without a cut, such as a trick from set-up to reveal, stays one beat and
one clip however many lines it spans: its `seconds` is the action's own running time, at most
the model's longest clip and 15 s when a route-B speaker talks in it, and only one character
speaks in it on route B, since a lipsync takes one visible speaker; a two-scene script is
usually 12–16), using `shot-list.md`. A beat that carries straight on from the one before,
with no jump in time, place or angle, is marked `continues`: it opens on the previous clip's
last frame (step 4) and gets no image of its own in step 2. That list is the plan the user
approves (**The plan first.**); every step below reads from it.

**1. Cast (free-ish, ~18 credits per character).** `list_characters` first (free): a
character already in the project's cast is reused — its first image is the portrait, and its
`voice_id` is the character's voice in step 3, with no voice to choose — and costs nothing
here. For each character not in the cast, `generate_image` a clean portrait in the chosen look
(with `image_urls` when the user gave a photo), then `wait_task` and note its link in the shot
list. Do the setting the same way, and each prop as a clean image of the object alone
(`placement_kind="prop"`, its name as `placement_group`).
Show the cast to the user; fix anyone who looks wrong before moving on. The portraits are
what the review checks every later shot against. Each new character the user keeps is saved
with `save_character` (name, portrait link) so the next video in this project reuses it; a
voice chosen in step 3 is saved the same way (`save_character` with the character's name,
`image_urls=[portrait link]` and `voice_id`), and the Frameo app's agent then uses that voice
too. A character `list_characters` shows with `in_library: false` has only a voice so far: its
portrait goes to `save_character` like a new one's, and its voice is kept.

**2. One image per beat (~18 credits each).** For every beat not marked `continues`,
`generate_image` with `image_urls` in this order, up to 14: the portraits of the characters in
it, the props it shows, the latest beat image in the same scene (it carries the extras, props
and positions across), and the setting last. The prompt is the action line in
the chosen look, naming each reference in words and nothing the beat does not name;
`aspect_ratio` from Ask first. Test: do
the first beat alone, `wait_task`, show it; then the rest. Then review every beat image
(`review-shots` steps 0, 1 and 4) before any clip is made from one.

**3. Test one line of dialogue, two ways (~65–130 credits).** Pick the first spoken line on a
beat with its own image (not a `continues` beat).
Route A — *native audio*: `generate_video` from that beat's image (`first_frame_url`), the
line and its delivery in the prompt, `generate_audio=true` (the quote must carry `generate_audio=true` as well — it is part of
the quote shape), on a model `list_models` marks as able to make sound. Route B — *speech + lipsync*: `search_voices` for the character in the
script's language, `generate_speech(text, voice_id)`, then
`generate_lipsync(video_url, audio_url, duration, audio_duration, aspect_ratio, resolution)`
over the beat's animated shot, at the size the clip was made at. The shot list's `speech s` column (the speech result's
`duration_seconds`, else 2.5 words per second in Hindi, 2.8 in English, plus 1 s headroom) is the lipsync's
`audio_duration`, and its `seconds` column (that estimate rounded up to whole seconds and
kept inside the model's range, and never above 15 s on a route-B beat — the lipsync's own
cap, whatever the video model allows) is the clip's `duration`. The two differ only when the range
clamps a long line, and then the lipsync needs the real speech length. The text says both
are approximate. Show both; the user
picks the voice and confirms the routes. The rule of thumb: native audio invents the voice afresh in every clip from the words
in the prompt, so it suits a character with a single line (cheaper, one step); a character who
speaks in more than one clip keeps one voice only on route B, with the same `voice_id` in every
`generate_speech`, and so does a language that is weak in the video model. A character
`list_characters` returned with a `voice_id` already has its voice, so its lines go route B with
it. The choice can be per character, and each route-B voice is saved with `save_character`
(step 1). A script with
narration instead of dialogue skips the lipsync: `generate_speech` for the narrator and mix
it in at the cut.

**4. All dialogue beats (the bulk: ~65 credits per 5 s clip, plus ~20 credits per second
of lipsync on route B).** Tell the user the batch total — `estimate_cost` is free, so quote each beat at its own
`seconds` from the shot list (plus its lipsync seconds on route B) and add them up — and
get one yes for that figure. Then, for each beat in script order:
`estimate_cost` for that clip, `generate_video` from its image (`first_frame_url`; a
`continues` beat opens on the previous beat's last frame, from `extract_frames` on the file that
goes into the cut — the lipsynced clip on route B — at that beat's `length`)
with that `quote_id` (route A: line in the prompt with audio on, quoted with
`generate_audio=true`; route B: silent, then `generate_speech` and `generate_lipsync`, each
with its own quote), `wait_task`. On Seedance models (the default) the portraits of the beat's
characters and the props it shows go along as `reference_image_urls`, which holds faces and
objects through the motion. The frame is the first reference image, so the prompt names the portraits
from the second image on, and it counts toward the model's cap; the quote carries
`first_frame=true` and `reference_count` = the number of `reference_image_urls`, not counting the
frame. A clip with a `continues` beat after it has its look checked (`review-shots` step 2,
free) before that beat is made from its last frame. The prompt names where the
people and the camera end up, since a clip with no end named can drift to an empty frame. The
batch approval is the
consent for each clip's `confirmed_by_user=true` — as long as that clip's quote, and the
running total, stay within the approved figure; when either goes over, stop and ask again
before the call. Retakes from the review are not part of this figure: they spend the plan's
separate retake allowance (`review-shots`).

Then review the clips (`review-shots` steps 2–4) before they are cut: a bad take found after
the cut costs the cut again. The look of every clip (steps 2 and 4) is checked now, since it is
free; clips with speech are analysed part by part in step 6, right before their part is cut,
because analyses share the cut's hourly run budget.

**5. Sound (free from the library; ~1 credit per 50 characters of prompt when generated).** The script's
cues (a crowd murmur, a door, a gasp) come from the free library first: `search_sound_effects`, then
`add_sound_to_library`; `generate_sound_effect` makes only the cues the library lacks. `generate_music` for one bed if the tone wants it — at most 120 s (the tool's cap), looped
at the join (`-stream_loop -1` on the bed input, `-shortest`) when the cut is longer.

**6. Cut (free).** `run_ffmpeg` takes at most 10 inputs and 4 outputs, so the cut is passes,
`wait_task` after each (its outputs are the next pass's inputs): per beat, mix that beat's
speech and cues over its clip into `beatN.mp4` at exactly the shot list's `length` —
`tpad=stop_mode=clone:stop_duration=<length>` on the video plus `apad` on the mixed audio
over-pad a short file (`stop_duration` is padding added, not a target) and `-t <length>`
then sets the final length; `-t` alone never extends anything —
up to four beats per call (one output each) and fewer when their files would pass ten
inputs: a clip plus every sound file counts, so a narrated beat with one cue is three
inputs and three such beats fill a call, and a beat with more than eight cue files gets its
cues pre-mixed into one stem first. Route-A and lipsynced beats already carry their sound
and go through the same pass with only their clip, for the length. Then concatenate the
beats in order in groups of at most 10 into `partN.mp4`, and repeat the grouping until the
final call — the parts plus the bed — fits; then that last call joins the parts, mixes the
bed under the dialogue and burns subtitles from a `.ass` sidecar built from the script's
lines (one cue per beat, timed from the `length` column of the shot list;
`sidecars=[{"name": "subs.ass", "content": ...}]` and the `ass` filter; the skeleton is in
`subtitles.md`; `timeout_seconds=600` on the group and final calls — a long join can pass the
300 s default). Output `final.mp4`. Sixteen beats is four mix calls, two group concats and
one final join — seven calls, a few more when cues shrink the mix batches. The budget is 30 runs per clock hour and every `run_ffmpeg` call counts, so a
part is at most 36 beats at one cue each — fewer with more sound — and the
next part waits for the next hour. With the review on, each analysed clip and each recheck is a
run too, taken from the same hour: a part's analyses, its expected rechecks and its cut calls
(the mix passes for its cue count, the concats and the join) add up to at most 30 — at one cue
per beat, a part whose beats all speak is at most 16 beats (16 analyses and 7 cut calls, room
for 7 rechecks and re-cuts together), and with more cues per beat fewer. Each part's clips are analysed right before
that part is cut, then the next part waits for the next hour.

**7. Hand over.** `wait_task` on the cut, then the review of the finished cut (`review-shots`
step 5); return the video link and `open_in_frameo`.

## Done

The user gets: the finished video, and a Frameo project with the cast portraits and
every shot on the canvas in script order, the dialogue clips, and the chat that records
each step — editable in the app from there. With it, the review note per shot: passed,
retaken and why, or kept with a known flaw. Say what was skipped (a beat that failed, a
voice the user may want to swap) rather than hiding it.

## Files

- `shot-list.md` — the cast and beat list template used in step 0.
- `subtitles.md` — the `.ass` sidecar skeleton and the timing rule for step 6.

## Worked example (a two-scene period drama, Hindi dialogue)

Cast: Chandrika, Arjun, the King of Avanti, Malti, Sudhan, an officer, a servant, a
minister; setting: the royal council hall. Two scenes, 15 beats: Chandrika's proposal over
the map (3), the King's approval and Arjun's reaction (3), Arjun turning on Malti (4), the
accusations (3), the banishment and Chandrika's silent close (2). Route for dialogue:
Chandrika, Arjun and the King speak more than once, so their lines go route B; step 3 tests
Chandrika's opening line both ways to pick her voice. Rough cost at 15 clips with a
5 s Seedance clip each and 12 spoken beats on route B: cast 8 × ~18 + beats 15 × ~18 + clips
15 × ~65 + lipsync 12 × 5 s × ~20 + sound ~5 = about 2,590 credits; a line on route A saves its
~100 credits of lipsync.

Referenced files: 4

subtitles-and-cut3.89 KB

View saved version →

---
name: subtitles-and-cut
description: "Trim, resize to 9:16 or 16:9, join clips, add a music bed, and burn captions from timings, a script, or a transcription of the footage — on footage the user already has. Free, except a transcription."
---

Needs: the user's video(s), uploaded through show_upload, imported from a web link with import_media_url, or already in Frameo; the caption text with timings, a script to fit, or nothing (the speech is transcribed)
Credits: 0 (ffmpeg is free; a music bed is ~1 per 50 characters of its prompt; transcribing the footage is a few credits per minute, charged after the run)
Time: 5–10 minutes

## When

Editing, not generating: cut a long take, make a vertical from a horizontal, join
takes, burn subtitles. Not for generating scenes.

## Ask first

1. **The footage** — links (uploaded or from `read_project`), each clip's length in
   seconds (`extract_frames` needs it; `read_project` reports `seconds` only for items the app has
   placed and saved, a generated clip's `wait_task` result has it as `duration_seconds`, and
   uploads carry none), and
   the target frame.
2. **Captions** — text with timings (SRT-like), a script to spread evenly, the footage's own
   speech (transcribed, a few credits per minute), or none.
3. **Sound** — keep the original, add a bed, or both.

## Steps

**Before anything.** `list_projects` (or `create_project`) for the `project_id` and, when the
project has several modules, the `module_id` — `extract_frames` and `run_ffmpeg` take that
pair. Nothing below spends credits. The one exception is a music bed the user wants
generated: one `generate_music` under the usual quote (`estimate_cost` with the same
project, told to the user, then the call with that `quote_id` and `confirmed_by_user=true`,
and `wait_task` for the link); a bed is at most 120 s (the tool's cap) and is looped at the
cut (`-stream_loop -1` on the bed input, `-shortest`) for longer footage.

**1. Look first (free).** `extract_frames(video_url, duration=<that clip's length>,
at_seconds=[…])` at a few timestamps to see what is in the footage before cutting; confirm the trim points with the user.

**2. Captions.** Write the `.ass` sidecar: one dialogue line per cue with start/end from
the user's timings. When only a script is given, spread the lines across the trimmed
length in proportion to their character counts and say the timing is approximate. When the
user wants the footage's own speech captioned, `analyze_media(media_url, question="Transcribe
the speech with start and end timestamps for every line", confirmed_by_user=true)` — the one
paid call with no quote: it is charged by clip length, a few credits per minute, the user is
told that and agrees first, and the finished run (`wait_task`) reports `credits_charged_estimate`.
Cues come from those timestamps, offset by the trim start.
Style: bottom-anchored, white with a dark outline, 15% of the height from the bottom
so phone UI does not cover it.

**3. The cut (free).** `run_ffmpeg` takes at most 10 inputs and 4 outputs: with more than
nine clips (plus the bed) concatenate in groups of at most 10 first — each input scaled, cropped,
`setsar=1` and set to one frame rate in that group call, since clips of different sizes do not join —
`wait_task`, and feed
those to the final call, repeating until the final call fits. That call: `-ss`/`-t` for the trim, `scale`+`crop` for 9:16
(centre crop unless the user points at a subject), `concat` for joins, `ass=subs.ass` for
captions, the bed mixed under with `amix`. Output `final.mp4`. `wait_task`, return the
link and `open_in_frameo`.

## Done

The cut on the Frameo project's canvas, shown with `show_generations` and `final: true`
where the chat app displays Frameo cards.
`open_in_frameo` on the result opens the project in the Frameo app, where every take sits on
the canvas, ready for retakes and edits.

## Files

- `ass-skeleton.md` — a minimal `.ass` file to fill in.

Referenced files: 3

talking-presenter6.77 KB

View saved version →

---
name: talking-presenter
description: "One person on camera delivers a script: a face (photo or designed), a fixed voice, lipsynced delivery, optional title card, cut to length."
---

Needs: a script; a photo of the presenter or a description; optionally a brand colour or logo
Credits: about 1,000–2,100 for a 30–60 second piece (lipsync is the main cost: ~20 credits per second of clip)
Time: 10–20 minutes

## When

Use for one person speaking to camera: an explainer, a founder message, a course intro, an
announcement, a customer-support answer. Any language the voice catalog covers.

Not for: several characters in scenes (`script-to-video`), or narration over other footage
(`faceless-story-video`).

## Ask first

1. **Who** — a photo (`show_upload` for files on the user's device, `import_media_url` for a web link; images attached to the chat do not reach Frameo, and `create_upload_url` is for clients that send the file themselves) or a description to design from.
2. **Voice** — the user's preference in words (warm, brisk, older, Hindi, British…); pick 3
   with `search_voices` and show their names.
3. **Frame** — 9:16 or 16:9, and whether a title card opens the piece.

Say the rough cost (the script's length decides it) in the same turn as the questions; the user's yes comes with the plan (Steps, **The plan first.**).

## Steps

**Before any paid step.** The project: `list_projects` (or `create_project`) gives the `project_id`
and, when the project has several modules, the `module_id`; every call below that takes a project —
`estimate_cost` included — gets that same pair; without it those tools answer `project_needed`. The
quote: one `estimate_cost(items=[…])` prices a stage in one call, with the same project, model,
size and number of `image_urls` or `reference_image_urls` (`reference_count`) as each generate call, and returns a `quote_id` per item plus the total. Each generate call
then passes its own item's `quote_id` and `confirmed_by_user=true`. A quote is single-use and lasts
15 minutes, so a long plan is priced stage by stage, right before each stage runs; a stage that
comes to more than the user approved is asked about again first. Generate calls return
`generation_ids`; `wait_task` returns the links, and `show_generations` shows each stage's running
and finished work in one card where the chat app displays Frameo cards: all the stage's ids at
once, before its first `wait_task`, and the finished result with `final: true`.

**The plan first.** Before the first paid call, the plan goes to the user in the chat as plain
text: what will be made, in order, one line per generation (for a script, the shot list; for a
set, each shot), with each line's credits and the total. The credits come from
`estimate_cost(items=[…])`, up to 10 items a call, so a long plan takes several calls; those
quotes may expire unused, since each stage is quoted again right before it runs. Nothing is
generated until the user says yes. The user can drop or change lines; a changed line is priced
again.

**Canvas rows.** Pass `shot_number` on every `generate_image` and `generate_video` of a shot (1, 2, 3… in story order; the same number for a retake and for that shot's video), so each shot gets its own row on the Frameo canvas. Cast, prop and location references take `placement_kind` (`character`, `prop` or `location`) and the subject's name as `placement_group` instead: they sit on their own board, and the shots built from them do not pile into their row.

**1. The presenter still (~18 credits).** `generate_image`: a mid-shot of the presenter facing
camera, neutral background unless the user wants one, `image_urls=[photo]` when there is a
photo. `wait_task`, show it; adjust once if asked.

**2. The voice (~1 credit per 50 characters).** `generate_speech(text=script, voice_id)` in one
call when the whole script fits one clip (15 s, the lipsync's cap, or the video model's max from
`list_models(kind="video")` when that is lower); otherwise split at sentence breaks — and at clause breaks (commas, dashes) when one
sentence alone is still too long — so each file fits one clip. `wait_task` gives the link; play it back to the user before spending on video. The result's
`duration_seconds`, the file's measured spoken length, is its raw estimate, kept as is for the
lipsync; a result without one is estimated at 2.8 words per second (2.5 for Hindi) plus one
second of headroom, and that estimate is approximate. When the user picks a presenter
`list_characters` already has in the project, its first image is the still and its `voice_id` the
voice. A new
presenter is saved with `save_character` under a name not already in the cast (the same name
replaces that character's voice), with `image_urls=[presenter still]` and the `voice_id`, so later
videos in this project reuse the same face and voice.

**3. The clip (~65 credits per 5 s + ~20 per second of lipsync).** For each speech file:
`generate_video` from the still (`first_frame_url`) with a prompt like "presenter speaking
to camera, subtle natural head movement, steady framing", `duration` = that file's raw
estimate rounded up to whole seconds, kept inside the video model's range
(`list_models(kind="video")`; 4–30 s on the default) and at most 15 s — step 2 sized the files to
that, and a file that measures a little longer still goes in whole as `audio_duration`,
`wait_task`, then `generate_lipsync(video_url, audio_url, duration, audio_duration=the raw
estimate, aspect_ratio, resolution)` — the raw one, so a clamped clip is still extended to
the speech; the size the clip was made at; and `duration` never above 15 s, the lipsync's
own cap, whatever the video model allows — and
`wait_task`. Test the first clip
before the rest.

**4. Title card (free).** If wanted: `render_motion_graphics` with `template="title_card"`, the title
(and a `subtitle`), the piece's `size` and a `duration` of 2 to 3 s; `wait_task` returns the card's clip,
joined in front in the cut. It counts against the 30 runs per clock hour.

**5. Cut (free).** With a title card, one pass of its own first: the card comes at full HD in its
`size`, so it is scaled to the clips' size (`scale`, then `setsar=1`) and given a bounded silent
track (`-f lavfi -t <card seconds> -i anullsrc=r=<the clips' sample rate>:cl=stereo`); an
unbounded `anullsrc` never ends and stalls the join. Then `run_ffmpeg` takes at most 10 inputs and 4 outputs:
concatenate the lipsynced clips in
order, in groups of at most 10 — the card counts as an input, so nine clips per group when
one is joined in front — and repeat until the final call fits (`wait_task` after each pass;
its outputs are the next pass's inputs). Output `final.mp4`. `wait_task`,
return the link and `open_in_frameo`.

## Done

The finished clip, the presenter still, and every clip in the Frameo project (chat and
canvas) for edits in the app.

## Files

None.

Referenced files: 2

thumbnail-set3.88 KB

View saved version →

---
name: thumbnail-set
description: "Three to five thumbnail concepts for a video: a face or product, a short title treatment, high contrast, delivered at upload size."
---

Needs: the video's title; a face photo or the product photos; the platform (YouTube 16:9, Instagram 1:1 or 4:5)
Credits: about 110–180 (3–5 images at ~36 each, the default model's 4K tier)
Time: 10 minutes

## When

After the video exists: a set to A/B. Not for the video itself.

## Ask first

1. **Title** — the 2–4 words that go on the image.
2. **Subject** — a face (`show_upload` for files on the user's device, `import_media_url` for a web link; images attached to the chat do not reach Frameo, and `create_upload_url` is for clients that send the file themselves) or the product photos.
3. **Platform** — sets the frame: YouTube 16:9, Instagram 1:1; for Instagram 4:5 generate at
   3:4 (the nearest supported ratio) and crop at the end.

## Steps

**Before any paid step.** The project: `list_projects` (or `create_project`) gives the `project_id`
and, when the project has several modules, the `module_id`; every call below that takes a project —
`estimate_cost` included — gets that same pair; without it those tools answer `project_needed`. The
quote: one `estimate_cost(items=[…])` prices a stage in one call, with the same project, model,
size and number of `image_urls` or `reference_image_urls` (`reference_count`) as each generate call, and returns a `quote_id` per item plus the total. Each generate call
then passes its own item's `quote_id` and `confirmed_by_user=true`. A quote is single-use and lasts
15 minutes, so a long plan is priced stage by stage, right before each stage runs; a stage that
comes to more than the user approved is asked about again first. Generate calls return
`generation_ids`; `wait_task` returns the links, and `show_generations` shows each stage's running
and finished work in one card where the chat app displays Frameo cards: all the stage's ids at
once, before its first `wait_task`, and the finished result with `final: true`.

**The plan first.** Before the first paid call, the plan goes to the user in the chat as plain
text: what will be made, in order, one line per generation (for a script, the shot list; for a
set, each shot), with each line's credits and the total. The credits come from
`estimate_cost(items=[…])`, up to 10 items a call, so a long plan takes several calls; those
quotes may expire unused, since each stage is quoted again right before it runs. Nothing is
generated until the user says yes. The user can drop or change lines; a changed line is priced
again.

**1. Concepts.** Propose 3–5 in one line each (big face + reaction, product + number, before/
after, a question, a bold object on colour); the user picks or edits.

**2. Images (~36 each at 4K).** `generate_image` per concept, each with its own quote,
`image_urls=[face or product]`,
prompt = the concept plus "high contrast, one focal point, space on the left third for
text, no small details", `aspect_ratio` from Ask first (1:1, 16:9, 9:16 or 3:4) and
`image_resolution` at the highest tier the model offers (`list_models(kind="image")`
names them, e.g. 4K on the default; the quote is exact for that model and tier). Leave the title off the image: the exact title is burned in at step 3, not generated.
`wait_task`, then show the set.

**3. Title and exact size (free).** One `run_ffmpeg` per pick: `scale` + `crop` to the
platform's pixel size (1280×720, 1080×1080, 1080×1350), and the title burned in from an `.ass`
sidecar with the `ass` filter (Noto Sans, large and bold, thick outline, in the free space) unless the
user prefers their own editor; ask which. `wait_task` for each
output link.

**4. Deliver.** Links and `open_in_frameo`.

## Done

The set on the canvas in one row.
`open_in_frameo` on the result opens the project in the Frameo app, where every take sits on
the canvas, ready for retakes and edits.

## Files

None.

Referenced files: 2

ugc-product-video7.61 KB

View saved version →

---
name: ugc-product-video
description: "Product-only UGC: handheld-style shots of the product with a voiceover, no creator on camera."
---

Needs: product photos; the one claim
Credits: about 390–470 for 30 seconds, plus a retake allowance of 15% of the image and clip total
Time: 10–20 minutes, plus about a minute per clip for the review

## When

When there is no creator to show, or the brand wants the product alone. Not for a person on
camera (`ugc-review-video`).

## Ask first

1. **The product** — photos (`show_upload` for files on the user's device, `import_media_url` for a web link; images attached to the chat do not reach Frameo, and `create_upload_url` is for clients that send the file themselves), the name, and the one thing the video
   must say about it.
2. **The voice** — the tone of the voiceover (warm, brisk, playful) and its language.
3. **Frame and length** — 9:16, 15 / 30 / 45 s.

Say the rough cost in the same turn as the questions; the user's yes comes with the plan (Steps, **The plan first.**).

## Steps

**Before any paid step.** The project: `list_projects` (or `create_project`) gives the `project_id`
and, when the project has several modules, the `module_id`; every call below that takes a project —
`estimate_cost` included — gets that same pair; without it those tools answer `project_needed`. The
quote: one `estimate_cost(items=[…])` prices a stage in one call, with the same project, model,
size and number of `image_urls` or `reference_image_urls` (`reference_count`) as each generate call, and returns a `quote_id` per item plus the total. Each generate call
then passes its own item's `quote_id` and `confirmed_by_user=true`. A quote is single-use and lasts
15 minutes, so a long plan is priced stage by stage, right before each stage runs; a stage that
comes to more than the user approved is asked about again first. Generate calls return
`generation_ids`; `wait_task` returns the links, and `show_generations` shows each stage's running
and finished work in one card where the chat app displays Frameo cards: all the stage's ids at
once, before its first `wait_task`, and the finished result with `final: true`.

**The plan first.** Before the first paid call, the plan goes to the user in the chat as plain
text: what will be made, in order, one line per generation (for a script, the shot list; for a
set, each shot), with each line's credits and the total. The credits come from
`estimate_cost(items=[…])`, up to 10 items a call, so a long plan takes several calls; those
quotes may expire unused, since each stage is quoted again right before it runs. Nothing is
generated until the user says yes. The user can drop or change lines; a changed line is priced
again.

**Canvas rows.** Pass `shot_number` on every `generate_image` and `generate_video` of a shot (1, 2, 3… in story order; the same number for a retake and for that shot's video), so each shot gets its own row on the Frameo canvas. Cast, prop and location references take `placement_kind` (`character`, `prop` or `location`) and the subject's name as `placement_group` instead: they sit on their own board, and the shots built from them do not pile into their row.

**Review (on by default).** Every still and clip is checked before it is built on, and the
finished cut before it is handed over, following `review-shots`: its checklist and steps come
from `get_skill("review-shots")`, loaded before step 1. The checks run on the step-1 still
before step 2 builds on it, after step 2 (`review-shots` steps 0, 1 and 4), after step 4 (steps 2 and 4) and after the cut (step 5); the
product's label, shape and colour are checked against the product photos in every shot, and a
failed shot is retaken once. No clip has speech on screen, so every check is free; a retake is
paid like the shot it replaces, from the retake allowance the plan carries as its own line for
the user to approve. The user can turn the review off.

**0. Script.** Write a voiceover script of the length chosen in Ask first: hook in the first
line, the one claim, a close. Show it; the user edits before anything is generated.

**1. Establishing still (~18).** `generate_image`: the product in a real setting, phone-video
framing, `image_urls=[product photos…]`; no person in frame. 

**2. Product-in-scene stills (~18 each).** `generate_image` with `image_urls=[product photos…]` for each beat that shows the product
(hero on a surface, in use, detail, packaging).

**3. Voice (~1 per 50 characters).** First the beat list, no tool call: estimate each beat
at 2.8 words per second (2.5 for Hindi) plus one second of headroom, say it is
approximate, and split until every piece is within it — more lines, more clips — before any speech or
clip is generated: a beat whose estimate is over the video model's max (30 s on the default) is
split, because there is no lipsync here to stretch the picture and the mix would run
past it. Then
`search_voices` for the tone and language, 3 candidates, then `generate_speech` per final beat (each
paid call, here and below, has its own quote); `wait_task` for the links. Each speech
result's `duration_seconds` is its measured spoken length; from here on the estimate means that
length when the result has one, and the words-per-second figure only when it has none. That estimate is the clip's `duration` in step 4.

**4. Clips (~65 per 5 s).** `generate_video` per final beat — one clip per piece, the first piece of a split beat on its source beat's still and each
later one on the previous piece's last frame (below) — with a handheld push-in or slow orbit in the prompt, silent; the voiceover is mixed at the cut. Every `generate_video` takes `duration` = that beat's speech estimate rounded up to whole
seconds and kept inside the model's range (`list_models(kind="video")` gives it; 4–30 s on
the default). `wait_task` each. Test the first beat before the rest.

On Seedance models (the default) the product photos go along as `reference_image_urls` on every
clip that shows the product, which holds the label through the motion. The frame is the first
reference image, so the prompt names the others from the second image on, and it counts toward the model's cap;
the quote carries `first_frame=true` and `reference_count` = the number of
`reference_image_urls`, not counting the frame. Each
later piece of a split beat opens on the previous piece's last frame (`extract_frames` on that
clip) once that piece's look has been checked, instead of the beat's still, so the join reads
as one take. Each prompt names where the clip ends (the product in frame) and nothing the beat
does not show.

**5. Cut (free).** `run_ffmpeg` takes at most 10 inputs and 4 outputs, so the cut is passes, `wait_task`
after each (its outputs are the next pass's inputs): per beat, mix its step-3 speech over its clip (`amix`, speech on top) into `beatN.mp4` at exactly
the clip's `duration` (`tpad`/`apad` over-pad, `-t <duration>` sets the length; captions are
timed from those durations), up to four beats per call (one output each) and fewer when their files would pass ten inputs
(a clip plus its speech is two inputs and four such beats fill a call); then concatenate the beats in order in groups of at most 10;
then one last call joins the groups and burns captions from an `.ass` sidecar in a casual
style. Output `final.mp4`, review the finished cut (`review-shots` step 5), and return the link
and `open_in_frameo`.

## Done

The UGC clip, the script, and every still and clip in the Frameo project, with the review note
per shot: passed, retaken and why, or kept with a known flaw.
`open_in_frameo` on the result opens the project in the Frameo app, where every take sits on
the canvas, ready for retakes and edits.

## Files

None.

Referenced files: 2

ugc-review-video9.29 KB

View saved version →

---
name: ugc-review-video
description: "The default UGC ad, a creator on camera reviewing a product in phone-video style: hook, one honest claim, product in hand, a close — lipsynced delivery."
---

Needs: product photos; the one claim; a creator photo or a description
Credits: about 500–650 for 30 seconds, plus a retake allowance of 15% of the image and clip total
Time: 15–25 minutes, plus about a minute per clip for the review

## When

The default UGC ask: a person talking about a product as if to their followers. Not for
product-only shots (`ugc-product-video`), unboxing (`ugc-unboxing-video`), wearing it
(`ugc-try-on-video`) or a how-to (`ugc-tutorial-video`).

## Ask first

1. **The product** — photos (`show_upload` for files on the user's device, `import_media_url` for a web link; images attached to the chat do not reach Frameo, and `create_upload_url` is for clients that send the file themselves), the name, and the one thing the video
   must say about it.
2. **The creator** — design one (age, vibe, setting), use the user's photo, or reuse a creator
   already in the project (`list_characters`: its first image is the step-1 still and its
   `voice_id` the voice in step 3, with nothing to generate or choose).
3. **Frame and length** — 9:16, 15 / 30 / 45 s.

Say the rough cost in the same turn as the questions; the user's yes comes with the plan (Steps, **The plan first.**).

## Steps

**Before any paid step.** The project: `list_projects` (or `create_project`) gives the `project_id`
and, when the project has several modules, the `module_id`; every call below that takes a project —
`estimate_cost` included — gets that same pair; without it those tools answer `project_needed`. The
quote: one `estimate_cost(items=[…])` prices a stage in one call, with the same project, model,
size and number of `image_urls` or `reference_image_urls` (`reference_count`) as each generate call, and returns a `quote_id` per item plus the total. Each generate call
then passes its own item's `quote_id` and `confirmed_by_user=true`. A quote is single-use and lasts
15 minutes, so a long plan is priced stage by stage, right before each stage runs; a stage that
comes to more than the user approved is asked about again first. Generate calls return
`generation_ids`; `wait_task` returns the links, and `show_generations` shows each stage's running
and finished work in one card where the chat app displays Frameo cards: all the stage's ids at
once, before its first `wait_task`, and the finished result with `final: true`.

**The plan first.** Before the first paid call, the plan goes to the user in the chat as plain
text: what will be made, in order, one line per generation (for a script, the shot list; for a
set, each shot), with each line's credits and the total. The credits come from
`estimate_cost(items=[…])`, up to 10 items a call, so a long plan takes several calls; those
quotes may expire unused, since each stage is quoted again right before it runs. Nothing is
generated until the user says yes. The user can drop or change lines; a changed line is priced
again.

**Canvas rows.** Pass `shot_number` on every `generate_image` and `generate_video` of a shot (1, 2, 3… in story order; the same number for a retake and for that shot's video), so each shot gets its own row on the Frameo canvas. Cast, prop and location references take `placement_kind` (`character`, `prop` or `location`) and the subject's name as `placement_group` instead: they sit on their own board, and the shots built from them do not pile into their row.

**Review (on by default).** Every still and clip is checked before it is built on, and the
finished cut before it is handed over, following `review-shots`: its checklist and steps come
from `get_skill("review-shots")`, loaded before step 1. The checks run on the step-1 still
before step 2 builds on it, after step 2 (`review-shots` steps 0, 1 and 4), after step 4 (steps 2 and 4) and after the cut (step 5); the
product's label, shape and colour are checked against the product photos in every shot, and a
failed shot is retaken once. The plan carries the retake allowance as its own line for the user
to approve with everything else. The paid check of clips with speech (`review-shots` step 3)
runs only when the user asks for it, since the talking beats are lipsynced to one fixed voice,
and its rough cost (a few credits per minute of clip) is said and agreed first; without it, the
note on each talking clip says its sound was not checked.
The user can turn the review off.

**0. Script.** Write a script of the length chosen in Ask first in the creator's voice: hook in the first line, the one
claim, a close. Show it; the user edits before anything is generated.

**1. Creator still (~18).** `generate_image`: the creator at home, phone held at arm's length, looking into the lens, phone-video framing,
`image_urls=[creator photo]` if given. 

**2. Product-in-scene stills (~18 each).** `generate_image` for each beat that shows the product
(in hand, on the table, close-up), with `image_urls=[step-1 creator still, product photos…]` when the creator or their hands are in
frame (the same person every time) and `image_urls=[product photos…]` when only the product
is; the prompt names nothing the beat does not show.

**3. Voice (~1 per 50 characters).** First the beat list, no tool call: estimate each beat
at 2.8 words per second (2.5 for Hindi) plus one second of headroom, say it is
approximate, and split any beat that is too long — the product beats at the video model's max
(`list_models(kind="video")`; 30 s on the default) because the mix at the cut cannot stretch
the picture, the talking beats at the smaller of that max and 15 s, the lipsync's cap —
until every piece is within it, so every clip's `duration` is its length on the timeline. Then
`search_voices` for the creator's age and language, 3 candidates, then `generate_speech` per final beat (each
paid call, here and below, has its own quote); `wait_task` for the links. The creator the user keeps is saved with `save_character` (a name,
`image_urls=[creator still]` and the chosen `voice_id`), so the next video in this project
reuses the same face and voice. Each speech
result's `duration_seconds` is its measured spoken length; from here on the estimate means that
length when the result has one, and the words-per-second figure only when it has none. That estimate is the clip's `duration` in step 4 and the `duration` / `audio_duration`
of its lipsync.

**4. Clips (~65 per 5 s + ~20 per second of lipsync).** For talking beats: `generate_video` from the creator still, then `generate_lipsync` with the step-3 speech for that beat. For product beats: `generate_video` from the product still with a small handheld move, silent. Every `generate_video` takes `duration` = that beat's speech estimate rounded up to whole
seconds and kept inside the model's range (`list_models(kind="video")` gives it; 4–30 s on
the default), 5 s for a beat with no speech; every `generate_lipsync` takes that same whole
`duration` (never above 15 s — the lipsync's own cap, whatever the video model allows),
`audio_duration` = the estimate, and the clip's `aspect_ratio` and `resolution`. `wait_task` each. Test the first beat before
the rest.

On Seedance models (the default) the product photos go along as `reference_image_urls` on every
clip that shows the product, and the creator still on every product clip the creator is in,
which holds the label and the face through the motion. The frame is the first reference image,
so the prompt names the others from the second image on, and it counts toward the model's cap;
the quote carries `first_frame=true` and `reference_count` = the number of
`reference_image_urls`, not counting the frame. Each
later piece of a split beat opens on the previous piece's last frame (`extract_frames` on the
clip that goes into the cut, the lipsync result on a talking beat) once that piece's look has
been checked, instead of the beat's still, so the join reads as one take. Each prompt names where
the clip ends and nothing the beat does not show: on a lipsynced beat the creator facing the
lens; on a silent beat the product in frame with the creator's eyes on it, never on the lens,
since the voiceover plays over lips that do not move.

**5. Cut (free).** `run_ffmpeg` takes at most 10 inputs and 4 outputs, so the cut is passes, `wait_task`
after each (its outputs are the next pass's inputs): per beat — lipsynced beats bring only their clip — mix its step-3 speech over its clip (`amix`, speech on top) into `beatN.mp4` at exactly
the clip's `duration` (`tpad`/`apad` over-pad, `-t <duration>` sets the length; captions are
timed from those durations), up to four beats per call (one output
each) and fewer when their files would pass ten inputs
(a clip plus its speech is two inputs and four such beats fill a call); then concatenate all the beats in
order in groups of at most 10; then one last call joins the groups and burns captions
from an `.ass` sidecar in a casual style. Output `final.mp4`, review the finished cut (`review-shots` step 5), and return the link
and `open_in_frameo`.

## Done

The UGC clip, the script, and every still and clip in the Frameo project, with the review note
per shot: passed, retaken and why, or kept with a known flaw.
`open_in_frameo` on the result opens the project in the Frameo app, where every take sits on
the canvas, ready for retakes and edits.

## Files

None.

Referenced files: 2

ugc-try-on-video9.3 KB

View saved version →

---
name: ugc-try-on-video
description: "Wearing and fit-check for clothing, jewellery, glasses or shoes: the creator in the item, a turn, a close-up, a verdict."
---

Needs: product photos; a creator photo (the person who wears it) or a description; size or fit notes
Credits: about 500–650 for 30 seconds, plus a retake allowance of 15% of the image and clip total
Time: 15–25 minutes, plus about a minute per clip for the review

## When

Wearable products. The creator's photo matters more here than anywhere else: the item has
to sit on a consistent person. Not for products that are held or used (`ugc-review-video`).

## Ask first

1. **The product** — photos (`show_upload` for files on the user's device, `import_media_url` for a web link; images attached to the chat do not reach Frameo, and `create_upload_url` is for clients that send the file themselves), the name, and the one thing the video
   must say about it.
2. **The creator** — design one (age, vibe, setting), use the user's photo, or reuse a creator
   already in the project (`list_characters`: its first image is the step-1 still and its
   `voice_id` the voice in step 3, with nothing to generate or choose).
3. **Frame and length** — 9:16, 15 / 30 / 45 s.

Say the rough cost in the same turn as the questions; the user's yes comes with the plan (Steps, **The plan first.**).

## Steps

**Before any paid step.** The project: `list_projects` (or `create_project`) gives the `project_id`
and, when the project has several modules, the `module_id`; every call below that takes a project —
`estimate_cost` included — gets that same pair; without it those tools answer `project_needed`. The
quote: one `estimate_cost(items=[…])` prices a stage in one call, with the same project, model,
size and number of `image_urls` or `reference_image_urls` (`reference_count`) as each generate call, and returns a `quote_id` per item plus the total. Each generate call
then passes its own item's `quote_id` and `confirmed_by_user=true`. A quote is single-use and lasts
15 minutes, so a long plan is priced stage by stage, right before each stage runs; a stage that
comes to more than the user approved is asked about again first. Generate calls return
`generation_ids`; `wait_task` returns the links, and `show_generations` shows each stage's running
and finished work in one card where the chat app displays Frameo cards: all the stage's ids at
once, before its first `wait_task`, and the finished result with `final: true`.

**The plan first.** Before the first paid call, the plan goes to the user in the chat as plain
text: what will be made, in order, one line per generation (for a script, the shot list; for a
set, each shot), with each line's credits and the total. The credits come from
`estimate_cost(items=[…])`, up to 10 items a call, so a long plan takes several calls; those
quotes may expire unused, since each stage is quoted again right before it runs. Nothing is
generated until the user says yes. The user can drop or change lines; a changed line is priced
again.

**Canvas rows.** Pass `shot_number` on every `generate_image` and `generate_video` of a shot (1, 2, 3… in story order; the same number for a retake and for that shot's video), so each shot gets its own row on the Frameo canvas. Cast, prop and location references take `placement_kind` (`character`, `prop` or `location`) and the subject's name as `placement_group` instead: they sit on their own board, and the shots built from them do not pile into their row.

**Review (on by default).** Every still and clip is checked before it is built on, and the
finished cut before it is handed over, following `review-shots`: its checklist and steps come
from `get_skill("review-shots")`, loaded before step 1. The checks run on the step-1 still
before step 2 builds on it, after step 2 (`review-shots` steps 0, 1 and 4), after step 4 (steps 2 and 4) and after the cut (step 5); the
product's label, shape and colour are checked against the product photos in every shot, and a
failed shot is retaken once. The plan carries the retake allowance as its own line for the user
to approve with everything else. The paid check of clips with speech (`review-shots` step 3)
runs only when the user asks for it, since the talking beats are lipsynced to one fixed voice,
and its rough cost (a few credits per minute of clip) is said and agreed first; without it, the
note on each talking clip says its sound was not checked.
The user can turn the review off.

**0. Script.** Write a script of the length chosen in Ask first in the creator's voice: hook in the first line, the one
claim, a close. Show it; the user edits before anything is generated.

**1. Creator still (~18).** `generate_image`: the creator full-length in front of a mirror or a
plain wall, wearing the item, phone-video framing. `image_urls=[creator photo, product
photos…]` in that order: the person is the base, the item the reference. With no creator
photo, the first product photo is the base and the prompt describes the person. 

**2. Product-in-scene stills (~18 each).** `generate_image` for each beat that shows the item on
the creator (front, a half turn, a detail) with `image_urls=[step-1 creator still, product
photos…]` — the same person as the base every time, the item as the reference.

**3. Voice (~1 per 50 characters).** First the beat list, no tool call: estimate each beat
at 2.8 words per second (2.5 for Hindi) plus one second of headroom, say it is
approximate, and split any beat that is too long — the fit beats at the video model's max
(`list_models(kind="video")`; 30 s on the default) because the mix at the cut cannot stretch
the picture, the talking beats at the smaller of that max and 15 s, the lipsync's cap —
until every piece is within it, so every clip's `duration` is its length on the timeline. Then
`search_voices` for the creator's age and language, 3 candidates, then `generate_speech` per final beat (each
paid call, here and below, has its own quote); `wait_task` for the links. The creator the user keeps is saved with `save_character` (a name,
`image_urls=[creator still]` and the chosen `voice_id`), so the next video in this project
reuses the same face and voice. Each speech
result's `duration_seconds` is its measured spoken length; from here on the estimate means that
length when the result has one, and the words-per-second figure only when it has none. That estimate is the clip's `duration` in step 4 and the `duration` / `audio_duration`
of its lipsync.

**4. Clips (~65 per 5 s + ~20 per second of lipsync).** Verdict beats: `generate_video` from the creator still, then `generate_lipsync` with the step-3 speech. Fit beats: `generate_video` from the try-on stills with 'slow half turn' / 'steps closer' in the prompt, silent. Every `generate_video` takes `duration` = that beat's speech estimate rounded up to whole
seconds and kept inside the model's range (`list_models(kind="video")` gives it; 4–30 s on
the default), 5 s for a beat with no speech; every `generate_lipsync` takes that same whole
`duration` (never above 15 s — the lipsync's own cap, whatever the video model allows),
`audio_duration` = the estimate, and the clip's `aspect_ratio` and `resolution`. `wait_task` each. Test the first beat before
the rest.

On Seedance models (the default) the product photos go along as `reference_image_urls` on every
clip that shows the product, and the creator still on every product clip the creator is in,
which holds the label and the face through the motion. The frame is the first reference image,
so the prompt names the others from the second image on, and it counts toward the model's cap;
the quote carries `first_frame=true` and `reference_count` = the number of
`reference_image_urls`, not counting the frame. Each
later piece of a split beat opens on the previous piece's last frame (`extract_frames` on the
clip that goes into the cut, the lipsync result on a talking beat) once that piece's look has
been checked, instead of the beat's still, so the join reads as one take. Each prompt names where
the clip ends and nothing the beat does not show: on a lipsynced beat the creator facing the
lens; on a silent beat the product in frame with the creator's eyes on it, never on the lens,
since the voiceover plays over lips that do not move.

**5. Cut (free).** `run_ffmpeg` takes at most 10 inputs and 4 outputs, so the cut is passes, `wait_task`
after each (its outputs are the next pass's inputs): per beat — lipsynced beats bring only their clip — mix its step-3 speech over its clip (`amix`, speech on top) into `beatN.mp4` at exactly
the clip's `duration` (`tpad`/`apad` over-pad, `-t <duration>` sets the length; captions are
timed from those durations), up to four beats per call (one output
each) and fewer when their files would pass ten inputs
(a clip plus its speech is two inputs and four such beats fill a call); then concatenate all the beats in
order in groups of at most 10; then one last call joins the groups and burns captions
from an `.ass` sidecar in a casual style. Output `final.mp4`, review the finished cut (`review-shots` step 5), and return the link
and `open_in_frameo`.

## Done

The UGC clip, the script, and every still and clip in the Frameo project, with the review note
per shot: passed, retaken and why, or kept with a known flaw.
`open_in_frameo` on the result opens the project in the Frameo app, where every take sits on
the canvas, ready for retakes and edits.

## Files

None.

Referenced files: 2

ugc-tutorial-video9.35 KB

View saved version →

---
name: ugc-tutorial-video
description: "A step-by-step how-to with \"Step N\" captions: setup, each step shown, the result."
---

Needs: product photos; the steps in order; a creator photo or a description
Credits: about 600–750 for 45 seconds, plus a retake allowance of 15% of the image and clip total
Time: 15–30 minutes, plus about a minute per clip for the review

## When

When the product needs showing how: setup, a routine, a recipe, an app flow. Not for a
review (`ugc-review-video`).

## Ask first

1. **The product** — photos (`show_upload` for files on the user's device, `import_media_url` for a web link; images attached to the chat do not reach Frameo, and `create_upload_url` is for clients that send the file themselves), the name, and the one thing the video
   must say about it.
2. **The creator** — design one (age, vibe, setting), use the user's photo, or reuse a creator
   already in the project (`list_characters`: its first image is the step-1 still and its
   `voice_id` the voice in step 3, with nothing to generate or choose).
3. **Frame and length** — 9:16, 15 / 30 / 45 s.

Say the rough cost in the same turn as the questions; the user's yes comes with the plan (Steps, **The plan first.**).

## Steps

**Before any paid step.** The project: `list_projects` (or `create_project`) gives the `project_id`
and, when the project has several modules, the `module_id`; every call below that takes a project —
`estimate_cost` included — gets that same pair; without it those tools answer `project_needed`. The
quote: one `estimate_cost(items=[…])` prices a stage in one call, with the same project, model,
size and number of `image_urls` or `reference_image_urls` (`reference_count`) as each generate call, and returns a `quote_id` per item plus the total. Each generate call
then passes its own item's `quote_id` and `confirmed_by_user=true`. A quote is single-use and lasts
15 minutes, so a long plan is priced stage by stage, right before each stage runs; a stage that
comes to more than the user approved is asked about again first. Generate calls return
`generation_ids`; `wait_task` returns the links, and `show_generations` shows each stage's running
and finished work in one card where the chat app displays Frameo cards: all the stage's ids at
once, before its first `wait_task`, and the finished result with `final: true`.

**The plan first.** Before the first paid call, the plan goes to the user in the chat as plain
text: what will be made, in order, one line per generation (for a script, the shot list; for a
set, each shot), with each line's credits and the total. The credits come from
`estimate_cost(items=[…])`, up to 10 items a call, so a long plan takes several calls; those
quotes may expire unused, since each stage is quoted again right before it runs. Nothing is
generated until the user says yes. The user can drop or change lines; a changed line is priced
again.

**Canvas rows.** Pass `shot_number` on every `generate_image` and `generate_video` of a shot (1, 2, 3… in story order; the same number for a retake and for that shot's video), so each shot gets its own row on the Frameo canvas. Cast, prop and location references take `placement_kind` (`character`, `prop` or `location`) and the subject's name as `placement_group` instead: they sit on their own board, and the shots built from them do not pile into their row.

**Review (on by default).** Every still and clip is checked before it is built on, and the
finished cut before it is handed over, following `review-shots`: its checklist and steps come
from `get_skill("review-shots")`, loaded before step 1. The checks run on the step-1 still
before step 2 builds on it, after step 2 (`review-shots` steps 0, 1 and 4), after step 4 (steps 2 and 4) and after the cut (step 5); the
product's label, shape and colour are checked against the product photos in every shot, and a
failed shot is retaken once. The plan carries the retake allowance as its own line for the user
to approve with everything else. The paid check of clips with speech (`review-shots` step 3)
runs only when the user asks for it, since the talking beats are lipsynced to one fixed voice,
and its rough cost (a few credits per minute of clip) is said and agreed first; without it, the
note on each talking clip says its sound was not checked.
The user can turn the review off.

**0. Script.** Write a script of the length chosen in Ask first in the creator's voice: hook in the first line, the one
claim, a close. Show it; the user edits before anything is generated.

**1. Creator still (~18).** `generate_image`: the creator at the place the steps happen, hands and product in frame, phone-video framing,
`image_urls=[creator photo]` if given. 

**2. Product-in-scene stills (~18 each).** `generate_image` for each beat that shows the product
(one per step, hands doing the step), with `image_urls=[step-1 creator still, product photos…]` when the creator or their hands are in
frame (the same person every time) and `image_urls=[product photos…]` when only the product
is; the prompt names nothing the beat does not show.

**3. Voice (~1 per 50 characters).** First the beat list, no tool call: estimate each beat
at 2.8 words per second (2.5 for Hindi) plus one second of headroom, say it is
approximate, and split any beat that is too long — the step beats at the video model's max
(`list_models(kind="video")`; 30 s on the default) because the mix at the cut cannot stretch
the picture, the talking beats at the smaller of that max and 15 s, the lipsync's cap —
until every piece is within it, so every clip's `duration` is its length on the timeline. Then
`search_voices` for the creator's age and language, 3 candidates, then `generate_speech` per final beat (each
paid call, here and below, has its own quote); `wait_task` for the links. The creator the user keeps is saved with `save_character` (a name,
`image_urls=[creator still]` and the chosen `voice_id`), so the next video in this project
reuses the same face and voice. Each speech
result's `duration_seconds` is its measured spoken length; from here on the estimate means that
length when the result has one, and the words-per-second figure only when it has none. That estimate is the clip's `duration` in step 4 and the `duration` / `audio_duration`
of its lipsync.

**4. Clips (~65 per 5 s + ~20 per second of lipsync for the intro and outro).** Intro and outro: `generate_video` from the creator still, then `generate_lipsync` with the step-3 speech. Steps: `generate_video` per final step beat — one clip per piece, the first piece of a split beat on its source beat's still and each
later one on the previous piece's last frame (below) — with the action in the prompt, silent, narration over it. The `.ass` sidecar at the cut carries a 'Step N — …' caption per step. Every `generate_video` takes `duration` = that beat's speech estimate rounded up to whole
seconds and kept inside the model's range (`list_models(kind="video")` gives it; 4–30 s on
the default), 5 s for a beat with no speech; every `generate_lipsync` takes that same whole
`duration` (never above 15 s — the lipsync's own cap, whatever the video model allows),
`audio_duration` = the estimate, and the clip's `aspect_ratio` and `resolution`. `wait_task` each. Test the first beat before
the rest.

On Seedance models (the default) the product photos go along as `reference_image_urls` on every
clip that shows the product, and the creator still on every product clip the creator is in,
which holds the label and the face through the motion. The frame is the first reference image,
so the prompt names the others from the second image on, and it counts toward the model's cap;
the quote carries `first_frame=true` and `reference_count` = the number of
`reference_image_urls`, not counting the frame. Each
later piece of a split beat opens on the previous piece's last frame (`extract_frames` on the
clip that goes into the cut, the lipsync result on a talking beat) once that piece's look has
been checked, instead of the beat's still, so the join reads as one take. Each prompt names where
the clip ends and nothing the beat does not show: on a lipsynced beat the creator facing the
lens; on a silent beat the product in frame with the creator's eyes on it, never on the lens,
since the voiceover plays over lips that do not move.

**5. Cut (free).** `run_ffmpeg` takes at most 10 inputs and 4 outputs, so the cut is passes, `wait_task`
after each (its outputs are the next pass's inputs): per beat — lipsynced beats bring only their clip — mix its step-3 speech over its clip (`amix`, speech on top) into `beatN.mp4` at exactly
the clip's `duration` (`tpad`/`apad` over-pad, `-t <duration>` sets the length; captions are
timed from those durations), up to four beats per call (one output
each) and fewer when their files would pass ten inputs
(a clip plus its speech is two inputs and four such beats fill a call); then concatenate all the beats in
order in groups of at most 10; then one last call joins the groups and burns captions
from an `.ass` sidecar in a casual style. Output `final.mp4`, review the finished cut (`review-shots` step 5), and return the link
and `open_in_frameo`.

## Done

The UGC clip, the script, and every still and clip in the Frameo project, with the review note
per shot: passed, retaken and why, or kept with a known flaw.
`open_in_frameo` on the result opens the project in the Frameo app, where every take sits on
the canvas, ready for retakes and edits.

## Files

None.

Referenced files: 2

ugc-unboxing-video9.26 KB

View saved version →

---
name: ugc-unboxing-video
description: "Unboxing / first-reaction / haul: the creator opens the package, reacts, shows what is inside."
---

Needs: product and packaging photos; a creator photo or a description
Credits: about 500–650 for 30 seconds, plus a retake allowance of 15% of the image and clip total
Time: 15–25 minutes, plus about a minute per clip for the review

## When

The package matters: first impressions, hauls, subscription boxes. Not for a plain
review (`ugc-review-video`).

## Ask first

1. **The product** — photos (`show_upload` for files on the user's device, `import_media_url` for a web link; images attached to the chat do not reach Frameo, and `create_upload_url` is for clients that send the file themselves), the name, and the one thing the video
   must say about it.
2. **The creator** — design one (age, vibe, setting), use the user's photo, or reuse a creator
   already in the project (`list_characters`: its first image is the step-1 still and its
   `voice_id` the voice in step 3, with nothing to generate or choose).
3. **Frame and length** — 9:16, 15 / 30 / 45 s.

Say the rough cost in the same turn as the questions; the user's yes comes with the plan (Steps, **The plan first.**).

## Steps

**Before any paid step.** The project: `list_projects` (or `create_project`) gives the `project_id`
and, when the project has several modules, the `module_id`; every call below that takes a project —
`estimate_cost` included — gets that same pair; without it those tools answer `project_needed`. The
quote: one `estimate_cost(items=[…])` prices a stage in one call, with the same project, model,
size and number of `image_urls` or `reference_image_urls` (`reference_count`) as each generate call, and returns a `quote_id` per item plus the total. Each generate call
then passes its own item's `quote_id` and `confirmed_by_user=true`. A quote is single-use and lasts
15 minutes, so a long plan is priced stage by stage, right before each stage runs; a stage that
comes to more than the user approved is asked about again first. Generate calls return
`generation_ids`; `wait_task` returns the links, and `show_generations` shows each stage's running
and finished work in one card where the chat app displays Frameo cards: all the stage's ids at
once, before its first `wait_task`, and the finished result with `final: true`.

**The plan first.** Before the first paid call, the plan goes to the user in the chat as plain
text: what will be made, in order, one line per generation (for a script, the shot list; for a
set, each shot), with each line's credits and the total. The credits come from
`estimate_cost(items=[…])`, up to 10 items a call, so a long plan takes several calls; those
quotes may expire unused, since each stage is quoted again right before it runs. Nothing is
generated until the user says yes. The user can drop or change lines; a changed line is priced
again.

**Canvas rows.** Pass `shot_number` on every `generate_image` and `generate_video` of a shot (1, 2, 3… in story order; the same number for a retake and for that shot's video), so each shot gets its own row on the Frameo canvas. Cast, prop and location references take `placement_kind` (`character`, `prop` or `location`) and the subject's name as `placement_group` instead: they sit on their own board, and the shots built from them do not pile into their row.

**Review (on by default).** Every still and clip is checked before it is built on, and the
finished cut before it is handed over, following `review-shots`: its checklist and steps come
from `get_skill("review-shots")`, loaded before step 1. The checks run on the step-1 still
before step 2 builds on it, after step 2 (`review-shots` steps 0, 1 and 4), after step 4 (steps 2 and 4) and after the cut (step 5); the
product's label, shape and colour are checked against the product photos in every shot, and a
failed shot is retaken once. The plan carries the retake allowance as its own line for the user
to approve with everything else. The paid check of clips with speech (`review-shots` step 3)
runs only when the user asks for it, since the talking beats are lipsynced to one fixed voice,
and its rough cost (a few credits per minute of clip) is said and agreed first; without it, the
note on each talking clip says its sound was not checked.
The user can turn the review off.

**0. Script.** Write a script of the length chosen in Ask first in the creator's voice: hook in the first line, the one
claim, a close. Show it; the user edits before anything is generated.

**1. Creator still (~18).** `generate_image`: the creator at a table with the sealed box in frame, phone-video framing,
`image_urls=[creator photo]` if given. 

**2. Product-in-scene stills (~18 each).** `generate_image` for each beat that shows the product
(box closed, box opening, contents laid out, one item held up), with `image_urls=[step-1 creator still, product photos…]` when the creator or their hands are in
frame (the same person every time) and `image_urls=[product photos…]` when only the product
is; the prompt names nothing the beat does not show.

**3. Voice (~1 per 50 characters).** First the beat list, no tool call: estimate each beat
at 2.8 words per second (2.5 for Hindi) plus one second of headroom, say it is
approximate, and split any beat that is too long — the opening beats at the video model's max
(`list_models(kind="video")`; 30 s on the default) because the mix at the cut cannot stretch
the picture, the talking beats at the smaller of that max and 15 s, the lipsync's cap —
until every piece is within it, so every clip's `duration` is its length on the timeline. Then
`search_voices` for the creator's age and language, 3 candidates, then `generate_speech` per final beat (each
paid call, here and below, has its own quote); `wait_task` for the links. The creator the user keeps is saved with `save_character` (a name,
`image_urls=[creator still]` and the chosen `voice_id`), so the next video in this project
reuses the same face and voice. Each speech
result's `duration_seconds` is its measured spoken length; from here on the estimate means that
length when the result has one, and the words-per-second figure only when it has none. That estimate is the clip's `duration` in step 4 and the `duration` / `audio_duration`
of its lipsync.

**4. Clips (~65 per 5 s + ~20 per second of lipsync).** Reaction beats: `generate_video` from the creator still before the box is opened, and after it
from the step-2 still of the opened product with the creator in frame, so the box stays open, then `generate_lipsync` with the step-3 speech for that beat. Opening beats: `generate_video` from the box stills with 'hands opening the box' in the prompt, silent. Every `generate_video` takes `duration` = that beat's speech estimate rounded up to whole
seconds and kept inside the model's range (`list_models(kind="video")` gives it; 4–30 s on
the default), 5 s for a beat with no speech; every `generate_lipsync` takes that same whole
`duration` (never above 15 s — the lipsync's own cap, whatever the video model allows),
`audio_duration` = the estimate, and the clip's `aspect_ratio` and `resolution`. `wait_task` each. Test the first beat before
the rest.

On Seedance models (the default) the product photos go along as `reference_image_urls` on every
clip that shows the product, and the creator still on every product clip the creator is in,
which holds the label and the face through the motion. The frame is the first reference image,
so the prompt names the others from the second image on, and it counts toward the model's cap;
the quote carries `first_frame=true` and `reference_count` = the number of
`reference_image_urls`, not counting the frame. Each
later piece of a split beat opens on the previous piece's last frame (`extract_frames` on the
clip that goes into the cut, the lipsync result on a talking beat) once that piece's look has
been checked, instead of the beat's still, so the join reads as one take. Each prompt names where
the clip ends and nothing the beat does not show: on a lipsynced beat the creator facing the
lens; on a silent beat the product in frame with the creator's eyes on it, never on the lens,
since the voiceover plays over lips that do not move.

**5. Cut (free).** `run_ffmpeg` takes at most 10 inputs and 4 outputs, so the cut is passes, `wait_task`
after each (its outputs are the next pass's inputs): per beat — lipsynced beats bring only their clip — mix its step-3 speech over its clip (`amix`, speech on top) into `beatN.mp4` at exactly
the clip's `duration` (`tpad`/`apad` over-pad, `-t <duration>` sets the length; captions are
timed from those durations), up to four beats per call (one output
each) and fewer when their files would pass ten inputs
(a clip plus its speech is two inputs and four such beats fill a call); then concatenate all the beats in
order in groups of at most 10; then one last call joins the groups and burns captions
from an `.ass` sidecar in a casual style. Output `final.mp4`, review the finished cut (`review-shots` step 5), and return the link
and `open_in_frameo`.

## Done

The UGC clip, the script, and every still and clip in the Frameo project, with the review note
per shot: passed, retaken and why, or kept with a known flaw.
`open_in_frameo` on the result opens the project in the Frameo app, where every take sits on
the canvas, ready for retakes and edits.

## Files

None.

Referenced files: 2

Package details

Publisher declarations from the archived package. These are separate from our research and the live service's terms.

Package author
Dashverse
Keywords
See publisher keywords

Declared capabilities

  • Read
  • Write

Package observed Oct 10, 2026.

Technical details
First seen
Sep 30, 2026 · 22:02 UTC
Last seen
Oct 10, 2026 · 18:00 UTC
Latest observed change
Oct 8, 2026 · 18:02 UTC
Collection status
Collected

plugin_asdk_app_6ab639b28e2081919b704b4ada7e478f

Download plugin data (JSON)

Before you connect Frameo

How do I connect it?

Open the publisher's marketplace listing to check current availability and follow its connection instructions. This directory does not install plugins. Check the requested access and any account requirements before connecting.

Check marketplace availability ↗

Does it require paid access?

We have not established the pricing or subscription requirements for this plugin. An absent price does not mean free access.

Compare researched pricing and access models →

How can I evaluate it?

Check the declared skills and available files, then try a small task whose result you can verify. Our archived descriptions and instructions establish publisher claims, not tested runtime quality. Review sources and coverage limits.