← Plugin catalog
Productivity
AI Producer by OpusClip
OpusClip v1.2.8
Publisher description
From the marketplace listing
Create an editable talking-head video in chat. Upload footage, direct cuts, captions, reframing, and motion graphics, then open the AI Producer preview to refine or export.
Language: English · Automatically detected from descriptions.
Files & skills
File archives
Plugin package28 files · 122 KBBrowse files →
Skill instructions
aip33.7 KB
---
name: aip
description: "Create or update an editable AI Producer video. Use when someone provides footage or an existing project and asks AI Producer to cut, caption, add motion graphics or music, restyle, export, or adapt a named Motion Library template, in any language (剪视频 / 剪辑 / 加字幕 / 做动效 / 配乐 / 导出成片 / 素材 / 插入图片 / 插一段视频). Settle the rough cut, fine cut and finishes rounds for a new project when the request leaves them open. Deliver the editable project link by default, export only on explicit request, and stop without inspecting the preview or output video."
---
# AI Producer preview
Create one editable AI Producer preview from the user's footage and creative brief. Shape the edit around the user's intent, the footage, and its audience.
## Targeted changes to an existing project
When the request targets an existing AI Producer project, begin from its current saved edit and make only the requested change. Do not upload the source again, restart project preparation, or reopen the rough cut, fine cut and finishes rounds unless the user asks to revisit one of them.
A natural request to use or adapt a Motion Library template, together with its `Project` and `Template` context line, is an existing-project change. Read the [Motion Library handoff](references/handoff.md) once and follow it before the new-project flow below. The handoff also owns compatibility with legacy Motion Library references.
## What to settle before you cut
The rough cut, the fine cut and the finishes decide the edit. Settle all three before you start editing, then edit without coming back for more.
Preparing the project is not editing, and it waits for no answer. In the turn the footage arrives, start it before you ask anything: mint the upload target with `sign_upload`, PUT the bytes, then call `start_transcribe_project`. Only then read the saved default and put up the first round. Ingesting and transcribing the source is the long part of the run and no answer in the rounds changes what it does, so a card raised first buys nothing and charges the user the whole wait. The three rounds are asked while preparation runs, and because the project exists by the time an answer comes back, each one is written down where it is given. Two things hold preparation back, and only these. A conversation with no computer behind it, which Known issues below covers: nothing is signed there. And the user asking to talk the edit through before anything is spent: that is an instruction, so tell them what preparing costs, settle what they want to settle, and start preparation when they say to go ahead.
Preparation spends, so whatever you say when you start it carries its price: the message that raises the first round, or your first reply when their opening settles every round and no card goes up. The price is read from `get_pricing` and never quoted from memory, and it reaches the user in plain words, as what getting their video ready costs.
Ask only for what the user has not already told you. Read their opening message against the rounds below; a round whose answer follows from what they wrote is settled, and you do not put it to them. Ask the rounds that are still open, in the order they are listed, one at a time. Every round settled - from their words or from their answers - is the signal to start production. Nothing else gates the work: there is no mode to pick and no feature list to shop.
### Start from the saved default
At the start of every new project, before the first round, read `get_editing_preset`.
- It returns a preset: put those settings to the user in plain words and ask them to confirm or change anything, in cards built from the same components the rounds use, each picker carrying the stored value as its `selected`. A text component takes no prefill - a card renders its field empty whatever you send - so a stored `editing_preference` is quoted back in your own message and the empty field is where a change goes. A card holds at most six server components and a full preset is five, so one card carries the whole of it, in round order, and you ask for the confirmation once at the end rather than after each setting. Confirmed as it stands, you edit with it and ask nothing further, and you write the confirmed values down with `record_choices` under the same card ids an answer would have used, so the project's own record says the rounds are settled and no later turn asks them again. Changed, you edit with the changed values and write those down the same way.
- It returns nothing: run the rounds.
Read it fresh on every project rather than remembering it between them. The user can change these settings away from the conversation, and this is the only place that change shows.
The default is the user's to set, and they set it on the settings page on the web, never through you: the server has a `save_editing_preset` tool, but it is not yours to call, and nothing in a conversation asks the user to keep their answers.
### The rounds
Each round is one card: call `present_choices` with that round's components, end your turn, and wait. Preparation is already under way by the first card, so `record_choices` has a project to write to and each answer lands as it arrives. A card needs no project behind it, so in the exceptional case where a round is raised before one exists - preparation refused, or the user answering a round you have not reached - the answer is carried and written down as soon as the project is created. When the user answers - on the card or in their own words - write it down with `record_choices` under that round's card id, then raise the next round - unless that round names something of its own that comes first. The latest recording for a card wins, so a user who changes their mind is recorded again, not argued with. `get_project` carries `choices`, what is answered, and `pending_choices`, what is not, in round order; a turn that picks up a conversation already in progress reads those two rather than guessing. A card's title and note are sentences the user reads, so they follow Talk to the user and name no card, component or tool.
| Round | Card id | Components | What you write down |
| --- | --- | --- | --- |
| Rough cut | `roughcut` | `rough_cut_boldness` | the value alone |
| Fine cut | `branding` | `editing_preference_example`, `editing_preference` | `<component id>=<value>`, one per answer they gave |
| Finishes | `finishing` | `caption_style`, `bgm_enabled`, `sfx_enabled` | the caption style's value alone, and the two toggles as `bgm_enabled=on`, `sfx_enabled=off` |
**Rough cut.** `off` means keep the original cut: no filler word, retake, false start or pause is removed. `off` is not a boldness and never travels on to anything that takes one. The other three - `conservative`, `balanced`, `aggressive` - say how far the clean-up goes. You make this cut yourself by writing the speaker clips, so its scope is yours to keep: on their own, these levels remove only filler words, verbatim restarts, abandoned false starts, failed retakes, and dead-air pauses. A fluent, complete sentence is content at every level, even when it seems less important, repetitive, or off topic; dropping it needs editorial scope from the user, and a scripted read with no disfluencies is left whole under cleanup alone. Without a requested duration, highlights, short version, or named removal, the speaker clips cover the rest of the source in order. Cut to the settled value and to the user's requested scope and duration, preserving the speaker's intended meaning. For a rough-cut-only or targeted speech-cleanup request, skip visual research and default visual treatments unless they are also requested. When revising the rough cut of an existing project, start from its current saved state: preserve choices outside the requested rough-cut change, and update affected captions, visuals, and audio timing to follow the revised cut. When the request includes both rough cut and fine cut, settle the content cut before finalizing timed treatments, then complete both. Pause for rough-cut review only when the user asks for it.
For a requested duration, highlights, or short version, select valuable, fluent speech and settle an engaging opening, the core argument, and a complete ending before fitting the length. Choose a self-contained final thought from the original speech, such as a conclusion, summary, or recommendation, with the context it needs. Do not fill the duration sequentially and truncate: avoid ending mid-sentence, mid-list, on a conjunction, or on a promise of more explanation. Keep the full final word and a short natural pause from the source; that closing pause is not disposable dead air. An approximate duration allows modest variation for a complete ending. For an exact duration or strict maximum, adjust earlier selections or choose a shorter complete ending instead of cutting the final sentence short.
**Fine cut.** This round asks how the user wants the footage edited: one worked example to start from, and their own words beside it. Both answers reach no tool. They are recorded, and applying them is yours - they are the direction you design the cut under. `editing_preference_example` offers `more-broll-image`, `more-motion-graphics` and `custom`; `custom` is the user saying the examples do not fit, so it carries nothing on its own and the words in `editing_preference` are the answer. `custom` with nothing typed is therefore not an answer: ask for their words before you record it, or the round reads as settled while holding no direction at all. An example on its own is a complete answer, and so is nothing at all - a user who wants neither is not blocked, so record what they did give and design the cut yourself.
The reference image is not on this card, because a card collects one line of text and an image is a file. So the card's note says nothing about a picture, and the ask is a message of its own. Look back over what the user has already sent. When they handed a picture over, put it into the project with `sign_workspace_upload` and write `reference_image=<path>` into the same `record_choices` call as the two answers, and ask them for nothing. When they sent none, record the two answers on their own, and then, before the finishes round goes up, ask them for the picture in plain words - your own sentence, not a card and not a line inside the note: whether they want to send a reference image to set the styling, and that answering "skip", or anything else, carries on without one. End the turn there and wait for the reply. An image comes back: upload it with `sign_workspace_upload` and record `branding` a second time - the two answers you already wrote down and `reference_image=<path>` beside them, because a recording replaces the whole entry for a card and an answer left out of the second call is gone - then raise the finishes round. "skip" or any other reply: raise the finishes round. Ask once for the whole project, and never a second time - not after a skip, and not when a turn picking the conversation back up finds `reference_image` already recorded under `branding` or the fine cut already behind it in `choices`. A user with no image to give is not blocked by the round.
**Finishes.** `caption_style` offers `no-caption` ("No captions") and eight caption patterns. Write the chosen token unprefixed, including `no-caption`, while the two toggles carry their component ids: `on` on its own identifies no question. A pattern goes to the caption skill without picking again. An omitted `caption_style` stays Auto: let the caption skill choose the pattern, rather than treating it as `no-caption`. Recording replaces the whole entry for a card, so write every finishing answer you hold each time you record one, or the ones you leave out are gone.
`no-caption` is a card answer, not a caption pattern: skip `build_captions` and omit the `narrator-captions` host from `render-engine/index.html`. For an existing captioned project, remove only that host, preserve transcript and editor artifacts, then upload the changed root and `commit_workspace`. Recording the answer alone does not change the edit. A saved preset that carries `no-caption` is read back unchanged.
`bgm_enabled=on` is the instruction to call `add_music` afterwards, and `off` is the instruction not to. `off` declines a new music run; it does not strip a bed already mixed into the cut, so do not offer to take music away with it. `on` is an answer, not a payload: it is not a value `add_music` takes, and that call's own inputs are stated where it states them. Music spends, and `get_pricing` is what says how much - never quote an amount from memory.
`sfx_enabled=on` is the instruction to call `add_sound_effects` on this project, and `off` is the instruction not to, so nothing is spent on sound effects. Where each effect lands is yours to decide: read the finished cut's timeline and place a hit where a sound earns its place - a cut, a beat of emphasis, the instant a motion graphic enters - in milliseconds on that timeline, never the source's. So call it once the cut and its timed treatments are settled, and when a later edit moves the cut, place the hits again. Send every hit in one call, 1 to 12 of them, because a call replaces the project's previous sound-effect track rather than adding to it. `on` is an answer, not a payload, exactly as with music. Sound effects spend, so read `get_pricing` before the call - never quote an amount from memory.
Not every install has the cards. When `present_choices` is not among your tools, the rounds are still asked, in plain text: each still-open round as one question of its own, in the user's language, offering the same options by name and in the order the table above lists them, one round per turn. Each round's options are named where that round is explained above, with one exception: the eight caption styles are not, so read them from the [dynamic caption skill](../aip-dynamic-caption/SKILL.md) before you ask the finishes round - with no card, the server is not the one offering them, and a style you invent is not one the caption layer can be built with. Write nothing down - `record_choices` is absent with it, and so are the preset tools, so skip the saved-default read - and carry the settled answers in the conversation instead. Everything else stands: preparation starts first and on the same terms, the reference image is still asked for as its own question, and a round the user's opening already settles is still not put to them.
## Source supporting visuals
Once the rounds are settled, apply the default editorial taste and research supporting material in the initial plan. Identify transcript beats that need real products, interfaces, devices, places, or event imagery. Use suitable supplied assets first; respect requests to avoid external sources.
- Review every supplied B-roll across its full time range: use representative frames for an overview, then inspect candidate action boundaries closely. Match each retained speech beat against operations, results, project lists, templates, feature entries, and tools, not just the first obvious actions. Finish that comparison before deciding coverage; finding a few matches is not a stopping rule. Match what the footage actually shows: a feature entry can illustrate availability, but does not demonstrate the operation or its result.
- Choose B-roll boundaries by action completeness, existing edits, topic changes, and dead time, not a fixed seconds threshold. Keep short complete shots and coherent pre-edited clips continuous; divide longer videos where distinct topics or operations warrant it. Preserve the necessary setup, action, and readable result. One demonstration may cover several consecutive related sentences. Avoid word-by-word cuts, source-time jumps within one operation, repeated close-up/wide switches, and brief presenter flashes between related demonstrations.
- A-roll follows the settled speech cut; fit B-roll to it without recutting speech to accommodate an asset's length. Trim irrelevant head or tail, adjust coverage over related speech, choose another excerpt, or return to the presenter when a clip does not fit. Do not default to loops, freezes, or speed changes to fill time. Let useful matches and continuity decide coverage, without a fixed clip count or coverage quota. Keep a concise local mapping of speech beats, source in/out times, and cut placements.
- For missing material, search the Web and open official or primary source pages to find images or clips. An empty `find_broll` result requires this lookup before falling back to diagrams. Use one batch of up to three targeted queries, plus at most one such follow-up batch for unresolved material.
- Batch page reads and downloads. Before authoring, inspect a downsized image sheet or representative source frames for subject match, usable detail, and crop suitability. Replace unsuitable candidates within the research limit.
- Use accepted assets in relevant beats with framing or motion that explains the speech. Unless the user names a trigger or a count, use each image or single-shot clip once at its best-fitting beat. A long video with multiple relevant moments can supply several distinct excerpts across the edit; do not restrict the source file to one use or repeat the same excerpt by default. A stated trigger such as "whenever X is mentioned" applies at every matching beat. Record source URLs locally; apply the [on-screen text rules](#on-screen-text) to visible labels.
## Delivery boundary
Deliver the editable project link by default. Treat "finished video" or "final cut" as a completed preview, not an explicit MP4 request. Export only when the user explicitly requests an MP4, rendered video file, or export. For an authorized export, wait for completion and return its link without opening, playing, sampling, or analyzing the output.
Do not inspect the generated preview or MP4, including browser playback, screenshots, contact sheets, frame extraction, audio analysis, or delegated inspection. Do not start a post-delivery review, repair, or aesthetic iteration. This stopping condition applies whether the deliverable is a preview or an explicitly requested export. Input-media analysis, including retrieved source assets, and local static contract checks before submission remain allowed.
Use only this package's [AI Producer composition contract](../aip-composition/SKILL.md) and [workspace reference](references/workspace.md) for integration. Read the workspace reference once before source upload or export, the composition contract once before authoring, and the [framing skill](../aip-framing/SKILL.md) once before deciding the speaker's frame for each beat; load the linked PIP example only when needed, and read the [dynamic caption skill](../aip-dynamic-caption/SKILL.md) once before the caption layer when it is listed. Do not search global HyperFrames or media-use skills, backend source, or repository documentation to make this video. Resolve these links relative to the skill actually loaded, never a remembered versioned cache path. If that path is stale, use the host's installed-plugin listing once to find the enabled package and its version.
## Talk to the user
Treat every commentary, confirmation, progress update, warning, refusal, and final reply as user-facing product copy. Match the user's language. Lead with the user's current state: what is ready, what is still happening, and what they need to do next. Keep routine progress concise.
Call the product AI Producer in everything the user reads, and never name the environment it runs against, the endpoint, the server, this plugin, or any part inside it.
Do not narrate implementation policy or expose skill names, tool names, IDs, tasks, digests, graphs, checkpoints, commits, upload signing, provider or runtime checks, internal files, or status codes. Keep those details inside tool calls and local logs. Translate a necessary limitation into its effect on the user's project and the next available step.
Plain language must remain accurate. Say that a project is ready only after the corresponding operation succeeds. For a refusal or failure, explain what could not be completed, how it affects the project, and what can happen next. Do not repeat a non-actionable internal warning unless it materially changes what the user should expect.
## Known issues
**The host reports that the MCP server requires OAuth reauthentication, or this plugin's tools are absent from the tool list with nothing else offered to explain it.** This is the client validating the credential it stores locally, not a service fault. No request leaves the client, so no service error exists to read and nothing can be retried or repaired from inside the conversation. Signing in clears it, but the session that hit it does not reliably reload the tools, so do not wait for them and do not deliberate over whether to. When the host reports a condition of its own, such as the service being unreachable or a network failure, that is a different problem and signing in does not address it: report what it means for the user's project instead of running this flow.
Both steps belong to the user, in this order.
1. Ask the user to sign in once to the `ai-producer` server in their own client. In Claude Code this is `/mcp`, then `plugin:aip:ai-producer`, which is how that host namespaces the same connection. In Codex, on the desktop app as much as in the CLI, it is `codex mcp login ai-producer` run by the user in their own terminal: the desktop app carries no sign-in control of its own for a plugin's connection, and the commands you run reach no network, so this is not something you can do for them. Name the entry for the host they are on, because the user selects it by name, and keep everything around it in product language. Do not assemble an authorization link, do not advise reinstalling the plugin, and do not modify MCP configuration. None of these restores the connection, and each can cost the user a working install.
2. Then ask the user to start a new chat and resume the work there. The hosts differ: in Codex the user @-mentions this plugin, and in Claude Code the user simply makes the request again, which loads it. State exactly what carries over: an existing project keeps its footage and its edit behind its project link, while anything the failed connection prevented from reaching AI Producer has to be sent again.
A plugin update can require signing in once more. That is expected, and it happens once.
**The conversation has no computer behind it: nothing you can run a command on, nowhere to write a file, and a video the user attached is something the host holds rather than a file you have.** In ChatGPT that is a Chat conversation. This plugin runs in Work and in Codex, each of which gives the conversation a computer of its own; Chat does not, and nothing on the AI Producer side changes that. The upload the service signs waits for bytes only a computer can send, and the edit is authored as files only a computer can write, so starting anyway ends in an error that reads as the plugin's or the video's when the cause is the kind of conversation. So do not start: no upload, no project, and no other route for the footage. Tell the user that their video is fine and so is AI Producer, and that editing it needs a Work chat or a Codex task: in ChatGPT on the web and in the desktop app, Work is the tab above the message box, and if it is not offered where they are, those two places have it. Ask them to start that conversation and mention AI Producer in it. What carries over is what already reached AI Producer: a project that exists keeps its footage and its edit behind its project link, so a follow-up on one takes the link there and not the footage again, while a video that only ever reached this conversation has to be attached again in the new one. Codex and Claude Code always have a computer, so this never applies to them.
## Show progress in Codex
In Codex, as soon as an AI Producer creation/preparation tool returns `project_id`, call `get_view_url`. In the same code-mode call, parse its single-use `url` and durable `agent_page_url` with `URL`. Require the same origin and the expected `/p/<project_id>` path. On the project URL, replace its query with `mcp=1&embed=1` followed by the original parameters other than `mode`, `mcp`, and `embed`; preserve `host` and other parameters. Set only the exchange URL's `next` parameter to that project path plus query and fragment; preserve its ticket and all other parameters. Pass the resulting exchange URL directly to `open_in_codex` with `target: {type: "browser", url: <exchange url>}` and `placement: "right"`, omitting `threadId`. Open during preparation; use the transformed `agent_page_url` for any later reopening in that browser. The exchange URL is single-use: do not fetch it separately or include it in messages or files.
For every user-facing project link, including the link given when a run starts, the final reply, and follow-up handoffs, parse `page_url` with `URL` and replace its query with `mcp=1&embed=0` followed by the original parameters other than `mode`, `mcp`, and `embed`; preserve `host` and other parameters. Never give the user the exchange URL or `agent_page_url`.
Opening is a status-only handoff to the user, not inspection. Use no external browser or computer-use action that returns page content. Keep the tab for live updates without reopening, refreshing, polling, playing, or inspecting it.
`queued` means pending. Retry an explicit failure at most once with the same tool, obtaining a fresh exchange URL only if the first may have been consumed. If unavailable or still failing, report it and continue editing; never claim the page opened.
## Default editorial taste
A settled round is the user's own instruction, so it outranks every default below, and a default never fills a round that is still open. A footage-only or vague opening is not licence to skip the rounds: settle them first, then plan from these defaults for everything the rounds and the user left unsaid.
Plan each new talking-head edit from its supplied media and publishing brief. The user's request, in the brief or in a later edit, is the highest instruction: it overrides any default below, and nothing listed or unlisted among your skills blocks it. A listed skill sets a default and the method for it; an unlisted one only drops that default, so a treatment the user asks for is still delivered with the tools the service exposes. Keep the rest of the defaults. Preserve established choices in follow-up edits unless the user requests changes.
### Editing
- Open with a sharp punch-in, then hold the framing; when sound effects are on for this project, a whoosh placed with `add_sound_effects` lands on the punch-in. Place massive, extra-bold text behind the presenter, spanning 80-90% of the frame width, partially occluded by the head but still readable. Reveal words individually; split or wrap instead of shrinking. Use high-contrast text, preserve the original background, and animate only the text after the punch-in.
- Use presenter close-ups and zooms for emphasis.
- For concrete subjects, prefer relevant real footage, then images, then vectors.
- Show the specific change described by the speech. Use generic geometric shapes, expanding cards, or checkmarks only when they clarify that change; prefer a direct demonstration when suitable footage or assets are available.
- Captions are on by default while aip-dynamic-caption is listed, and they stay on during visual moments; the framing skill places the band for each layout. Apply the finishing round's choice.
- For sarcasm, self-deprecation, or brief reminders, use a monochrome or desaturated presenter close-up for at most one sentence, then restore color.
- Use smooth effects and layout transitions; minimize hard cuts.
- No background music.
### Visual style
- Use a pure black canvas with near-black foreground surfaces, fills only, no strokes; preserve source-media colors unless a specific treatment is intended.
- For abstract relationships, mechanisms, and processes, prefer 3D animation, then SVG or vector.
- Every visual must animate purposefully from entry to exit; each motion reveals a step, relationship, or change; transitions and camera moves don't count.
- No titles, headings, or decorative copy. If a beat has no meaningful visual without its copy, show the presenter, with captions only when enabled.
- Center the focal point within each visual area. Keep all graphics inset at least 10% from the edges throughout the animation; avoid dense repetition and unnecessary layers so graphics stay large.
### On-screen text
These rules govern extra authored copy; spoken captions remain a separate layer governed by the caption skill.
- Authored vectors, diagrams, and animations default to no secondary text. Keep core values and units attached to the quantities they depict. Add an object, category, or relationship label only when its absence creates a specific ambiguity that the visual and simultaneous speech do not resolve. Omit labels that merely name a recognizable object or restate what the visual and speech already convey.
- Sourced B-roll or images may carry one short source or example label, only when the material could be mistaken for the speaker's own footage or demonstration. Authored graphics carry no provenance labels, "illustration" disclaimers, or explanatory footnotes.
- Before publishing each effect, apply a deletion test to its local composition source: if removing an extra label preserves the meaning in the context of the speech, remove it. This is a source-text check before submission, not inspection of the generated preview.
## Make, submit, stop
Inspect the supplied media and plan once. When the user ties a visual to an on-screen cue such as a gesture, a prop, or an action, locate that moment from representative source frames as well as the transcript; this is input-media analysis, not output inspection. Batch independent metadata reads and material downloads in one script. Fetch a transcription only after its task completes; do not repeatedly fetch an unfinished transcript. When the available tool schema supports `wait_seconds`, use it with the supported limit; otherwise obey the returned polling interval. Waiting is not a reason to load more skills or inspect backend code.
Plan the whole edit locally. When the current schema exposes `authoring`, set it to `true` on a new project's creation/preparation request or an existing project's first `sign_workspace_upload` request. Set it to `true` on any intermediate commit and `false` on the final commit, so the service can retain and then close external-authoring status. This uses existing calls only; never add a provider call, model subtask, heartbeat, or polling round for status. Omit the field when the schema lacks it.
For a fresh project in Codex, use the [progressive checkpoint flow](references/progressive-publication.md). Plan the complete edit once. Publish the first effect as soon as it is ready. Then author the remaining effects in batches of up to two while the preceding batch publishes. Keep drafts isolated and let the helper publish each effect in order using the AI Producer MCP tools already loaded in the current task. Preserve the planned effects, timing, and transitions; batching controls scheduling only. Keep unfinished beats full-length with normal presenter framing, and publish without intentional delays. Do not launch another Codex CLI, app server, task, or model subtask. The last commit closes authoring status: the final batch of a caption-free edit, otherwise the caption commit that follows the last batch.
Run the batch helper in the execution that writes its drafts. It handles signing, uploading, committing, held waiting, and receipt acceptance without separate model coordination or receipt-collection turns. The final receipt includes earlier warnings. If the host cannot keep an execution alive while authoring continues, await each batch without claiming overlap. Other hosts, existing effect graphs, and zero-effect edits use ordinary complete-graph publication. Measure actual tokens and duration; overlapping work is not a cost or quality guarantee.
The progressive helper runs the local contract and timeline checks, freezes the hashes the current task supplies to `commit_workspace`, and advances only from that commit's accepted receipt. Use the plan it returns once; do not recompute its digest or expected files in the host. For ordinary publication, run [preflight.py](scripts/preflight.py) as described in the workspace reference, or skip it and let admission report contract errors if no Python 3.9+ interpreter is available; batch signatures and uploads with [upload_batch.py](scripts/upload_batch.py) or the host's batch HTTP tools, then commit with the current `base_digest`. Every commit remains all-or-nothing. A signing, upload, commit, task, or receipt failure leaves the current checkpoint closed without automatically retrying a mutation or adopting another writer's digest. Report it with the last successful effect count; do not restart the full sequence. Tool descriptions own authentication, limits, and task state.
After the final requested graph is accepted, deliver the durable project link (`page_url`, retrieved if missing and transformed for the user as above) in the final reply, then stop unless an export was explicitly requested. Give the user that link in the same reply that says the run has started, not only at the hand-back: the project exists and is watchable from the moment it is created, and a run is long. A view-ticket url is never that link - it is single-use and signs whoever opens it into your own account, so it does not go to the user. Translate refusals and material warnings under the user-facing language contract; do not repeat internal codes or mechanisms. Do not claim visual quality, audio synchronization, or export compatibility. Close by handing final confirmation to the user: state what was applied, then ask them to confirm the result on the project page, already open beside the conversation in Codex, otherwise at the link in this reply. Word this as a completed delivery waiting on the user's review.
Referenced files: 17
aip-composition11.8 KB
---
name: aip-composition
description: "AI Producer-specific composition and editor contract. Read when using the AI Producer plugin to author its HTML workspace; not the general HyperFrames creation workflow."
---
# HyperFrames for an AI Producer project
<!-- SPDX-License-Identifier: Apache-2.0 -->
Copyright 2026 HeyGen, Inc. Modifications Copyright 2026 OpusClip. Modified by OpusClip for AIP; see [source attribution, license, and changes](../../THIRD_PARTY_NOTICES.md).
HTML is the video: a composition is an HTML element with `data-*` timing attributes and one paused GSAP timeline the player drives. AI Producer's editor plays your HTML with its own player, and the export renders it with AI Producer's own pinned HyperFrames; they are two programs with separate playback and editing contracts. This page is the interface those two read, nothing more. What to draw, how it moves, where captions sit and what a moment looks like are your decisions.
## What AI Producer accepts
- The project lives under `render-engine/`: `index.html`, `compositions/*.html`, and assets under `public/`. AI Producer stages `public/source.mp4` (the recording) and `public/source.mp3` (its audio) and puts GSAP at `public/vendor/gsap.min.js`; load GSAP from that path.
- The accepted files and upload limits are described in the [AI Producer workspace format](../aip/references/workspace.md). Scripts from outside the project are refused, so a composition's logic is inline.
- A user's image is staged at `public/images/<name>` and a user's clip at `public/videos/<stem>.mp4`, `<name>` being the filename reduced to ASCII letters, digits, `-`, `_` and `.` (a space becomes `-`, anything else is dropped), because other characters change the URL the browser asks for. The hand-back check confirms every file a composition references through `src` or `href` is in the tree; a file referenced any other way (`data-pip-src`, a CSS `url(...)`) is not seen by it, so stage it in the same upload batch as the composition and confirm it in the listing after the commit.
## What the editor recognises in index.html
- The root: `<div id="stage" data-composition-id="finecut-root" data-start="0" data-duration="<total s>" data-width="W" data-height="H">`. Place this root directly in the document body, not inside a `<template>`. The id is fixed: the editor seeks the root timeline as `window.__timelines["finecut-root"]`, so a root zoom or transition lives on that timeline, and the root creates the registry before any composition registers:
```html
<script>
window.__timelines = window.__timelines || {};
const rootTl = gsap.timeline({ paused: true });
...root tweens, e.g. on #speaker...
window.__timelines["finecut-root"] = rootTl;
</script>
```
- The speaker: `<video id="speaker" class="clip" src="public/source.mp4" muted playsinline data-volume="0">` and `<audio id="speaker-audio" class="clip" src="public/source.mp3" data-volume="1">`. Keep unique media IDs. Speaker video is silent; the source audio is the audible leader. Generic extra `<audio class="clip">` tags are not AI Producer sound-effect tracks: AI Producer audio tracks require its editing document and derived `track-audio` elements. After a cut, use N video/audio pairs over the same two files. Each element has `class="clip speaker-clip"`, a unique `id`, and `data-start`, `data-duration`, `data-media-start`, and `data-track-index`. A pair shares its `data-hf-id` and timing; give each pair a distinct `data-hf-id`. Keep `speaker` and `speaker-audio` on the first pair. DOM order is output order on each speaker track, so write the spans in playback order with consecutive output starts. The EditingScript AV track must carry the same clip IDs and source/output times; the [workspace helper](../aip/references/workspace.md) synchronizes it and caption timing before publication.
- A visual moment: one `<div class="visual-host clip" data-composition-id="<id>" data-composition-src="compositions/<file>.html" data-start data-duration data-track-index data-width data-height>` per moment. The host div is the moment the editor shows and lets the user move; content written straight into the root is not a moment. The host may carry the caption policy for its window: `data-hide-captions="false"` keeps the running caption on under the moment (absent or `"true"` hides it, the default), `data-caption-position="<integer percent of frame height>"` moves the caption band's anchor for that window, and `data-caption-ink="#<six hex digits>"` gives the caption type that colour under the moment, for a ground this composition painted light. The editor, the service's edit paths, and the export re-derive the caption layer's windows from these attributes, so a moment moved or deleted carries its policy with it; the [dynamic caption skill](../aip-dynamic-caption/SKILL.md) passes the same values to the caption build.
- Captions: a host div the same way, `data-composition-id="narrator-captions"` pointing at `compositions/narrator_captions.html`. That id is how the editor and the export find the caption layer, whoever wrote it: the editor reads the `.caption-band` element inside it, and both read the words from one inline `const WORDS = [...];` JSON array, one object per word with `word`, `start` and `end` (seconds on the cut) and an integer `group_id` shared by the words of one phrase; a file declaring that id without the array is refused at commit. The service's caption build writes that file and `compositions/transcript-src.json` and returns the host div; the [dynamic caption skill](../aip-dynamic-caption/SKILL.md) owns the style pick when it is listed. Mount that host div as returned. Its `.caption-band` is a full-canvas container, and the phrase inside it sits at the position the build reports, already clamped to the pattern's safe band; an offset on the band or the host from `index.html` (`top`, `inset`, `transform`, or any override) moves the whole canvas and puts every word off screen. To move the captions, call the build again with a different `position`. An edit without captions needs none of these files.
- A speaker seat or PIP that belongs to an editable visual moment lives inside that moment: `<div class="speaker-pip-frame" data-aip-editable="speaker-pip-frame"><video id="<unique>" data-pip-src="public/source.mp4" data-start="<the host's data-start>" data-duration="<the host's data-duration>" data-media-start="<source seconds at that start>" muted playsinline data-volume="0"></video></div>`, with no `src`. Paint an opaque moment ground over the root speaker, and animate the frame inside the child timeline when it should travel from full frame into its seat. The one exception to the timing table is that this video's `data-start` is the host's global start, not 0, because the export re-anchors it on a cut project by setting its media offset to that value; the editor maps it through the cut itself. Keeping the picture and payload under the same host makes a timeline move or delete apply to both. A speaker cutout the service staged through `frame_speaker` (`public/speaker_<slug>.webm`, alpha video cut to its window) is wired the same way with `data-media-start="0"`, over a `public/source.mp4` view of the same window at identical geometry; the [framing skill](../aip-framing/SKILL.md) owns when and where.
- Footage other than the recording (a user clip, a stock clip): default to a root video clip on its own track in `index.html` for both PIP and split, `<video id="<unique>" class="clip" data-track-index="<n>" src="public/videos/<stem>.mp4" data-start data-duration data-media-start muted playsinline data-volume="0">`, with the moment that frames it (a mask, a frame, a title) drawing over it. A nested alternative is the same tag inside the moment's composition, `<video id="<unique>" class="clip" src="public/videos/<stem>.mp4" data-start="<the host's data-start>" data-duration data-media-start muted playsinline data-volume="0">`, timed on the host's clock like a PIP. The root clip is what the user can move on the timeline; the nested one moves with its moment.
- An element the user should move and restyle as one object carries `data-aip-editable="<token>"` where the token is one of that element's own classes; the editor also treats `.list-item`, `.glass-card`, `.media-card`, `.speaker-pip-frame` and `.item-bar` as such objects. Anything else inside a moment is reachable only as text or image leaves.
Do not implement a seat or PIP owned by a visual moment as absolute-time geometry tweens on the root speaker. Moving or deleting the host does not rewrite arbitrary root-timeline tweens, so that shape would remain at its old time. Root speaker geometry is only for a camera treatment that is not owned by an independently editable visual moment. The optional [PIP example](references/pip-transition.md) shows the effect-owned full-frame/PIP transition.
## A sub-composition file
```html
<template>
<div data-composition-id="<id>" data-width="W" data-height="H">
...content, with its <style> and <script> inside the template...
<script>
const root = document.querySelector('[data-composition-id="<id>"]');
const tl = gsap.timeline({ paused: true });
...tweens on root.querySelector(...) targets...
window.__timelines["<id>"] = tl;
</script>
</div>
</template>
```
The inner div's `data-composition-id` equals the host's. Bind the root exactly as the snippet does, with `[data-composition-id="<id>"]` and nothing else in that selector, and reach every element through descendant queries on that root. The export mount strips `data-composition-id` and the other `data-*` attributes from the inner div, so at export the host is the only element carrying the id: a selector that also requires a class, an id, an attribute, or `:not(.visual-host)` on the same element, or a `:scope >` child selector, matches nothing at export, the script throws before it registers, and the export refuses the moment. Scope styles and element queries to the composition's content so multiple moments can coexist in the same document. Keep the composition root visible at its base pose; animate an inner wrapper or its children, not the composition root or host. The editor mounts only what is inside `<template>`, and the export mounts the template itself: a `<script>` written after `</template>` is not part of the moment in either, so write nothing after `</template>`; the file needs no mount script of its own. Timeline time 0 is the host's `data-start`, and the player seeks the timeline rather than playing it. Create and register the paused timeline synchronously after its DOM exists. Render-critical changes belong on that timeline, not in timers, requestAnimationFrame loops, playback callbacks, or CSS animations. Do not call `tl.play()` or start media playback yourself; the player owns the clock. Keep animation repeats finite and randomness seeded so a seek reaches a deterministic state. The host attributes own the duration and active window; do not add empty tweens to set duration or manually nest the child timeline into the root timeline.
## Timing attributes
| attribute | on | value |
| --- | --- | --- |
| `data-start` | every clip | seconds, relative to the parent composition (a PIP or footage clip inside a composition: the host's global start, see above) |
| `data-duration` | every clip | seconds, the clip's own length |
| `data-track-index` | every clip | integer; clips on one track cannot overlap |
| `data-media-start` | video and audio | offset into the source file, seconds |
| `data-volume` | registered media | playback gain, 0 to 1; does not enroll arbitrary audio in AI Producer mixing |
| `data-width`, `data-height` | compositions | the canvas, px |
## Preview acceptance
The [AI Producer skill](../aip/SKILL.md) owns delivery and the no-inspection stopping condition. Keep this contract separate from creative direction. A successful local animation sample or accepted upload does not establish AI Producer editor or export correctness.
Referenced files: 2
aip-dynamic-caption11 KB
---
name: aip-dynamic-caption
description: "Pick one of the eight preset dynamic caption styles for an AI Producer project and have the service build the caption layer: how to choose the pattern from the video's content, which knobs the host controls, and the one call. Listed, it makes captions the default."
---
# Dynamic captions for an AI Producer project
The caption layer is one composition, `compositions/narrator_captions.html`, with `compositions/transcript-src.json` beside it. The service builds both from a subtitle pattern you pick and the project's transcript; you do not write caption HTML, CSS, or word data. Your job is the pick and its few knobs, made from what you know about the video.
## What you decide
1. **The pattern.** One per video, from the table below. Read the transcript at `get_transcript` with `detail: segments` (a tenth of the bytes) for register and pace, and look at the source for brightness, contrast, where the face sits, and whether a brand accent exists. Those are the inputs the table keys on.
2. **Emphasis words** (optional). At most one per phrase: the payload word a viewer should feel land. Take `word_id` values from `get_transcript` with `detail: words`. Every pattern has a designed treatment for an emphasis word; a phrase without one renders in the quiet register. Leave the list out and the service picks them with one caption-plan call.
3. **Position** (optional). Leave `position` out unless the brief places the captions; the service then uses the pattern's default. If you set it, it is the band's vertical anchor as a percent of frame height: 50 is mid-frame, 80 is the lowest the text-safe zone allows, and the service clamps anything past that, tighter for a pattern whose block grows downward. The canvas decides the default and the pattern: on a 1080x1920 canvas the band sits at 70 and every pattern is available; on a 1920x1080 or 1080x1080 canvas the speaker's torso reaches the frame's floor, so the band sits at 78 (the bottom band), the type is smaller (44 to 56px) and holds more words per phrase, and a pattern whose block grows downward (`blur-ladder`, `editorial-stack`, `lead-in-flare`, `tilt-slam`, `scribble-subtitle`, `inverted-stack`) is rendered as `pace-adaptive`; the task result's `pattern` is what was baked and `substituted_from` names what you asked for. Pick `pace-adaptive` or `solo-word-punch` outright for a landscape or square video. `placement` is `adaptive` (default: the band moves off the speaker where the project's person samples exist; the task result's `placement` says whether it did) or `fixed`. Horizontal composition is part of each pattern's design and is not a knob; change the pattern to change it. The task result's `position` is the anchor actually baked; when it differs from what you passed, the pattern's safe band clamped it. The fix is another call with a position inside the band, never a CSS offset on the host div or the `.caption-band` in `index.html`: the band is the full canvas, so an offset pushes every word off screen.
4. **Hide windows** (optional). `hide_intervals` is a list of `{start, end}` seconds where the band stays hidden. Captions stay on during visual moments by default; hide only for a `headline` overlay or a `statement` full cover (the spoken phrase drawn big), a `source-text` or `photo` full cover whose relevant detail must occupy the band and cannot move, or a stretch the brief wants silent. An empty list keeps captions on for the whole video and never hides anything by itself.
5. **Font** (optional). Leave `font` out unless the brief names a typeface; the captions then render in Geist with Fraunces for the emphasis word. If the brief names one, pick the closest pairing register and pass it as `font`: `standard` (Geist, Fraunces), `quiet-luxury` (Plus Jakarta Sans, Newsreader), `avant-garde` (Bricolage Grotesque, Fraunces), `heavy-hitter` (Schibsted Grotesk, Spectral), `pure-editorial` (Familjen Grotesk, Source Serif 4), `warm` (Hanken Grotesk, Crimson Pro). The sans is every caption word and the serif is the emphasis word; there is no free family name. The task result's `font` names the pairing baked.
6. **Caption windows** (optional). `caption_windows` is a list of `{start, end, position?, ink?}` on the cut, one per visual moment whose band differs from the whole-video caption. `position` is that window's anchor as a percent of frame height; the [framing skill](../aip-framing/SKILL.md) gives the value per layout and canvas (`Where the caption sits`). `ink` is a six-digit hex the caption type takes in that window, for a moment whose ground the composition painted light: a full cover on a pale canvas, or the canvas or card under the band while the speaker sits in a seat. Pick a colour with a large lightness gap against that ground (`#0a0c12` on a pale ground), never a mid grey; the type stays white everywhere else, and an overlay rides footage, which moves under the words, so it never takes an ink. Each window names at least one of the two, and where two windows overlap the earlier start wins. Write the same values you pass on that moment's host in `index.html` (`data-hide-captions="false"`, `data-caption-position`, `data-caption-ink`), because the editor and the export re-derive the caption layer from those attributes after any edit, and a moment host without `data-hide-captions="false"` hides the captions there at the first editor rebuild; the call bakes the first copy. The task result's `windows` has one row per window: `position` as you passed it, `placed` where the pattern's safe band lets it land, `band`, the pixel box of that landing (`anchor`, `top_px`, `bottom_px`), and `ink` as you passed it; when `placed` differs from `position`, keep a seat or a graphic clear of `band`, not of the number you asked for. The top-level `band` is the pixel box the whole-video anchor produces, the region a seat or a graphic keeps clear.
7. **Accent** (optional). Three patterns paint an accent colour: `lead-in-flare` (the emphasis word's type), `scribble-subtitle` (its marks), and `tilt-slam` (the slab under the loudest word). Pass `accent` as a six-digit hex when the brand has one; leave it out and those patterns render their white fallback. The service never picks a colour for a host project, so any colour in the caption is the `accent` you passed. The task result's `accent` is the hex painted, or null.
## The call
Call `build_captions` with `project_id`, `caption_style` (the pattern), and any of `position`, `placement`, `hide_intervals`, `caption_windows`, `emphasis_word_ids`, `font`, `accent`. It returns a task; read it with `get_task` (its `poll_args` say when). The finished task's `result` names the files the service staged (`files`, one `{path, sha256}` per file; `staged` lists the same paths), the pattern and position it baked, the emphasis it used as `word_id` values, and `host_div`, the one element for `index.html`:
```html
<div
class="visual-host clip"
data-composition-id="narrator-captions"
data-composition-src="compositions/narrator_captions.html"
data-start="0"
data-duration="<total s>"
data-track-index="<n>"
data-width="W"
data-height="H"
></div>
```
Place `host_div` as written (its duration and canvas come from the root you committed; placeholders mean commit the root first) inside the root, last, on its own track, then upload your tree and `commit_workspace`. When that commit names `expected_files`, add the result's `files` entries to the list exactly as returned; a batch that leaves them out keeps the caption files staged and the round reports them in `retained_staged`. A commit that names no batch promotes them with everything staged. The editor finds the running caption by the `.caption-band` element inside that host. A file you stage yourself at `compositions/narrator_captions.html` or `compositions/transcript-src.json` replaces the service's. A caption layer you write yourself follows the caption entry of the [composition contract](../aip-composition/SKILL.md). Hand back the editor link; there is no local inspection step. One `build_captions` run is one flat charge, whether or not you named the emphasis words.
## Patterns
| Choose when the video is | pattern | Rotate away when |
| --------------------------------------------------------------------------------------------------------------------------------------------- | ------------------- | -------------------------------------------------------------------------------------- |
| A tutorial or a punchy talking head where each sentence should feel hand-cut; face outside the 50-73% band | `blur-ladder` | captions should be a quiet legibility layer; the footage is busy |
| A piece where the running caption is the design itself; darkish source; face outside the 50-68% band | `editorial-stack` | captions should stay quiet under the speaker or under visual moments |
| A dark speaker cut against bright b-roll, where the caption should read as grade rather than overlay; real tonal contrast in the caption band | `inverted-stack` | flat mid-tone footage |
| **The default**: every sentence wants a visible payoff beat; a brand `accent` you can pass (the only pattern that colors caption type) | `lead-in-flare` | delivery so fast that a big tail every sentence lags the voice |
| Dense, fast, or information-heavy; captions should stay uniform under the speaker | `pace-adaptive` | a punchier per-word emphasis or a calmer whole-line read fits better |
| Creator confessional, story-time, a scrappy hand-made brand register | `scribble-subtitle` | a formal or corporate register; dense small on-screen text |
| Short imperative narration, step-throughs, hype delivery, where the caption is a beat | `solo-word-punch` | reflective or information-dense narration |
| Vlogs, hot takes, sticker energy, loud and playful; a brand `accent` you can pass fills the slab under the loudest word | `tilt-slam` | a formal or corporate register; dense small on-screen text; captions should stay quiet |
When captions share the frame with visual moments for most of the video, prefer a quiet pattern (`pace-adaptive`, `blur-ladder`): `tilt-slam` and `solo-word-punch` compete with an animation for the same beat, and `inverted-stack` cannot take an ink (its difference blend would invert the colour). One pattern per video. The pick is yours; the pattern's motion and layout are the service's, and its typeface is the pairing `font` names (Geist with Fraunces italic emphasis by default), staged with the composition under `public/fonts/`.
Referenced files: 1
aip-framing22.4 KB
--- name: aip-framing description: "Frame the speaker for an AI Producer project. Read once before authoring the root layout: pick the canvas first, then who owns the frame for each beat (the speaker, the footage, or the canvas), and for each of the three layouts (seat, overlay, and full cover) the forms it offers, where that canvas lets it sit, what it keeps, and how it is built, with the speaker measured, not guessed." --- # Framing for an AI Producer project Framing decides who owns the frame for each beat: the speaker, the footage, or the canvas. It is one decision per beat, made from the canvas, the transcript, and the footage, and it is yours. How anything looks (color, corners, shadows, type) is the visual style's; this page owns each layout's forms and the geometry it must keep. The [composition contract](../aip-composition/SKILL.md) owns what the editor and the export read. Apply the user's brief and the [AI Producer text rules](../aip/SKILL.md#on-screen-text) before choosing a form that contains text. The forms below describe geometry; their label, headline, and statement examples do not authorize extra copy. ## What you decide 1. **The canvas, first.** Read it from the brief's platform: 1080x1920 for Reels, TikTok, and Shorts; 1080x1080 for square feed posts; otherwise the source aspect, usually 1920x1080. The canvas decides which forms and positions each layout has (tables below); a layout drawn for one canvas is not reused on another. 2. **The register, per beat.** One of the four in the table below; its chapter owns the forms, the positions its canvas allows, what it keeps, and how it is built. 3. **The form, per moment.** Seat, overlay, and full cover each offer a short menu of forms in their chapters below. Name the register and the form in the plan. 4. **The measurement, before you place.** Call `frame_speaker` with every window you will seat, overlay, cut out, or reframe, in as few calls as the limit allows: each window's `start` and `end` on the cut, a `slug`, `matte` (true only for a cutout), and the `slot` the speaker will occupy (the seat's rect, a circle's side twice, an aperture's hole, or the canvas for an overlay, a full-frame reframe, and a cutout). Read the task with `get_task`. Each window in its `result.windows` carries `presence` (skip a seat or an overlay where `seat_ok` is false and a cutout where `matte_ok` is false), `face` and `head` as source fractions, and for the slot `object_position` (paste its `css` onto the seat's `data-pip-src` view, or onto the root clip for a full-frame reframe), `head_in_slot`, and `fits`. When `fits.height` is false the slot is wider than the source's aspect allows at that height, so `object-fit: cover` is scaling the source to the slot's width and the head is drawn at that width: narrow the slot at the same height and the head shrinks with it, or give it at least `fits.min_slot_height`. Measure the revised slot before placing it: if `min_slot_height` did not fall, the slot is height-bound and only deepening helps, and a slot narrowed until `fits.width` turns false has gone too far. Never deepen the crop instead. For a canvas slot, `head_in_slot` is the head's box on the canvas, the area an overlay keeps clear. A cutout window takes and returns more; the [cutout reference](references/cutout.md) owns those fields. The call is free and takes at most 8 windows, with matte windows totalling at most 60 s a call; when the plan has more, split the windows across calls, keep every `slug` unique across them, and read each call's task. 5. **The windows.** A register that shows the speaker is legal only over a window where the speaker is on camera. When the source cuts away to something the creator chose to show, let it play raw rather than covering it. 6. **The caption band, per moment.** Captions stay on during a visual moment: the moment gives way to the band, not the reverse. Decide where the band sits for each moment from its layout chapter below (`Where the caption sits`), declare it on the moment host (`data-hide-captions="false"` with `data-caption-position`, and `data-caption-ink` only where the ground under the band is one the composition painted light), and pass the same windows to the caption call, which the [dynamic caption skill](../aip-dynamic-caption/SKILL.md) owns. Hide the band under a moment only for a `headline` overlay or a `statement` full cover (the spoken phrase drawn big), a `source-text` or `photo` full cover whose relevant detail must occupy the band and cannot move, or a stretch the brief wants silent; those windows go to `hide_intervals`. One whole-video `position` cannot serve a PIP and a full cover at once, so the per-moment anchor is what lets one video hold both. ## Registers | register | what it is | | --- | --- | | `speaker` | the root speaker at full frame, optionally with a zoom that stays on the root timeline | | `seat` | the speaker moved into a declared shape at a declared position on an opaque ground, with the payload drawn in the band the seat leaves; forms and positions in its chapter | | `overlay` | the footage stays full-bleed and sharp, and one payload group at a time rides it in the area the measured head leaves clear | | `full-cover` | the canvas owns the frame: one opaque ground across the whole page carries the payload, and the speaker is covered but still heard | ## Video assets A video asset the speech puts on screen is seated in one of two layouts. Both are seats and keep everything a seat keeps. - PIP: the asset is the ground, and the speaker is a small inset placed clear of the asset's important subjects, action, and text. - Top/bottom split: the asset above, the speaker in a low-centered `card` below it. The asset may run the full width; the card keeps its side margins, because with portrait footage a full-width speaker panel draws the head at full source scale and crops a close-up. - The asset keeps its aspect ratio and its important content; panel proportions and inset placement adjust to it. - The caption band sits between the asset and the speaker panel, as under any portrait seat (`Where the caption sits under a seat`); the panel starts below the band. ## The three layouts Seat, overlay, and full cover are written the same way below: what it is, its forms, where each canvas lets it sit, what it keeps, and how it is built. A `speaker` beat needs none of this: it is the root speaker at full frame, and its zoom stays on the root timeline. ## Seat The speaker moves into a declared shape at a declared position on an opaque ground, and the payload is drawn in the band the seat leaves. Each seated framing declares its shape, position, and resting rect once. Beats that share that framing keep it, and a different moment may declare a different framing. What never happens is re-deriving a rect inside a run to fit a long line or a wider figure: a payload that does not fit the band needs a different framing, not a nudged seat. ### Seat forms Every seat is one of these shapes at one of the positions its canvas allows. Name the shape and the position in the plan; the numbers follow. | shape | what it is | | --- | --- | | `card` | a rounded window clear of the frame's edges, an object resting on the ground | | `stratum` | the speaker flush to one or more frame edges, a layer of the page rather than an object on it | | `circle` | an equal-sided window with `border-radius: 50%`, the head centered in it | | `aperture` | an opaque ground with a hole cut in it; the live speaker shows through the hole, which sits on the head, so its position comes from the measured face, never from the layout grid | | `cutout` | the pop: the room recedes into a full-width card and the speaker, cut free of it, stands proud of the card; the headroom above the head hosts the payload; the [cutout reference](references/cutout.md) owns its measurement and build | ### Seat positions by canvas The canvas decides where a seat may sit. A left or right column is a landscape layout; a portrait seat stays horizontally centered and clear of both side edges, with its payload above it, except for video-asset PIP and the cutout, whose full-width card is measured by `frame_speaker` and owned by its reference. | canvas | seats that fit | do not use | | --- | --- | --- | | portrait 9:16 (1080x1920) | `card` low-centered at about 78% of the width with the payload above, `card` mid-centered at the same width with payload above and below, `circle` low-centered, `aperture` centered on the head | a full-width `card` or `stratum` speaker window (the cutout's card is the one exception: its silhouette stands above the card, measured): with portrait footage a seat as wide as the canvas draws the head at full source scale, so a close-up head needs more than half the canvas or loses its crown and chin; a left or right column at any width: the band beside it is too narrow for a payload and the head reads small, so a portrait payload sits above or below the seat; corner circles below 30% of the width | | square 1:1 (1080x1080) | `card` low-centered or upper-centered, `stratum` as a bottom or top band, `circle` low-centered or in a lower corner at 30% to 36% of the width, `aperture` centered | side columns; stacked seats taller than half the frame | | landscape 16:9 (1920x1080) | `card` as a left or right column at 30% to 40% of the width with the payload beside it, `stratum` as a side column flush to three edges, `circle` in a lower corner at 22% to 28% of the height, `aperture` on the head, `card` centered with the payload split to both sides | low-centered cards with the payload above: the band is a thin strip; visuals-above/presenter-below stacks that crop the head to a band | ### Where the caption sits under a seat Look at the seat the speaker actually occupies, not the source frame: measure the head in the seat, then give the band its own region outside the seat and outside the payload. The whole-video caption never shrinks into the seat, and a seat that cannot leave the band its head intact takes another framing (narrow it, or change register). The pixel rows below are the 1080x1920 and 1920x1080 zoning the design settled on; `data-caption-position` takes the first value for a pattern that centres on its anchor and the second for one that grows downward (`lead-in-flare`, `blur-ladder`, `editorial-stack`). | canvas | the band | the payload | the seat | | --- | --- | --- | --- | | portrait 9:16 | between the payload and the seat, about y 920 to 1110: `data-caption-position` 53, or 50 | above the band, about y 265 to 770 | below the band, from about y 1240 (a `card` about 780x460) | | landscape 16:9 | the bottom band, about y 800 to 890: `data-caption-position` 78 (the service renders a single-row pattern on this canvas, so the band is one row) | beside the seat, ending above y 700 | a side column ending above y 700 | ### What a seat keeps - **One audible speaker.** The root track-0 video and its paired audio remain the canonical speaker. An editable seated moment uses one muted `data-pip-src` view inside its composition while an opaque ground hides the root picture; it never adds audio or a second independently playing source. This ownership view moves and disappears with the moment, and the editor maps it through the cut. - **The head stays whole.** Crown to jaw with room to breathe: a seat that cuts the forehead or shaves the chin has failed at its one job. When the head cannot fit, narrow the seat or deepen it; never choose which edge to sever. The view fills its seat at cover scale and no further: a scale on the view draws the head larger than the measurement saw. - **Center the head, measured, not guessed.** `object-position` is the `frame_speaker` result's `object_position` for that window and slot, pasted as returned onto the view that shows the speaker: it centers the measured face and holds the crown, jaw and cheeks inside the slot. `50% 50%` is a guess that crops a high-framed head at the crown; a narrow column crops harder, so every shape gets its own slot in the call, and on cut media every speaker clip its own window. - **The seat is an object on a ground.** For video-asset PIP, the asset fills the ground behind the speaker inset while its important content stays visible. Other seated beats paint an opaque ground across the frame and place the moment-owned speaker view in its seat; an aperture clips that view to the measured hole. In those layouts, nothing tucks under, overlaps into, or straddles the seat to buy room; content low in the band clears the seat's width as well as its top edge. - **A separate payload lives in the band the seat leaves.** The payload is composed inside that band, not against the whole frame. - **A seat arrives once per run of beats.** Adjacent beats that share a framing keep the seat where it is: it does not re-enter, re-settle, or re-announce itself; only the payload turns over. A new framing arrives at a beat boundary, either from the full frame or as a morph from the previous seat. ### How a seat is built - Every `seat` move owned by a visual moment lives inside that moment's paused child timeline. Put a muted `data-pip-src="public/source.mp4"` video in a `.speaker-pip-frame`, keep its global start and duration aligned with the host, and set its media start to the source time at the host's anchor. The composition contract's [editable speaker PIP](../aip-composition/references/pip-transition.md) is the single source for the snippet. - Animate the frame's box (`left`, `top`, `width`, `height`), crop (`object-position` with `object-fit: cover`) and corner radius together over the same interval. A circle uses equal sides and `borderRadius: "50%"`; an aperture clips the same moment-owned view to its measured hole. The opaque composition ground hides the unchanged root picture. - Never couple an independently editable host to absolute-time speaker geometry on `finecut-root`. The editor moves and deletes the host and its child content as one unit, but it does not discover or rewrite unrelated GSAP tweens at the old global time. ## Overlay The footage owns the frame and one payload rides it. The root speaker keeps playing full-bleed and sharp underneath, and the moment draws only what the payload needs: no page, no wash, no copy of the source. ### Overlay forms Every overlay is one of these forms. Name the form and where it sits in the plan. | form | what it is | | --- | --- | | `annotation` | drawn graphics, a short label, or an image placed in the area clear of the head | | `headline` | the phrase just spoken, drawn big as the beat's whole payload; its window goes to `hide_intervals`, because the captions would otherwise repeat the line | ### Overlay positions by canvas The measured head decides where an overlay may sit: the payload stays outside `head_in_slot` for that window, with room to breathe around it. | canvas | overlays that fit | do not use | | --- | --- | --- | | portrait 9:16 (1080x1920) | `headline` in the lower third, or above the head where the head box leaves real headroom; `annotation` in the band above or below the head at full width | anything beside the face: the flank is too narrow to read, so a portrait overlay sits above or below the head | | square 1:1 (1080x1080) | `headline` in the lower third; `annotation` beside the head on the flank the head box leaves open, or below the head | payloads on both flanks at once; a payload that crosses the head box | | landscape 16:9 (1920x1080) | `headline` in the lower third or on the open flank; `annotation` beside the head on the flank an off-center face leaves open | a band above the head: it is a thin strip; a payload that crosses the head box | ### Where the caption sits under an overlay The head, the payload, and the band avoid one another. The footage stays full-bleed; the payload takes the empty region outside the head box, and the band takes another band of its own, so the payload can animate without the caption following it. Only avoiding the face is not enough: the band must not cross the payload's key change, and the payload must not enter the band as it moves. When only one empty region exists, shrink or move the payload, or change to a seat or a full cover; the caption never follows the face word by word. | canvas | the band | the payload | | --- | --- | --- | | portrait 9:16 | the lower third, about y 1400 to 1600: `data-caption-position` 78, or 72 for a growing pattern | below the head, about y 860 to 1250 | | landscape 16:9 | the bottom band, about y 800 to 890: `data-caption-position` 78 (the service renders a single-row pattern on this canvas, so the band is one row) | the open flank, ending above y 620 | ### What an overlay keeps - **The overlay rides sharp footage.** The composition never pauses, scales, or reframes the source under an overlay. - **The head stays clear.** Name the overlay's window in the `frame_speaker` call with the canvas as its `slot`, and keep every text, plate, and image outside the returned `head_in_slot`; a thin connector line may reach past it toward what it points at. On a reframed canvas the box holds while the root clip carries the same result's `object_position`. The box describes the unzoomed frame: when a root zoom runs during the overlay's window, clear the box as it stands at the zoom's largest scale in that window, each edge pushed away from the zoom's origin by that scale, or end the zoom before the overlay starts. - **Payloads take turns.** An overlay holds one payload group at a time; the next arrives only as the previous yields. ### How an overlay is built - An overlay is an ordinary visual moment whose composition stays transparent: no `background` on the composition root or on any full-frame wrapper, the payload inside one positioned wrapper, and its entrance and exit on the moment's own paused child timeline. The composition contract's [editable speaker PIP](../aip-composition/references/pip-transition.md) example minus its ground and its speaker view is an overlay. - It embeds no video: no `data-pip-src` view and no copy of the source, because the root speaker underneath is the picture. Leave the root speaker's geometry to the root timeline. - Declare the window's caption policy on the host (`data-hide-captions="false"` and its `data-caption-position`) and pass it to the caption call as a caption window. A `headline` is the exception: pass its window as `hide_intervals`, because the captions would otherwise repeat the phrase the headline draws. ## Full cover The canvas owns the frame. The moment paints one opaque ground across the whole canvas and everything sits on it; nobody is seated, and the root speaker keeps playing underneath, unseen and still heard. ### Full-cover forms Every full cover is one of these forms. Name the form in the plan. | form | what it is | | --- | --- | | `figure` | a drawn diagram, chart, or mechanism at full scale | | `photo` | a real image staged on the ground as content; a video asset is seated as `Video assets` above says, not covered | | `statement` | one oversized line, a few words at the largest size the page allows | | `source-text` | a crop of dense source material, enlarged until the relevant detail reads | ### Full-cover positions by canvas The canvas decides how a page is laid out. | canvas | pages that fit | do not use | | --- | --- | --- | | portrait 9:16 (1080x1920) | one centered focal element, or a vertical stack of two: the figure above and its one-line label below | side-by-side columns: each is too narrow to read | | square 1:1 (1080x1080) | one centered focal element, or a stack of two rows | columns narrower than half the frame; more than two rows | | landscape 16:9 (1920x1080) | one centered focal element, or two columns side by side: a figure and its label, a before and an after | a vertical stack of three or more rows: each is a thin strip | ### Where the caption sits on a full cover Settle the band first, then compose the payload in the space it leaves: the animation recentres, scales, and moves inside that space, and the band must not cross an arrow's end, a compared result, a value, or a legend through the whole entry and exit. Check the frames where the speaker shows briefly at the switch. Only where the ground under the band is one this composition painted light (a pale panel, a light style the brief asked for) add `data-caption-ink` with a hex that contrasts with that ground (`#0a0c12` on a pale ground); the default black canvas keeps the white type, so the ink is a per-moment call, never a full-cover default, and never over footage or a photo, which move under the words. A `statement` hides the captions, and so does a `source-text` or `photo` whose relevant detail must occupy the band and cannot move up. | canvas | the band | the payload | | --- | --- | --- | | portrait 9:16 | about y 1390 to 1590: `data-caption-position` 78, or 72 for a growing pattern | recentred above the band, about y 360 to 1060 | | landscape 16:9 | the bottom band, about y 800 to 890: `data-caption-position` 78 (the service renders a single-row pattern on this canvas, so the band is one row) | above the band, about y 160 to 630 | ### What a full cover keeps - **One ground, wholly owned.** The composition paints one opaque ground across the whole canvas, and everything sits on it. An image is staged on the ground as content, never stretched into the ground itself. - **Nobody is seated.** No `data-pip-src` view, no window onto the speaker, and no copy of the source. The cutout is the one exception, and it is a seat form with its own [reference](references/cutout.md). - **The speaker underneath does not move.** The root speaker keeps its geometry for the whole cover; a root zoom is released before the cover starts. ### How a full cover is built - A full cover is an ordinary visual moment whose composition root paints the opaque ground (`position: absolute; inset: 0` with a solid `background`), with the payload's entrance and exit on the moment's own paused child timeline. The composition contract's [editable speaker PIP](../aip-composition/references/pip-transition.md) example minus its speaker view is a full cover. - It embeds no speaker view. The root speaker and its audio keep playing underneath at their ordinary geometry; release a root zoom before the full cover starts. - Declare the window's caption policy on the host and pass it to the caption call as a caption window; a `statement`, or a payload whose relevant detail must occupy the band, goes to `hide_intervals` instead. Reference every image through `src` so the hand-back check sees it. ## The speaker beat - Root zooms on a `speaker` beat that is not owned by a visual moment may remain root tweens on every active track-0 speaker clip (`#stage > video[data-track-index="0"]`). Release that camera treatment before the next seated moment; a root zoom never lands inside a seat.
Referenced files: 2
Package details
Publisher declarations from the archived package. These are separate from our research and the live service's terms.
- Package author
- OpusClip
Package observed Oct 2, 2026.
Technical details
- First seen
- Sep 30, 2026 · 22:02 UTC
- Last seen
- Oct 2, 2026 · 06:00 UTC
- Collection status
- Collected
plugin_asdk_app_6aab97bee30881919b757db1f5d91f99
Download plugin data (JSON)