← AI Producer by OpusClipCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to AI Producer by OpusClip
Snapshot Sep 30, 2026 · 22:59 UTC · version 1.2.8
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "aip-framing",
"description": "Frame the speaker for an AI Producer project. Read once before authoring the root layout: pick the canvas first, then who owns the frame for each beat (the speaker, the footage, or the canvas), and for each of the three layouts (seat, overlay, and full cover) the forms it offers, where that canvas lets it sit, what it keeps, and how it is built, with the speaker measured, not guessed.",
"included_files": [
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 167
},
{
"relative_path": "references/cutout.md",
"size_in_bytes": 5789
}
],
"skill_md_contents": "---\nname: aip-framing\ndescription: \"Frame the speaker for an AI Producer project. Read once before authoring the root layout: pick the canvas first, then who owns the frame for each beat (the speaker, the footage, or the canvas), and for each of the three layouts (seat, overlay, and full cover) the forms it offers, where that canvas lets it sit, what it keeps, and how it is built, with the speaker measured, not guessed.\"\n---\n\n# Framing for an AI Producer project\n\nFraming decides who owns the frame for each beat: the speaker, the footage, or the canvas. It is one decision per beat, made from the canvas, the transcript, and the footage, and it is yours. How anything looks (color, corners, shadows, type) is the visual style's; this page owns each layout's forms and the geometry it must keep. The [composition contract](../aip-composition/SKILL.md) owns what the editor and the export read.\n\nApply the user's brief and the [AI Producer text rules](../aip/SKILL.md#on-screen-text) before choosing a form that contains text. The forms below describe geometry; their label, headline, and statement examples do not authorize extra copy.\n\n## What you decide\n\n1. **The canvas, first.** Read it from the brief's platform: 1080x1920 for Reels, TikTok, and Shorts; 1080x1080 for square feed posts; otherwise the source aspect, usually 1920x1080. The canvas decides which forms and positions each layout has (tables below); a layout drawn for one canvas is not reused on another.\n2. **The register, per beat.** One of the four in the table below; its chapter owns the forms, the positions its canvas allows, what it keeps, and how it is built.\n3. **The form, per moment.** Seat, overlay, and full cover each offer a short menu of forms in their chapters below. Name the register and the form in the plan.\n4. **The measurement, before you place.** Call `frame_speaker` with every window you will seat, overlay, cut out, or reframe, in as few calls as the limit allows: each window's `start` and `end` on the cut, a `slug`, `matte` (true only for a cutout), and the `slot` the speaker will occupy (the seat's rect, a circle's side twice, an aperture's hole, or the canvas for an overlay, a full-frame reframe, and a cutout). Read the task with `get_task`. Each window in its `result.windows` carries `presence` (skip a seat or an overlay where `seat_ok` is false and a cutout where `matte_ok` is false), `face` and `head` as source fractions, and for the slot `object_position` (paste its `css` onto the seat's `data-pip-src` view, or onto the root clip for a full-frame reframe), `head_in_slot`, and `fits`. When `fits.height` is false the slot is wider than the source's aspect allows at that height, so `object-fit: cover` is scaling the source to the slot's width and the head is drawn at that width: narrow the slot at the same height and the head shrinks with it, or give it at least `fits.min_slot_height`. Measure the revised slot before placing it: if `min_slot_height` did not fall, the slot is height-bound and only deepening helps, and a slot narrowed until `fits.width` turns false has gone too far. Never deepen the crop instead. For a canvas slot, `head_in_slot` is the head's box on the canvas, the area an overlay keeps clear. A cutout window takes and returns more; the [cutout reference](references/cutout.md) owns those fields. The call is free and takes at most 8 windows, with matte windows totalling at most 60 s a call; when the plan has more, split the windows across calls, keep every `slug` unique across them, and read each call's task.\n5. **The windows.** A register that shows the speaker is legal only over a window where the speaker is on camera. When the source cuts away to something the creator chose to show, let it play raw rather than covering it.\n6. **The caption band, per moment.** Captions stay on during a visual moment: the moment gives way to the band, not the reverse. Decide where the band sits for each moment from its layout chapter below (`Where the caption sits`), declare it on the moment host (`data-hide-captions=\"false\"` with `data-caption-position`, and `data-caption-ink` only where the ground under the band is one the composition painted light), and pass the same windows to the caption call, which the [dynamic caption skill](../aip-dynamic-caption/SKILL.md) owns. Hide the band under a moment only for a `headline` overlay or a `statement` full cover (the spoken phrase drawn big), a `source-text` or `photo` full cover whose relevant detail must occupy the band and cannot move, or a stretch the brief wants silent; those windows go to `hide_intervals`. One whole-video `position` cannot serve a PIP and a full cover at once, so the per-moment anchor is what lets one video hold both.\n\n## Registers\n\n| register | what it is |\n| --- | --- |\n| `speaker` | the root speaker at full frame, optionally with a zoom that stays on the root timeline |\n| `seat` | the speaker moved into a declared shape at a declared position on an opaque ground, with the payload drawn in the band the seat leaves; forms and positions in its chapter |\n| `overlay` | the footage stays full-bleed and sharp, and one payload group at a time rides it in the area the measured head leaves clear |\n| `full-cover` | the canvas owns the frame: one opaque ground across the whole page carries the payload, and the speaker is covered but still heard |\n\n## Video assets\n\nA video asset the speech puts on screen is seated in one of two layouts. Both are seats and keep everything a seat keeps.\n\n- PIP: the asset is the ground, and the speaker is a small inset placed clear of the asset's important subjects, action, and text.\n- Top/bottom split: the asset above, the speaker in a low-centered `card` below it. The asset may run the full width; the card keeps its side margins, because with portrait footage a full-width speaker panel draws the head at full source scale and crops a close-up.\n- The asset keeps its aspect ratio and its important content; panel proportions and inset placement adjust to it.\n- The caption band sits between the asset and the speaker panel, as under any portrait seat (`Where the caption sits under a seat`); the panel starts below the band.\n\n## The three layouts\n\nSeat, overlay, and full cover are written the same way below: what it is, its forms, where each canvas lets it sit, what it keeps, and how it is built. A `speaker` beat needs none of this: it is the root speaker at full frame, and its zoom stays on the root timeline.\n\n## Seat\n\nThe speaker moves into a declared shape at a declared position on an opaque ground, and the payload is drawn in the band the seat leaves. Each seated framing declares its shape, position, and resting rect once. Beats that share that framing keep it, and a different moment may declare a different framing. What never happens is re-deriving a rect inside a run to fit a long line or a wider figure: a payload that does not fit the band needs a different framing, not a nudged seat.\n\n### Seat forms\n\nEvery seat is one of these shapes at one of the positions its canvas allows. Name the shape and the position in the plan; the numbers follow.\n\n| shape | what it is |\n| --- | --- |\n| `card` | a rounded window clear of the frame's edges, an object resting on the ground |\n| `stratum` | the speaker flush to one or more frame edges, a layer of the page rather than an object on it |\n| `circle` | an equal-sided window with `border-radius: 50%`, the head centered in it |\n| `aperture` | an opaque ground with a hole cut in it; the live speaker shows through the hole, which sits on the head, so its position comes from the measured face, never from the layout grid |\n| `cutout` | the pop: the room recedes into a full-width card and the speaker, cut free of it, stands proud of the card; the headroom above the head hosts the payload; the [cutout reference](references/cutout.md) owns its measurement and build |\n\n### Seat positions by canvas\n\nThe canvas decides where a seat may sit. A left or right column is a landscape layout; a portrait seat stays horizontally centered and clear of both side edges, with its payload above it, except for video-asset PIP and the cutout, whose full-width card is measured by `frame_speaker` and owned by its reference.\n\n| canvas | seats that fit | do not use |\n| --- | --- | --- |\n| portrait 9:16 (1080x1920) | `card` low-centered at about 78% of the width with the payload above, `card` mid-centered at the same width with payload above and below, `circle` low-centered, `aperture` centered on the head | a full-width `card` or `stratum` speaker window (the cutout's card is the one exception: its silhouette stands above the card, measured): with portrait footage a seat as wide as the canvas draws the head at full source scale, so a close-up head needs more than half the canvas or loses its crown and chin; a left or right column at any width: the band beside it is too narrow for a payload and the head reads small, so a portrait payload sits above or below the seat; corner circles below 30% of the width |\n| square 1:1 (1080x1080) | `card` low-centered or upper-centered, `stratum` as a bottom or top band, `circle` low-centered or in a lower corner at 30% to 36% of the width, `aperture` centered | side columns; stacked seats taller than half the frame |\n| landscape 16:9 (1920x1080) | `card` as a left or right column at 30% to 40% of the width with the payload beside it, `stratum` as a side column flush to three edges, `circle` in a lower corner at 22% to 28% of the height, `aperture` on the head, `card` centered with the payload split to both sides | low-centered cards with the payload above: the band is a thin strip; visuals-above/presenter-below stacks that crop the head to a band |\n\n### Where the caption sits under a seat\n\nLook at the seat the speaker actually occupies, not the source frame: measure the head in the seat, then give the band its own region outside the seat and outside the payload. The whole-video caption never shrinks into the seat, and a seat that cannot leave the band its head intact takes another framing (narrow it, or change register). The pixel rows below are the 1080x1920 and 1920x1080 zoning the design settled on; `data-caption-position` takes the first value for a pattern that centres on its anchor and the second for one that grows downward (`lead-in-flare`, `blur-ladder`, `editorial-stack`).\n\n| canvas | the band | the payload | the seat |\n| --- | --- | --- | --- |\n| portrait 9:16 | between the payload and the seat, about y 920 to 1110: `data-caption-position` 53, or 50 | above the band, about y 265 to 770 | below the band, from about y 1240 (a `card` about 780x460) |\n| landscape 16:9 | the bottom band, about y 800 to 890: `data-caption-position` 78 (the service renders a single-row pattern on this canvas, so the band is one row) | beside the seat, ending above y 700 | a side column ending above y 700 |\n\n### What a seat keeps\n\n- **One audible speaker.** The root track-0 video and its paired audio remain the canonical speaker. An editable seated moment uses one muted `data-pip-src` view inside its composition while an opaque ground hides the root picture; it never adds audio or a second independently playing source. This ownership view moves and disappears with the moment, and the editor maps it through the cut.\n- **The head stays whole.** Crown to jaw with room to breathe: a seat that cuts the forehead or shaves the chin has failed at its one job. When the head cannot fit, narrow the seat or deepen it; never choose which edge to sever. The view fills its seat at cover scale and no further: a scale on the view draws the head larger than the measurement saw.\n- **Center the head, measured, not guessed.** `object-position` is the `frame_speaker` result's `object_position` for that window and slot, pasted as returned onto the view that shows the speaker: it centers the measured face and holds the crown, jaw and cheeks inside the slot. `50% 50%` is a guess that crops a high-framed head at the crown; a narrow column crops harder, so every shape gets its own slot in the call, and on cut media every speaker clip its own window.\n- **The seat is an object on a ground.** For video-asset PIP, the asset fills the ground behind the speaker inset while its important content stays visible. Other seated beats paint an opaque ground across the frame and place the moment-owned speaker view in its seat; an aperture clips that view to the measured hole. In those layouts, nothing tucks under, overlaps into, or straddles the seat to buy room; content low in the band clears the seat's width as well as its top edge.\n- **A separate payload lives in the band the seat leaves.** The payload is composed inside that band, not against the whole frame.\n- **A seat arrives once per run of beats.** Adjacent beats that share a framing keep the seat where it is: it does not re-enter, re-settle, or re-announce itself; only the payload turns over. A new framing arrives at a beat boundary, either from the full frame or as a morph from the previous seat.\n\n### How a seat is built\n\n- Every `seat` move owned by a visual moment lives inside that moment's paused child timeline. Put a muted `data-pip-src=\"public/source.mp4\"` video in a `.speaker-pip-frame`, keep its global start and duration aligned with the host, and set its media start to the source time at the host's anchor. The composition contract's [editable speaker PIP](../aip-composition/references/pip-transition.md) is the single source for the snippet.\n- Animate the frame's box (`left`, `top`, `width`, `height`), crop (`object-position` with `object-fit: cover`) and corner radius together over the same interval. A circle uses equal sides and `borderRadius: \"50%\"`; an aperture clips the same moment-owned view to its measured hole. The opaque composition ground hides the unchanged root picture.\n- Never couple an independently editable host to absolute-time speaker geometry on `finecut-root`. The editor moves and deletes the host and its child content as one unit, but it does not discover or rewrite unrelated GSAP tweens at the old global time.\n\n## Overlay\n\nThe footage owns the frame and one payload rides it. The root speaker keeps playing full-bleed and sharp underneath, and the moment draws only what the payload needs: no page, no wash, no copy of the source.\n\n### Overlay forms\n\nEvery overlay is one of these forms. Name the form and where it sits in the plan.\n\n| form | what it is |\n| --- | --- |\n| `annotation` | drawn graphics, a short label, or an image placed in the area clear of the head |\n| `headline` | the phrase just spoken, drawn big as the beat's whole payload; its window goes to `hide_intervals`, because the captions would otherwise repeat the line |\n\n### Overlay positions by canvas\n\nThe measured head decides where an overlay may sit: the payload stays outside `head_in_slot` for that window, with room to breathe around it.\n\n| canvas | overlays that fit | do not use |\n| --- | --- | --- |\n| portrait 9:16 (1080x1920) | `headline` in the lower third, or above the head where the head box leaves real headroom; `annotation` in the band above or below the head at full width | anything beside the face: the flank is too narrow to read, so a portrait overlay sits above or below the head |\n| square 1:1 (1080x1080) | `headline` in the lower third; `annotation` beside the head on the flank the head box leaves open, or below the head | payloads on both flanks at once; a payload that crosses the head box |\n| landscape 16:9 (1920x1080) | `headline` in the lower third or on the open flank; `annotation` beside the head on the flank an off-center face leaves open | a band above the head: it is a thin strip; a payload that crosses the head box |\n\n### Where the caption sits under an overlay\n\nThe head, the payload, and the band avoid one another. The footage stays full-bleed; the payload takes the empty region outside the head box, and the band takes another band of its own, so the payload can animate without the caption following it. Only avoiding the face is not enough: the band must not cross the payload's key change, and the payload must not enter the band as it moves. When only one empty region exists, shrink or move the payload, or change to a seat or a full cover; the caption never follows the face word by word.\n\n| canvas | the band | the payload |\n| --- | --- | --- |\n| portrait 9:16 | the lower third, about y 1400 to 1600: `data-caption-position` 78, or 72 for a growing pattern | below the head, about y 860 to 1250 |\n| landscape 16:9 | the bottom band, about y 800 to 890: `data-caption-position` 78 (the service renders a single-row pattern on this canvas, so the band is one row) | the open flank, ending above y 620 |\n\n### What an overlay keeps\n\n- **The overlay rides sharp footage.** The composition never pauses, scales, or reframes the source under an overlay.\n- **The head stays clear.** Name the overlay's window in the `frame_speaker` call with the canvas as its `slot`, and keep every text, plate, and image outside the returned `head_in_slot`; a thin connector line may reach past it toward what it points at. On a reframed canvas the box holds while the root clip carries the same result's `object_position`. The box describes the unzoomed frame: when a root zoom runs during the overlay's window, clear the box as it stands at the zoom's largest scale in that window, each edge pushed away from the zoom's origin by that scale, or end the zoom before the overlay starts.\n- **Payloads take turns.** An overlay holds one payload group at a time; the next arrives only as the previous yields.\n\n### How an overlay is built\n\n- An overlay is an ordinary visual moment whose composition stays transparent: no `background` on the composition root or on any full-frame wrapper, the payload inside one positioned wrapper, and its entrance and exit on the moment's own paused child timeline. The composition contract's [editable speaker PIP](../aip-composition/references/pip-transition.md) example minus its ground and its speaker view is an overlay.\n- It embeds no video: no `data-pip-src` view and no copy of the source, because the root speaker underneath is the picture. Leave the root speaker's geometry to the root timeline.\n- Declare the window's caption policy on the host (`data-hide-captions=\"false\"` and its `data-caption-position`) and pass it to the caption call as a caption window. A `headline` is the exception: pass its window as `hide_intervals`, because the captions would otherwise repeat the phrase the headline draws.\n\n## Full cover\n\nThe canvas owns the frame. The moment paints one opaque ground across the whole canvas and everything sits on it; nobody is seated, and the root speaker keeps playing underneath, unseen and still heard.\n\n### Full-cover forms\n\nEvery full cover is one of these forms. Name the form in the plan.\n\n| form | what it is |\n| --- | --- |\n| `figure` | a drawn diagram, chart, or mechanism at full scale |\n| `photo` | a real image staged on the ground as content; a video asset is seated as `Video assets` above says, not covered |\n| `statement` | one oversized line, a few words at the largest size the page allows |\n| `source-text` | a crop of dense source material, enlarged until the relevant detail reads |\n\n### Full-cover positions by canvas\n\nThe canvas decides how a page is laid out.\n\n| canvas | pages that fit | do not use |\n| --- | --- | --- |\n| portrait 9:16 (1080x1920) | one centered focal element, or a vertical stack of two: the figure above and its one-line label below | side-by-side columns: each is too narrow to read |\n| square 1:1 (1080x1080) | one centered focal element, or a stack of two rows | columns narrower than half the frame; more than two rows |\n| landscape 16:9 (1920x1080) | one centered focal element, or two columns side by side: a figure and its label, a before and an after | a vertical stack of three or more rows: each is a thin strip |\n\n### Where the caption sits on a full cover\n\nSettle the band first, then compose the payload in the space it leaves: the animation recentres, scales, and moves inside that space, and the band must not cross an arrow's end, a compared result, a value, or a legend through the whole entry and exit. Check the frames where the speaker shows briefly at the switch. Only where the ground under the band is one this composition painted light (a pale panel, a light style the brief asked for) add `data-caption-ink` with a hex that contrasts with that ground (`#0a0c12` on a pale ground); the default black canvas keeps the white type, so the ink is a per-moment call, never a full-cover default, and never over footage or a photo, which move under the words. A `statement` hides the captions, and so does a `source-text` or `photo` whose relevant detail must occupy the band and cannot move up.\n\n| canvas | the band | the payload |\n| --- | --- | --- |\n| portrait 9:16 | about y 1390 to 1590: `data-caption-position` 78, or 72 for a growing pattern | recentred above the band, about y 360 to 1060 |\n| landscape 16:9 | the bottom band, about y 800 to 890: `data-caption-position` 78 (the service renders a single-row pattern on this canvas, so the band is one row) | above the band, about y 160 to 630 |\n\n### What a full cover keeps\n\n- **One ground, wholly owned.** The composition paints one opaque ground across the whole canvas, and everything sits on it. An image is staged on the ground as content, never stretched into the ground itself.\n- **Nobody is seated.** No `data-pip-src` view, no window onto the speaker, and no copy of the source. The cutout is the one exception, and it is a seat form with its own [reference](references/cutout.md).\n- **The speaker underneath does not move.** The root speaker keeps its geometry for the whole cover; a root zoom is released before the cover starts.\n\n### How a full cover is built\n\n- A full cover is an ordinary visual moment whose composition root paints the opaque ground (`position: absolute; inset: 0` with a solid `background`), with the payload's entrance and exit on the moment's own paused child timeline. The composition contract's [editable speaker PIP](../aip-composition/references/pip-transition.md) example minus its speaker view is a full cover.\n- It embeds no speaker view. The root speaker and its audio keep playing underneath at their ordinary geometry; release a root zoom before the full cover starts.\n- Declare the window's caption policy on the host and pass it to the caption call as a caption window; a `statement`, or a payload whose relevant detail must occupy the band, goes to `hide_intervals` instead. Reference every image through `src` so the hand-back check sees it.\n\n## The speaker beat\n\n- Root zooms on a `speaker` beat that is not owned by a visual moment may remain root tweens on every active track-0 speaker clip (`#stage > video[data-track-index=\"0\"]`). Release that camera treatment before the next seated moment; a root zoom never lands inside a seat.\n"
}SHA-256: 309e16cdc70cd5b7680a475632949ab059b5bf3aae91b92f0b0d32a567f34a99