← AI Producer by OpusClipCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to AI Producer by OpusClip
Snapshot Sep 30, 2026 · 22:59 UTC · version 1.2.8
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "aip-dynamic-caption",
"description": "Pick one of the eight preset dynamic caption styles for an AI Producer project and have the service build the caption layer: how to choose the pattern from the video's content, which knobs the host controls, and the one call. Listed, it makes captions the default.",
"included_files": [
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 137
}
],
"skill_md_contents": "---\nname: aip-dynamic-caption\ndescription: \"Pick one of the eight preset dynamic caption styles for an AI Producer project and have the service build the caption layer: how to choose the pattern from the video's content, which knobs the host controls, and the one call. Listed, it makes captions the default.\"\n---\n\n# Dynamic captions for an AI Producer project\n\nThe caption layer is one composition, `compositions/narrator_captions.html`, with `compositions/transcript-src.json` beside it. The service builds both from a subtitle pattern you pick and the project's transcript; you do not write caption HTML, CSS, or word data. Your job is the pick and its few knobs, made from what you know about the video.\n\n## What you decide\n\n1. **The pattern.** One per video, from the table below. Read the transcript at `get_transcript` with `detail: segments` (a tenth of the bytes) for register and pace, and look at the source for brightness, contrast, where the face sits, and whether a brand accent exists. Those are the inputs the table keys on.\n2. **Emphasis words** (optional). At most one per phrase: the payload word a viewer should feel land. Take `word_id` values from `get_transcript` with `detail: words`. Every pattern has a designed treatment for an emphasis word; a phrase without one renders in the quiet register. Leave the list out and the service picks them with one caption-plan call.\n3. **Position** (optional). Leave `position` out unless the brief places the captions; the service then uses the pattern's default. If you set it, it is the band's vertical anchor as a percent of frame height: 50 is mid-frame, 80 is the lowest the text-safe zone allows, and the service clamps anything past that, tighter for a pattern whose block grows downward. The canvas decides the default and the pattern: on a 1080x1920 canvas the band sits at 70 and every pattern is available; on a 1920x1080 or 1080x1080 canvas the speaker's torso reaches the frame's floor, so the band sits at 78 (the bottom band), the type is smaller (44 to 56px) and holds more words per phrase, and a pattern whose block grows downward (`blur-ladder`, `editorial-stack`, `lead-in-flare`, `tilt-slam`, `scribble-subtitle`, `inverted-stack`) is rendered as `pace-adaptive`; the task result's `pattern` is what was baked and `substituted_from` names what you asked for. Pick `pace-adaptive` or `solo-word-punch` outright for a landscape or square video. `placement` is `adaptive` (default: the band moves off the speaker where the project's person samples exist; the task result's `placement` says whether it did) or `fixed`. Horizontal composition is part of each pattern's design and is not a knob; change the pattern to change it. The task result's `position` is the anchor actually baked; when it differs from what you passed, the pattern's safe band clamped it. The fix is another call with a position inside the band, never a CSS offset on the host div or the `.caption-band` in `index.html`: the band is the full canvas, so an offset pushes every word off screen.\n4. **Hide windows** (optional). `hide_intervals` is a list of `{start, end}` seconds where the band stays hidden. Captions stay on during visual moments by default; hide only for a `headline` overlay or a `statement` full cover (the spoken phrase drawn big), a `source-text` or `photo` full cover whose relevant detail must occupy the band and cannot move, or a stretch the brief wants silent. An empty list keeps captions on for the whole video and never hides anything by itself.\n5. **Font** (optional). Leave `font` out unless the brief names a typeface; the captions then render in Geist with Fraunces for the emphasis word. If the brief names one, pick the closest pairing register and pass it as `font`: `standard` (Geist, Fraunces), `quiet-luxury` (Plus Jakarta Sans, Newsreader), `avant-garde` (Bricolage Grotesque, Fraunces), `heavy-hitter` (Schibsted Grotesk, Spectral), `pure-editorial` (Familjen Grotesk, Source Serif 4), `warm` (Hanken Grotesk, Crimson Pro). The sans is every caption word and the serif is the emphasis word; there is no free family name. The task result's `font` names the pairing baked.\n6. **Caption windows** (optional). `caption_windows` is a list of `{start, end, position?, ink?}` on the cut, one per visual moment whose band differs from the whole-video caption. `position` is that window's anchor as a percent of frame height; the [framing skill](../aip-framing/SKILL.md) gives the value per layout and canvas (`Where the caption sits`). `ink` is a six-digit hex the caption type takes in that window, for a moment whose ground the composition painted light: a full cover on a pale canvas, or the canvas or card under the band while the speaker sits in a seat. Pick a colour with a large lightness gap against that ground (`#0a0c12` on a pale ground), never a mid grey; the type stays white everywhere else, and an overlay rides footage, which moves under the words, so it never takes an ink. Each window names at least one of the two, and where two windows overlap the earlier start wins. Write the same values you pass on that moment's host in `index.html` (`data-hide-captions=\"false\"`, `data-caption-position`, `data-caption-ink`), because the editor and the export re-derive the caption layer from those attributes after any edit, and a moment host without `data-hide-captions=\"false\"` hides the captions there at the first editor rebuild; the call bakes the first copy. The task result's `windows` has one row per window: `position` as you passed it, `placed` where the pattern's safe band lets it land, `band`, the pixel box of that landing (`anchor`, `top_px`, `bottom_px`), and `ink` as you passed it; when `placed` differs from `position`, keep a seat or a graphic clear of `band`, not of the number you asked for. The top-level `band` is the pixel box the whole-video anchor produces, the region a seat or a graphic keeps clear.\n7. **Accent** (optional). Three patterns paint an accent colour: `lead-in-flare` (the emphasis word's type), `scribble-subtitle` (its marks), and `tilt-slam` (the slab under the loudest word). Pass `accent` as a six-digit hex when the brand has one; leave it out and those patterns render their white fallback. The service never picks a colour for a host project, so any colour in the caption is the `accent` you passed. The task result's `accent` is the hex painted, or null.\n\n## The call\n\nCall `build_captions` with `project_id`, `caption_style` (the pattern), and any of `position`, `placement`, `hide_intervals`, `caption_windows`, `emphasis_word_ids`, `font`, `accent`. It returns a task; read it with `get_task` (its `poll_args` say when). The finished task's `result` names the files the service staged (`files`, one `{path, sha256}` per file; `staged` lists the same paths), the pattern and position it baked, the emphasis it used as `word_id` values, and `host_div`, the one element for `index.html`:\n\n```html\n<div\n class=\"visual-host clip\"\n data-composition-id=\"narrator-captions\"\n data-composition-src=\"compositions/narrator_captions.html\"\n data-start=\"0\"\n data-duration=\"<total s>\"\n data-track-index=\"<n>\"\n data-width=\"W\"\n data-height=\"H\"\n></div>\n```\n\nPlace `host_div` as written (its duration and canvas come from the root you committed; placeholders mean commit the root first) inside the root, last, on its own track, then upload your tree and `commit_workspace`. When that commit names `expected_files`, add the result's `files` entries to the list exactly as returned; a batch that leaves them out keeps the caption files staged and the round reports them in `retained_staged`. A commit that names no batch promotes them with everything staged. The editor finds the running caption by the `.caption-band` element inside that host. A file you stage yourself at `compositions/narrator_captions.html` or `compositions/transcript-src.json` replaces the service's. A caption layer you write yourself follows the caption entry of the [composition contract](../aip-composition/SKILL.md). Hand back the editor link; there is no local inspection step. One `build_captions` run is one flat charge, whether or not you named the emphasis words.\n\n## Patterns\n\n| Choose when the video is | pattern | Rotate away when |\n| --------------------------------------------------------------------------------------------------------------------------------------------- | ------------------- | -------------------------------------------------------------------------------------- |\n| A tutorial or a punchy talking head where each sentence should feel hand-cut; face outside the 50-73% band | `blur-ladder` | captions should be a quiet legibility layer; the footage is busy |\n| A piece where the running caption is the design itself; darkish source; face outside the 50-68% band | `editorial-stack` | captions should stay quiet under the speaker or under visual moments |\n| A dark speaker cut against bright b-roll, where the caption should read as grade rather than overlay; real tonal contrast in the caption band | `inverted-stack` | flat mid-tone footage |\n| **The default**: every sentence wants a visible payoff beat; a brand `accent` you can pass (the only pattern that colors caption type) | `lead-in-flare` | delivery so fast that a big tail every sentence lags the voice |\n| Dense, fast, or information-heavy; captions should stay uniform under the speaker | `pace-adaptive` | a punchier per-word emphasis or a calmer whole-line read fits better |\n| Creator confessional, story-time, a scrappy hand-made brand register | `scribble-subtitle` | a formal or corporate register; dense small on-screen text |\n| Short imperative narration, step-throughs, hype delivery, where the caption is a beat | `solo-word-punch` | reflective or information-dense narration |\n| Vlogs, hot takes, sticker energy, loud and playful; a brand `accent` you can pass fills the slab under the loudest word | `tilt-slam` | a formal or corporate register; dense small on-screen text; captions should stay quiet |\n\nWhen captions share the frame with visual moments for most of the video, prefer a quiet pattern (`pace-adaptive`, `blur-ladder`): `tilt-slam` and `solo-word-punch` compete with an animation for the same beat, and `inverted-stack` cannot take an ink (its difference blend would invert the colour). One pattern per video. The pick is yours; the pattern's motion and layout are the service's, and its typeface is the pairing `font` names (Geist with Fraunces italic emphasis by default), staged with the composition under `public/fonts/`.\n"
}SHA-256: 01e04f028b62b33291ac61e48b22a9cb658eaf50901357d7505a2058ac1876ca