{"id":6788,"plugin_id":"plugin_asdk_app_6a61e06bd2d08191ab2faa84dfd37b96","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T22:49:07.317Z","digest":"ed81d85d2e8f0db2732829deb1a38c48ac0a81d88707620515cdca8ac5c8dc83","against":null,"payload":{"name":"create-moment","description":"Create an editable moment (clip) from an ingested video via the Montage MCP tools and deliver an editor link. Use after video ingestion when the user wants a clip/moment created from their video.","included_files":[],"skill_md_contents":"---\nname: create-moment\ndescription: Create an editable moment (clip) from an ingested video via the Montage MCP tools and deliver an editor link. Use after video ingestion when the user wants a clip/moment created from their video.\n---\n\n# Create a moment from an ingested video\n\nEnd-to-end flow: read the ingestion outputs, pick clip boundaries, create the\nmoment, wait for clip preprocessing (downscale + speaker detection), and hand\nthe user an editor link. All tools are on the Montage MCP server.\n\n## 1. Locate the project/video\n\n- `list_projects` / `get_project` to find the project the user means.\n- If the video was just uploaded with `upload_video`, poll\n  `get_upload_status(workflow_id)` until `completed` — it returns the\n  `project_id` and `file_id`. Ingestion must be complete before creating a\n  moment (transcription and video validation results are required).\n- Never pass a video URL to `create_moment` — it takes `project_id`/`file_id`\n  and resolves the video itself. On paid plans ingestion repoints the project's\n  video at the noise-removed copy, so the moment is cut from cleaned audio\n  automatically. Creating a moment before ingestion finishes can catch the\n  pre-correction video, which is a second reason to wait for `completed`.\n\n## 2. Understand the content\n\nRead all four ingestion outputs before picking anything — each answers a\ndifferent question, and a moment chosen from the transcript alone will cut\nacross bad footage.\n\n- `get_video_validation(project_id)` — technical ground truth: `duration_seconds`\n  (the hard upper bound for every segment), `width`/`height`/`fps`, codecs,\n  `is_valid` + `failed_checks`/`error_summary`, and the probe rollups\n  `speech_detection_info`, `audio_quality_info`, `person_detection_info`. Read\n  it first: it tells you whether the footage even has usable speech, audio and\n  people on camera.\n- `get_transcript(project_id)` — what was said, with timestamps and\n  `speaker_id`. Paginate for full text. This is where segment boundaries come\n  from.\n- `get_video_analysis(project_id)` — the vision-model read: scene breakdown,\n  clip candidates with hook/payoff reasoning, `category_name`,\n  `max_people_count`, reframe instructions. Treat its clip candidates as\n  proposals to validate against the transcript, not as final boundaries.\n- `get_video_corpus(project_id, view=\"rollup\")` — on-screen reality: coverage\n  fraction, subject roster with speaking/tracked seconds, hard-veto ranges,\n  screen-text inventory. Only fall back to `view=\"full\"` when the rollup is\n  genuinely not enough — it returns the whole index and is large.\n- `get_moments(project_id)` to see what already exists — don't duplicate.\n\nLow `coverage_fraction` means those windows were never analyzed, not that\nnothing happens there. Say so instead of reporting absence as a finding.\n\n## 3. Pick segments\n\nChoose one or more `[start, end]` second-pairs, and make every one of them\nsurvive all four reads:\n\n- Boundaries come from **transcript** timestamps — cut on sentence edges, not\n  mid-word.\n- The story beat comes from **video analysis** scenes/clip candidates — a\n  moment should be one coherent beat (or a hook plus its payoff).\n- **Corpus** hard-veto ranges are exclusions: do not cut inside them, and\n  prefer windows where the speaking subject is also on camera.\n- **Validation** bounds it: `0 <= start < end <= duration_seconds`, and if\n  `is_valid` is false or speech/person detection is empty, say so before\n  creating anything.\n\nMultiple pairs compose one moment. Write a short `title` (and optionally a\n`summary`) describing the moment.\n\n## 4. Create the moment\n\n```\ncreate_moment(project_id, title, segments, utterances, summary?, kind?, file_id?)\n```\n\n- `utterances` (required for speaker detection): the transcript segments from\n  step 2 that overlap the chosen `segments`, as\n  `[{\"start\": <sec>, \"end\": <sec>, \"text\": \"...\", \"speaker\": \"<speaker_id>\"}, ...]`\n  — map `get_transcript` fields `start_time`→`start`, `end_time`→`end`,\n  `speaker_id`→`speaker`, using absolute video seconds\n  (`timestamp_format=\"seconds\"`). Cover every chosen segment window; without\n  utterances, speaker detection is skipped.\n- `kind`: content (default), hook_montage, ad, or sponsor_read.\n- `file_id` only when the project has multiple files.\n- Requires the `moments:write` scope.\n\nThe response contains `moment.id`, `preprocessing.{status, workflow_id}`, and\n`editor_url`.\n\n## 5. Wait for preprocessing\n\n- If `preprocessing.status` is `\"processing\"`: poll\n  `get_workflow_status(preprocessing.workflow_id)` every ~30s until `status`\n  is `completed` (typically 2–10 minutes). On `failed`/`timed_out`, report it\n  — the moment still exists, but the editor may lack the downscaled preview.\n- `\"ready\"` means enrichment already exists; `\"skipped\"` means preprocessing\n  could not run (`preprocessing.reason`: `no_video_url` / `no_input_codec` /\n  `no_utterances`) — nothing to poll, tell the user why.\n\n## 6. Deliver\n\nGive the user the `editor_url` from the create_moment response verbatim:\n\n```\nhttps://studio.montage.app/project/{project_id}?view=clipEditing&clipId={moment_id}\n```\n\n(`clipId` is the moment id.)\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}