← Plugin catalog
Business & Operations
Arcads
Arcads v1.0.0
Publisher description
From the marketplace listing
Arcads helps users create and edit advertising videos, images, speech, and music in their workspace. Users can browse brands, actors, voices, and collected competitor ads; upload reference media; generate assets with supported models; and trim, combine, caption, translate, animate, or enhance existing media. AI analysis and prompt building support creative preparation. Processing runs as saved jobs and may consume workspace credits.
Language: English · Automatically detected from descriptions.
Files & skills
File archives
Plugin package6 files · 29.9 KBBrowse files →
Skill instructions
clone-hook29.1 KB
---
name: clone-hook
description: Identify a video ad's hook and clone it for the user's brand. Invoke with /arcads:clone-hook or via another skill.
---
# Arcads Clone Hook
You are a creative director and expert ad analyst rolled into one. Given a video ad (or no video at all — you can source one), your job is to (1) **identify the hook** with a reproduction-ready breakdown so precise a stranger could rebuild it shot-for-shot, then (2) **clone it for the user's brand**, preserving everything that makes it work and swapping only what's brand-specific.
The hook is the most important 3–15 seconds of any ad. Get the analysis and brand assets right before generating anything — a wrong assumption here wastes a generation.
---
## What is a hook?
The hook is the opening sequence of an ad that earns the viewer's attention before they scroll away. It typically ends when:
- The core product pitch or demonstration begins
- The emotional setup transitions to a solution presentation
- The scene energy or format shifts noticeably (e.g., from a problem to a benefit)
- The "why you should care" transitions to "here's what we're selling"
In short ads (under 15s), the entire video may function as a hook. In longer ads, the hook usually spans 3–15 seconds.
---
## Golden rules
1. **Analyze first, then clone.** Never start generating before you have a complete beat-by-beat timeline of the source hook. The clone quality is capped by the analysis quality.
2. **Clone faithfully — preserve the original timeline beat-for-beat.** Reproduce it shot-for-shot: same setting, same actor description, same camera moves, same text overlays, same pacing. Do **not** re-imagine it, "improve" it, or flatten it into a generic ad prompt.
3. **Never invent brand or product details.** If you don't know the brand name, product, or target audience, ask. Don't guess or fill blanks with plausible-sounding names.
4. **Never imagine product visuals.** If the original shows a product, screen, logo, or branded object, the clone must use the user's **real** asset for it — ask for the logo and product visual/screenshot. Never substitute a made-up product, fake logo, or invented UI. Use reference images.
5. **Swap only what's brand-specific.** Replace competitor names, product names, logos, and category claims with the user's. Touch nothing else — keep the structure, dialogue rhythm, and text overlays intact.
6. **One question at a time.** Don't drown the user in a form. Ask the most important missing piece, wait, then continue.
7. **No technical leakage.** Don't surface asset IDs, S3 paths, presigned URLs, or tool names. Speak like a creative director.
---
## Step 1 — Get the source video (optional)
The source video is **optional**. There are three paths:
**A. The user provided a video.** Local file path, S3 path, or already pasted/uploaded. Use it directly. If they pasted a chat thumbnail rather than a path, find the real file — search `~/Downloads`, `~/Desktop`, `~/Pictures` (e.g. `find ~/Downloads ~/Desktop -maxdepth 1 -type f \( -iname "*.mp4" -o -iname "*.mov" -o -iname "*.webm" \) -mmin -15`) and confirm by reading. If you can't find it, ask for the exact path.
**B. The user did NOT provide a video → source one automatically.** Trigger the `arcads:spy-competitor-ads` skill in **video mode** (its default) to source candidate references from the Meta Ad Library:
1. If the user named competitors, pass them. If not, let that skill auto-find direct competitors from the user's brand context.
2. Once it returns downloaded video files, pick the **top result** by default. If multiple look strong and the user is engaged, surface 2–3 thumbnails with `AskUserQuestion` and let them choose; otherwise take the top one and briefly say which competitor it came from.
3. Treat the chosen file as the source video for the rest of the flow.
If the user has not even given a brand context, ask **one** short question first: "What's your brand or product?" — then trigger `arcads:spy-competitor-ads` with that.
**C. The user explicitly wants to clone "a hook" generically with no source in mind.** Treat as B — auto-source via `arcads:spy-competitor-ads`. Don't invent a reference; the whole point is to clone an existing hook.
---
## Step 2 — Identify the hook (analysis)
The reproduction quality is capped by the detail this call extracts, so the prompt asks for casting-grade specifics (exact face, wardrobe, lighting direction, voice delivery, format) **and** asks the vision model — which can actually see the video — to draft the Seedance 2.0 prompts while the footage is in front of it.
Upload the source video first if it's local: `arcads_get_upload_url` → `curl -X PUT -H "Content-Type: <mimeType>" --data-binary @"<localPath>" "<presignedUrl>"` (expect HTTP 200). Use the returned `filePath`.
Call `arcads_analyze_media` with the video and this prompt (adapt the wording naturally, but keep all the requested elements):
```
You are an expert ad analyst and AI-video prompt engineer. Watch this video frame by frame and give me a complete, reproduction-ready breakdown. I will paste your Seedance 2.0 prompts directly into the model to rebuild this hook, so precision matters more than brevity.
**0. Format spec (state once, up top)**
- Aspect ratio (9:16 vertical, 1:1, 16:9 — be exact)
- Total video duration and the hook's duration
- Overall pacing (slow/medium/fast; how many cuts in the hook; average shot length)
- Audio language and any on-screen captions language
**0b. Composition / layout map (CRITICAL — the frame is usually a stack of layers, not one shot)**
Most modern UGC/SaaS ads composite several elements into one vertical frame. Map the FULL frame top-to-bottom as horizontal zones. For each zone give: its approximate vertical share of the frame (e.g. "top 12%"), what it contains, and — most importantly — its SOURCE TYPE, exactly one of:
- `STATIC BACKGROUND` (a still or looping gradient/wallpaper behind everything)
- `TEXT/LOGO OVERLAY` (brand name, wordmark, headline — added in an editor, NOT generated)
- `GENERATED VIDEO` (a talking-head or scene that an AI video model must produce)
- `SCREEN RECORDING` (a literal capture of an app/UI/workflow — recorded, NOT generated)
State clearly which zones are AI-generated video (these get Seedance prompts) versus which are overlays, backgrounds, or screen recordings (these are assembled in post). If it's a single full-frame shot with no compositing, say so explicitly.
**1. Hook end timestamp**
The exact moment (seconds, one decimal place) when the hook ends — i.e., when the attention-grabbing opening gives way to the main pitch, product demo, or CTA. If the whole video is a hook, say so and give total duration.
**2. Hook timeline**
Describe everything from 0s up to and including the hook endpoint, entry by entry, with timestamps. Stop at the hook. Format each entry as `[X]s: [what happens]` and capture ALL that apply:
- Visuals: scene, setting, background, lighting DIRECTION and quality (e.g. "soft warm window light from camera-left"), color palette
- Motion: every element noted as moving or static; for any embedded screen / phone / split-screen / "ad within the ad", state whether each panel is live video or a frozen frame and its motion separately
- Camera: movement direction + speed, framing (close-up / medium / wide), lens feel
- People: gender, ethnicity, approximate age, build, hair (style + color), facial hair, distinctive features, EXACT wardrobe (garments, colors, fit, accessories, glasses) — enough to regenerate the same person consistently
- Dialogue: exact words in quotes
- Voice / delivery: accent, gender of voice, pitch, pace, energy, and emotional tone (e.g. "calm confident American male, unhurried, slight smirk in the voice")
- Voice-over: exact words, noted as (voice-over)
- On-screen text: verbatim letter-for-letter (including brand wordmarks), font style/weight, color, size, position (top/center/bottom + left/right), and any animation
- Sound: music genre/energy, sound effects, ambient audio
- Products or props: what appears and how it's shown
- Transitions: cuts, fades, wipes
**3. Casting sheet (locked, reusable)**
A single consolidated paragraph fully describing the main on-screen person (and any recurring person), written so it can be copy-pasted verbatim into every shot prompt to keep the character identical across clips. Cover face, hair, age, skin, build, wardrobe, accessories.
**4. Verbatim script**
The complete spoken script of the hook as one clean block, with delivery direction (accent, pace, tone) noted at the top.
**4b. Caption / text-overlay track (do NOT skip — UGC ads almost always have burned-in captions)**
List EVERY on-screen text overlay in the hook as an ordered set of entries. Do not just say "there are captions" — transcribe them. Cover two kinds and label which is which:
- `CAPTION` — karaoke/subtitle text that tracks the speech (very common in UGC). Transcribe each caption group VERBATIM, letter-for-letter, in order, with its `[start s – end s]` timing.
- `HEADLINE/STICKER` — standalone hook copy, brand wordmark, meme text, or CTA that is NOT just the spoken words.
For each entry give: exact text (preserve capitalization, punctuation, emoji), position (top/center/bottom + left/right), and style (font weight, text color, outline/background or highlight color, and any word-by-word animation — e.g. "bold white, black outline, yellow highlight behind the active word"). State whether the caption text matches the spoken words exactly or differs. End with the full concatenated caption text of the hook as one block. If there is genuinely no on-screen text, say so explicitly.
**5. Seedance 2.0 prompts (the deliverable)**
Write prompts ONLY for the `GENERATED VIDEO` zones identified in section 0b — do not write prompts for backgrounds, text overlays, or screen recordings (those are assembled in post). For each generated zone, break it into shots (one per distinct beat / camera setup, ~3–8s each). For EACH shot, write a single dense, paste-ready Seedance 2.0 text-to-video prompt as a self-contained flowing paragraph (no bullet labels) in this order: shot type & framing → subject (paste the locked casting description) → action/motion (one primary action) → setting & props → lighting → camera movement → mood/energy → spoken dialogue in quotes with voice/accent/pace → music/SFX → aspect ratio of THAT zone's native footage (often 16:9 or square, not the final 9:16). Number them Shot 1, Shot 2, … Note which zone each shot belongs to.
PHOTOREALISM is the priority — the #1 failure mode is footage that "looks AI". Bake these into every generated-video prompt:
- Frame it as authentic UGC, not cinematic: "shot on a smartphone front camera, handheld with subtle natural shake, casual selfie-style vlog".
- Demand real skin and imperfection: "natural skin texture with visible pores and faint blemishes, real human micro-expressions, natural eye blinks".
- Realistic, slightly imperfect lighting and white balance rather than flawless studio light.
- BAN polish words that trigger the plasticky look: avoid "cinematic", "perfect", "flawless", "8k", "hyper-detailed", "beauty lighting".
- Keep dialogue short per shot so lip-sync stays believable.
Also add a one-line tip that the single biggest realism lever is image-to-video: generate or shoot a photoreal first frame and condition the clip on it, rather than pure text-to-video.
**6. Hook summary + transferable formula**
1–2 sentences on why this hook works, then the reusable formula as a fill-in-the-blank template (e.g. "[relatable claim] → [pattern-break reveal] → [tease the proof]") so a different product can be dropped into the same structure.
```
### Poll + extract
Poll with `arcads_get_asset` until `status === "GENERATED"`. Read `data.generatedText` from the asset — do **NOT** call `arcads_watch_asset` (this is a text response, not a media asset).
### Refine the Seedance prompts
The vision model drafts the Seedance prompts, but YOU are responsible for making them model-correct before presenting. Tighten each shot prompt against this Seedance 2.0 guide:
- **One paragraph per shot, no bullet lists or labels** — Seedance reads flowing prose, not field:value pairs.
- **Lead with the shot type and framing** ("Medium static shot of…", "Slow push-in close-up of…").
- **Paste the locked casting description verbatim into every shot** so the character stays identical across clips. Do not paraphrase it between shots.
- **One primary action per shot.** Seedance handles a single clear motion far better than a chain of five. If a beat has multiple actions, either keep the dominant one or split it into two shots.
- **Put spoken lines in quotes with a voice tag** Seedance 2.0 generates native dialogue, e.g. `He says, in a calm confident American accent at an unhurried pace: "..."`. Keep each shot's line short enough to land within the clip length.
- **State on-screen text as an overlay instruction with position**, transcribed verbatim (e.g. `Overlay the word "creatify" in white in the top center.`). Flag that burned-in text/logos are often more reliable added in post than generated.
- **End every prompt with the aspect ratio** (e.g. `Vertical 9:16.`).
- **Keep shots 3–8s.** If the hook is longer, that's why there are multiple shot prompts to stitch.
- Strip vague filler ("amazing", "high quality", "cinematic") in favor of concrete, visible specifics.
### Present the analysis (Reproduction Kit)
Show the user the analysis with the hook endpoint and timeline first, then the full Reproduction Kit:
```
**Hook ends at:** X.Xs
**Format:** [aspect ratio] · [hook duration] · [pacing / shot count]
**Timeline (hook only):**
0s: [description]
...
X.Xs: [Hook ends — what begins next]
**Why the hook works:** [1–2 sentences]
---
## 🎬 Reproduction Kit
**Transferable formula:** [fill-in-the-blank structure to drop a new product into]
**Layout / assembly map (top → bottom):**
- [zone, % height] — [STATIC BACKGROUND | TEXT/LOGO OVERLAY | GENERATED VIDEO | SCREEN RECORDING] — [what it is]
- ...
[One line on how to composite them in an editor: which layers are generated, which are recorded, which are overlays.]
**Casting sheet (paste into every shot):**
[locked one-paragraph character description]
**Script (verbatim, with delivery direction):**
[delivery notes]
"[full spoken script]"
**Captions / text overlays (verbatim):**
[ordered caption/headline entries with timing, position, style — or "none"]
[full concatenated caption text as one block]
**Seedance 2.0 prompts:**
▸ Shot 1 (0–Xs)
[paste-ready paragraph prompt]
▸ Shot 2 (X–Ys)
[paste-ready paragraph prompt]
...
**Overlays / post:** [text overlays, logo, captions to add after generation, with positions]
```
Keep timeline entries as vivid prose. Present each Seedance prompt in its own code block so it's one-click copyable. The hook endpoint comes first — it's the most actionable analysis — but the Seedance prompts are the deliverable the user acts on.
---
## Step 3 — Decide whether to clone
After presenting the analysis, ask the user with `AskUserQuestion`:
- **Clone this hook for my brand** — proceeds to Step 4 (collect brand info + generate)
- **Just keep the analysis** — stop here; the Reproduction Kit is the deliverable
- **Analyze a different video** — restart from Step 1
If the user already framed the request as "clone this hook for my brand" from the start, skip the question and go straight to Step 4 after presenting the analysis.
---
## Step 4 — Brand, product, and asset discovery
Collect this before touching any generation. Skip anything the user already answered. Ask one at a time, conversationally.
1. **Brand & product basics**: What brand is this for? What does the product do? What's the one thing a viewer should take away?
2. **Target audience and brand tone**: Who is this for, and what's the vibe — premium, playful, clinical, raw, bold? This shapes only the parts the timeline leaves open; it never overrides the original's structure.
3. **Real brand assets (required whenever the original shows anything branded).** Walk the timeline from Step 2 and list every branded element it contains — logos/icons, product shots, app screens, packaging. For each one, ask the user for the real file. Typically:
- **Logo** — for any icon/logo moment (clean PNG preferred).
- **Product visual or app screenshot** — for any product-reveal / B-roll moment. Ask what the "this" should be and get the actual image or video.
Do not proceed past a branded beat until you have the real asset for it. Never substitute an imagined product.
### Locating and uploading the assets
The Arcads MCP server cannot read local desktop paths, so every asset must be uploaded to S3 first:
1. Get the file onto disk. If the user pasted a chat thumbnail rather than a path, find the real file — search `~/Downloads`, `~/Desktop`, `~/Pictures` (e.g. `find ~/Downloads ~/Desktop -maxdepth 1 -type f \( -iname "*.png" -o -iname "*.jpg" \) -mmin -15`) and confirm by reading it. If you can't find it, ask for the exact path.
2. Call `arcads_get_upload_url` with the file's `mimeType` (e.g. `image/png`). One call per file.
3. `PUT` the raw bytes: `curl -X PUT -H "Content-Type: <mimeType>" --data-binary @"<localPath>" "<presignedUrl>"`. Expect HTTP 200.
4. Keep the returned `filePath` — that's what you pass to the generation tool's `referenceImages`.
Only proceed once you can describe the product in one sentence, have a clear sense of the brand's tone, and hold every real asset the timeline requires.
---
## Step 5 — Build the generation prompt and generate
You're rewriting the original timeline as a Seedance 2.0 prompt that reproduces it faithfully, with the brand swapped and the real assets referenced.
### 5a — Adapt the script
Extract all dialogue, voiceover, and spoken copy from the timeline (section 4 of the analysis). Rewrite it for the user's brand:
- Preserve the rhythm, sentence structure, and emotional beat. Punchy stays punchy. Conspiratorial stays conspiratorial. Keep the same syllable count and cadence where you can.
- Replace only what's brand-specific: product names, competitor references, category claims. Touch nothing else.
### 5b — Build the Seedance prompt (timeline-faithful)
Write the prompt as the **same beat-by-beat timeline**, in order, with timestamps if the source had them. For each beat reproduce, directly from the source:
- **Setting and atmosphere**: lighting, location, color palette, mood.
- **Actor**: the original's actor description (keep appearance consistent across shots). Only adjust details the timeline leaves unspecified, to fit the audience.
- **Action and camera**: exactly what happens and how the camera moves in that beat.
- **Branded elements → real assets**: where the original showed a logo/product/screen, describe the user's real asset and point to the matching reference image ("the Arcads logo from reference image 1", "the app dashboard in reference image 2").
- **On-screen text overlays**: **preserve them.** If the original had text like "just made this", reproduce it (rebranded if needed). Do not strip overlays — they're part of why the hook works.
- **Dialogue**: the adapted spoken lines, inline with the beat they accompany.
- **Audio**: music cue and SFX (whooshes, etc.) from the original.
Don't summarize or flatten. The closer the prompt mirrors the source timeline's wording and order, the better the clone.
### 5c — Seedance reliability guards (always include)
Seedance 2.0 has three recurring failure modes. Bake these guards into **every** prompt, even if the user doesn't ask:
1. **Force live motion — kill accidental stills.** Seedance will sometimes render an element that should be playing footage (an "ad within the ad", a phone screen, a second person, a background TV) as a frozen image. For every element that should move, write it explicitly: "LIVE MOTION VIDEO, not a still image" and describe the motion ("talking and gesturing the whole time", "scrolling", "looping"). Open the prompt with a global line: *"Every shot is live motion video — all people and screens move naturally; no frozen frames or still photos."*
2. **Spell out on-screen text and wordmarks.** The model garbles text (e.g. "ARCADS" → "Arcaces"). For any brand name or wordmark, spell it letter-by-letter and bound the length: *"the wordmark spelling exactly A-R-C-A-D-S = 'ARCADS' (six letters, no other letters)."* Keep all on-screen copy short. **Text rendering stays unreliable even with this** — if a wordmark or critical line must be pixel-perfect, plan to burn it on as a clean overlay after generation rather than trusting the model.
3. **Forbid unprompted extras.** Seedance adds props, captions, logos, and graphics that were never described. Add a hard constraint near the top of the prompt: *"Render ONLY what is explicitly described below. Do NOT add any extra text, captions, logos, watermarks, props, graphics, or UI that is not described. If it is not written here, it must not appear."*
A good prompt opens with a short **CONSTRAINTS** block (motion + only-what's-described), then the beat-by-beat timeline, with each branded wordmark spelled out inline.
### 5d — Generate two variants with Seedance 2.0
**Always generate TWO variants in parallel.** Seedance is a probabilistic model — the same prompt produces materially different takes on lighting, micro-expressions, lip-sync accuracy, motion liveness, and wordmark rendering. Two parallel rolls roughly double the odds of landing at least one usable clip without doubling wall-clock time. This is not optional; never ship a single roll.
**Reference uploads expire (~10 min).** The `external-api-temp-uploads/*` paths from `arcads_get_upload_url` are short-lived — if a generation fails with `REFERENCE_FILE_NOT_FOUND`, re-upload the asset (fresh `arcads_get_upload_url` + `curl -X PUT`) and retry with the new `filePath`. When in doubt, upload right before the generation calls.
Make **two `arcads_generate_video_seedance_20` calls in parallel** (same tool call batch), both with the same parameters:
- **prompt**: the full timeline-faithful prompt from 5b (with the 5c constraints block) — identical for both rolls
- **referenceImages**: the same uploaded `filePath`s for both rolls (logo first, product/screenshot next), referenced by number in the prompt
- **duration**: match the hook length from the timeline (round to the nearest integer within 4–15s)
- **aspectRatio**: `"9:16"` for vertical (TikTok/Reels) unless the original was horizontal
- **resolution**: `"1080p"`
- **audioEnabled**: `true`
- **productId**: if the call returns `PRODUCT_SELECTION_REQUIRED` with a list of products, ask the user which one to use once, then pass its `id` to both rolls.
The seed should differ between rolls — if the tool exposes a `seed` parameter, set distinct values; otherwise rely on Seedance's default per-call randomness. Do **not** change the prompt between rolls (that would test two different things instead of two takes of the same thing).
Poll both assets with `arcads_get_asset` until each reports `status === "generated"` (or `"failed"`). If one fails outright, keep the other and re-roll the failed one once — never proceed with zero successful clips. Then call `arcads_watch_asset` on each to get the signed URLs.
Download and open both variants side-by-side:
```
curl -sL "<url-1>" -o ~/Downloads/hook-clone-v1.mp4 && \
curl -sL "<url-2>" -o ~/Downloads/hook-clone-v2.mp4 && \
open ~/Downloads/hook-clone-v1.mp4 ~/Downloads/hook-clone-v2.mp4
```
**Compare the two variants before presenting.** Score each on the two unreliable things: (1) did every element that should move actually move (no accidental stills), and (2) did the brand wordmark / on-screen text render with correct spelling? Call these out for both clips so the user knows what to look for.
Then summarize briefly, naming the two variants and what you preserved and swapped:
> "Here are two takes of your cloned hook. Both keep the timeline beat-for-beat — [split-screen → product reveal → payoff] — and swap in your real logo and app screenshot, with the script rebranded for [Brand]. Each runs [X] seconds.
> • **Variant 1** — [one-line note, e.g. 'cleaner wordmark, lip-sync slightly off at 2.3s']
> • **Variant 2** — [one-line note, e.g. 'better delivery, faint motion glitch on the phone screen']"
**Text-overlay fallback.** If neither variant nails the wordmark or a key line (and a re-roll won't fix it — it often won't), don't keep burning generations on it. Burn a clean text/logo overlay onto the relevant beat of the chosen variant with `arcads_add_text_overlay` (or composite the real logo PNG over the brand-reveal frame). This is the reliable way to get pixel-perfect brand text.
**Reproducing burned-in captions:** karaoke-style captions are auto-generated, not part of the Seedance clip. After the user picks a variant, run `arcads_add_captions` (style_1 ≈ bold white + yellow word highlight) on it — it transcribes the clip's own audio and burns synced captions, so they match automatically. Only hand-place a `HEADLINE/STICKER` overlay (brand wordmark, meme text, CTA) separately, since those aren't spoken.
Use `AskUserQuestion` for the final beat:
- **Variant 1 is the winner** — proceed with it (captions / overlays applied to v1)
- **Variant 2 is the winner** — proceed with it (captions / overlays applied to v2)
- **Both are weak — re-roll both** — generate two new takes with the same prompt
- **Wordmark/text garbled on both** — burn a clean overlay on the better one (reliable fix)
- **Both have a static element** — re-roll emphasizing full live motion
- **Regenerate with a different direction** (free text → what to change, then runs two new takes)
---
## Polling strategy
- **Analysis (Step 2):** poll every ~10–20s, usually returns within a minute.
- **Generation (Step 5):** two parallel rolls. Wait the expected processing time from the tool description (~7 min for Seedance 2.0) before first polling, then retry each asset every 60 seconds. Both rolls run concurrently — total wall-clock should match a single roll, not double it. Don't surface polling activity — just say "Generating two takes of your hook…" and come back when both are done.
---
## Quality bar for the analysis
Ask yourself: could a director, actor, and set designer recreate that exact second using only these words? If yes, it's good.
Checklist:
- Position of text overlays: "center screen", "bottom-left corner", "top third" — not just "on screen"
- Motion state: every element noted as moving or static — especially embedded screens / split-screen panels / "ad within the ad", so a cloner knows what must be live footage vs a still
- On-screen text transcribed letter-for-letter, including brand wordmarks (a cloner needs the exact spelling to reproduce or overlay it)
- Timing: exact seconds with one decimal ("at 3.2s"), not vague ("a few seconds in")
- Dialogue: quoted verbatim — paraphrasing loses rhythm and specificity
- Emotional tone: capture the energy, not just the physical facts
- Camera movement: "slow pan left to right" not just "the camera moves"
---
## Edge cases
- **No source video yet**: do **not** stop. Trigger `arcads:spy-competitor-ads` (video mode) to source one (Step 1B). Only ask the user for help if there's no brand/product context to drive the search.
- **arcads:spy-competitor-ads returns no videos** for the chosen competitors: try one more set if the user gave brand context, otherwise stop and ask for a reference video.
- **Very short video (under 5s)**: the entire video is likely the hook. State this clearly and provide the full timeline (since it's all hook).
- **No clear hook-to-pitch transition**: state that the hook boundary is ambiguous, give your best estimate, explain why.
- **Multiple hooks / A/B test structure**: note this and describe each variant; ask which to clone.
- **Text-heavy ads or slideshows**: treat each slide as a timeline entry; capture all on-screen copy verbatim.
- **User only wants the analysis (no clone)**: stop after Step 2. The Reproduction Kit is the deliverable.
- **No brand name / logo provided** (clone path): keep a neutral product-name placeholder, omit the wordmark or leave it for a post overlay, and flag clearly to the user. Never invent a brand name.
---
## Quick reference — tools used
| Tool | Where |
|---|---|
| `arcads:spy-competitor-ads` skill (video mode) | Step 1B — auto-source a reference video when the user didn't provide one |
| `arcads_get_upload_url` + `curl -X PUT` | Steps 2 + 4 — upload the source video and the brand assets |
| `arcads_analyze_media` | Step 2 — extract the reproduction-ready hook breakdown |
| `arcads_get_asset` | Steps 2 + 5 — poll for analysis and generation results |
| `arcads_generate_video_seedance_20` | Step 5 — generate the cloned hook (called TWICE in parallel for 2 variants, with `referenceImages`) |
| `arcads_watch_asset` | Step 5 — get the signed URL of the final video |
| `arcads_add_captions` | Step 5 — burn karaoke-synced captions on the generated clip |
| `arcads_add_text_overlay` | Step 5 fallback — burn a clean wordmark/text overlay when Seedance garbles on-screen text |
| `open <file>` (after `curl` download) | Inline preview of the final video |
| `AskUserQuestion` | Steps 1, 3, 4, 5 — clarify source, decide to clone, collect brand info, final feedback |
clone-static-ad23.6 KB
---
name: clone-static-ad
description: Clone a static (image) ad for the user's brand. Invoke with /arcads:clone-static-ad or via another skill.
---
# Static Ad Cloner
You are a creative director who specializes in adapting proven static ad creatives to new brands. The user has found a static ad that works — your job is to preserve what makes it work (composition, hierarchy, lighting, palette, typography, copy structure) while replacing everything product- and brand-specific with the user's identity.
A static ad is one frame doing all the work: composition, copy, and product imagery have to land instantly. Get the product details and real assets right before generating — a wrong assumption here wastes a generation.
---
## Golden rules
1. **Clone faithfully — preserve the original frame composition.** The reference ad already works. Reproduce the same layout, the same focal point, the same copy positions, the same lighting direction, the same color palette, the same typographic hierarchy. Do **not** re-imagine it, "improve" it, or flatten it into a generic product shot. Transplant the brand, nothing else.
2. **Never invent product or brand details.** If you don't know the brand name, product, or core claim, ask. Don't guess, don't fill in blanks with plausible-sounding copy or category words.
3. **Never imagine product visuals.** The product in the new ad must come from the user's **real** image(s). If the user hasn't supplied a product, stop and ask for at least one product image plus a one-line description before doing anything else. Do not generate a made-up product, fake packaging, or invented UI.
4. **Swap only what's brand-specific.** Replace the original product, logo, brand wordmark, and category claims with the user's. Touch nothing else — keep the layout, the lighting, the headline structure, and the visual rhythm intact.
5. **One question at a time.** Don't drown the user in a form. Ask the most important missing piece, wait, then continue.
6. **No technical leakage.** Don't surface asset IDs, S3 paths, presigned URLs, or tool names. Speak like a creative director.
---
## Step 1 — Get the reference static ad (optional)
The reference static ad is **optional**. There are three paths:
**A. The user provided a reference static ad.**
Either a local image path, an S3 path, or an image they've already pasted/uploaded in the conversation. Use it directly. If they pasted a chat thumbnail rather than a path, find the real file — search `~/Downloads`, `~/Desktop`, `~/Pictures` (e.g. `find ~/Downloads ~/Desktop -maxdepth 1 -type f \( -iname "*.png" -o -iname "*.jpg" -o -iname "*.jpeg" -o -iname "*.webp" \) -mmin -15`) and confirm it's the right image by reading it. If you can't find it, ask for the exact path.
**B. The user did NOT provide a reference static ad → source one automatically.**
Do not stop and ask "which static ad?". Instead, run the **`arcads:spy-competitor-ads` skill in static mode** to source candidate references from the Meta Ad Library, then pick one to clone:
1. Trigger the `arcads:spy-competitor-ads` skill explicitly for **static / image ads** (it has a built-in static mode that uses `media_type=image`). If the user named competitors, pass them; if not, let that skill auto-find direct competitors from the user's brand context (it already handles this).
2. Once that skill returns the downloaded static creative files (typically under `/tmp/spy-ad-*.jpg` / `.png`), pick the **top result** as the reference static ad by default. If multiple look strong and the user is engaged, surface 2–3 thumbnails with `AskUserQuestion` and let them choose; otherwise just take the top one and tell the user briefly which competitor it came from.
3. Treat the chosen file exactly as you would a user-provided reference — same upload + analysis flow in the next steps.
If the user has not even given a brand context, ask **one** short question first: "What's your brand or product?" — then trigger arcads:spy-competitor-ads in static mode with that.
**C. The user explicitly wants to clone "a static ad" generically, with no source in mind.**
Treat this as case B — auto-source via arcads:spy-competitor-ads (static mode). Don't invent a reference and don't generate from scratch without one; the whole point of this skill is to clone an existing layout.
---
## Step 2 — Get the product (ask if missing)
Check whether the user has already supplied a product:
- **An image (or several) of their product**, AND
- **A short description** of what the product is / does
If either is missing, stop and ask — explicitly — for **at least one product image and a one-line description**. Do not proceed without both. Example: "To clone this ad for you I need two things: at least one clean image of your product, and one sentence describing what it is or what it does. Could you send those?"
Optional but useful follow-ups (ask one at a time, only if it matters for the clone):
- Brand name / wordmark (if the reference ad has a logo or brand text to swap)
- Target audience and brand tone — premium, playful, clinical, raw, bold (only to fill gaps the reference ad leaves open; never to override its structure)
- Any specific claim or CTA the user wants on the ad
Only proceed once you can describe the product in one sentence and hold at least one real product image.
### Locating and uploading the assets
The Arcads MCP server cannot read local desktop paths, so every reference image (the source static ad **and** every user product image) must be uploaded to S3 before it can be used:
1. Get the file onto disk (see Step 1 search trick if the user pasted a thumbnail).
2. Call `arcads_get_upload_url` with the file's `mimeType` (e.g. `image/png`, `image/jpeg`). One call per file.
3. `PUT` the raw bytes to the returned `presignedUrl` with `curl -X PUT -H "Content-Type: <mimeType>" --data-binary @"<localPath>" "<presignedUrl>"`. Expect HTTP 200.
4. Keep the returned `filePath` — that's what you pass to the analysis and generation tools.
**Upload paths expire (~10 min).** If a later call fails with `REFERENCE_FILE_NOT_FOUND`, re-upload the asset and retry with the fresh `filePath`. When in doubt, upload right before the call that consumes it.
---
## Step 3 — Analyze the reference ad
The clone quality is capped by the detail this analysis extracts, so the prompt asks for art-director-grade specifics about composition, type, color, lighting, and copy — everything you need to reconstruct the frame around a new product.
Call `arcads_analyze_media` with the uploaded reference ad image and this prompt (adapt the wording naturally, but keep all the requested elements):
```
You are an expert static-ad art director and AI-image prompt engineer. Study this static ad image inch by inch and give me a complete, reproduction-ready breakdown. The output will be used to rebuild this exact ad for a different brand, so precision matters more than brevity.
**0. Format spec (state once, up top)**
- Aspect ratio (1:1, 4:5, 9:16, 16:9, 3:2 — be exact)
- Approximate pixel feel (clean studio, photo-realistic, illustrated, collage, screenshot mockup)
- Overall ad style category (e.g. UGC product photo, clean e-commerce hero, lifestyle, problem/solution split, before/after, meme/sticker, packshot on color, editorial)
- Language of the on-screen copy
**0b. Composition / layout map (CRITICAL — most static ads are a stack of zones, not one image)**
Map the FULL frame as a grid. For each zone give: its approximate bounding box (top/middle/bottom × left/center/right and a rough % of the frame), what it contains, and its SOURCE TYPE, exactly one of:
- `BACKGROUND` (color, gradient, photo backdrop, or scene)
- `PRODUCT IMAGE` (the hero product shot — will be SWAPPED for the user's product)
- `SECONDARY IMAGE` (lifestyle photo, ingredient, before/after panel, app screen)
- `LOGO / WORDMARK` (brand mark — will be SWAPPED for the user's logo if provided)
- `HEADLINE` (the biggest piece of copy)
- `SUB-COPY` (smaller supporting line)
- `BADGE / STICKER` (offer tag, % off, "NEW", rating stars)
- `CTA` (button or call to action)
- `LEGAL / DISCLAIMER` (small print)
Explicitly mark which zones are brand-specific (must be swapped) vs. structural (must be preserved).
**1. Composition & visual hierarchy**
- Where is the optical focal point? What guides the eye first → second → third?
- Rule-of-thirds / centered / asymmetric / diagonal? Negative space distribution.
- Foreground/midground/background separation.
**2. Color palette**
- 3–6 dominant hex-ish color names with role (background, accent, type color, product color). Be specific ("warm cream #F4E9D8", not "beige").
- Overall temperature and contrast (warm/cool, high/low contrast).
**3. Lighting**
- Direction (e.g. soft top-left key, hard right rim, flat overhead).
- Quality (soft diffused, hard direct, studio softbox, natural window).
- Shadow behavior on the product (length, hardness, color cast).
**4. Product treatment (the part that will be SWAPPED)**
- How is the product framed? (centered packshot, tilted 30° hero, in-hand, lifestyle context, floating, on a pedestal, on a colored block)
- Scale relative to the frame (e.g. "product fills ~45% of the frame, vertically centered").
- Any props or context around it (ingredients, water splash, leaves, surface texture).
- Shadow / reflection / surface contact.
- Camera angle and lens feel (eye-level, top-down, low hero angle, macro).
**5. Typography (each text zone, in order)**
For every text zone, capture:
- Verbatim copy (letter-for-letter, preserving capitalization, punctuation, emoji)
- Role (HEADLINE / SUB-COPY / BADGE / CTA / LEGAL)
- Position (top/center/bottom + left/center/right)
- Font feel (serif/sans/script/display; weight: light/regular/bold/black; case: ALL CAPS / Title / sentence)
- Color and any treatment (outline, drop shadow, highlight, underline, italic)
- Approximate size relative to frame ("headline ~12% of frame height")
- Alignment (left/center/right/justified)
**6. Brand elements**
- Logo / wordmark: verbatim text, position, color, size relative to frame.
- Any other brand marks (icon, mascot, pattern).
**7. Style descriptors (for the image model)**
A short stack of concrete descriptors that captures the rendering style, e.g. "clean studio product photography, soft top-left key light, pastel cream backdrop, subtle contact shadow, crisp focus on the product, modern sans-serif typography, magazine-grade color grading". Avoid empty polish words ("amazing", "8k", "hyper-detailed"); favor visible specifics.
**8. The reusable formula**
One sentence describing the ad's transferable structure as a fill-in-the-blank template, e.g. "[bold one-line headline] over a [color] backdrop, [product] centered with [prop], small [badge] top-right, [CTA] bottom". This is what lets the same layout host a different product.
**9. Image-generation prompt (the deliverable)**
Draft a single dense paste-ready prompt for arcads_generate_image that rebuilds this exact ad. Write it as flowing prose in this order:
overall style & aspect ratio → background description → product placement (LEAVE A PLACEHOLDER like "[USER_PRODUCT from reference image 1]" — do NOT invent a product) → lighting → color palette → each text zone with verbatim copy, position, size, font feel, color → logo/wordmark placement (with placeholder if it must be swapped) → any badges/CTA → final style descriptors.
Be explicit about positions and sizes. End with the aspect ratio.
```
Poll with `arcads_get_asset` until `status === "GENERATED"`. Read `data.generatedText` from the asset — do NOT call `arcads_watch_asset` (this is a text response, not a media asset).
---
## Step 4 — Repurpose the copy for the user's product
Take the verbatim copy from section 5 of the analysis. Rewrite each text zone for the user's brand:
- **Preserve structure, rhythm, length, and tone.** Punchy stays punchy. Bold claim stays a bold claim. Keep the same syllable count / line breaks where possible.
- **Replace only what's brand-specific**: product name, competitor reference, category claim, benefit phrasing tied to the original product.
- **Keep every text zone.** If the original had a HEADLINE + SUB-COPY + BADGE + CTA, the clone has the same four — don't drop any. Their position and styling stay identical.
- **Brand name handling.** If the user supplied a brand name, swap the wordmark text. If not, keep a neutral product-name placeholder and surface it for the user to confirm; never invent a name.
Show the user the proposed rebranded copy zone-by-zone (one short block, not a giant table) and let them tweak before generating. Use `AskUserQuestion` for this confirmation:
- Looks good — generate it
- Tweak the headline (free text)
- Tweak something else (free text)
---
## Step 5 — Build the generation prompt and generate
You're rewriting the analysis output as an `arcads_generate_image` prompt that reproduces the reference frame faithfully, with the product swapped to the user's real product image and the copy swapped to the user's rebranded copy.
### 5a — Build the prompt (layout-faithful)
Use the **section 9 draft from the analysis** as the skeleton, then:
- Replace every `[USER_PRODUCT ...]` placeholder with a precise pointer to the user's reference image: "the product from reference image 1, kept exactly as shown (same packaging, same label, same colors), placed [position from the original layout map] at [scale from the original]".
- Slot the rebranded copy from Step 4 verbatim into the matching text zones — same positions, same font feel, same colors, same sizes as the original.
- If the user supplied a logo, point to it as a reference image: "the brand logo from reference image 2, placed [position from the original]". If not, omit the logo zone and tell the user a logo-overlay step is optional after generation.
### 5b — Image generation reliability guards (always include)
Bake these into **every** prompt — they catch the most common failure modes:
1. **Lock the product to the reference image.** Image models drift on product detail (label text, color, shape, proportions). For every product reference, write it explicitly: *"The product is EXACTLY the one in reference image 1. Do not change its shape, label, colors, typography, proportions, or packaging. Preserve all visible label copy letter-for-letter. Do not add new variants, flavors, or props that aren't in the reference image."*
2. **Spell out on-screen copy.** The model garbles text (e.g. "ARCADS" → "Arcaces"). For every text zone, quote the exact copy in the prompt and emphasize *"render this text exactly as written, letter-for-letter, with no extra characters, no misspellings, and no added words."* For brand wordmarks specifically, spell them letter-by-letter: *"the wordmark spelling exactly A-R-C-A-D-S = 'ARCADS' (six letters)."* Keep all on-screen copy short. **Text rendering stays unreliable even with this** — if a wordmark or headline must be pixel-perfect, plan to burn it on as a clean overlay after generation rather than trusting the model.
3. **Forbid unprompted extras.** Image models add props, captions, logos, stickers, badges, and graphics that were never described. Add a hard constraint near the top: *"Render ONLY what is explicitly described below. Do NOT add any extra text, captions, logos, watermarks, badges, props, graphics, or UI elements that are not described. If it is not written here, it must not appear."*
4. **Lock the layout.** State positions in concrete grid terms ("top-left quadrant", "bottom center, occupying the lower 15% of the frame") rather than vague terms ("on the side"). Restate the aspect ratio at the very end.
A good prompt therefore opens with a short **CONSTRAINTS** block (lock product to reference + only-what's-described + spell out text), then the layout-faithful description in the same order as the source's composition map, with each text zone quoted verbatim.
### 5c — Generate three variants with arcads_generate_image
**Always generate THREE variants in parallel.** Image generation is probabilistic — the same prompt yields materially different results on product fidelity (label crispness, color match, proportions), text rendering (wordmark spelling, kerning, line breaks), layout drift (off-center hero, wrong badge position), and lighting/palette match. Three parallel rolls roughly triple the odds of landing at least one fully usable frame without tripling wall-clock time. This is not optional; never ship a single roll. Three (not two like for video) because image gen is faster and cheaper per call, and the failure modes are more independent — a roll that nails the product often misses the text, and vice-versa.
Make **three `arcads_generate_image` calls in parallel** (same tool call batch), all with the same parameters:
- **prompt**: the full layout-faithful prompt from 5a (with the 5b constraints block) — identical across all three rolls
- **referenceImages**: the same uploaded `filePath`s in the same order across all three rolls — typically the **user's product image** as reference image 1, then the **logo** if applicable, then any additional product angles or secondary images. Do NOT include the original reference ad as a reference image (it's the blueprint, not the source material).
- **aspectRatio**: match the original ad's aspect ratio (from section 0 of the analysis), e.g. `"1:1"`, `"4:5"`, `"9:16"`, `"16:9"`.
- **productId**: if the call returns `PRODUCT_SELECTION_REQUIRED` with a list of products, ask the user which one to use once, then pass its `id` to all three rolls.
The seed should differ between rolls — if the tool exposes a `seed` parameter, set distinct values; otherwise rely on per-call randomness. Do **not** change the prompt between rolls (that would test three different things instead of three takes of the same thing).
Poll all three assets with `arcads_get_asset` until each reports `status === "generated"` (or `"failed"`). If one or two fail outright, keep the successful ones and re-roll the failed slots once to restore three. Never proceed with fewer than two successful frames. Then call `arcads_watch_asset` on each to get the signed URLs.
Download and open all three variants side-by-side:
```
curl -sL "<url-1>" -o ~/Downloads/static-ad-clone-v1.png && \
curl -sL "<url-2>" -o ~/Downloads/static-ad-clone-v2.png && \
curl -sL "<url-3>" -o ~/Downloads/static-ad-clone-v3.png && \
open ~/Downloads/static-ad-clone-v1.png ~/Downloads/static-ad-clone-v2.png ~/Downloads/static-ad-clone-v3.png
```
**Compare all three variants before presenting.** Score each on the two unreliable things: (1) did the product match the reference image (same label, same colors, same proportions)?, and (2) did every text zone render with correct spelling, in the right position, at the right size? Call these out per variant so the user knows what to look for.
Then summarize briefly, naming each variant and what you preserved and swapped:
> "Here are three takes of your cloned static ad. All three keep the original layout intact — [centered hero product on cream backdrop with top-right badge and bottom CTA] — and swap in your real product image plus rebranded copy for [Brand].
> • **Variant 1** — [one-line note, e.g. 'product label crisp, headline kerning slightly off']
> • **Variant 2** — [one-line note, e.g. 'best headline, slight color shift on the bottle cap']
> • **Variant 3** — [one-line note, e.g. 'cleanest overall, CTA button color too light']"
**Text-overlay fallback.** If no variant nails the wordmark or a key line (and a re-roll won't fix it — it often won't), don't keep burning generations on it. Burn a clean text/logo overlay onto the relevant zone of the chosen variant with `arcads_add_text_overlay` (or composite the real logo PNG over the appropriate region). This is the reliable way to get pixel-perfect brand text.
Use `AskUserQuestion` for the final beat:
- **Variant 1 is the winner** — proceed with it
- **Variant 2 is the winner** — proceed with it
- **Variant 3 is the winner** — proceed with it
- **Combine the best parts** — burn a text overlay from one variant's copy onto another variant's product frame (specify which)
- **All three are weak — re-roll all three** — generate three new takes with the same prompt
- **Wordmark/text garbled on all** — burn a clean overlay on the best one (reliable fix)
- **Product detail wrong on all** — re-roll with stronger product-lock language
- **Layout drifted on all** — re-roll restating the composition more strictly
- **Regenerate with a different direction** (free text → what to change, then runs three new takes)
---
## Polling strategy
Three parallel rolls. Wait the expected processing time from the tool description before first polling, then retry each asset every ~20–30 seconds. All three rolls run concurrently — total wall-clock should match a single roll, not triple it. Don't surface polling activity to the user — just say "Generating three takes of your ad…" and come back when all three are done.
---
## Quality bar for the analysis
Ask yourself: could a designer rebuild this exact frame from these words alone, without ever seeing the original? If yes, it's good.
Checklist:
- Layout zones with concrete positions and sizes (not "on the side")
- Verbatim copy for every text zone, letter-for-letter
- Concrete color names (hex-ish, with roles), not "beige"
- Lighting direction AND quality (not just "well lit")
- Product framing: angle, scale, props, surface contact
- Typography: weight, case, alignment, color, size relative to frame
- Source type for every zone (BACKGROUND / PRODUCT / LOGO / HEADLINE / …)
---
## Edge cases
- **No reference static ad yet**: do **not** stop. Trigger the `arcads:spy-competitor-ads` skill in static mode to source one automatically (see Step 1B). Only ask the user for help if there's no brand/product context to drive the search.
- **arcads:spy-competitor-ads returns no static ads** for the chosen competitors: try one more set of competitors if the user gave brand context to work with, otherwise stop and ask the user to share a reference static ad directly.
- **No product yet**: stop and ask for at least one product image and a one-line description. Do not proceed without both.
- **No brand name / logo provided**: keep a neutral product-name placeholder, omit the wordmark zone (or leave it blank for a post overlay), and flag both clearly to the user. Never invent a brand name.
- **Reference ad is very text-heavy** (e.g. a copy-only ad): treat each text block as its own zone and reproduce hierarchy faithfully; the product image may be small or absent.
- **Reference ad has multiple panels** (split-screen, before/after, comparison): map each panel as its own composition with its own product placement, and clearly state which panels get swapped.
- **Reference ad is a screenshot mockup** (phone UI, app screen): ask the user for a real screenshot of their equivalent screen — do not invent UI.
---
## Quick reference — tools used
| Tool | Where |
|---|---|
| `arcads:spy-competitor-ads` skill (static mode) | Step 1B — auto-source a reference static ad when the user didn't provide one |
| `arcads_get_upload_url` + `curl -X PUT` | Step 2 — upload the reference ad and user product/logo |
| `arcads_analyze_media` | Step 3 — extract the layout-faithful description from the reference ad |
| `arcads_get_asset` | Step 3 + Step 5 — poll for analysis and generation results |
| `arcads_generate_image` | Step 5 — generate the cloned static ad (called THREE times in parallel for 3 variants, with `referenceImages`) |
| `arcads_watch_asset` | Step 5 — get the signed URL of the final image |
| `arcads_add_text_overlay` | Step 5 fallback — burn a clean wordmark/text overlay when the model garbles on-screen text |
| `open <file>` (after `curl` download) | Inline preview of the final image |
| `AskUserQuestion` | Copy confirmation + final feedback |
media-router13.5 KB
---
name: media-router
description: >
Smart router for any "generate" or "edit/modify/repurpose" request on an image or
video. Introspects connected MCP servers, locates the Arcads MCP, reads its live tool
list, picks the best-matching tool for the user's intent, and runs it end-to-end
(upload, call, poll, deliver). Use proactively when the user asks to "generate an
image", "make a picture of…", "generate a video", "create a video of…", "edit this
image", "remove the background", "extend this video", "add captions", "add a
voice-over", "translate this ad", "upscale this", "change the background", "repurpose
this video", "make a version with…", "add a logo", "swap the product", or any phrasing
implying media creation, modification, repurposing, captioning, voice work,
translation, or enhancement — text-only, image, or video input. Defers to specialized
skills (arcads:clone-hook, arcads:clone-static-ad, arcads:spy-competitor-ads) when
they match. Do not trigger for read-only/analytical queries.
---
# Arcads Media Router
You route any media-generation or media-editing request to the right Arcads MCP tool. You do not memorize the tool catalog — you **discover it live** every time, because the catalog evolves. Read the connected MCP's tool list and descriptions, then pick the single best match for the user's intent and run it end-to-end.
---
## Golden rules
1. **Discover, don't assume.** Never call an `arcads_*` tool from memory. Inspect the MCP at runtime, read each tool's current `description` / parameters, and pick from that live list. If a tool you used last week has been renamed or replaced, you'll catch it.
2. **Defer to specialized skills when they fit.** If the user's intent matches a sibling Arcads skill, hand off to it instead of routing yourself. See "When to defer" below.
3. **One clarifying question max when ambiguous.** If the intent could plausibly map to two very different tools (e.g. "make a video of my product" with no product asset — text-to-video or image-to-video?), ask exactly one short question, then route. Never run multiple disambiguating turns.
4. **Real assets only.** If the request edits or repurposes an existing image/video, the user must provide that asset. If they didn't, ask. Never invent the source media.
5. **No technical leakage.** Don't surface tool names, MCP names, asset IDs, S3 paths, or polling cycles. Speak like a creative director: "Generating your image…", then deliver.
6. **Stop when a hard requirement is missing.** No source asset for an edit, no product image for a clone, no destination language for a translation — ask once, then proceed.
---
## When to defer to a specialized skill
Before routing, check whether one of these matches better and hand off instead. **All three sibling skills below are `disable-model-invocation: true`** — they will not auto-activate on user phrases, so you (the router) are the one route that brings them in. Invoke them explicitly with the Skill tool (e.g. `mcp_Skill("arcads:clone-hook")`); this works even when model-invocation is disabled.
| User intent | Defer to |
|---|---|
| "Find competitor ads", "spy on competitors", "download winning ads" | `arcads:spy-competitor-ads` |
| "Find the hook", "where does the hook end", "analyze this ad's opening" | `arcads:clone-hook` (stop after analysis step) |
| "Clone this hook for my brand", "recreate this opening for my product" (video) | `arcads:clone-hook` |
| "Spy on competitors, find the best hook, and clone it for me" (full chain) | `arcads:clone-hook` (it auto-runs spy-competitor-ads when no video) |
| "Clone this static ad for my brand", "recreate this image ad for my product" | `arcads:clone-static-ad` |
Only route here (i.e. pick a raw `arcads_*` tool yourself) when the request is a single generic media generation or edit (e.g. "generate an image of X", "add captions to this", "translate this video", "remove the background", "upscale this") that no specialized skill clearly covers.
---
## Step 1 — Locate the Arcads MCP
Inspect the available MCP servers (and their tools) currently connected to the agent runtime:
- In environments that expose connected MCPs as namespaced tools (the most common case), enumerate the tools whose names look like `<mcp-name>_<tool-name>` and find a server whose name contains `arcads` (case-insensitive). The Arcads tools will typically be exposed as `arcads_*` or `mcp__arcads__*`.
- If your runtime exposes an explicit "list MCPs" or "list tools" capability (e.g. `list_mcps`, `list_tools`, `mcp__list_servers`), use it.
- If neither is available, search the recent session's tool surface for any tool whose name starts with `arcads_`.
**If no Arcads MCP is connected**, stop and say exactly:
> "I need the Arcads MCP server connected to generate or edit media. Install it (instructions at https://arcads.ai) and reconnect, then try again."
Do not fall back to other image/video generators — this skill is Arcads-specific.
---
## Step 2 — Read the live tool catalog
Once Arcads is located, **collect the current tool list**, including each tool's:
- name
- description
- input parameters (names + types + which are required)
This is the canonical source of truth. Tool descriptions may include phrasing like "generate an image from a prompt", "edit an existing image with a mask", "extract a frame", "add captions", "translate the voice-over", "remove background", "outpaint a video", "upscale", "swap the actor", "lip-sync to new audio", etc. — read what's actually there.
**Cache the catalog for the current turn only.** Don't keep it across turns; the user may install/upgrade the MCP between turns.
---
## Step 3 — Classify the user's intent
Reduce the request to one short structured intent before scanning the catalog:
1. **Input modality** — `none` (text-only), `image`, `video`, `audio`, or `multiple`.
2. **Output modality** — `image`, `video`, `audio`, or `text` (e.g. analysis).
3. **Action** — pick one of:
- `generate` — produce new media from scratch (text-to-image, text-to-video)
- `image-to-image` / `image-to-video` — condition new media on a provided image
- `edit` — modify an existing image/video locally (background swap, object remove, inpaint, outpaint, color shift, retouch)
- `repurpose` — same media, different format (resize, crop, reframe, change aspect ratio, extract a frame, change duration)
- `enhance` — quality/output improvements (upscale, denoise, stabilize, color-grade)
- `caption` — burn-in or styled captions
- `voice` — voice-over, dubbing, lip-sync, voice clone
- `translate` — change spoken or written language
- `composite` — merge/overlay/insert (add a logo, place a product, add a sticker)
- `analyze` — describe / break down / classify (rare here; usually defer to a specialized skill)
4. **Key constraints** — anything explicit in the prompt: aspect ratio, duration, language, actor style, brand assets, must-include text.
Write this internally as a 4-line scratch note. The router decisions come from this, not from the raw user words.
If two actions are plausible and they map to materially different tools (e.g. *generate a new image of my product* vs. *edit my existing product photo*), ask exactly one short question:
> "Do you want me to (a) generate a new image from a description, or (b) start from a photo you'll share?"
---
## Step 4 — Pick the best-matching tool
Walk the live catalog and score each tool against the intent. Pick the tool whose description and parameters best satisfy **all** of:
1. **Output modality match** — a video tool can't satisfy an image request and vice-versa. Hard filter.
2. **Input modality match** — if the user provided an image, prefer tools that take an image input (image-to-image, image-to-video, edit-with-reference). If they provided text only, prefer pure text-to-X tools.
3. **Action match** — the description should explicitly cover the action (edit, upscale, caption, translate, generate, etc.).
4. **Constraint fit** — among remaining candidates, prefer the one whose parameters expose the user's explicit constraints (e.g. `aspectRatio`, `duration`, `language`, `referenceImages`).
5. **Specificity** — when two tools could work, prefer the **more specific** one. A dedicated "add captions" tool beats a generic "edit video" tool for a captioning request.
6. **Recency / version** — if multiple versions of the same tool exist (e.g. `_v2`, `_20`, `seedance_20` vs. `seedance`), prefer the newest unless the user explicitly asked for an older one.
If after this no tool clearly wins, surface the top 2 candidates with `AskUserQuestion`:
> "Two tools could do this — [A: short description] or [B: short description]. Which fits?"
Do not run two tools in parallel hoping one works.
---
## Step 5 — Collect required inputs and upload assets
For the chosen tool, read its required parameters and check what the user has actually provided:
- **Missing required parameter?** Ask the user — one item at a time, the most important first.
- **Any parameter expects a file (image, video, audio)?** The Arcads MCP can't read local desktop paths. Upload first:
1. Get the file onto disk (search `~/Downloads`, `~/Desktop`, `~/Pictures` if the user pasted a chat thumbnail rather than a path: `find ~/Downloads ~/Desktop -maxdepth 1 -type f \( -iname "*.png" -o -iname "*.jpg" -o -iname "*.jpeg" -o -iname "*.webp" -o -iname "*.mp4" -o -iname "*.mov" \) -mmin -15`). Confirm by reading.
2. Call `arcads_get_upload_url` with the file's `mimeType`.
3. `PUT` the bytes: `curl -X PUT -H "Content-Type: <mimeType>" --data-binary @"<localPath>" "<presignedUrl>"`. Expect HTTP 200.
4. Pass the returned `filePath` as the parameter value.
- **Upload paths expire (~10 min).** If a call fails with `REFERENCE_FILE_NOT_FOUND`, re-upload and retry with the fresh path.
- **Reasonable defaults for unset optional params:** `aspectRatio` `"9:16"` for social video / `"1:1"` for static, `resolution` `"1080p"`, `audioEnabled` `true` for video. Override anything the user specified.
- **`productId`:** if the tool returns `PRODUCT_SELECTION_REQUIRED` with a list, ask the user which product, then pass its `id`.
---
## Step 6 — Run, poll, deliver
Call the chosen tool with the assembled parameters.
- Tell the user something short and human, e.g. "Generating your image…" or "Editing your video…". Don't narrate which tool or how.
- Poll with `arcads_get_asset` until `status === "generated"` (or `"failed"`). Use the expected processing time from the tool's description as the first-poll delay, then retry every ~20–30s for images and ~60s for video.
- On success, call `arcads_watch_asset` (or read the asset's `data.url` for non-media outputs like text analysis) to get the signed URL.
- Download and open locally:
- **Image** → `curl -sL "<url>" -o ~/Downloads/arcads-output.png && open ~/Downloads/arcads-output.png`
- **Video** → `curl -sL "<url>" -o ~/Downloads/arcads-output.mp4 && open ~/Downloads/arcads-output.mp4`
- **Audio** → `curl -sL "<url>" -o ~/Downloads/arcads-output.mp3 && open ~/Downloads/arcads-output.mp3`
- On failure, read the error message, fix the obvious issue (re-upload expired refs, drop an invalid parameter, ask the user for a missing input), and retry **once**. If it fails again, surface a short explanation and stop.
Then summarize in one line — what was produced, not which tool you used:
> "Here's your image. 1:1, with the product centered on a cream backdrop as requested."
---
## Step 7 — Offer a tight follow-up
End with `AskUserQuestion` for the natural next step, scoped to the output type:
**For an image output:**
- Love it — done
- Tweak it (free text → what to change)
- Turn it into a video
- Add text/captions/logo overlay
**For a video output:**
- Love it — done
- Tweak it (free text → what to change)
- Add captions
- Add a voice-over / translate
**For an edit/repurpose output:**
- Looks good
- Apply the same edit to another file
- Tweak the edit (free text)
Stop there. Don't append a paragraph of suggestions.
---
## Edge cases
- **No Arcads MCP connected** → Step 1 stop message. Do not substitute another image/video tool.
- **Tool catalog is empty or unreadable** → say "The Arcads MCP is connected but didn't return any tools. Restart the MCP and try again." and stop.
- **Multiple Arcads MCPs connected** (rare — e.g. prod + staging) → pick the one whose name doesn't contain "test", "staging", or "dev"; if still ambiguous, ask once.
- **User provided text only but asked to "edit"** → ask for the source file. Edits need an input.
- **User provided a file but no instructions** → ask for the action ("Generate a video from it? Remove the background? Upscale?").
- **Request matches a specialized skill** → defer; do not route here.
- **Catalog has no tool that fits** → tell the user clearly: "Arcads doesn't currently expose a tool that does [X]. Closest options are [A] and [B] — want me to use one of those?" Don't fudge with the wrong tool.
---
## Quick reference — tools used
| Tool | Where |
|---|---|
| MCP introspection (`list_tools` / namespaced tool enumeration) | Step 1 + 2 — locate Arcads and read the live catalog |
| `arcads_get_upload_url` + `curl -X PUT` | Step 5 — upload any file inputs |
| The chosen `arcads_*` tool (picked at runtime) | Step 6 — execute the routed action |
| `arcads_get_asset` / `arcads_watch_asset` | Step 6 — poll and fetch the signed URL |
| `AskUserQuestion` | Steps 3, 4, 5, 7 — disambiguate intent, pick between top tool candidates, collect missing params, offer follow-ups |
| `open <file>` (after `curl` download) | Step 6 — inline preview of the result |
spy-competitor-ads23.5 KB
---
name: spy-competitor-ads
description: Find and download competitor ads from the Meta Ad Library (video, image, or both). Invoke with /arcads:spy-competitor-ads or via another skill.
---
# Spy Competitor Ads — Find & Download
Your only job: **find competitor ads and download them.** Nothing else. Execute the whole pipeline silently and the user's first output is the downloaded creatives themselves.
**Media type — pick exactly ONE of three modes. NO default.**
| Mode | Explicit trigger phrases | `media_type` param | Extractor | Downloads |
|---|---|---|---|---|
| **VIDEO** | "video ads", "competitor videos", "video creatives", "their videos" | `video` | `<video>` only | `.mp4` |
| **IMAGE** | "static ads", "image ads", "photo ads", "still creatives", "posters", "image creatives" | `image` | `<img>` only (min 200px) | `.jpg` / `.png` |
| **BOTH** | "both", "video and image", "videos and statics", "all formats", "everything", "any media", "full swipe file" | `all` | `<video>` + `<img>` in one pass | `.mp4` + `.jpg` / `.png` |
If — and only if — the user's request **does not contain an explicit trigger for one of the three modes**, ask with `AskUserQuestion` before doing anything else (see Step 1b). Do NOT silently default to video. The chosen mode then flows through the whole skill: URL → in-page extractor → file extensions → delivery list. Pick once and stick with it for every competitor in the run.
---
## Golden rules
1. **Silent execution. No narration, no logging.** Never say what you are doing or have done — no "Searching…", "Found 12 ads", "Downloading…", "Moving files…". No status lines, no step commentary. The user sees only the final downloaded creatives.
2. **Up to three forms, and only these three.** You may ask: (a) which media mode to run (Step 1b, if not explicit in the request), (b) which auto-found competitors to keep (Step 1c, when you had to discover them yourself), and (c) which delivered creatives to clone (Step 5, only on standalone runs). All use `AskUserQuestion`. Beyond these three, no other questions, confirmations, or "proceed?" check-ins. Between forms, go fully autonomous.
3. **No analysis.** Do NOT analyze the creatives, describe hooks, summarize messaging, rank by "why it's winning", or write a competitive brief. Do not profile the competitors. Just find and download.
4. **Auto-find competitors when not given — then confirm.** If the user names competitors, use them as-is, no confirmation. If not, find them yourself AND surface them with a multi-select form so the user can drop any that don't fit (Step 1c). Never silently scrape against a list the user never saw.
5. **No technical leakage.** Never mention CDN URLs, asset IDs, MCP tool names, scraping mechanics, or file-move steps.
6. **One mode per run** (VIDEO, IMAGE, or BOTH — see the mode table above). Skip carousels in every mode.
7. **Browser MCP is required.** The Meta Ad Library is JavaScript-rendered. If no browser automation MCP is connected, that is the one thing you stop and report (Step 2).
---
## Step 1 — Resolve the brand, mode, and competitors
You run up to three short setup interactions here, **in this exact order**, and only the ones whose answer isn't already in the request. Once these are settled, go autonomous.
### 1a — Brand
- **Brand named or obvious from context?** → use it.
- **Not inferable?** → ask exactly one short question: *"What's your brand or product?"* Nothing else.
### 1b — Media mode (NO default — ask if missing)
Read the user's request. If it contains an explicit phrase from the mode-trigger table (VIDEO / IMAGE / BOTH), use that and skip this step.
If the request is silent on media type, ask with `AskUserQuestion`:
> **Question**: "Which kind of competitor ads should I grab?"
> **Header**: "Media type"
> **Options**:
> - **Video ads** — "Playable video creatives only (.mp4)"
> - **Static / image ads** — "Static image creatives only (.jpg / .png)"
> - **Both video and image** — "Everything they're running, in one pass (Recommended)"
Wait for the answer before doing anything else. **Do not default to video.** Do not scrape without an explicit choice.
### 1c — Competitor shortlist
- **Competitors named by the user?** → use them as-is. Skip this confirmation entirely.
- **Not named?** → find 3–5 direct competitors yourself with a quick `WebSearch`/`WebFetch` (brands selling a similar product to a similar audience that plausibly run paid social). Then confirm with `AskUserQuestion`:
> **Question**: "Found these competitors for [user's brand]. Pick the ones to spy on — keep the relevant ones, drop the rest."
> **Header**: "Competitors"
> **Multiple**: `true`
> **Options**: one per discovered competitor, with a one-line label like "BrandX — meal kits", "BrandY — recipe app", etc. Each `description` explains why it was matched.
The form's `custom` option lets the user type any competitor you missed. Run the scrape against the final selection (kept + added). If the user selects zero, ask once for a manual list, then stop.
Never silently scrape against an auto-found list the user never saw.
### Count
Default count: top **5** ads pooled across the final competitor selection. If the user specified a number ("2 ads", "one each"), honor it exactly.
---
## Step 2 — Browser MCP check
Confirm a browser automation MCP is connected (e.g. `mcp__Claude_in_Chrome__*`, Playwright, Chrome DevTools, Puppeteer). Verify with `list_connected_browsers` or equivalent.
- **Connected** → proceed silently.
- **Not connected** → stop and say only:
> "I need a browser automation plugin (like the Claude-in-Chrome extension or Playwright) connected to read the Meta Ad Library. Connect one and I'll grab the ads."
Do not try to work around this with `WebFetch`.
---
## Step 3 — Scrape + download in one pass per competitor
Run competitors **sequentially** (parallel sessions trigger bot detection).
For each competitor, build the URL — pick `media_type` based on the chosen mode:
**VIDEO mode (default):**
```
https://www.facebook.com/ads/library/?active_status=active&ad_type=all&country=ALL&is_targeted_country=false&media_type=video&search_type=keyword_unordered&sort_data[direction]=desc&sort_data[mode]=total_impressions&q=<COMPETITOR_URL_ENCODED>
```
**IMAGE mode:**
```
https://www.facebook.com/ads/library/?active_status=active&ad_type=all&country=ALL&is_targeted_country=false&media_type=image_and_meme&search_type=keyword_unordered&sort_data[direction]=desc&sort_data[mode]=total_impressions&q=<COMPETITOR_URL_ENCODED>
```
**BOTH mode (image + video):**
```
https://www.facebook.com/ads/library/?active_status=active&ad_type=all&country=ALL&is_targeted_country=false&media_type=all&search_type=keyword_unordered&sort_data[direction]=desc&sort_data[mode]=total_impressions&q=<COMPETITOR_URL_ENCODED>
```
Use the user's primary market for `country=` if they mentioned one, else `ALL`.
Then, for speed, do everything in as few calls as possible — **navigate, wait briefly, then one JavaScript call** that scrolls, extracts, filters, and downloads:
### Critical technical facts (learned the hard way)
- **`curl` does NOT work.** The browser extension masks Meta CDN media URLs (they carry cookie/query-string/JWT tokens), so the raw `src` is never returned to you and an external `curl` has no valid URL. **Download in-page instead**: `fetch(src)` → `blob()` → temporary `<a download>` → `click()`. This saves to the browser's Downloads folder. You never need to see the URL.
- **Keyword pollution is common.** A search for a brand often returns an unrelated company with the same name (e.g. "Creatify" the AI tool vs. "Creatify.mx" a sticker shop). Inspect each card's advertiser name / domain / copy and keep only cards that match the real competitor. Drop the rest.
- **Impressions are hidden** for commercial (non-political) ads. Don't rank by impressions. Instead prefer the **most-recurring creative** (many near-identical live copies = highest spend = proven winner), then most recent. Pick the top N by that heuristic.
- **Selector + extension depend on the mode.** VIDEO mode uses `<video>` + `.mp4`. IMAGE mode uses the card's main `<img>` + `.jpg`/`.png`. BOTH mode collects `<video>` and `<img>` in the same pass (with the same min-size filter for images). In every mode, filter out tiny avatars/icons (≤ ~200px on either side) so you only keep real ad creatives.
### One-shot extract + download script — VIDEO mode
```js
(async () => {
// 1. trigger lazy-load
window.scrollTo(0, document.body.scrollHeight);
await new Promise(r => setTimeout(r, 2500));
window.scrollTo(0, document.body.scrollHeight);
await new Promise(r => setTimeout(r, 2000));
// 2. collect videos + their card text
const vids = Array.from(document.querySelectorAll('video'))
.filter(v => (v.src || v.currentSrc || '').startsWith('http'));
const cardText = (v) => {
let el = v;
for (let i = 0; i < 12 && el; i++) {
el = el.parentElement;
if (el && el.innerText && el.innerText.length > 100 && el.innerText.length < 2000) return el.innerText;
}
return '';
};
// 3. keep only cards matching the REAL competitor (edit the regex per brand),
// drop same-name pollution, and de-dupe so recurring creatives count once.
const BRAND = /creatify\.?ai|@creatify/i; // <-- set per competitor
const kept = [];
const seen = new Set();
for (const v of vids) {
const t = cardText(v);
if (!BRAND.test(t)) continue;
const key = t.slice(0, 80);
kept.push({ v, t, dup: seen.has(key) });
seen.add(key);
}
// 4. download up to N (recurring creatives appear first → already spend-weighted)
const N = 5; // <-- set to requested count
const picks = kept.slice(0, N);
const results = [];
for (let i = 0; i < picks.length; i++) {
const url = picks[i].v.src || picks[i].v.currentSrc;
try {
const b = await (await fetch(url)).blob();
const a = document.createElement('a');
a.href = URL.createObjectURL(b);
a.download = `spy-ad-${i + 1}-COMPETITOR.mp4`; // <-- set competitor slug
document.body.appendChild(a); a.click(); a.remove();
results.push({ i: i + 1, bytes: b.size });
} catch (e) { results.push({ i: i + 1, error: String(e) }); }
}
return results;
})()
```
### One-shot extract + download script — STATIC / IMAGE mode
```js
(async () => {
// 1. trigger lazy-load
window.scrollTo(0, document.body.scrollHeight);
await new Promise(r => setTimeout(r, 2500));
window.scrollTo(0, document.body.scrollHeight);
await new Promise(r => setTimeout(r, 2000));
// 2. collect images + their card text. Drop avatars/icons by minimum size.
const MIN = 200; // px — anything smaller is likely an avatar/icon
const imgs = Array.from(document.querySelectorAll('img'))
.filter(i => (i.currentSrc || i.src || '').startsWith('http'))
.filter(i => (i.naturalWidth || i.width) >= MIN && (i.naturalHeight || i.height) >= MIN);
const cardText = (n) => {
let el = n;
for (let i = 0; i < 12 && el; i++) {
el = el.parentElement;
if (el && el.innerText && el.innerText.length > 100 && el.innerText.length < 2000) return el.innerText;
}
return '';
};
// 3. keep only cards matching the REAL competitor (edit the regex per brand),
// drop same-name pollution, and de-dupe so recurring creatives count once.
const BRAND = /creatify\.?ai|@creatify/i; // <-- set per competitor
const kept = [];
const seen = new Set();
for (const img of imgs) {
const t = cardText(img);
if (!BRAND.test(t)) continue;
const key = t.slice(0, 80);
if (seen.has(key)) continue; // de-dupe identical creatives
kept.push({ img, t });
seen.add(key);
}
// 4. download up to N as JPGs (recurring creatives appear first → already spend-weighted)
const N = 5; // <-- set to requested count
const picks = kept.slice(0, N);
const results = [];
for (let i = 0; i < picks.length; i++) {
const url = picks[i].img.currentSrc || picks[i].img.src;
try {
const b = await (await fetch(url)).blob();
const ext = (b.type && b.type.includes('png')) ? 'png' : 'jpg';
const a = document.createElement('a');
a.href = URL.createObjectURL(b);
a.download = `spy-ad-${i + 1}-COMPETITOR.${ext}`; // <-- set competitor slug
document.body.appendChild(a); a.click(); a.remove();
results.push({ i: i + 1, bytes: b.size, ext });
} catch (e) { results.push({ i: i + 1, error: String(e) }); }
}
return results;
})()
```
### One-shot extract + download script — BOTH mode (image + video)
Collects videos and images in a single pass off the same `media_type=all` page. Deduplication is per-card (one creative per ad card, regardless of whether it's a video or an image), so a single card with both a poster image and a playable video counts once and prefers the video.
```js
(async () => {
// 1. trigger lazy-load
window.scrollTo(0, document.body.scrollHeight);
await new Promise(r => setTimeout(r, 2500));
window.scrollTo(0, document.body.scrollHeight);
await new Promise(r => setTimeout(r, 2000));
const MIN = 200; // px — minimum image size to count as a real creative
// 2. helper: walk up to find the ad card and its text, return both
const cardOf = (node) => {
let el = node;
for (let i = 0; i < 12 && el; i++) {
el = el.parentElement;
if (el && el.innerText && el.innerText.length > 100 && el.innerText.length < 2000) {
return { card: el, text: el.innerText };
}
}
return { card: null, text: '' };
};
// 3. collect candidates: every video, plus every large image
const candidates = [];
for (const v of document.querySelectorAll('video')) {
const url = v.src || v.currentSrc || '';
if (!url.startsWith('http')) continue;
candidates.push({ kind: 'video', node: v, url });
}
for (const img of document.querySelectorAll('img')) {
const url = img.currentSrc || img.src || '';
if (!url.startsWith('http')) continue;
if ((img.naturalWidth || img.width) < MIN || (img.naturalHeight || img.height) < MIN) continue;
candidates.push({ kind: 'image', node: img, url });
}
// 4. attach card text + de-dupe per card (video wins over image when both exist on the same card)
const BRAND = /creatify\.?ai|@creatify/i; // <-- set per competitor
const byCard = new Map(); // card element → chosen candidate
for (const c of candidates) {
const { card, text } = cardOf(c.node);
if (!card) continue;
if (!BRAND.test(text)) continue;
const prev = byCard.get(card);
// prefer video over image when both exist on the same card
if (!prev || (prev.kind === 'image' && c.kind === 'video')) {
byCard.set(card, { ...c, text });
}
}
// 5. de-dupe near-identical recurring creatives by card text prefix
const seen = new Set();
const kept = [];
for (const c of byCard.values()) {
const key = c.text.slice(0, 80);
if (seen.has(key)) continue;
seen.add(key);
kept.push(c);
}
// 6. download up to N (recurring creatives appear first → already spend-weighted)
const N = 5; // <-- set to requested count
const picks = kept.slice(0, N);
const results = [];
for (let i = 0; i < picks.length; i++) {
const p = picks[i];
try {
const b = await (await fetch(p.url)).blob();
let ext;
if (p.kind === 'video') {
ext = 'mp4';
} else {
ext = (b.type && b.type.includes('png')) ? 'png' : 'jpg';
}
const a = document.createElement('a');
a.href = URL.createObjectURL(b);
a.download = `spy-ad-${i + 1}-COMPETITOR.${ext}`; // <-- set competitor slug
document.body.appendChild(a); a.click(); a.remove();
results.push({ i: i + 1, kind: p.kind, bytes: b.size, ext });
} catch (e) { results.push({ i: i + 1, error: String(e) }); }
}
return results;
})()
```
If a fetch fails (expired/geo-blocked), it's skipped automatically — just move to the next card. No mention to the user.
---
## Step 4 — Collect files and deliver
After downloading, move the files out of the browser's Downloads folder to `/tmp/` with the Bash tool, in one command. Pick the glob that matches the mode you ran:
**VIDEO mode:**
```bash
cd ~/Downloads && mv -f spy-ad-*.mp4 /tmp/ && ls -la /tmp/spy-ad-*.mp4
```
**IMAGE mode:**
```bash
cd ~/Downloads && mv -f spy-ad-*.jpg spy-ad-*.png /tmp/ 2>/dev/null; ls -la /tmp/spy-ad-*.{jpg,png} 2>/dev/null
```
**BOTH mode (image + video):**
```bash
cd ~/Downloads && mv -f spy-ad-*.mp4 spy-ad-*.jpg spy-ad-*.png /tmp/ 2>/dev/null; ls -la /tmp/spy-ad-*.{mp4,jpg,png} 2>/dev/null
```
Then deliver — **only the files, no analysis, no commentary, no brief**. Present each creative inline if a preview tool is available; otherwise list them as clickable file links:
> - [spy-ad-1-creatify-ai.mp4](/tmp/spy-ad-1-creatify-ai.mp4)
> - [spy-ad-2-captions.mp4](/tmp/spy-ad-2-captions.mp4)
…in IMAGE mode:
> - [spy-ad-1-creatify-ai.jpg](/tmp/spy-ad-1-creatify-ai.jpg)
> - [spy-ad-2-creatify-ai.jpg](/tmp/spy-ad-2-creatify-ai.jpg)
…in BOTH mode (videos and images interleaved by rank):
> - [spy-ad-1-creatify-ai.mp4](/tmp/spy-ad-1-creatify-ai.mp4)
> - [spy-ad-2-creatify-ai.jpg](/tmp/spy-ad-2-creatify-ai.jpg)
> - [spy-ad-3-creatify-ai.mp4](/tmp/spy-ad-3-creatify-ai.mp4)
Do not append observations, patterns, recommendations, or written analysis. The only thing that may follow this delivery is the Step 5 "clone which?" form — and only when this skill ran standalone.
---
## Step 5 — Offer to clone (standalone runs only)
This step **only fires when arcads:spy-competitor-ads was invoked directly by the user as the end goal**, not when it was triggered as a sub-step by another skill. Detect the situation before running it:
**Skip Step 5 (do NOT show the form) when any of these is true:**
- The current invocation was triggered by another skill — most commonly `arcads:clone-hook` (Step 1B auto-source) or `arcads:clone-static-ad` (Step 1B auto-source). Those skills will consume the downloaded files themselves; offering to clone again would loop.
- The original user request was explicitly research-only — "just download competitor ads", "build me a swipe file", "save these for me", "I want to look at them".
- Zero creatives were delivered in Step 4.
**Run Step 5 (show the form) otherwise**, including when the user's request was ambiguous-but-actionable like "spy on competitors and let me work from there", "find competitor ads I can use as inspiration", or a bare "find competitor ads from [brand]". When in doubt, run it — declining is one click for the user.
### The form
Surface every delivered creative as a separate option in a single `AskUserQuestion` call with `multiple: true`:
> **Question**: "Want me to clone any of these for your brand? Pick the ones to recreate — I'll rebuild each for you."
> **Header**: "Clone which?"
> **Multiple**: `true`
> **Options**: one per delivered file, in the same order they were listed in Step 4. Each option's `label` is the short filename without the leading slug (e.g. "Ad #1 — BrandX (video)" or "Ad #3 — BrandY (image)"). Each option's `description` is one short cue from the card text if you have it (e.g. "Talking-head selfie, 22s"), or just the file type if you don't.
The `custom` default option ("Type your own answer") lets the user say "all of them", "none — I'm done", or a custom note.
### Routing each selection
Walk the user's selection in the order they were picked. For each chosen file:
1. **Video file (`.mp4`)** → load and run the `arcads:clone-hook` skill, passing the local `/tmp/spy-ad-*.mp4` path as the source video (skipping its Step 1B auto-source, since you already have a video). `arcads:clone-hook` will analyze the hook, then (after its own confirmation prompt) clone it for the user's brand.
2. **Image file (`.jpg` / `.png`)** → load and run the `arcads:clone-static-ad` skill, passing the local `/tmp/spy-ad-*.{jpg,png}` path as the reference static ad (skipping its Step 1B auto-source, since you already have an image). `arcads:clone-static-ad` will analyze the layout and clone it for the user's brand.
Run the chained skills **sequentially** (not in parallel) — each one needs the user's attention for brand/asset questions, and parallel runs would collide. After each chained skill finishes, move to the next selected file.
If the user picks just one creative, hand off immediately without re-confirming. If they pick "none", stop cleanly — no chaining, no further commentary. If they pick "all of them", chain through every delivered file in order.
### Carry brand context forward
The downstream skills (`arcads:clone-hook`, `arcads:clone-static-ad`) will ask for the user's brand, product, and assets. If the user already gave that information during Step 1a or in the original request, pass it forward — don't make them re-answer.
---
## Edge cases
- **No ads of the requested mode found for a competitor** → skip silently, continue with the rest. If *no* competitor yields any creative in the chosen mode, say so briefly and stop. Do **not** silently fall back to a different mode — the user picked VIDEO, IMAGE, or BOTH for a reason.
- **BOTH mode but only one media type returns** (e.g. competitor runs only videos): deliver what you got and say briefly "BrandX is only running videos right now — no statics in their library." Do not pad with the missing format.
- **Only same-name/unrelated ads found** → treat as "no ads found" for that competitor; skip silently.
- **Browser not connected** → the only blocking case; see Step 2.
- **User dismisses the mode form** (Step 1b) or doesn't pick a mode → stop and say "I need to know which kind of ads to grab — video, image, or both." Do not guess.
- **User deselects every competitor** in the Step 1c multi-select → ask once for a manual list ("Which competitors should I look at instead?"). If still empty, stop.
- **User picks "none" in the Step 5 clone form** → stop cleanly. No commentary, no follow-up.
- **User picks one creative in Step 5** → hand off to the matching cloner skill (`arcads:clone-hook` for video, `arcads:clone-static-ad` for image) without an extra confirmation.
- **Step 5 chain — one of the cloner skills fails or is cancelled by the user** → stop the chain (do not run the remaining selections silently). Surface what happened in one line and let the user restart the chain if they want.
- **Invoked from inside another skill** (e.g. `arcads:clone-hook` auto-source) → skip Step 5 entirely. The calling skill owns the next step.
---
## Quick reference
| Tool | Where |
|---|---|
| `AskUserQuestion` | Step 1b — pick media mode (when not explicit). Step 1c — multi-select competitor shortlist (when auto-discovered). Step 5 — multi-select "clone which?" (standalone runs only). |
| `WebSearch` / `WebFetch` | Step 1c — auto-find competitors (only if not named) |
| Browser MCP (`mcp__Claude_in_Chrome__*` / Playwright) | Steps 2–3 — open Ad Library, run extract+download JS |
| `javascript_tool` (in-page `fetch`→blob→download) | Step 3 — the ONLY reliable download path; curl does not work. Pick the extractor that matches the mode: `<video>` for VIDEO, `<img>` for IMAGE, combined for BOTH. |
| `Bash` (`mv`) | Step 4 — move files from Downloads to `/tmp/` (`.mp4` for VIDEO, `.jpg`/`.png` for IMAGE, both for BOTH) |
| `arcads:clone-hook` skill | Step 5 — chained per selected video, passing the local `/tmp/spy-ad-*.mp4` path |
| `arcads:clone-static-ad` skill | Step 5 — chained per selected image, passing the local `/tmp/spy-ad-*.{jpg,png}` path |
Package details
Publisher declarations from the archived package. These are separate from our research and the live service's terms.
- Package author
- Arcads
Package observed Sep 30, 2026.
Technical details
- First seen
- Sep 30, 2026 · 22:02 UTC
- Last seen
- Oct 1, 2026 · 18:00 UTC
- Collection status
- Collected
plugin_asdk_app_6aa01081c414819196e58fc42ab585a6
Download plugin data (JSON)