{"id":5743,"plugin_id":"plugin_asdk_app_6a19d5012d308191a00a48780f7dcdcc","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T22:46:14.911Z","digest":"20eae14c233bbd98b97fe39f3adbe1f62586da7b9a57da8068276041d226a6b8","against":null,"payload":{"name":"fal-models-catalog","description":"Navigate fal.ai model families by media modality and production role. Use when the user asks which model, endpoint, or model family is appropriate for image, video, audio, 3D, editing, training, restoration, try-on, or analysis.\n","included_files":[],"skill_md_contents":"---\nname: fal-models-catalog\ndescription: >\n  Navigate fal.ai model families by media modality and production role. Use\n  when the user asks which model, endpoint, or model family is appropriate for\n  image, video, audio, 3D, editing, training, restoration, try-on, or analysis.\n---\n\n# fal.ai Models Catalog\n\nThis is a curated routing layer for MCP clients. It should guide endpoint selection,\nthen execution should happen through the fal.ai MCP tools exposed by this\nplugin. Do not shell out to genmedia CLI.\n\nUse `model-routing` first for common production defaults. Use this skill when\nthe question is broader: comparing modalities, finding a category, choosing a\nfamily, or explaining tradeoffs.\n\n## Modality Map\n\n- Text to image: campaign visuals, product stills, character concepts,\n  editorial photography, posters, UI mockups, image typography.\n- Image to image: edits, inpainting, reference preservation, style transfer,\n  product placement, background replacement, upscaling, restoration.\n- Text to video: cinematic clips, product reveals, narrative shots, social\n  concepts, motion drafts.\n- Image to video: animate an approved still, product hero motion, b-roll,\n  talking-head source frames, first-frame continuity.\n- Reference to video: stronger continuity from characters, products, or style\n  references where supported.\n- Text/audio to talking head: spokesperson, UGC creator, avatar, lip sync.\n- Text to audio: narration, TTS, music, sound effects.\n- Audio to text: transcription, subtitles, diarization, audio cleanup.\n- Image/text to 3D: objects, characters, game assets, GLB/OBJ/PLY outputs.\n- Image to text / vision: OCR, captioning, segmentation, detection, analysis.\n- Training: LoRA and fine-tune style workflows, only when a dataset exists.\n\n## Selection Pattern\n\n1. Identify the artifact role: final commercial, draft, utility transform,\n   analysis, training, or intermediate step.\n2. If the user did not name a specific endpoint, call `recommend_model` before\n   execution and use catalog search when the recommended list is too generic.\n3. Choose modality and endpoint family.\n4. Inspect schema before assuming fields such as `image_url`, `image_urls`,\n   `reference_image_url`, `duration`, `aspect_ratio`, `seed`, `quality`,\n   `audio_url`, or `enable_rigging`.\n5. Check pricing when the job is long, high resolution, batched, video, audio,\n   3D, or likely to be repeated.\n6. For uncertain categories, use MCP catalog search and docs search, then\n   choose from verified endpoints.\n\n## Production Defaults\n\n- Text-heavy stills: `openai/gpt-image-2`.\n- Product stills and campaign heroes: `openai/gpt-image-2`,\n  `fal-ai/nano-banana-pro`, `fal-ai/nano-banana-2`.\n- Product/reference edits: `fal-ai/nano-banana-pro/edit`,\n  `openai/gpt-image-2/edit`.\n- Final-quality video: Seedance 2.0 text/image/reference video endpoints.\n- Fast video drafts: Grok Imagine Video endpoints.\n- Talking head: `veed/fabric-1.0`, `veed/fabric-1.0/text`,\n  `fal-ai/creatify/aurora`, `fal-ai/sync-lipsync/v2`.\n- Background removal: use catalog search for current Bria/background endpoints\n  and inspect schema.\n- 3D: prefer Meshy v6 class endpoints for rigging/animation when available.\n\n## Avoid\n\n- Choosing a model only because it is popular when the artifact role is clear.\n- Using text-to-image when the user supplied a reference that must be preserved.\n- Using cheap draft endpoints for final brand/product assets.\n- Generating readable legal, medical, financial, or claim text unless the user\n  supplies exact wording and the chosen model supports text well.\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}