{"id":17571,"plugin_id":"plugins_6a7a0fafc7e881918b67a46ce717faa1","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:14:16.036Z","digest":"4410ae3188e4155ae3febf7845a29d3e75fffffa59f77a7e2899b546310a9ced","against":null,"payload":{"description":"Compare Codex model, reasoning-effort, and Standard/Fast combinations using the live Codex catalog, official OpenAI pricing and evaluations, and optional matched local measurements. Use when a user asks which Codex model or effort to choose, what a setting gains or loses, whether Fast mode is worth it, which combinations are supported, asks to turn Model Compass on or off, requests its companion view, asks for an evidence-backed interactive comparison, or requests a controlled model duel on their own task.","included_files":[{"relative_path":"agents/openai.yaml","size_in_bytes":270},{"relative_path":"assets/delta-racetrack.fragment.html","size_in_bytes":32964},{"relative_path":"assets/effort-field.fragment.html","size_in_bytes":14192},{"relative_path":"assets/neon-constellation.fragment.html","size_in_bytes":34909},{"relative_path":"assets/tradeoff-lab.fragment.html","size_in_bytes":18032},{"relative_path":"references/companion-mode.md","size_in_bytes":1305},{"relative_path":"references/evidence-policy.md","size_in_bytes":2332},{"relative_path":"references/openai-baseline.json","size_in_bytes":5857},{"relative_path":"scripts/get-codex-catalog.mjs","size_in_bytes":3912}],"name":"compare-model-tradeoffs","skill_md_contents":"---\nname: compare-model-tradeoffs\ndescription: Compare Codex model, reasoning-effort, and Standard/Fast combinations using the live Codex catalog, official OpenAI pricing and evaluations, and optional matched local measurements. Use when a user asks which Codex model or effort to choose, what a setting gains or loses, whether Fast mode is worth it, which combinations are supported, asks to turn Model Compass on or off, requests its companion view, asks for an evidence-backed interactive comparison, or requests a controlled model duel on their own task.\n---\n\n# Compare Model Tradeoffs\n\nProduce a decision receipt that separates what OpenAI establishes, what the\nuser's own runs establish, what is derived, and what remains unknown.\n\n## Display state\n\nModel Compass is quiet by default. Maintain a conversational preference named\n`model_compass_visual` for the current thread:\n\n- Default `model_compass_visual` to `off` in every new or forked thread.\n- `Show Model Compass`, `Visualize this comparison`, and clear equivalents are\n  one-shot requests. Render once without changing the preference.\n- `Enable Model Compass for this thread`, `Turn Model Compass on`, and clear\n  equivalents set the preference to `on`. Acknowledge the change; render\n  immediately only when the same turn also asks for a comparison.\n- `Disable Model Compass for this thread`, `Turn Model Compass off`, and clear\n  equivalents set it to `off`. The latest explicit instruction wins.\n- Treat a bare `turn off` as Model Compass only when no competing target is\n  active in the conversation. Otherwise ask one short clarification.\n- When resuming the same thread, recover the latest explicit preference from\n  its history. If history does not establish a value, use `off`.\n- When writing a handoff or compaction summary, preserve exactly one state\n  line: `Model Compass visual: on` or `Model Compass visual: off`.\n\nThis preference is transcript-backed rather than durable host state. Never\nclaim stronger persistence than the available thread history establishes.\n\nWhen the preference is `off`, provide an ordinary model-choice answer as a\nconcise text decision receipt without an inline visualization. End with one\nunobtrusive sentence offering `Show Model Compass` when a visual would help.\nWhen it is `on`, render only on relevant model, effort, speed, cost, or latency\ndecision turns. A one-turn `text only` request does not change the preference.\nDisabling affects future responses; do not claim to remove visuals already in\nthe transcript.\n\n`Open Model Compass companion view` is a one-shot request. Render the\nnarrow-friendly delta racetrack, then explain once that the desktop app can pop\nout the active chat and optionally keep it on top. Do not change\n`model_compass_visual` unless the user also explicitly enables it.\n\nDo not persist the display preference to files or other threads. Do not\nautomatically open, move, resize, pin, or keep a window on top. Do not render a\nvisual on unrelated turns merely because the preference is on.\n\nPets are a separate user-controlled ChatGPT desktop feature. A plugin cannot\ndetect that the app was minimized, wake or tuck away a pet, replace the pet\nactivity tray, or put this visualization inside the pet overlay. If the user\nasks for pet integration, explain that an awake pet can show cross-chat status\nand return them to ChatGPT, but `/pet`, **Wake Pet**, and **Tuck Away Pet** stay\nunder user control.\n\n## Workflow\n\n1. Read `references/evidence-policy.md` and\n   `references/companion-mode.md`.\n2. Infer the task shape, billing surface, current choice, and hard constraints\n   from the conversation. Use GPT-5.6 Sol + medium + Standard as the documented\n   Power baseline when the user gives no baseline.\n3. Discover valid choices from the live Codex catalog:\n\n   ```bash\n   node scripts/get-codex-catalog.mjs\n   ```\n\n   The helper starts the local Codex app server and calls only its read-only\n   `model/list` method. It does not read credentials or Codex files directly.\n   If the sandbox blocks that exact helper command, do not request broader\n   filesystem access. Use current host tool metadata when it exposes model\n   combinations; otherwise use the dated packaged baseline and mark live\n   compatibility as unknown. Never run Codex debug-dump commands or parse\n   Codex SQLite, rollout logs, `models_cache.json`, global state, or\n   application bundles.\n4. Refresh material official facts through the OpenAI developer-docs tools.\n   Use only OpenAI-operated sources. Start from\n   `references/openai-baseline.json` when remote docs are unavailable, and show\n   its `asOf` date. Re-check at least model guidance, Speed, pricing, and the\n   latest model launch when freshness could change the decision.\n5. Join catalog choices to official facts by exact model ID. Let the catalog\n   determine supported efforts and service tiers. Keep desktop presets separate\n   from model-level defaults. Present Ultra as multi-agent delegation, not as a\n   scalar effort value.\n6. Classify every metric as `official-exact`, `official-proxy`, `local`,\n   `derived`, or `unknown`. Preserve benchmark version, evaluation effort when\n   known, source, and date.\n7. When the current display state permits a visual, render one compact\n   interactive comparison. Choose the view that best matches the decision:\n   - `assets/neon-constellation.fragment.html` for a broad model landscape and\n     Pareto-style exploration. Use it when a one-shot `Show Model Compass`\n     request does not identify a concrete comparison.\n   - `assets/effort-field.fragment.html` for inspecting the complete\n     OpenAI-published GeneBench-Pro model × effort matrix.\n   - `assets/delta-racetrack.fragment.html` for a selected-versus-baseline\n     decision receipt. Prefer this view when the user is comparing one concrete\n     choice.\n   - `assets/tradeoff-lab.fragment.html` remains the compact general-purpose\n     fallback.\n   - Copy it to the thread-scoped visualization directory under a new\n     lower-case hyphenated filename.\n   - Patch the embedded catalog and official snapshot before displaying it.\n   - Keep all data inline; do not fetch from the fragment.\n   - Use the exact `::codex-inline-vis{file=\"name.html\"}` directive.\n   - Keep controls to model, effort, and Standard/Fast unless another input is\n     essential to the user's stated decision.\n   - For a Companion one-shot request, prefer\n     `assets/delta-racetrack.fragment.html`; it is designed to remain legible\n     in a narrow popped-out chat.\n8. Lead with a tradeoff receipt:\n   - selected and baseline configurations;\n   - exact price, credit, availability, and Fast facts;\n   - closest comparable published proxy;\n   - domain-specific effort evidence, if relevant;\n   - local evidence and sample size;\n   - explicit unknowns;\n   - cheapest useful next experiment.\n\n## Matched measurement\n\nRun a model duel only after the user explicitly asks to measure configurations\nand accepts any material credit or API cost.\n\n- Pin prompt, inputs, repo state, tools, permissions, timeout, and success\n  criteria.\n- Use isolated worktrees or disposable copies for code-changing trials.\n- Interleave configuration order and use at least three trials when practical.\n- Record wall time, output and reasoning tokens, tool time when observable,\n  retries, failures, and deterministic test or artifact results.\n- Treat fewer than five trials as exploratory.\n- Use a blinded rubric only when deterministic grading is not possible.\n- Keep raw proprietary prompts and code out of persistent telemetry by default.\n\nDo not automatically apply model, effort, or Fast settings. Offer the exact\nconfiguration command or setting after the comparison and require confirmation\nbefore changing persistent defaults.\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}