← Empire LLM for CodexCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to Empire LLM for Codex
Snapshot Sep 30, 2026 · 23:13 UTC · version 1.7.2
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"description": "Fetch current Artificial Analysis language-model benchmarks, join exact OpenRouter model identities, apply transparent weighted scoring and modality requirements, and render a sandboxed vertical comparison chart in supported Codex or ChatGPT surfaces. Use for current best-model questions, agentic or coding benchmark comparisons, route-fit visualizations, and benchmark-weighted model selection; do not use it to dispatch a model completion.",
"included_files": [
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 313
}
],
"name": "empire-benchmarks",
"skill_md_contents": "---\nname: empire-benchmarks\ndescription: Fetch current Artificial Analysis language-model benchmarks, join exact OpenRouter model identities, apply transparent weighted scoring and modality requirements, and render a sandboxed vertical comparison chart in supported Codex or ChatGPT surfaces. Use for current best-model questions, agentic or coding benchmark comparisons, route-fit visualizations, and benchmark-weighted model selection; do not use it to dispatch a model completion.\n---\n\n# Empire Benchmarks\n\nCreate a current, attributable benchmark view while Codex remains the analyst\nand routing authority. This skill is read-only: it may call model metadata APIs,\nbut it must never dispatch an inference request, reserve project budget, send\nrepository evidence, or expose credentials to the generated UI.\n\n## Build the benchmark view\n\n1. Translate the request into a profile (`agentic`, `coding`, `balanced`, or\n `value`), requested input/output modalities, top-model count, and any explicit\n scoring weights. Preserve the user's date wording in the explanation. The API\n returns current data; do not claim it reconstructs a historical snapshot.\n2. Resolve this skill's directory as `EMPIRE_BENCHMARK_ROOT`, then resolve the\n shared runner at `../../scripts/empire_benchmarks.py`.\n3. For a current route-qualified agentic view with Pro/Commercial identity\n evidence, run:\n\n ```bash\n python3 \"$EMPIRE_BENCHMARK_ROOT/../../scripts/empire_benchmarks.py\" rank \\\n --profile agentic \\\n --route-qualified-only \\\n --top 8 \\\n --output /tmp/empire-benchmarks.json\n ```\n\n Add `--input-modality image` or another requirement only when the user needs\n it. Add repeatable `--weight agentic=45` style overrides when the user states\n priorities. The runner normalizes weights to 100% and reports them.\n Artificial Analysis Free omits `openrouter_api_id` and modality metadata.\n For that tier, omit `--route-qualified-only` to create an explicitly unverified\n research chart; never claim a qualified route or infer missing IDs. Explain\n that Pro/Commercial identity evidence is needed for exact joins. Do not\n upgrade a subscription or change the user's cost policy automatically.\n4. The live command reads the Artificial Analysis credential and optional\n OpenRouter credential through Empire's existing environment/keyring boundary.\n If a key is missing or inaccessible, use `$empire-settings`; never ask the\n user to paste a key into chat or put one in a command argument.\n5. Treat the returned JSON as the source of truth. Verify its source receipt,\n tier, index version, retrieval time, weights, evidence coverage, modality\n state, and exact OpenRouter match before recommending a route.\n\n## Render inside Codex\n\nWhen the current surface supports in-conversation visualizations, choose a\nwritable thread visualization path and run:\n\n```bash\npython3 \"$EMPIRE_BENCHMARK_ROOT/../../scripts/empire_benchmarks.py\" render \\\n --input /tmp/empire-benchmarks.json \\\n --output /absolute/thread/visualization/empire-model-benchmarks.html\n```\n\nReturn the generated fragment with the host's visualization content reference.\nThe fragment is self-contained, embeds only sanitized benchmark fields and local\n14 px model icons, performs no browser-side network requests, and lets users\nadjust the scoring weights. Its follow-up button may ask Codex to compare the\nleaders, but it may not dispatch an external model.\n\nIf the surface cannot render visualizations, rerun `rank` with `--view markdown`\nand return the accessible Markdown table. Do not claim that Codex CLI or an IDE\nextension rendered an interactive chart.\n\n## Evidence and routing boundaries\n\n- Artificial Analysis `/api/v2/language/models` is used for Pro or Commercial\n access. `auto` falls back to the documented `/api/v2/language/models/free`\n endpoint only after a tier-level `403`; every page is validated and bounded.\n- Current benchmark indices, pricing, and performance come from Artificial\n Analysis. Empire's weighted score is an inference, not an Artificial Analysis\n rank. Always keep the visible attribution link.\n- [Artificial Analysis API documentation](https://artificialanalysis.ai/data-api/docs)\n defines which fields each tier exposes. Missing identity or modality fields\n are a capability limitation, not permission to substitute fuzzy matches.\n- OpenRouter is queried only through its authenticated user-visible model\n catalog. A model becomes route eligible only when\n `openrouter_api_id` matches an OpenRouter model ID exactly and required\n modalities pass. Fuzzy names may be discussed as unverified, never routed.\n- Missing evidence reduces the weighted score; it is not silently converted to\n a zero benchmark. Incompatible modalities fail route eligibility.\n- OpenAI models may be included with `--include-openai` as research baselines,\n but they remain ineligible as Empire partner models because Codex is the\n OpenAI lead.\n- A chart is a read-only comparison. Before any later external completion,\n rerun the normal Empire route/review authorization flow and respect its cost,\n privacy, freshness, and project-budget gates.\n- Artificial Analysis requires attribution, and API access or redistribution\n rights depend on the user's tier and agreement. Never package a live API\n response or benchmark snapshot into a release artifact.\n\n## Useful prompts\n\n- “Give me the best agentic-model benchmarks as of September 2026 and show the\n top eight in a vertical weighted chart.”\n- “Compare current coding models with 50% coding, 25% agentic, 15% modality,\n and 10% cost weight. Require text and image input.”\n- “Show the best-value non-OpenAI partners available to my OpenRouter account,\n then explain why the top two complement Codex.”\n"
}SHA-256 of public snapshot: 969572540b8dc32a12f8432cf5b96dce9b0cce641f6be8554a4fedb866569c49