← Files Model CompassARCHIVED FILE
skills/compare-model-tradeoffs/references/evidence-policy.md
2.28 KB · Oct 2, 2026 · 00:30 UTC
# Evidence policy Use this policy whenever a comparison presents a number or recommendation. ## Evidence classes 1. **Official exact** — live catalog fields, token prices, credit rates, Fast multipliers, supported efforts, and context limits. 2. **OpenAI-published measurement** — a named benchmark with its version, harness notes, evaluation setting when known, source URL, and publication date. 3. **Locally measured** — matched runs on the user's task, with sample size, elapsed time, token usage, grader, failures, and configuration. 4. **Derived estimate** — arithmetic based on published or local inputs. Show the formula, assumptions, and an interval when uncertainty is material. 5. **Unknown** — no defensible numeric value. Keep it visible as unknown. ## Non-negotiable rules - Never turn the model-card 1–5 capability or speed icons into percentages. - Never invent a universal intelligence score or average unrelated benchmarks into one. - Never infer an effort-specific score when the evaluation does not identify its effort. - Never assume quality is monotonic with effort. - Never convert 1.5× model speed into one-third less wall-clock task time unless the model-time share is known. - Never mix ChatGPT credits, API dollars, included plan messages, or legacy per-message rates. - Never compare benchmark rows across different versions or harnesses without an explicit comparability note. - Treat fewer than five local trials as exploratory. - Preserve a last-known-good official snapshot if refresh parsing fails, and show its retrieval date. ## Configuration semantics - Treat the desktop **Power** preset separately from model-catalog defaults. Current docs define Power as GPT-5.6 Sol + medium reasoning. - Treat **Max** as more single-agent reasoning time. - Treat **Ultra** as a topology change that delegates to multiple agents, not another scalar effort level. - Treat **Fast** as a service tier. Read supported tiers from the live catalog and keep internal aliases such as `fast` and `priority` out of the user-facing label. ## Recommendation rule Recommend a configuration only when the evidence supports the user's task and constraints. If missing evidence can change the choice, recommend the smallest matched experiment that resolves it instead.
SHA-256: 3c106f755fd397fa63e0a60e0da5c2371b24b69a1b6f8a589b13d789e7259ef2