{"id":18677,"plugin_id":"plugins_6a8b1e75fc008191a65fa89587954dc6","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:14:53.300Z","digest":"3e4e0e904bc530c3750a48ffc89eb46ff8f50d979fe19d8dc08f031bbfab7e36","against":null,"payload":{"description":"Use when the user asks for a Raw vs CompText verdict, quality regression check, efficiency delta, or evidence-based benchmark comparison.","included_files":[{"relative_path":"agents/openai.yaml","size_in_bytes":309}],"name":"compare-benchmark","skill_md_contents":"---\nname: compare-benchmark\ndescription: Use when the user asks for a Raw vs CompText verdict, quality regression check, efficiency delta, or evidence-based benchmark comparison.\n---\n\n# Compare benchmark\n\n- Prefer `proof-summary.json`; regenerate it from `result.json` with bundled `scripts/proof-summary.mjs` if absent or stale.\n- Confirm the receipt is bound to the expected manifest/result digest before relying on it.\n- Keep quality and efficiency separate and surface every regression flag.\n- Read the full `result.json` only when the receipt is insufficient; read events only for event-level evidence.\n- Lead with verdict, then quality deltas, efficiency deltas, regressions, and artifact/evidence references.\n- Never hide failed arms or incompatible manifests by averaging.\n- Load `../../references/benchmark-policy.md` only for ambiguity or live-run policy questions."},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}