{"id":18681,"plugin_id":"plugins_6a8b1e75fc008191a65fa89587954dc6","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:14:53.346Z","digest":"9dd7fbebe952fc04b2441abc7ea816b47facb3fefbc23e42b798a8f5b9af619b","against":null,"payload":{"description":"Use when the user asks for benchmark status, failures, evidence, active/completed arms, or whether a rerun is required.","included_files":[{"relative_path":"agents/openai.yaml","size_in_bytes":307}],"name":"inspect-benchmark","skill_md_contents":"---\nname: inspect-benchmark\ndescription: Use when the user asks for benchmark status, failures, evidence, active/completed arms, or whether a rerun is required.\n---\n\n# Inspect benchmark\n\n- Resolve the exact run ID first and prefer immutable Benchmark Lab artifacts over process guesses.\n- For `smoke-orion-001`, prefer `proof-summary.json`; regenerate it from `result.json` with bundled `scripts/proof-summary.mjs` if absent or stale.\n- Read `events.jsonl` only when event-level failure evidence is required.\n- For other runs, verify their actual data directory before reading it.\n- Report status, runner, manifest digest, arm state, failures, evidence refs, metrics-so-far, and rerun decision.\n- Never expose credentials, environment secrets, or unrelated raw logs.\n- Load `../../references/benchmark-policy.md` only when policy interpretation is needed."},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}