← CompText BenchmarkCONTENT HISTORY

Update to CompText Benchmark

Snapshot Sep 30, 2026 · 23:14 UTC · version 0.1.5

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "inspect-benchmark",
  "description": "Use when the user asks for benchmark status, failures, evidence, active/completed arms, or whether a rerun is required.",
  "included_files": [
    {
      "relative_path": "agents/openai.yaml",
      "size_in_bytes": 307
    }
  ],
  "skill_md_contents": "---\nname: inspect-benchmark\ndescription: Use when the user asks for benchmark status, failures, evidence, active/completed arms, or whether a rerun is required.\n---\n\n# Inspect benchmark\n\n- Resolve the exact run ID first and prefer immutable Benchmark Lab artifacts over process guesses.\n- For `smoke-orion-001`, prefer `proof-summary.json`; regenerate it from `result.json` with bundled `scripts/proof-summary.mjs` if absent or stale.\n- Read `events.jsonl` only when event-level failure evidence is required.\n- For other runs, verify their actual data directory before reading it.\n- Report status, runner, manifest digest, arm state, failures, evidence refs, metrics-so-far, and rerun decision.\n- Never expose credentials, environment secrets, or unrelated raw logs.\n- Load `../../references/benchmark-policy.md` only when policy interpretation is needed."
}

SHA-256: 9dd7fbebe952fc04b2441abc7ea816b47facb3fefbc23e42b798a8f5b9af619b