← CompText BenchmarkCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to CompText Benchmark
Snapshot Sep 30, 2026 · 23:14 UTC · version 0.1.5
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "inspect-benchmark",
"description": "Use when the user asks for benchmark status, failures, evidence, active/completed arms, or whether a rerun is required.",
"included_files": [
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 307
}
],
"skill_md_contents": "---\nname: inspect-benchmark\ndescription: Use when the user asks for benchmark status, failures, evidence, active/completed arms, or whether a rerun is required.\n---\n\n# Inspect benchmark\n\n- Resolve the exact run ID first and prefer immutable Benchmark Lab artifacts over process guesses.\n- For `smoke-orion-001`, prefer `proof-summary.json`; regenerate it from `result.json` with bundled `scripts/proof-summary.mjs` if absent or stale.\n- Read `events.jsonl` only when event-level failure evidence is required.\n- For other runs, verify their actual data directory before reading it.\n- Report status, runner, manifest digest, arm state, failures, evidence refs, metrics-so-far, and rerun decision.\n- Never expose credentials, environment secrets, or unrelated raw logs.\n- Load `../../references/benchmark-policy.md` only when policy interpretation is needed."
}SHA-256: 9dd7fbebe952fc04b2441abc7ea816b47facb3fefbc23e42b798a8f5b9af619b