← CompText BenchmarkCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to CompText Benchmark
Snapshot Sep 30, 2026 · 23:14 UTC · version 0.1.5
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"description": "Use when the user asks for a Raw vs CompText verdict, quality regression check, efficiency delta, or evidence-based benchmark comparison.",
"included_files": [
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 309
}
],
"name": "compare-benchmark",
"skill_md_contents": "---\nname: compare-benchmark\ndescription: Use when the user asks for a Raw vs CompText verdict, quality regression check, efficiency delta, or evidence-based benchmark comparison.\n---\n\n# Compare benchmark\n\n- Prefer `proof-summary.json`; regenerate it from `result.json` with bundled `scripts/proof-summary.mjs` if absent or stale.\n- Confirm the receipt is bound to the expected manifest/result digest before relying on it.\n- Keep quality and efficiency separate and surface every regression flag.\n- Read the full `result.json` only when the receipt is insufficient; read events only for event-level evidence.\n- Lead with verdict, then quality deltas, efficiency deltas, regressions, and artifact/evidence references.\n- Never hide failed arms or incompatible manifests by averaging.\n- Load `../../references/benchmark-policy.md` only for ambiguity or live-run policy questions."
}SHA-256 of public snapshot: 3e4e0e904bc530c3750a48ffc89eb46ff8f50d979fe19d8dc08f031bbfab7e36