← CompText BenchmarkCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to CompText Benchmark
Snapshot Sep 30, 2026 · 23:14 UTC · version 0.1.5
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"description": "Use when the user asks to run or rerun a Raw vs CompText benchmark, validate token reduction, or check quality regressions.",
"included_files": [
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 302
}
],
"name": "run-benchmark",
"skill_md_contents": "---\nname: run-benchmark\ndescription: Use when the user asks to run or rerun a Raw vs CompText benchmark, validate token reduction, or check quality regressions.\n---\n\n# Run benchmark\n\nFor the built-in offline fixture run exactly:\n\n`node \"${CODEX_HOME:-$HOME/.codex}/plugins/cache/comptext-marketplace/comptext-benchmark/0.1.5/scripts/smoke-receipt.mjs\" /root/comptext/apps/comptext-benchmark-lab`\n\nTreat its single JSON line as the primary evidence receipt. Do not run `find`, list run files, inspect bundled script source, or read full `result.json` unless that command fails or the user asks for raw evidence.\n\nFor live runs, reuse the Benchmark Lab, freeze one manifest, keep non-experimental variables identical across arms, and load `../../references/benchmark-policy.md`. MCP is optional.\n\nNever claim an efficiency win when the receipt reports a quality regression. Return status, digests, runner/model, arm state, quality, efficiency, verdict, and artifact paths."
}SHA-256 of public snapshot: 612c1ec47cf5b361eabde5309ec6694e3815d54349ccba6345b87ba85cc9f50c