← Plugin catalog
Developer Tools

CompText Benchmark

CompText Labs v0.1.5

Publisher description

From the marketplace listing

Run, inspect, and compare Raw versus CompText experiments while keeping quality and efficiency separate and preserving evidence.

Language: English · Automatically detected from descriptions.

Files & skills

File archives

Plugin package17 files · 4.58 KBBrowse files →
Skill instructions
compare-benchmark875 Bytes

View saved version →

---
name: compare-benchmark
description: Use when the user asks for a Raw vs CompText verdict, quality regression check, efficiency delta, or evidence-based benchmark comparison.
---

# Compare benchmark

- Prefer `proof-summary.json`; regenerate it from `result.json` with bundled `scripts/proof-summary.mjs` if absent or stale.
- Confirm the receipt is bound to the expected manifest/result digest before relying on it.
- Keep quality and efficiency separate and surface every regression flag.
- Read the full `result.json` only when the receipt is insufficient; read events only for event-level evidence.
- Lead with verdict, then quality deltas, efficiency deltas, regressions, and artifact/evidence references.
- Never hide failed arms or incompatible manifests by averaging.
- Load `../../references/benchmark-policy.md` only for ambiguity or live-run policy questions.

Referenced files: 1

inspect-benchmark855 Bytes

View saved version →

---
name: inspect-benchmark
description: Use when the user asks for benchmark status, failures, evidence, active/completed arms, or whether a rerun is required.
---

# Inspect benchmark

- Resolve the exact run ID first and prefer immutable Benchmark Lab artifacts over process guesses.
- For `smoke-orion-001`, prefer `proof-summary.json`; regenerate it from `result.json` with bundled `scripts/proof-summary.mjs` if absent or stale.
- Read `events.jsonl` only when event-level failure evidence is required.
- For other runs, verify their actual data directory before reading it.
- Report status, runner, manifest digest, arm state, failures, evidence refs, metrics-so-far, and rerun decision.
- Never expose credentials, environment secrets, or unrelated raw logs.
- Load `../../references/benchmark-policy.md` only when policy interpretation is needed.

Referenced files: 1

run-benchmark970 Bytes

View saved version →

---
name: run-benchmark
description: Use when the user asks to run or rerun a Raw vs CompText benchmark, validate token reduction, or check quality regressions.
---

# Run benchmark

For the built-in offline fixture run exactly:

`node "${CODEX_HOME:-$HOME/.codex}/plugins/cache/comptext-marketplace/comptext-benchmark/0.1.5/scripts/smoke-receipt.mjs" /root/comptext/apps/comptext-benchmark-lab`

Treat its single JSON line as the primary evidence receipt. Do not run `find`, list run files, inspect bundled script source, or read full `result.json` unless that command fails or the user asks for raw evidence.

For live runs, reuse the Benchmark Lab, freeze one manifest, keep non-experimental variables identical across arms, and load `../../references/benchmark-policy.md`. MCP is optional.

Never claim an efficiency win when the receipt reports a quality regression. Return status, digests, runner/model, arm state, quality, efficiency, verdict, and artifact paths.

Referenced files: 1

Package details

Publisher declarations from the archived package. These are separate from our research and the live service's terms.

Package author
CompText Labs
Keywords
benchmark, context-compression, evaluation, comptext

Package observed Oct 2, 2026.

Technical details
First seen
Sep 30, 2026 · 22:02 UTC
Last seen
Oct 2, 2026 · 18:00 UTC
Collection status
Collected

plugins_6a8b1e75fc008191a65fa89587954dc6

Download plugin data (JSON)