Oct 9, 2026 · 2 saved observations
Discoverability changed from “UNLISTED” to “LISTED”.
Metadata evidence →Listing evidence →CompText Labs v0.1.5
From the marketplace listing
Run, inspect, and compare Raw versus CompText experiments while keeping quality and efficiency separate and preserving evidence.
Language: English · Automatically detected from descriptions.
Search terms declared by the publisher.
Oct 9, 2026 · 2 saved observations
Discoverability changed from “UNLISTED” to “LISTED”.
Metadata evidence →Listing evidence →--- name: compare-benchmark description: Use when the user asks for a Raw vs CompText verdict, quality regression check, efficiency delta, or evidence-based benchmark comparison. --- # Compare benchmark - Prefer `proof-summary.json`; regenerate it from `result.json` with bundled `scripts/proof-summary.mjs` if absent or stale. - Confirm the receipt is bound to the expected manifest/result digest before relying on it. - Keep quality and efficiency separate and surface every regression flag. - Read the full `result.json` only when the receipt is insufficient; read events only for event-level evidence. - Lead with verdict, then quality deltas, efficiency deltas, regressions, and artifact/evidence references. - Never hide failed arms or incompatible manifests by averaging. - Load `../../references/benchmark-policy.md` only for ambiguity or live-run policy questions.
Referenced files: 1
--- name: inspect-benchmark description: Use when the user asks for benchmark status, failures, evidence, active/completed arms, or whether a rerun is required. --- # Inspect benchmark - Resolve the exact run ID first and prefer immutable Benchmark Lab artifacts over process guesses. - For `smoke-orion-001`, prefer `proof-summary.json`; regenerate it from `result.json` with bundled `scripts/proof-summary.mjs` if absent or stale. - Read `events.jsonl` only when event-level failure evidence is required. - For other runs, verify their actual data directory before reading it. - Report status, runner, manifest digest, arm state, failures, evidence refs, metrics-so-far, and rerun decision. - Never expose credentials, environment secrets, or unrelated raw logs. - Load `../../references/benchmark-policy.md` only when policy interpretation is needed.
Referenced files: 1
---
name: run-benchmark
description: Use when the user asks to run or rerun a Raw vs CompText benchmark, validate token reduction, or check quality regressions.
---
# Run benchmark
For the built-in offline fixture run exactly:
`node "${CODEX_HOME:-$HOME/.codex}/plugins/cache/comptext-marketplace/comptext-benchmark/0.1.5/scripts/smoke-receipt.mjs" /root/comptext/apps/comptext-benchmark-lab`
Treat its single JSON line as the primary evidence receipt. Do not run `find`, list run files, inspect bundled script source, or read full `result.json` unless that command fails or the user asks for raw evidence.
For live runs, reuse the Benchmark Lab, freeze one manifest, keep non-experimental variables identical across arms, and load `../../references/benchmark-policy.md`. MCP is optional.
Never claim an efficiency win when the receipt reports a quality regression. Return status, digests, runner/model, arm state, quality, efficiency, verdict, and artifact paths.Referenced files: 1
Publisher declarations from the archived package. These are separate from our research and the live service's terms.
Package observed Oct 10, 2026.
plugins_6a8b1e75fc008191a65fa89587954dc6
Download plugin data (JSON)Open the publisher's marketplace listing to check current availability and follow its connection instructions. This directory does not install plugins. Check the requested access and any account requirements before connecting.
Check marketplace availability ↗
We have not established the pricing or subscription requirements for this plugin. An absent price does not mean free access.
Compare researched pricing and access models →
Check the declared skills and available files, then try a small task whose result you can verify. Our archived descriptions and instructions establish publisher claims, not tested runtime quality. Review sources and coverage limits.