← CavemanCONTENT HISTORY

Update to Caveman

Snapshot Sep 30, 2026 · 23:17 UTC · version 2.7.0

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "caveman-manage",
  "description": "Use when the user explicitly asks to start, approve, cancel, promote, or roll back a Caveman experiment.",
  "included_files": [
    {
      "relative_path": "agents/openai.yaml",
      "size_in_bytes": 225
    }
  ],
  "skill_md_contents": "---\nname: caveman-manage\ndescription: Use when the user explicitly asks to start, approve, cancel, promote, or roll back a Caveman experiment.\n---\n\n# Manage eval-gated experiments\n\nTreat every lifecycle change as a production control action. Read current state\nand results, then report one supported recommendation or block.\nCurrent agent MCP is intentionally read-only: control-api does not yet enforce a\ncomplete lifecycle transition table and evidence gate atomically.\n\n## Non-negotiable gates\n\n1. A request to review, inspect, explain, or recommend authorizes reads only.\n2. Never approve an experiment whose results are pending, whose required\n   guardrails are absent, or whose evidence reports a breach.\n3. Never convert experiment lift into `verified_savings`. Only active real\n   traffic plus provider-causal, provider-complete ledger evidence can do that.\n4. Never supply an organization id. Project and tenant scope come from the\n   logged-in Caveman identity and server RBAC.\n5. Never execute a lifecycle mutation, even after user approval. Exact\n   `<action>:<experiment_id>` strings are agent-generatable and are not proof of\n   human intent.\n6. Unknown states and server errors fail closed. Report exact\n   `cave_snake_code`.\n\n## Step 1 — Load project and experiment\n\nPrefer MCP:\n\n```text\ncaveman_context {}\ncaveman_experiment_get {\"action\":\"get\",\"experiment_id\":\"<id>\"}\ncaveman_experiment_get {\"action\":\"results\",\"experiment_id\":\"<id>\"}\n```\n\nUse `{\"action\":\"list\"}` when the user has not named an id.\n\nCLI fallback:\n\n```bash\ncaveman cloud experiments list\ncaveman cloud experiments show <id>\ncaveman cloud experiments results <id>\n```\n\nStop if login, project, experiment, or results are unavailable.\n\n## Step 2 — Evaluate evidence\n\nReport:\n\n- current lifecycle state and safety class;\n- control and candidate sample sizes;\n- quality or eval result;\n- latency, error, cost, retry, drop, and escalation guardrails when present;\n- evidence cost;\n- rollback or hold reason;\n- whether result is pending, failed, promotable, or active.\n\nAbsence is not a pass. If a required field is absent, state\n`evidence incomplete` and do not propose approval.\n\n## Step 3 — Propose one action\n\nAllowed actions:\n\n- `start` — only from a startable draft or queued state with configured graders;\n- `approve` — only with complete passing evidence and a safety class the\n  current role may approve;\n- `cancel` — stop a non-active experiment the user no longer wants;\n- `rollback` — revert an active or harmful change through the server's linked\n  policy path. Current deployments may reject this honestly with\n  `cave_not_implemented`; never describe that response as a rollback.\n\nShow recommendation and id:\n\n```text\nProposed action: approve experiment 7f...\nReason: candidate passed quality and every configured guardrail.\nExecution: blocked until server-authoritative lifecycle and evidence gates ship.\n```\n\nDo not treat earlier generic statements such as \"manage it\" or \"do what is best\"\nas mutation approval.\n\n## Step 4 — Block unsafe execution\n\nDo not emit or run an executable lifecycle command. Explain that current server\ndoes not yet enforce every evidence/state transition atomically. CLI and MCP\nagent surfaces therefore expose experiment reads only.\n\n## Step 5 — Re-read after external operator action\n\nIf operator says they executed command, read detail and results again. Report\nserver-observed post-state, audit or result response, and any policy-delivery\nstatus returned. Never infer success from operator intent alone.\n\nUse this close:\n\n```text\nAction: <action> <experiment-id>\nBefore: <state>\nServer response: <status and cave_snake_code if any>\nAfter: <re-read state>\nBasis: experiment evidence only. Verified savings unchanged unless the signed\nledger independently records active, provider-causal real-traffic savings.\n```\n"
}

SHA-256: 79c17b4761f3dea737fc674c743c3939eb6484ad5d50a0b1f96721da0699d803