{"id":17614,"plugin_id":"plugins_6a7a2f9e3968819187e30eaee8da1435","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:14:17.476Z","digest":"0c3409b378ca65db53854fefd4a708dfc7ed429ffbe1fae30a58075cda03604e","against":null,"payload":{"description":"Score a finished OKR cycle against the rubric, challenge inflated and sandbagged scores, extract three lessons into okrdev/LESSONS.md, and close the cycle as scored (or abandoned). Use when a cycle is ending or has ended, someone says \"run our retro\", \"score the quarter\", \"close out this cycle\", or a dead cycle needs an honest burial before planning the next one.","included_files":[],"name":"retro","skill_md_contents":"---\nname: retro\ndescription: Score a finished OKR cycle against the rubric, challenge inflated and sandbagged scores, extract three lessons into okrdev/LESSONS.md, and close the cycle as scored (or abandoned). Use when a cycle is ending or has ended, someone says \"run our retro\", \"score the quarter\", \"close out this cycle\", or a dead cycle needs an honest burial before planning the next one.\n---\n\n# Cycle retro\n\nYou are running okrdev's retro: score every KR against the rubric, make the scores survive\nscrutiny, turn the cycle into exactly three lessons, and close the file. Budget ~60 minutes\nof human time. Like check-ins, you pre-draft everything — the humans' time goes to judgment,\nnot arithmetic. The ritual script is in docs/rituals.md; the scoring rules in docs/method.md\n(both ship with this plugin).\n\n## 1. Preflight\n\n1. **Is okrdev installed?** Check for `okrdev/config.md` in the repo root. Missing → say so\n   and point at `/okrdev:install`. Stop.\n2. Read `okrdev/config.md` frontmatter for `level` and `cycle_length`.\n3. **Level 0?** There are no cycles at Level 0 — nothing to retro. Say so and point at\n   `/okrdev:plan`, which handles the upgrade to Level 1 when they're ready.\n4. **Find the cycle.** Look in `okrdev/okrs/` for the file with `status: active` (or the\n   cycle the human named).\n   - No active cycle and no file at all → nothing to score; point at `/okrdev:plan`.\n   - Only `scored` or `abandoned` files → the last cycle is already closed. Summarize its\n     LESSONS.md block in two lines and point at `/okrdev:plan`.\n   - Active cycle whose `end` date is still weeks away → confirm intent: \"Scoring now closes\n     the cycle early — right call if it's truly done, otherwise wait.\" Proceed only on a\n     clear yes.\n\n## 2. The abandoned path — offer it when it's honest\n\nIf the cycle is dead — check-ins stopped weeks ago, the team pivoted, nobody can say what the\nnumbers are — don't force a scoring theater. Offer to close it unscored:\n\n1. Flip the cycle file to `status: abandoned`.\n2. Append one dated line to `okrdev/LESSONS.md`: the cycle id, `abandoned`, and the reason in\n   the team's own words.\n3. Commit, and point straight at `/okrdev:plan`.\n\nTwo minutes, no inquisition. Systems die by silent decay, not by decision — an honest\nabandonment is a decision, and it beats a zombie cycle blocking the next real one.\n\n## 3. Pre-draft the scoring sheet — before engaging the humans\n\nRead, compute, and assemble everything first:\n\n- **The cycle file**: every KR with its type, DRI, baseline, target, milestone anchors\n  (`Notes:`), `Revised:` blocks, and `Status: dropped` markers.\n- **Every check-in in `okrdev/checkins/<cycle>/`**: build a per-KR confidence history\n  (needed for sandbag detection and for confronting score/confidence mismatches), pull\n  evidence from \"What moved,\" and collect the cycle's Judgment calls — overrides, emergency\n  post-hoc reviews, mid-cycle revisions.\n- **Check-in adherence**: check-ins held ÷ weeks in the cycle (e.g. `11/13`). This goes in\n  LESSONS.md; a low number is usually the first lesson writing itself.\n- **Proposed scores**: pre-fill wherever the actual is already in evidence. For each metric\n  KR you'll need the actual number — pull it from check-ins if recorded, otherwise flag it\n  for the DRI to bring.\n- **Cycle-wide tallies**: emergency count (more than 2 in a cycle gets said out loud — the\n  one unaudited escape hatch is where all gaming funnels), side-quest box-hours spent, and\n  maintenance share if computable from PR/commit `KR:` tags.\n\n## 4. Score, KR by KR\n\nWalk the sheet with the room (solo mode: with the one human — you argue the other side).\nFor each KR, the DRI states the actual; you apply the rubric:\n\n- **Metric KRs**: `score = clamp((actual − baseline) / (target − baseline), 0, 1)`.\n  The formula exists so retros don't degenerate into vibes. No actual number available →\n  that's not a scoring problem, it's an instrumentation lesson; score conservatively from\n  what evidence exists and record why in `Notes:`.\n- **Milestone KRs**: score the highest anchor fully reached — 0.3, 0.7, or 1.0 as defined at\n  planning. No partial credit between anchors; the anchors were agreed precisely so nobody\n  has to negotiate 0.55 versus 0.6 today.\n- **Dropped KRs** (`Status: dropped`): not scored. One line on why they were dropped —\n  they're reviewed in step 6, not averaged into anything.\n- Record each score in the KR's `Score:` field.\n\nWhile scoring, challenge in both directions:\n\n- **Inflation.** A score must trace to evidence. \"0.8 — what's the actual?\" is the whole\n  move. Milestone claims replay their anchors: \"0.7 claimed — show the thing in its 0.7\n  state,\" in its own medium — a preview for code; the signed contract, the published page,\n  the hire started for everything else. The anchors from planning are the demo script.\n  Confidence history is your mirror: a KR that sat at 0.9 all cycle and scores 0.4\n  (or the reverse) means the check-ins were theater — name it, kindly.\n- **Sandbagging — aspirational KRs only.** The signals: target hit before 60% of the cycle\n  had elapsed, plus confidence ≥0.9 flat from week one. Flag it as an input to next\n  planning's target-setting, not as an accusation. **Never flag a committed KR for scoring\n  1.0** — hitting a commitment is the job, and naive 1.0-flagging just teaches people to\n  score 0.93.\n- **Committed misses.** Any committed KR under 1.0 gets a root-cause note in its `Notes:`\n  before the retro moves on. Root-cause the plan and the system, not the person — policed\n  people game classifications; coached people use them.\n\n## 5. Report the results — separately\n\nReport committed and aspirational results as two lists, never averaged together. A blended\nnumber is meaningless: 1.0 is the passing grade for one type and evidence of sandbagging\nrisk for the other.\n\nCalibration to say out loud:\n\n- **Committed** should sit at or near 1.0. Anything under it is the headline of this retro.\n- **Aspirational** lands well around 0.6–0.7. All 1.0s → targets were too soft (see the\n  sandbag flags). Everything ≤0.3 → the plan was fantasy; next planning should assume less\n  throughput, and LESSONS.md is how it will know.\n\n## 6. Revisions and judgment calls review\n\nWalk every `Revised:` block and every dropped KR, one question each: was it the right call,\nmade early enough — or did the revision quietly convert a miss into a win? Re-scoping to\ndeclare victory is inflation with paperwork.\n\nThen the cycle's judgment calls, assembled in step 3: overrides (what pattern do they show?),\nemergencies (were they? what did they protect? recurring emergencies mean something upstream\nis broken), side-quests (did the box-hours budget hold?), and evidence re-class lines,\neach replayed against the cycle's end state: \"expected evidence was the signed contract —\ndid it arrive?\" This is the framework keeping\nits core promise — the coach never blocks, it remembers, and the retro is where the memory\ngets read.\n\n## 7. Extract exactly three lessons\n\nThree, not five — scarcity forces ranking. Lessons are about the system, not people. The\ntest for each: would it change what next cycle's planning session does? If not, it's an\nobservation, not a lesson. Good ones sound like: \"we can't score what we don't instrument —\nbaseline KRs first,\" or \"our aspirational average is 0.45; plan for 60% of the throughput\nwe feel like we have.\"\n\n## 8. Write it down\n\nAppend a dated block to `okrdev/LESSONS.md` (append-only — never edit prior blocks; they're\nthe planning record):\n\n```markdown\n## 2026-Q3 — scored 2026-10-01\nCommitted: KR2.1 1.0, KR2.2 0.8 — one miss, root cause in cycle file.\nAspirational: KR1.1 0.7, KR1.2 0.4 — avg 0.55. (KR1.3 dropped W33.)\nCheck-in adherence: 11/13 weeks.\nRevisions: KR1.1 target raised in W31 (early sandbag flag) — right call.\nLessons:\n1. <lesson>\n2. <lesson>\n3. <lesson>\n```\n\nThen close the cycle file: every `Score:` filled, root-cause notes on committed misses in\nplace, and `status: active` → `status: scored`.\n\nCommit both files. If the repo runs cycle-file changes through PRs (it did at planning),\nopen one titled `Retro: <cycle>` — same audit-trail rationale, and it can merge immediately:\nthe retro conversation was the review. Otherwise commit directly to main. With non-technical\nhumans, narrate the step in plain words as you do it.\n\n## 9. Roll forward\n\n- Point at `okrdev/PARKING_LOT.md`'s Promoted section — those items plus the fresh\n  LESSONS.md block are next planning's inputs, and they're ready now.\n- Propose `/okrdev:plan`. Momentum matters: a scored cycle with no successor is how teams\n  drift back to unexamined work. If the team needs a breather, fine — but get the planning\n  session on the calendar before the room empties.\n\n## Async mode\n\nWhen the team can't meet: interview each DRI separately about their KRs (actuals, proposed\nscores, root causes), merge into one scoring sheet, flag any KR where your rubric result and\nthe DRI's proposal disagree, and resolve those with the objective's DRI before writing\nanything. The LESSONS.md block notes it was run async — a retro nobody attended together is\nstill a retro, but the record should say so.\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}