← Files Compound EngineeringARCHIVED FILE
skills/ce-babysit-pr/references/report.md
5.8 KB · Oct 3, 2026 · 06:34 UTC
# Step 4: report / summary Write every summary through the `ce-noslop` skill. The rules below are what this skill adds on top. Every stop — and every checkpoint tick — ends with a summary. Below the first line, write it however reads cleanly; the format is yours. What matters is that it hits these goals, because each counters a specific way these summaries fail: - **Outcome first, unmissable — open with one status line.** Emoji, state, then one clause of evidence composed from the final snapshot (quiet time, CI, remaining backlog, parked residuals — your wording, real values), so the state is scannable instead of buried in prose. Only the state phrases are fixed: - `✅ Looks merge-ready — <evidence>. Your call to merge.` — for a confirmed managed stack `✅ Ready as the next PR in the stack — <evidence>.`, for a manual dependency chain `✅ Ready relative to its parent — <evidence>.` - `🟡 Cautiously looks ready — <stalled-reviewer evidence>. Your call to merge.` Other stops follow the same shape with an emoji that states the condition — `🎉 Merged`, `🚫 Closed`, `⛔ Blocked`, `⏱️ Budget exhausted`, `⏸️ Paused` are the common ones. A ready declaration never opens with anything but ✅ or 🟡, and no other state may open with those two. - **PR state first in live updates.** Say what changed for the PR and what remains. Treat detector mechanics such as a wake, snapshot, re-arm, or head as internal implementation detail; mention them only when they explain a failure or required user action. - **Action receipts are historical.** When the host asks for an action receipt, including an `ACTIONS` trailer, list only mutations actually performed or observed this tick. Planned, next, or pending transitions stay in the narrative and never count as actions. - **A run recap at every true stop.** An hour-long watch resolves feedback and fixes CI the user never watched happen; the stop summary is the only place that work becomes visible, and omitting it is this skill's most common reporting failure — a merge-ready stop that states only the current PR state (CI green, no threads) has skipped this goal. After the status line, recap the run in a few short lines: what the feedback was about and how it settled (grouped by theme, with counts), what CI broke and the nature of each fix (one clause each), what was pushed, how long the watch ran, and what remains parked. The test: the reader could decide whether to merge and explain the PR's journey without scrolling back. Bare counts fail it — "resolved 11 threads" without what they concerned tells the user nothing — and so does the opposite extreme, a per-thread or per-check transcript. Build the recap from what survives in session context, verified and gap-filled from the PR's own remote record — resolved review threads, the PR's commit list, check runs, via `gh` — plus the state dir's parked items and the snapshot's elapsed time; never from conversation memory alone. A long watch has usually outlived the context that saw its early rounds, and the state dir deliberately forgets handled work (resolved threads leave the fetch, a new head clears dispatched checks), so the PR's remote record is the durable source for what the run actually did. If the watch changed nothing, one line says so. - **Escalations are prominent and complete.** When the canonical `needs_human_residuals` set is non-empty, add `## Needs your decision` immediately after the outcome line. Pair every unchanged payload with its `decision_id` from the snapshot's `human_decisions` view, then render `quoted_feedback`, `investigation`, `decision_reason`, every `options` entry and tradeoff, `recommendation` when non-null, and every `thread_urls` link. Return the residual objects unchanged. Never derive this section from a route-specific counter or summary, and never resolve covered threads. - **Chain scope is explicit.** For a managed stack, state the active layer's position, the run posture (`target` / `stack-ready` / `stack-land`), whether it is ready as next, whether stack-wide continuation was accepted or declined (under `target`), the next transition/draft/human boundary, and any `upstack_needs_rebase` residuals. Under `target`/`stack-ready` when ready-as-next and not auto-merging, print the exact `gh stack merge <N> --yes --squash` command. For a manual dependency chain, name the parent/dependent PRs and qualify readiness relative to the parent. Never imply that target-local success made the whole chain healthy. - **Surface the judgment calls, not the routine fixes.** Where the loop (through its delegates) did something other than the literal ask — a fix implemented differently than the reviewer suggested, feedback declined or rebutted as wrong, or a call a human steered mid-loop — name it in one line with the *why*. These are the calls a reasonable person would want to know were made on their behalf. Skip the routine "reviewer asked, we fixed it" items; those stay in the aggregate count. If a human decision or a stated preference shaped how an item went, reflect that so the record shows why the call landed where it did. If nothing non-routine was decided, say nothing — do not manufacture calls to look thorough. - **Honest about settledness.** If it looks ready, say how long it has been quiet and that it is your call to merge. Never imply "safe to merge." - **Disclose a stalled reviewer succinctly.** Name the reviewer when identifiable, otherwise name the observed signal; say how long no additional review progress was observed, and state that the lifecycle never produced its normal completion marker. Give the host-rendered resume invocation as the resume path and mention a known manual review trigger only when the repository exposes one. - **Checkpoint mode ends with the resume path.** State plainly that monitoring is paused and give the exact command to run the next tick.
SHA-256: 39f2671663759344960c4d5ad7c6319286cbea1635c12f67e66431f1a06b3096