← Files Compound EngineeringARCHIVED FILE
skills/ce-babysit-pr/references/settle.md
16.1 KB · Oct 4, 2026 · 12:33 UTC
# Stop conditions, settle, and the merge-ready judgment **"Looks ready" is a judgment, not just an elapsed timer.** It is never enough that CI is green and the PR has been quiet for a while — see the review-still-coming gate below. **The in-progress signal gates only the *merge-ready declaration* — never the work.** Keep resolving open feedback as it arrives even while a review is in progress: **do not wait for the 👀 to clear before acting on the comments it has already posted.** Waiting for the review to *finish* before addressing feedback it already left would serialize the exact way waiting for a full CI run before addressing comments would — the same mistake the core principle forbids. Act on every open item continuously; the *only* thing the in-progress signal withholds is the "looks ready" call. (The engine reports only that the PR went quiet and for how long; whether a review is still coming is yours to judge at the settle decision, under the gate in Step 3.) ## Step 3: Stop conditions Otherwise (interactive), classify each condition as either a **true stop** (the watch ends and hands back) or a **standing residual** (surfaced to the user and blocking a *merge-ready declaration*, but which the self-sustaining watch **keeps running around** — it does **not** end the loop). **Terminal**, **looks-merge-ready**, **budget**, and **`blocked-external-drained`** are true stops for the active PR (under `stack-ready`/`stack-land` a quiescent layer is a transition before any of these — `references/stack.md`); an initial **`blocked-external`** observation enters the bounded review drain below instead of stopping. **`needs-human`**, **`blocked-failing`**, `stack-sync-needed` / chain-probe uncertainty, and an unresolved **semantic conflict** are standing residuals. A stale dependent above a target that is itself current is a stack-health residual, not a blocker to calling the target ready *as the next PR*. In **checkpoint mode** the single tick ends regardless of class; the distinction only changes what you report and whether you print a resume command. **True stops — the watch ends:** - **Terminal** — PR `MERGED` or `CLOSED`, except when this run just completed an authorized `stack-land` merge on that PR: that MERGED outcome is a layer transition (see stack-land land step above), not a run-level Terminal stop. - **Looks merge-ready (settled)** — GitHub itself reports it mergeable: `mergeability_certain`, `mergeable == "MERGEABLE"`, and `merge_state_status == "CLEAN"` (this defers required-check and required-review policy to GitHub after the merge computation is bound to the independently proven current base and current head), `base_ref_blocker == null`, `checks_terminal` is true (nothing still running), there is **zero actionable backlog** — `counts.threads == 0` **and** `counts.comments == 0` (no unresolved inline threads and no un-acted top-level/review-body feedback) — **and `open_needs_human == 0`** (a thread or comment you deferred for a human decision means it is *not* ready — surface it, do not call it merge-ready), **and `branch_currency_blocker == null`** (no open, claimed, or parked base-movement item), **and** `quiet_seconds` has reached the settle threshold **and the review-still-coming gate below is clear** — either no review looks to be on its way, or the bounded wait it allows has run out and you say so. "Mergeable against current base" does not mean the PR head contains the latest base commit: `BEHIND`, `DIRTY`, branch-protection requirements, or an explicitly selected always-current policy may require maintenance, while ordinary base movement with GitHub reporting `CLEAN` does not. Chain state further qualifies that result: a managed target requires `stack_blocker == null` (the snapshot defers to GitHub's certain `MERGEABLE`/`CLEAN` read for plain trunk drift); call it **"ready as the next PR in the stack"**, never independently ready, and list stale upstack entries separately. In a manual dependency chain, call it **"ready relative to its parent"** and name any open parent that must land first. Base identity `race`, `mergeability-pending`, or `probe-error`, chain `probe-error`, or an uncleared unknown managed freshness (`stack_blocker` present) blocks readiness only until a fresh observation proves the result. GitHub can leave its test merge on the old base indefinitely after main moves; when head and live base have been stable for the engine's stale window it emits `base.identity == "stale-computation"` with `merge_computation_stale: true`, which counts as certain — declare ready if the rest of the gate holds, and disclose in the report that GitHub's cached merge computation was stale (a fresh recompute happens at merge time). The settle threshold is the script's **300s default**; the *only* thing that ever widens it is the re-arm after a rejected `merge-ready` wake (the wake protocol below) — never pre-widen the initial arm, review bots or not. The settle window is a *cooling-off* signal — evidence the PR stopped moving, **not** a guarantee no further review is coming. **Before you report it ready, reflect on PR-description freshness (the final checkpoint).** A watch full of incremental commits — fixes, new behavior, resolved feedback, a base-into-head merge — routinely leaves the *original* PR description describing a PR that no longer exists. If what the PR now does has materially drifted from its description (new/removed behavior, a changed approach, resolved caveats), **refresh it autonomously: invoke `ce-commit-push-pr` in description-update mode, non-interactively (`mode:pipeline` — it rewrites and applies via `gh pr edit` directly, no preview prompt). Do not ask — a current description is part of leaving a PR merge-ready**, and updating it is one of `ce-commit-push-pr`'s own functions. If the description still reflects the change, leave it untouched. Report an independent PR as "looks ready — your call to merge," never "safe to merge." For a confirmed managed stack, apply the transition paragraph above before treating this layer stop as the whole run's stop. In **checkpoint mode** you cannot enforce elapsed time between manual re-runs, so if it is otherwise clean but `quiet_seconds` is under the threshold, say "green now, re-run in ~5 min to confirm it stayed quiet before merging." - **Blocked on external CI approval, after draining review** — `checks_awaiting_approval > 0` with no actionable backlog means a workflow is **awaiting a base-repo maintainer's approval to run** (GitHub's fork-PR security gate). Neither you nor the loop can trigger it; **never auto-approve** the run. CI is blocked, but review may still move independently, so the first `blocked_external` observation is not an interactive stop: - **Interactive self-sustaining watch:** without asking, apply the same review-still-coming judgment and re-arm at the normal active cadence (`--interval 150`) with `--blocked-external-drain-seconds 300` when no review looks to be on its way, or `--blocked-external-drain-seconds 900` when one does. The helper persists a narrow head-scoped clock: new or edited external feedback, review submission/signal movement, or a new head resets it; an unchanged approval gate, the loop's own replies, check/base/stack movement, and disposition-only bookkeeping do not. Incoming feedback still wakes ahead of the gate — resolve it immediately, then re-arm the drain on the new evidence. If approval clears, return to the ordinary CI watch. - **Drain expiry:** `blocked-external-drained` is the decision wake. At 900 quiet seconds, concrete prior current-head timing from the same reviewer may justify one re-arm with `--blocked-external-drain-seconds 1800`; it may never shorten the 900-second floor or extend beyond 1800 on a merely missing completion signal. When no review looked to be coming, 300 seconds is terminal — but apply the review-still-coming judgment once more before handing back. A reviewer can announce itself after the tier was chosen, and nothing in the drain observes that on its own, so the expiry wake is the only place that movement can still be caught. Only an explicit user request to keep watching the approval gate for a longer stated duration overrides these defaults, and it remains capped by this invocation's original budget. - **Handback:** once the selected drain expires with the gate still present, stop without asking another question. Report that all observed feedback was handled, how long the current head was review-quiet, that CI never ran because maintainer approval is still required, and give the host-rendered resume invocation. A later notification or explicit re-invocation starts the next bounded watch. - **Pipeline / unattended:** do **not** drain, ask, or spin — return a `blocked-external` residual with the run URL and terminate. Its bounded orchestrator contract explicitly does not wait on human review or approval. - **Checkpoint:** process the current tick, report the gate, state that monitoring is paused, and give the resume invocation; it cannot enforce elapsed drain time itself. - **Budget exhausted** — active `invocation_elapsed_seconds` reaches the fixed invocation budget (default 8h of active watch time, or the duration the user selected at entry), **or** raw wall-clock reaches the 3-calendar-day backstop, a round-count cap the user set, or the user aborts. The `max-runtime` wake names which ceiling fired via `max_runtime_ceiling` (`active-budget` or `backstop`). Watch re-arms and confirmed-managed-stack layer transitions must match the original ID, start, and budget; they can neither reset nor extend the cap and preserve accumulated dead time. The deadline's final refresh may still report `terminal` or an already-settled `merge-ready`; both stop immediately and start no new work. Otherwise `max-runtime` outranks actionable/residual work so no additional agent round begins. On `max-runtime`, report the emitted `invocation_elapsed_seconds` and budget, never `persisted_state_age_seconds`, and do not automatically mint another invocation. This is the blunt cost floor beneath the trajectory-driven non-convergence stop above — it catches a runaway that never trips the convergence trigger, not the normal way a stuck PR ends. A bounded invocation hands control back; only a later explicit user/orchestrator invocation starts another budget. **Standing residuals — surface, then keep watching (these do NOT end the self-sustaining loop):** - **`needs-human`** — accumulated typed decisions from `ce-resolve-pr-feedback` or `ce-debug`, or a semantic merge conflict Step 2 could not resolve mechanically. A managed upstack conflict follows Step 7's narrower rule: abort the manager transaction and surface it rather than deciding another PR layer's semantics. Persist each complete residual through the shared atomic mark, surface its full decision payload, and leave every covered review thread open. A parked decision blocks merge-ready, but in interactive continuous mode it does not end the watch: keep handling independent review and CI streams. The snapshot invalidates the grouped residual and re-actionizes every surviving source when any covered source changes or disappears; never maintain a separate route-specific decision or reopen path. - **`blocked-failing`** — a dispatched check `ce-debug` left terminally red (`has_failing_checks` with `counts.ci == 0`, nothing new to dispatch). Same shape: surface the red residual, it blocks merge-ready, but a later commit or head SHA may clear it — **keep watching**, and the detector will not re-wake on the same red residual (arm-time baseline). Only a **true stop** above, or the user, ends the loop. - **`stack-blocked`** — the target is manager-stale or of unknown managed freshness while GitHub does not report it certain `MERGEABLE`/`CLEAN` (plain trunk drift under a CLEAN read is not a blocker), manager discovery failed, or Step 7 could not complete its clean upstack transaction. Surface the `stack_blocker` and relevant `pr_chain` entries; continue review/CI work, but do not perform an ordinary base update or declare the target ready. A later manager sync, successful probe, or successful Step 7 maintenance clears the residual. A remaining stale upstack entry is a stack-health residual even when the target can be reported ready as next. **Is a review still coming? (part of the looks-ready gate).** The settle window tells you the PR stopped moving. It does not tell you a review is not on its way — the window can elapse before a backgrounded reviewer even starts. Judge that separately at the settle decision, from one cheap `gh` look at the current head, covering every surface an announcement can arrive on: reactions on the PR body, top-level comments, the check runs on that head, and reviews against it. A resumed run has no memory of an interim comment it already classified, and a dispatched one is absent from the snapshot, so the lookup is the only place it can still be seen. A reviewer that announces itself — a 👀 reaction, a "reviewing…" note — has told you it **started**. Nothing obliges it to retract that when it finishes, so a standing announcement is not evidence that work continues. Look instead for what that same reviewer produced on this head, and ask whether it accounts for the review it announced. Something carrying that reviewer's verdict — its review, its comment, or a check run that *is* its review — having reached a terminal conclusion means it stopped: say which, and never read a timeout or a skip as approval. Anything of its still running means it is still working. Output that does not account for the announced review leaves the review outstanding, so treat that reviewer like one that produced nothing: an unrelated check from the same app finishing while its review has not appeared is not the review finishing. A reviewer that announced itself and has nothing to show for it either way is the genuinely undecidable case. An announcement is not the only evidence a review is coming. A reviewer that reviewed an **earlier** head and not this one has told you the same thing through its own history on this PR, and many reviewers work that way without ever announcing anything — in some repos no reviewer announces at all. Treat that reviewer as one that announced with nothing to show yet, and it clears the same way: when it reviews this head, or when the bounded wait runs out. Keep the asymmetry in mind: a **present** signal is informative, an **absent** one tells you nothing, since many reviewers announce nothing and simply post. So never wait terminally for a done signal, and never let a missing one block on its own. Wait while you can see work in progress, and wait a bounded while when a reviewer announced itself with nothing to show for it. Re-arm by re-running the arm command with a longer `--settle-seconds` — passing it is what makes the re-arm real, since quiet already exceeds the 300s default and the ordinary arm would re-fire on the next poll. Use: about 900, extended once to no more than 1800 when concrete prior timing for that reviewer justifies it. Those are ceilings, not targets — announcing reviewers observed here finish in four to six minutes, while a slow reviewer that never announces had a 21.6-minute p90 over 157 reviews and a much longer tail beyond it. Once 1800 quiet seconds pass on unchanged evidence, stop and say plainly what you could not confirm; never re-arm past it on the same unchanged evidence. A widened window is your judgment about the evidence you looked at, so when that evidence moves the watch wakes you on `review-evidence-moved` rather than sleeping the window out against a state that has changed. Re-judge from the current head as at any settle decision: if what moved accounts for the review you were waiting on, the ordinary window decides again, and if it does not, re-arm as before. The engine reports only that something moved; whether it resolved anything is yours. Readiness is a recommendation to the user and never authorization to merge, except under `posture:stack-land`, where selecting that posture was the user's advance yes and a settled prefix does land. So calling it early usually costs only a disclosure, while under `stack-land` it costs landing sooner than the reviewer finished — and holding forever costs the PR in every posture.
SHA-256: 8725f46dd07f7bfa7bf486ac218efcbec3431fecc2d6a0e0e2ab2de240d273a2