← Files Compound EngineeringARCHIVED FILE

skills/ce-babysit-pr/references/tick.md

27.2 KB · Oct 2, 2026 · 00:33 UTC

↓ Download file

# The tick: snapshot, watch, marks, ordering

Read this before the first snapshot. Shell state does not persist between tool calls; every command re-sets its own variables.

## Step 2: Run one tick

A tick is fully resumable from disk, so any re-invocation drives it — a scheduler, `/loop`, or the user re-running the skill an hour later. Set `SKILL_DIR` to the directory containing this SKILL.md, then snapshot both streams in one batch:

```bash
SKILL_DIR="<absolute path of the directory containing the SKILL.md you just read>";
SCRATCH_ROOT="/tmp/compound-engineering-$(id -u)"; [ ! -L "$SCRATCH_ROOT" ] && (umask 077; mkdir -p "$SCRATCH_ROOT") 2>/dev/null && [ ! -L "$SCRATCH_ROOT" ] && [ -O "$SCRATCH_ROOT" ] && [ -w "$SCRATCH_ROOT" ] || SCRATCH_ROOT="${TMPDIR:-/tmp}/compound-engineering-$(id -u)"; [ ! -L "$SCRATCH_ROOT" ] && (umask 077; mkdir -p "$SCRATCH_ROOT") && [ ! -L "$SCRATCH_ROOT" ] && [ -O "$SCRATCH_ROOT" ] && chmod 700 "$SCRATCH_ROOT" || { echo "unsafe scratch root: $SCRATCH_ROOT" >&2; exit 1; };
STATE_DIR="$SCRATCH_ROOT/ce-babysit-pr/<host>-<owner>-<repo>-<N>";
(umask 077; mkdir -p "$STATE_DIR") || exit 1; chmod 700 "$STATE_DIR" || exit 1;
PY="$(for c in python3 python py; do command -v "$c" >/dev/null 2>&1 && "$c" -c '' >/dev/null 2>&1 && { echo "$c"; break; }; done)"; [ -n "$PY" ] || { echo "no working Python 3 interpreter on PATH" >&2; exit 1; };
"$PY" "$SKILL_DIR/scripts/pr-snapshot" snapshot --pr <N> --repo <[host/]owner/repo> --state-dir "$STATE_DIR" --start-invocation --invocation-budget-seconds <seconds>
```

This is the only command that may start a budget. Use the user's requested duration when supplied; otherwise use the fixed **8-hour default** (`28800`). The budget is spent in **active watch-capability time**, not raw wall-clock: while the in-session watch runs, a span where the whole process was suspended (a closed laptop) is excluded from `invocation_elapsed_seconds`, so time the agent could not watch does not drain the cap. Detection is coarse — an activity gap wider than a threshold well above the poll interval is charged to dead time; ordinary polls, agent ticks, and human-blocked waits keep counting. A separate **3-calendar-day wall-clock backstop** caps every invocation regardless of excluded dead time (the stale-PR / zombie-watch ceiling). Checkpoint mode and the durable/cron path have no continuous poll cadence, so they retain wall-clock accounting. Record the output's `invocation_id`, `invocation_started_at`, and `invocation_budget_seconds` as `RUN_INVOCATION_ID`, `RUN_STARTED_AT`, and `RUN_BUDGET_SECONDS`. Require `invocation_elapsed_seconds <= 60`; otherwise fail before arming a watcher. Durable PR dispositions, dedup, and trajectory survive a new invocation, but its budget clock does not. Every later snapshot and watch arm must present all three recorded values; the helper rejects a missing/mismatched token, anchor, or budget. A managed-stack layer transition additionally uses `--continue-invocation`. Never use `--start-invocation` after this first snapshot: re-arms, mutations, retries, review/CI rounds, and stack transitions share one non-rolling budget, and a re-arm preserves accumulated dead time rather than resetting it.

Treat every fresh `snapshot` as the canonical source of truth for review-thread state; its bundled fetch paginates the full thread connection. Never replace it with a one-shot `reviewThreads(first:N)` result. If a direct diagnostic query is genuinely necessary, follow `pageInfo` until `hasNextPage == false` before drawing a count or unresolved-state conclusion.

**In the self-sustaining watch, back the tick with the background change-detector.** `pr-snapshot watch` runs that same fetch→diff on an interval with **no agent tokens** and prints a single `BABYSIT_WAKE {reason,url,...}` line *only* when there's work to inspect (`actionable` for an unresolved thread or failed CI; `feedback-candidate` for a non-thread body that still needs resolver judgment) or a stop/residual condition (`terminal` / `blocked-external` / `blocked-external-drained` / `blocked-failing` / `base-ref-blocked` / `unrequested-base-merge` / `downstack-actionable` (a lower managed-stack layer re-opened; `references/stack.md`) / `stack-blocked` / `needs-human` / `merge-ready` after the settle window / `max-runtime` / `stop-signal` / `invocation-superseded`) — then exits. A `feedback-candidate` wake is not a detector claim that a fix or reply is required: a resolver pass that silent-drops the body is a normal classification outcome, not a false positive. Keep the session active while waiting for that line with the harness's tools (Step 1); on the sentinel, run the tick below:

```bash
SKILL_DIR="<absolute path of this skill's directory>"; SCRATCH_ROOT="/tmp/compound-engineering-$(id -u)"; [ ! -L "$SCRATCH_ROOT" ] && (umask 077; mkdir -p "$SCRATCH_ROOT") 2>/dev/null && [ ! -L "$SCRATCH_ROOT" ] && [ -O "$SCRATCH_ROOT" ] && [ -w "$SCRATCH_ROOT" ] || SCRATCH_ROOT="${TMPDIR:-/tmp}/compound-engineering-$(id -u)"; [ ! -L "$SCRATCH_ROOT" ] && (umask 077; mkdir -p "$SCRATCH_ROOT") && [ ! -L "$SCRATCH_ROOT" ] && [ -O "$SCRATCH_ROOT" ] && chmod 700 "$SCRATCH_ROOT" || { echo "unsafe scratch root: $SCRATCH_ROOT" >&2; exit 1; }; STATE_DIR="$SCRATCH_ROOT/ce-babysit-pr/<host>-<owner>-<repo>-<N>"; (umask 077; mkdir -p "$STATE_DIR") || exit 1; chmod 700 "$STATE_DIR" || exit 1; RUN_INVOCATION_ID="<invocation_id>"; RUN_STARTED_AT="<invocation_started_at>"; RUN_BUDGET_SECONDS="<invocation_budget_seconds>";
PY="$(for c in python3 python py; do command -v "$c" >/dev/null 2>&1 && "$c" -c '' >/dev/null 2>&1 && { echo "$c"; break; }; done)"; [ -n "$PY" ] || { echo "no working Python 3 interpreter on PATH" >&2; exit 1; };
"$PY" "$SKILL_DIR/scripts/pr-snapshot" watch --pr <N> --repo <[host/]owner/repo> --state-dir "$STATE_DIR" --interval 150 --invocation-id "$RUN_INVOCATION_ID" --session-started-at "$RUN_STARTED_AT" --invocation-budget-seconds "$RUN_BUDGET_SECONDS"
```

In a confirmed managed stack under `stack-ready`/`stack-land`, append `--downstack-pr <N>` for every open lower layer to that arm so the watcher probes below (`references/stack.md`).

Watch ownership is **latest-valid-watcher-wins**. A newer invocation first cancels any older invocation still preflighting, but does not disturb the active watcher; only after a successful first snapshot does it atomically supersede and gracefully terminate that active process. Every wake and snapshot carries `watch_generation`. On delivery, compare the wake's generation with one fresh snapshot; a stale wake is discarded and coalesced into that current read, and a current wake whose attention set already cleared is also a no-op rather than another tick. An `invocation-superseded` wake means another explicit invocation now owns the durable state: end the old loop without acting or re-arming it. Re-arming with the same invocation token preserves `last_change_at`, `invocation_started_at`, and `invocation_budget_seconds`; it cannot restart or extend either timer.

Do **not** pass `--settle-seconds` or `--blocked-external-drain-seconds` on the ordinary arm. The script's 300s default is the initial merge-ready settle window; Step 3 alone sets `--settle-seconds` after a rejected `merge-ready` wake and `--blocked-external-drain-seconds` after an approval-gate wake begins the bounded review drain.

**Shell state does not persist between separate tool calls.** `SKILL_DIR` and `STATE_DIR` are set only for the command they appear in; the later `mark` calls (Steps 3 and 5) run as their own invocations, so re-set both inline in each of those commands — or pass the absolute paths directly. A bare `$SKILL_DIR` in a fresh call is empty and resolves to the wrong path.

**`<host>` in `STATE_DIR` is load-bearing for GitHub Enterprise.** Derive it from the PR URL's host (or `gh repo view --json url`); use the same value in every `mark`. Keying only by `<owner>-<repo>-<N>` would let two PRs with the same `owner/repo#N` on *different* hosts (github.com + a GHE instance) share one `state.json`, so one host's dispositions/dispatched CI would silence or contaminate the other's actionable set. On plain github.com the host segment is just `github.com`. **Pass the same host in `--repo <host>/<owner>/<repo>`** (the documented `[HOST/]OWNER/REPO` selector) so `pr-snapshot`'s first `gh pr view` — which runs before it parses the URL host — queries the right host instead of the checkout's default `github.com`.

The snapshot emits the **attention set** — unresolved threads you have not yet acted on, **non-thread feedback candidates** (top-level PR comments + review-submission bodies) you have not yet classified, and failing checks on the current head you have not yet dispatched — plus the exact current `branch_currency` item and its `attention` route. It also emits `pr_state`, `mergeable`, `merge_state_status`, `base`, `base_ref_blocker`, `host_branch_update_capability`, `branch_currency_blocker`, `review_decision`, `head_sha`, `head_changed`, `quiet_seconds`, `invocation_elapsed_seconds`, `invocation_remaining_seconds`, `persisted_state_age_seconds`, `checks_awaiting_approval` / `blocked_external`, and the head-scoped `blocked_external_first_seen_at`, `blocked_external_review_last_activity_at`, `blocked_external_review_quiet_seconds`, and `blocked_external_review_moved_this_tick` review-drain facts (see Step 3), plus a `pr_chain` block and a `trajectory` block (cross-tick facts: `check_recur_max`, `recurring_checks`, `unresolved_trend`, `new_threads_this_tick`, `stream_alternations`, `heads_since_progress`, `invariant_rounds`). `base.historical_oid` is GitHub's historical `baseRefOid`; it is diagnostic and is not the current base tip. Current-base identity requires the independent exact Git ref (`base.oid`) to match the PR `baseRef.target.oid` (`base.graphql_oid`). For a mergeable result, the generated `potentialMergeCommit` must also name that current base and the observed PR head as its two parents. Only that proven binding emits `base.identity == "current"`; a base movement race emits `race`, temporary merge-commit generation emits `mergeability-pending`, and a failed or malformed probe emits `probe-error`. These transient blockers disable `mergeability_certain` and re-poll. A `DIRTY` / `CONFLICTING` result may omit `potentialMergeCommit`; matching current-base observations still make the conflict result usable. Invocation time and persisted-state age are separate; never report one as the other. `pr_chain` carries the two independent axes: `manager_status` (`confirmed|absent|probe-error`) and `relationship_status` (`dependent|independent|probe-error`), plus manager source, target/upstack freshness, ordered entries, and ordinary parent/dependent PRs when available. The JSON field remains `actionable.comments` for the claim→act→confirm protocol, but its members are candidates awaiting semantic classification, not detector-proven action items. For non-thread feedback, the deterministic fetch excludes only empty bodies — including the PR author's own comments, since an author asking for a change on their agent-opened PR is ordinary feedback; loop prevention is the `dispatched` mark, not identity. It does **not** decide from content, bot identity, or comment-vs-review surface whether an external message is valid feedback; `ce-resolve` applies that judgment. The snapshot **never** marks a surfaced item handled just from observing it; an item stays in the attention set until you confirm you acted or classified it (`mark`) or remote truth removes it (a resolved thread drops out of the fetch). Every `mark` write must present the same `RUN_INVOCATION_ID`, `RUN_STARTED_AT`, and `RUN_BUDGET_SECONDS`; a stale resolver tick must fail before it can silence work in a replacement invocation. So a crashed, failed, or superseded resolve pass leaves its items in the set next tick. The state schema and the claim→act→confirm protocol are in `references/watch-loop.md`; read it before acting if this session has not already loaded it.

**The `trajectory` is facts, not a verdict — you hand it to the leaves, they judge convergence.** When it crosses a trigger (`check_recur_max >= 2`, `stream_alternations >= 3`, a rising `unresolved_trend` with `new_threads_this_tick > 0` across passes, `heads_since_progress >= 2`, or any `invariant_rounds[].rounds >= 2` — a key at 2 recorded rounds means the next fix would be its third), pass the trajectory to that tick's `ce-debug`/`ce-resolve-pr-feedback` invocation as **mandatory input** **before** that leaf mutates and let it decide whether this is ordinary progress or genuine non-convergence (a leaf may then return a `needs-human` residual that parks the *whole stream*, e.g. an emergent CI trade-off, a wrong-approach nitpick cluster, or a third invariant round). Never declare non-convergence yourself. The trigger→route→park→re-open protocol is in `references/watch-loop.md` (**Non-convergence**); read it before acting on this if this session has not already loaded it.

## Ordering invariant (full text)

**The ordering invariant (this is the whole point):**

1. **Terminal check first.** If `pr_state` is `MERGED` or `CLOSED`, stop and report — the loop is done — **except** when this run just completed an authorized `stack-land` merge on that PR: treat that MERGED outcome as a managed-stack layer transition (see Step 3's stack-land land step), not a run-level Terminal stop.
2. **Capture the head SHA now** (`git rev-parse HEAD` or the snapshot's `head_sha`) so you can tell later whether the comment pass pushed.

**Managed-stack pre-push baseline.** Before invoking a delegate that may push the active target in a confirmed managed stack, record a recoverable baseline from a fresh `gh stack view --json`: the manager-ordered open branches at or above the target (target plus open dependents) and each branch's current remote-tracking OID on the tracking remote. Require a clean worktree and still-confirmed manager membership for the target/current branch. If either precondition fails, this is a true stop for the active invocation in every mode: do not invoke a delegate, run another tick, or arm/re-arm a watcher; state the residual and give the host-rendered resume invocation. Do not stop for missing atomic multi-ref push proof — current `gh stack push` may update branches non-atomically (`github/gh-stack#216`); prefer all-or-none when an installed manager later proves atomic push, but always re-probe after push rather than assuming it.

3. **Feedback before CI.** If the attention set has **either** unresolved threads **or** non-thread feedback candidates (`counts.threads > 0` or `counts.comments > 0`), invoke `ce-resolve-pr-feedback` **once**, passing the resolved PR ref — the base `[HOST/]OWNER/REPO#N` or the full PR URL from the snapshot's `url` (so a fork→upstream PR resolves against the **upstream base**, not the fork checkout's `origin`, which would query the wrong PR namespace) — in full mode **with `mode:pipeline`** (non-interactive: it parks any `needs-human` on the thread and returns it as a structured residual instead of pausing on a blocking user question, which would stall the autonomous watch — the same reason Step 2 step 5 invokes `ce-debug mode:pipeline`); it re-fetches and judges *all* feedback — inline threads, review bodies, and top-level comments — and is idempotent on empty. The `actionable.comments` field contains the top-level/review-body candidates the resolver would otherwise not know the loop cares about — a Changes-Requested review body or a bare top-level "please rename X" with **no inline thread** must still trigger a pass. **When the review trigger above is crossed (rising backlog, new-item arrivals, a repeating cluster, or any `invariant_rounds[].rounds >= 2`), pass the `trajectory`** so it can judge a treadmill / wrong-approach nitpick cluster and return one approach-level `needs-human` instead of fixing forever — **and, when the recurring items are *valid* and share one root and fix, request a bounded-class assessment** so it consolidates the equivalent sites this PR touched into a single fix rather than dripping one per head (`references/watch-loop.md`, Non-convergence). An `invariant_rounds[].rounds >= 2` trigger routes that trajectory **before** this pass may fix/commit/push. One resolve pass per tick — never fan out multiple.

For any delegate result, process its typed `needs-human` residuals through one boundary. Immediately render every complete payload under `## Needs your decision`, preserve it unchanged for caller return, and write it to an OS-temp JSON file. Persist each residual once: the snapshot validates the full schema, freezes every source's current observation, and publishes one decision under the locked state write. It emits the unchanged payload in canonical `needs_human_residuals` and its answer-routing ID in `human_decisions`; source dispositions remain ordinary open/dispatched facts. Any covered observation changing or disappearing invalidates the whole decision and reactivates its surviving sources, but remote activity is never an answer. Continue independent work without resolving covered threads.

After the resolver, reconcile every **comment you passed that is not covered by a returned typed residual**. A top-level comment or review body never drops out of the fetch on its own, and `ce-resolve` may silently drop boilerplate, status noise, or other non-actionable feedback after applying agent judgment. Mark every comment you passed as `dispatched` unless the shared residual mark already parked it; never add a route-specific `needs-human` mark. Leaving an uncovered comment unmarked keeps it in the attention set, so the loop can never settle:

```bash
SKILL_DIR="<absolute path of this skill's directory>"; SCRATCH_ROOT="/tmp/compound-engineering-$(id -u)"; [ ! -L "$SCRATCH_ROOT" ] && (umask 077; mkdir -p "$SCRATCH_ROOT") 2>/dev/null && [ ! -L "$SCRATCH_ROOT" ] && [ -O "$SCRATCH_ROOT" ] && [ -w "$SCRATCH_ROOT" ] || SCRATCH_ROOT="${TMPDIR:-/tmp}/compound-engineering-$(id -u)"; [ ! -L "$SCRATCH_ROOT" ] && (umask 077; mkdir -p "$SCRATCH_ROOT") && [ ! -L "$SCRATCH_ROOT" ] && [ -O "$SCRATCH_ROOT" ] && chmod 700 "$SCRATCH_ROOT" || { echo "unsafe scratch root: $SCRATCH_ROOT" >&2; exit 1; }; STATE_DIR="$SCRATCH_ROOT/ce-babysit-pr/<host>-<owner>-<repo>-<N>"; (umask 077; mkdir -p "$STATE_DIR") || exit 1; chmod 700 "$STATE_DIR" || exit 1;
PY="$(for c in python3 python py; do command -v "$c" >/dev/null 2>&1 && "$c" -c '' >/dev/null 2>&1 && { echo "$c"; break; }; done)"; [ -n "$PY" ] || { echo "no working Python 3 interpreter on PATH" >&2; exit 1; };
"$PY" "$SKILL_DIR/scripts/pr-snapshot" mark --pr <N> --repo <[host/]owner/repo> --state-dir "$STATE_DIR" --invocation-id "$RUN_INVOCATION_ID" --session-started-at "$RUN_STARTED_AT" --invocation-budget-seconds "$RUN_BUDGET_SECONDS" --disposition needs-human --residual-file <path-containing-one-exact-typed-residual>
"$PY" "$SKILL_DIR/scripts/pr-snapshot" mark --state-dir "$STATE_DIR" --invocation-id "$RUN_INVOCATION_ID" --session-started-at "$RUN_STARTED_AT" --invocation-budget-seconds "$RUN_BUDGET_SECONDS" --comment <ID> --disposition dispatched --invariant-key <key-returned-by-the-leaf-on-a-fix-outcome>
```

Passing `--pr`/`--repo` on the shared residual mark is load-bearing: `mark` re-reads every covered thread's current last comment (including your just-posted replies) before the atomic write, so later reviewer activity re-opens the whole group instead of being swallowed. Publish the decision only when every covered thread was re-read; a missing observation leaves the complete source set actionable. Covered comments and review bodies retain their snapshot edit identities for the same reason. A **dispatched comment** mark needs no baseline: it stays silenced until an explicit `mark --disposition open` and is never auto-reactivated by a body edit, because status bots rewrite their bodies on every push. A genuinely new request arrives as a review thread or a new comment, both still surfaced. On a **fix** outcome, persist the leaf-returned `invariant_key` on that same dispatched mark (`--thread` or `--comment` plus `--invariant-key`); omit the flag when the leaf returned none. The leaf does not run `pr-snapshot`.

When the user answers, map their response to the displayed `decision_id`, preserve the exact response in a file, and record the shared answer transition before acting on it. The answer is consumed only when this mark succeeds. Read-only prohibits executing the transition, not rendering it. To complete a read-only envelope, return the literal command below as the sole pending transition; substitute exact known values and retain explicit placeholders for unavailable invocation metadata or the answer-file path. That rendered command must include the literal `--answer-decision` and `--answer-file` flags. A prose paraphrase or an in-memory state move is incomplete.

```bash
SKILL_DIR="<absolute path of this skill's directory>"; SCRATCH_ROOT="/tmp/compound-engineering-$(id -u)"; [ ! -L "$SCRATCH_ROOT" ] && (umask 077; mkdir -p "$SCRATCH_ROOT") 2>/dev/null && [ ! -L "$SCRATCH_ROOT" ] && [ -O "$SCRATCH_ROOT" ] && [ -w "$SCRATCH_ROOT" ] || SCRATCH_ROOT="${TMPDIR:-/tmp}/compound-engineering-$(id -u)"; [ ! -L "$SCRATCH_ROOT" ] && (umask 077; mkdir -p "$SCRATCH_ROOT") && [ ! -L "$SCRATCH_ROOT" ] && [ -O "$SCRATCH_ROOT" ] && chmod 700 "$SCRATCH_ROOT" || { echo "unsafe scratch root: $SCRATCH_ROOT" >&2; exit 1; }; STATE_DIR="$SCRATCH_ROOT/ce-babysit-pr/<host>-<owner>-<repo>-<N>"; (umask 077; mkdir -p "$STATE_DIR") || exit 1; chmod 700 "$STATE_DIR" || exit 1; RUN_INVOCATION_ID="<invocation_id>"; RUN_STARTED_AT="<invocation_started_at>"; RUN_BUDGET_SECONDS="<invocation_budget_seconds>";
PY="$(for c in python3 python py; do command -v "$c" >/dev/null 2>&1 && "$c" -c '' >/dev/null 2>&1 && { echo "$c"; break; }; done)"; [ -n "$PY" ] || { echo "no working Python 3 interpreter on PATH" >&2; exit 1; };
"$PY" "$SKILL_DIR/scripts/pr-snapshot" mark --state-dir "$STATE_DIR" --invocation-id "$RUN_INVOCATION_ID" --session-started-at "$RUN_STARTED_AT" --invocation-budget-seconds "$RUN_BUDGET_SECONDS" --answer-decision <decision_id> --answer-file <path-containing-the-exact-human-response>
```

The next snapshot returns that record in `answered_human_decisions` and makes every still-matching source ordinary actionable work. Apply the answer, then use the existing dispatched/confirmed marks; the answered record retires when its sources are handled or their observations move. Never convert a thread reply, comment edit, check rerun, head move, or currency change into an answer implicitly.

Surface those decisions in Step 4 and continue independent work. Once none remains, `mode:pipeline` returns the canonical set as its decision handoff; interactive continuous mode keeps watching around the parked sources. Also retain the resolver's **non-routine verdicts** — a fix done differently than the reviewer suggested (`fixed-differently`), feedback it declined (`declined`) or rebutted as wrong (`not-addressing`) — for the Step 4 summary; a plain `fixed` is routine and not worth carrying.
4. **Stale-SHA cancellation.** Compare the current head SHA to the one captured in step 2. If it **changed**, the comment pass (or someone) pushed — the CI failures in this snapshot are against a dead SHA, so **do not act on them**; the new run will surface next tick. If it did **not** change, continue to CI.
5. **CI on the current head.** Aggregate *all* actionable failing checks into one remediation pass — do not dispatch per check. Classify from metadata:
   - **Flaky/infra** (known-flaky job, infrastructure/timeout signal) → extract the run ID **and the full base repo including host** from the failing check's `details_url` (`https://<host>/<owner>/<repo>/actions/runs/<run-id>/…`) and `gh run rerun <run-id> --failed -R <host>/<owner>/<repo>`. Passing the run ID is load-bearing unattended: omitting it drops `gh run rerun` to an interactive run-picker menu that blocks `mode:pipeline`. Passing the host-qualified `-R <host>/<owner>/<repo>` is load-bearing for fork→upstream and GitHub Enterprise PRs: the run lives in the **base** repo on its own host, so a bare `-R <owner/repo>` (or no `-R`) targets the fork or the default `github.com` and 404s. On plain github.com the host segment is optional but harmless.
   - **Real test/build failure** → invoke `ce-debug mode:pipeline` once, seeded with the failing jobs and their log tails — **and, when the CI trigger above is crossed, the `trajectory` (`recurring_checks`, `check_recur_max`, `heads_since_progress`) so it can judge oscillation vs ordinary progress.** Its structured return `status` is exactly one of `fixed-and-pushed`, `fixed-not-pushed`, `flaky-infra`, `diagnosed-no-fix`, or `needs-human`; do not invent `infra-retry` or `stale`. Handle convergent fixes and retries as before. A `needs-human` result returns complete typed residuals whose `check` sources enter the canonical set through the shared boundary above; park those checks with the same residual file. `diagnosed-no-fix` and failed pushes remain ordinary red residuals. Never weaken a test, rebase, reset, or force-push to clear them.
   The shared boundary above already parks every check covered by a `needs-human` residual. Record each other check you acted on so it is not re-dispatched at this head (re-set the vars inline):

   ```bash
   SKILL_DIR="<absolute path of this skill's directory>"; SCRATCH_ROOT="/tmp/compound-engineering-$(id -u)"; [ ! -L "$SCRATCH_ROOT" ] && (umask 077; mkdir -p "$SCRATCH_ROOT") 2>/dev/null && [ ! -L "$SCRATCH_ROOT" ] && [ -O "$SCRATCH_ROOT" ] && [ -w "$SCRATCH_ROOT" ] || SCRATCH_ROOT="${TMPDIR:-/tmp}/compound-engineering-$(id -u)"; [ ! -L "$SCRATCH_ROOT" ] && (umask 077; mkdir -p "$SCRATCH_ROOT") && [ ! -L "$SCRATCH_ROOT" ] && [ -O "$SCRATCH_ROOT" ] && chmod 700 "$SCRATCH_ROOT" || { echo "unsafe scratch root: $SCRATCH_ROOT" >&2; exit 1; }; STATE_DIR="$SCRATCH_ROOT/ce-babysit-pr/<host>-<owner>-<repo>-<N>"; (umask 077; mkdir -p "$STATE_DIR") || exit 1; chmod 700 "$STATE_DIR" || exit 1;
   PY="$(for c in python3 python py; do command -v "$c" >/dev/null 2>&1 && "$c" -c '' >/dev/null 2>&1 && { echo "$c"; break; }; done)"; [ -n "$PY" ] || { echo "no working Python 3 interpreter on PATH" >&2; exit 1; };
   "$PY" "$SKILL_DIR/scripts/pr-snapshot" mark --state-dir "$STATE_DIR" --invocation-id "$RUN_INVOCATION_ID" --session-started-at "$RUN_STARTED_AT" --invocation-budget-seconds "$RUN_BUDGET_SECONDS" --check "<key>"
   ```

   (A new head SHA clears these automatically.)

6. **Branch currency & conflicts (the third stream — after comments and CI).** Consume the exact current `branch_currency` item; no item means no base-into-head mutation. Full route, claim lifecycle, and the `BEHIND`/`DIRTY` mechanisms: `references/branch-currency.md`.
7. **After an authorized target-head push in a confirmed managed stack, preserve the upstack before resuming the watch** (`gh stack rebase "<first-dependent-branch>" --upstack --no-trunk` + `gh stack push`, abort and park on conflict): `references/stack.md`.
8. **After any mutation, re-snapshot** at the start of the next tick, passing the same `--invocation-id "$RUN_INVOCATION_ID" --session-started-at "$RUN_STARTED_AT" --invocation-budget-seconds "$RUN_BUDGET_SECONDS"` — the head SHA and CI universe have changed, but the invocation-wide budget has not. Do not run a second `snapshot` mid-tick to re-derive CI; that is what caused stale-SHA confusion.

During accepted managed-stack continuation there is one watcher (probing lower layers via `--downstack-pr`) and one mutated target; the transition condition, the return rule, and the landing gate are owned by `references/stack.md`.

SHA-256: 1b82874642000c97255fe5273d84699d8650eae9b164f491e162b33d1816dcd7