← Files Compound EngineeringARCHIVED FILE
skills/ce-babysit-pr/references/setup.md
10.3 KB · Oct 3, 2026 · 06:34 UTC
# Step 1 detail: forge check, PR resolution, chain classification, checkout, watch mode ## Prerequisites The loop runs `gh`, `git`, and a bundled Python helper against a local checkout with filesystem access. A harness without those (some sandboxed GUI environments) cannot run this skill — say so and stop rather than half-running. ## Step 1: Confirm GitHub, resolve the PR, pick an execution mode **GitHub only.** This skill and everything it delegates to speak GitHub's API (`gh`, review threads, Actions). First confirm the repo is on GitHub: `gh repo view` succeeding is the positive signal (it also covers GitHub Enterprise that `gh` is configured for). If it fails, inspect the remote — `git remote get-url origin` pointing at a `gitlab.*` host means GitLab, `bitbucket.*` means Bitbucket. On any non-GitHub forge (or if `gh` can't resolve the repo at all), **stop and tell the user ce-babysit-pr is GitHub-only** and that GitLab/other forges are not yet supported. Do not proceed into `gh` calls that will spray confusing errors. Then resolve the target PR from the argument (number/URL) or the current branch. If no open PR exists, report and stop. Resolve draft state with the PR. For an automatic calling-skill handoff without an explicit user watch-mode token, this check must be the stateless pre-bootstrap read `gh pr view --json isDraft` — never `snapshot --start-invocation`, which mints a new invocation and would supersede a watch a user explicitly authorized on that draft — and a draft target reports its draft status and stops here, before any bootstrap or watcher, per the "Draft PRs are opt-in" boundary. On user-invoked runs the first snapshot's emitted `pr_is_draft` serves as the ongoing signal. **Automatically classify the target's PR chain; never rely on the user to announce a stack.** The first snapshot and every later poll probe the read-only local manager with `gh stack view --json`, accepting it only when its branch list contains the target PR. If that cannot prove membership, the helper uses a read-only GraphQL fallback. A successful null stack means `pr_chain.manager_status == "absent"`. The specific stack-field schema-unavailable response also means `"absent"` only when a separate read-only lookup resolves the repository's default branch; auth, transport, rate-limit, malformed, other GraphQL, or failed default-branch probes mean `"probe-error"`. When no manager is confirmed, ordinary open-PR base/head relationships distinguish an independent PR from a manual dependency chain. Discovery never runs `gh stack checkout`, imports a stack, switches branches, or changes remote state. **Only when the fresh snapshot has `manager_status == "confirmed"` may stack-wide continuation activate; no other classification authorizes it.** A manual dependency chain never activates stack-wide continuation: keep it target-local even when its base/head topology resembles the manager's ordered branches. `probe-error` also stays target-local and mutation-conservative until a later snapshot positively confirms the manager. Discovery still runs for every babysit — posture does not disable confirmed-manager detection or Step 7 upstack maintenance. For a confirmed managed stack, inspect the manager's ordered entries once before choosing the active layer; this is read-only orientation, not multi-PR monitoring. Resolve **posture** per the table above before semantic work. If posture is still `target` and the requested middle PR has an unsettled downstack layer, offer once to begin at the lowest unsettled non-draft layer and proceed upward (`stack-ready`), with target-only as the alternative; do not silently redirect semantic work to another PR. If posture is already `stack-ready` or `stack-land` and the requested PR has an unsettled downstack layer, begin at the lowest unsettled non-draft layer without asking (downstack-to-upstack). If all downstack layers are settled, begin on the requested PR. When the requested PR already looks ready or later settles under `target`, offer once to continue to the immediate open non-draft upstack layer if it needs work (accepting selects `stack-ready` for the rest of the run). That one-time offer expands semantic babysit scope on an already confirmed managed stack — it is **not** a proactive suggestion to create or adopt PR stacks. An explicit request to babysit the managed stack counts as `stack-ready` acceptance, so do not ask redundantly. In `mode:pipeline`, which cannot ask, continue beyond the requested PR only when the invocation already supplied `posture:stack-ready`, `posture:stack-land`, or equivalent stack-wide scope; otherwise return the next candidate as a residual. Once `stack-ready` or `stack-land` is in effect, that posture authorizes sequential semantic babysitting through the confirmed managed stack without asking again at each layer. Keep one active PR target and one watcher: revalidate manager membership and ordered state at each transition, stop the old watcher, switch/check out the next immediate layer, then initialize its own snapshot state with `--continue-invocation` and the same three recorded values on the flags the first snapshot used — `--invocation-id "$RUN_INVOCATION_ID" --session-started-at "$RUN_STARTED_AT" --invocation-budget-seconds "$RUN_BUDGET_SECONDS"` (the anchor flag is `--session-started-at`, not `--invocation-started-at`) — plus `--continue-dead-time-seconds <prior layer's `invocation_dead_time_seconds`>` so the shared active-time budget carries the suspended time already excluded on earlier layers (each layer's state dir accumulates its own dead time, so without this the new layer would count that prior suspend as active) — **and re-state the same `posture:` value on the continue invocation**. The invocation budget is not renewed per layer. Never skip past a draft or enter it unless the user explicitly included that draft; never advance past a layer with a `needs-human` blocker. Stop at the first draft outside scope, human-blocked layer, end of the stack, budget, or user stop. Reconfirm `manager_status == "confirmed"` before every cross-PR transition — loss of positive confirmation ends stack-wide continuation rather than degrading into manual-chain behavior. **Verify the local checkout is the PR's head *branch* before any delegated mutation.** `ce-resolve-pr-feedback` and `ce-debug` commit and push the **currently checked-out branch** — so a checkout that isn't the PR's head branch makes their fixes fail to push or land on the wrong branch. A matching `HEAD` **SHA is not sufficient**: a detached HEAD or a *different* local branch that happens to point at the PR head SHA passes a SHA check yet still can't push the PR's branch. So verify the checkout is actually on the PR's head **ref with a matching upstream**: resolve `gh pr view <ref> --json headRefName,headRefOid,isCrossRepository`, and confirm `git branch --show-current` equals `headRefName` (and the upstream tracks the PR head repo). The robust default is to **just run `gh pr checkout <ref>` before mutating** (it checks out the head branch and sets tracking, and handles fork heads it can push to). If you cannot — **no push access to the PR's head ref** (you have it when the head repo is yours, when you have write access to it, or on someone else's fork when `maintainerCanModify` is true) or a dirty checkout — **stop and tell the user to checkout the PR's branch** rather than mutating the wrong one. Switching a *clean* checkout to the PR's branch is not a reason to ask; do it. Babysitting the current branch's own PR (the common case) already satisfies this. Then establish **how the watch sustains itself** — a skill can't be re-invoked by magic once its turn ends, so *you* set up the loop. **The default is a self-sustaining, in-session watch: you do not do one tick and hand back a resume command.** Read `references/watch-loop.md` once here for the mechanics; later steps cite its sections rather than re-reading it. Then: **User-runnable resume syntax.** Whenever this skill prints or copies a resume invocation, default to `/ce-babysit-pr <url>` and, when the run posture is not `target`, append the same `posture:stack-ready` or `posture:stack-land` token so checkpoint / durable / session re-entry keeps stack scope. Use `$ce-babysit-pr <url> [posture:…]` only when the active host is Codex or explicitly documents dollar-prefixed skill invocation. Render only the invocation as inline code and output one form only. - **Self-sustaining in-session watch (default).** Keep monitoring in the current session until a stop condition is met. Run `pr-snapshot watch`, the deterministic change detector, and wait for its `BABYSIT_WAKE` output using the harness's tools. The detector polls without agent tokens; each wake returns control to this agent for one tick of judgment and sub-skill calls, followed by re-arming the detector. Never collapse those judgments into a shell script. `references/watch-loop.md` owns scheduling and wake handling. - **Checkpoint.** Use checkpoint mode only when the user requests it or the harness cannot keep the session active while waiting for the detector’s output. Run exactly one tick, persist, report, and print the resume invocation. Say monitoring is paused. Never fake a loop with a foreground `sleep` or by ending the turn and promising to continue. - **Pipeline** (`mode:pipeline`, set by an orchestrator like `lfg`) — run **bounded synchronous ticks in-line**: the orchestrator is the scheduler, so loop ticks yourself (snapshot → act → re-snapshot) until the **pipeline stop** (Step 3), then return. Fully non-interactive. See "Pipeline mode" below for the deltas — a different stop condition, native residual surfacing, and a structured return — and read `references/watch-loop.md` for its bound. **Durability.** The in-session watch is session-bound; if the session closes, re-invoking with the host-rendered resume syntax resumes cleanly (state is fully persisted on disk). For an unattended watch that must outlive the session (days), escalate to a durable scheduler where one exists — Grok `scheduler_create --durable`, or a cron running `<harness-cli> exec '<host-rendered resume invocation>'` — accepting that a fresh headless run reconstructs from disk and loses this conversation's context (persist consequential decisions so it does not re-litigate). If the user passed a mode, honor it; otherwise pick per harness capability, state it in one line, and proceed.
SHA-256: a6649a318a43432ecfc13064f7555c2286168a8450ccda484a9014ece7eaa0dc