← Files WebMCP KitARCHIVED FILE
skills/implement/references/interactive.md
10.8 KB · Oct 5, 2026 · 18:31 UTC
# Interactive loop
Phase D is reviewed live in the Explorer instead of in chat. The state folder `<workspace>/.webmcp/` is the source of truth; the browser is a view and input surface, never proof that Codex handled an action. Write state files for the watcher to render. Keep NDJSON entries on one line and use ISO-8601 `ts` values (`date -u +%Y-%m-%dT%H:%M:%SZ`).
## State folder
- `.run.json` is server-owned runtime metadata, schema v1: `{"version":1,"run_id","capability","workspace","port","pid","started_at"}`. `workspace` is canonical. The default server start rotates the run; `--resume` reuses a crashed run only for recovery in the same agent task/session. A clean Done teardown clears liveness. Do not commit runtime metadata or infer a current run from an old port. Automatic hook binding is immutable for a run and never transfers to another Codex task.
- `plan.json` is `{"journey","site","suggestions":[{"id","name","description","why","status","params":[{"name","type","description"}]}]}`. `site` is a short page/site label; `journey` is the narrative. `why` is one short, plain sentence naming the user need without repeating the tool name. Keep status honest at every transition: proposed -> building -> review -> approved, or `declined` when omitted.
- `_status.ndjson` carries phase transitions `{"ts","phase"}` (propose | build | review | verify | done), suggestion steps `{"ts","suggestion","step","state"}` (code | verify; start | done), and optional run steps. A finished verify step includes `"outcome":"verified|failed|could-not-verify"`. Phase and step are different shapes.
- `_chat.ndjson` carries replies and the Propose handoff: `{"ts","from":"claude","re":"<suggestion id|null>","text"}`. Text is markdown-lite: blank-line paragraphs, `- ` bullets, `**bold**`, and inline backticks only.
- `_feedback.ndjson` is server-owned. New entries are `{"event_id","run_id","order","type","ts","payload"}`; old lines without the envelope remain readable for old runs. Never write this journal. Browser requests are `{"request_id","type","payload"}` and are not durable until the server returns `{"type":"recorded","request_id","event_id","run_id","order"}`. If a disconnect loses that response, the Explorer checks a newer current-run snapshot for the exact type/payload and a post-send timestamp/order. It keeps submit/approval disabled during a short reconciliation window, drains any coalesced journal update before deciding, accepts a delayed matching append, and offers an explicit retry only after the fresh journal still proves the action absent.
- `_delivery.ndjson` is server/hook-owned delivery evidence (`waiting | claimed | timeout | conflict | error`). A claim is not model receipt and is not handling. An eventless `error`/`conflict` after a run-level waiting window replaces that window for later unacked events; event-specific delivery evidence still takes precedence for its event.
- `_ack.ndjson` is the portable handled trail: `{"ts","run_id","event_id","status":"handled"}`. Only a matching handled ack means Codex completed the event's effect.
- `<id>.code.md` is the Explorer review copy of the real tool module. Generate it mechanically from `<srcroot>/webmcp/<id>.<ext>`; never retype it.
## Startup
0. Quietly run `command -v bun`. If missing, offer **install Bun** and browser review, or **review in chat** with `--no-interactive-loop`. Install only with consent. Missing Bun never implies `--non-interactive`.
1. Start `bun <this skill's dir>/interactive/server.ts <workspace>` as a background task whose stdout remains readable. Do not add `--resume` for a new run: a normal start rotates run identity.
2. Use the complete Explorer URL printed by that process. It includes `/?capability=...&run_id=...`; do not rebuild it from a remembered port. Open that URL unless a robot/headless client is driving.
3. Treat the matching `.run.json` as the endpoint record. Health and shutdown are run-aware: `/healthz?capability=<capability>&run_id=<run_id>` and `/shutdown?capability=<capability>&run_id=<run_id>`.
4. Choose the available wake path:
- **Codex:** inspect `/hooks` first. If it lists this plugin's bundled `PreToolUse` and `Stop` hooks, use the normal trust flow and ask the developer to review them there when needed; never use a trust-bypass flag. Automatic delivery is bound to this Codex task/session for the lifetime of the run; a second task cannot inherit or steal it. Listed hooks may use `PLUGIN_ROOT`/`PLUGIN_DATA`, but they never interpret product phases or mark an event handled. Their `Stop` path guards `stop_hook_active` so it cannot continue itself forever. If `/hooks` does not list this plugin, the current host is not exposing that wake capability: make no automatic-delivery promise and use same-task manual replay.
- **Claude Code, where Monitor is supported:** arm a persistent Monitor on the exact `ws://localhost:<port>/ws?role=claude&capability=<capability>&run_id=<run_id>` from the run metadata. Keep it armed through phase Done.
- **Fallback:** absent, timed-out, disabled, or untrusted hooks and an unavailable Monitor do not strand the run. The Explorer says only that the action is safely saved; ask the developer to resume or message the original Codex task bound to this run. On that turn, scan the journal manually. A fresh task may also scan and reconcile unacked events manually, but it must not claim the old task's automatic binding. Mention `/hooks` only if it lists this plugin in the bound task.
## Handling a recorded event
Hooks and Monitor notifications are wakeups, not workflow authorities. On every automatic or manual wake:
1. Read `.run.json`, then re-read the `_feedback.ndjson` envelope matching both `run_id` and `event_id`. Do not act from notification prose or socket delivery alone. Reject malformed/wrong-run candidates.
2. If `_ack.ndjson` already has matching `run_id` + `event_id` + `status:"handled"`, skip it. Otherwise reconcile any partial effect left by an interrupted earlier attempt before doing more work.
3. Handle the event using the phase judgment below and the skill's normal rules. Hooks must not encode or infer phases.
4. Persist the event's effect first. Then run `bun <this skill's dir>/interactive/ack-event.ts <workspace> <run_id> <event_id>`. Never hand-write `_ack.ndjson`, and never ack a socket send, delivery claim, intention, or failed effect.
5. When scanning manually, process every valid unacked envelope for this run in ascending numeric `order`, one at a time. Re-read the ack trail between items. A burst is server-ordered, not timestamp-ordered.
For `approve`, durably set shipped suggestions to `approved` and append phase `verify`, then ack immediately before the lengthy verification/PR work. On resume, that durable transition proves approval was accepted; continue the ladder without applying approval twice.
## Phases
1. **Propose.** After A-C, write `plan.json`, then append one `_chat.ndjson` handoff (`re:null`) under 200 words: one sentence with the tool count and why, one `- ` bullet per tool, only genuine decisions needing attention, then the available actions. Append phase `propose`, point the developer to the printed Explorer URL, and end the turn.
2. **Decide.** For each `comment`, append a short reply and revise `plan.json` when needed. Picks are already durable but are not actionable, so the model does not handle or ack them. End the turn after the comment effect and its ack.
3. **Build.** `submit` is Phase-D approval. First set omitted suggestions to `declined` and append phase `build`; that is the durable acceptance boundary. Start the dev server and watch-mode typechecker. Build approved suggestions in order: status building -> code start -> write the real module -> mechanically regenerate `<id>.code.md` -> code done -> verify start -> scoped static checks -> verify done with outcome -> status review. Write the entry module last, after the boot baseline. Handle mid-build feedback at the next safe pause. A `cancel` finishes the in-progress unit, preserves built tools at review, declines never-started tools, explains the result in chat, appends phase review, then acks; it is not Done.
4. **Review.** Plain `comment` is plan-level feedback. `feedback` for a built suggestion revises the real module and review copy, reruns scoped checks, appends a fresh verify outcome, and replies with the exact change. Stay in review. A declined suggestion stays terminal unless the developer explicitly restores it to proposed; building it still needs a fresh submit or unambiguous go.
5. **Verify.** After the durable approval transition and ack, run the full Phase-F ladder against the warmed server. Append fresh per-suggestion verify start/done lines; build-time static outcomes were not runtime verdicts. Track actual docs/PR work with run-level start/done lines. The PR includes the durable `.webmcp/` plan and approval trail, never live runtime plumbing.
6. **Done.** Only after the PR exists, append phase `done`, then tear down. Never mark Done while verification or PR creation remains.
## Rules and recovery
- A `comment` beginning `Tool request:` asks for a capability. During propose, add/revise its suggestion. After submit, add it as proposed and require explicit go before building.
- `.webmcp/` plan/review state is the ADR-0003 carve-out from "nothing before approval". Tool modules, dependencies, and branches remain gated on submit.
- **Resume in the same agent task/session:** read and validate `.run.json`, call the run-aware health URL, and reprint its capability/run URL when live. If the recorded process is dead, start `bun <this skill's dir>/interactive/server.ts <workspace> --resume`, use its new stdout URL/metadata, restore that task's wake path, then scan valid unacked events in order and continue from durable state. If `.run.json` is absent or invalid, there is no identity the server can safely resume; report that and rotate only when the developer asks for a new run. Use `--resume` only for same-task crashed-server recovery.
- **Continue from a fresh agent task/session:** first read the old journals and reconcile valid current unacked events manually in numeric `order`; the old automatic hook binding does not transfer. Manual reconciliation may finish the old run. If the developer instead wants continued automatic browser interaction, explicitly rotate: cleanly call the run-aware shutdown endpoint when the old server is live, or abandon its dead runtime metadata when it is not, then start `bun <this skill's dir>/interactive/server.ts <workspace>` without `--resume`. Use only the new printed URL/run metadata and establish the new task's own wake path. Never describe this as resuming or transferring the old hook session.
- **Teardown:** after Done, stop a Claude Monitor if one exists, POST the run-aware shutdown URL, and confirm the recorded port no longer listens. Codex hooks are plugin lifecycle hooks, not per-run background jobs to disable. A clean shutdown clears runtime liveness; journals remain as the review trail.
SHA-256: e7f382875f7506786f03c37d739bdcceacd55aa48978c06e04c11e88537a56cb