← Files Compound EngineeringARCHIVED FILE

skills/ce-work/references/cross-model-execution.md

30 KB · Oct 3, 2026 · 06:34 UTC

↓ Download file

# Cross-Model Execution Contract

Load this reference only after the cross-model engine is selected or recovery of an existing external run is activated. It defines the fixed-route, authority, fallback, identity, receipt, and serial transaction contract. The host drives the bundled controller, detached runner, and adapter; no worker response or process exit can substitute for controller and Git evidence.

## Resolve one requested route

Use only these targets: `codex`, `claude`, `grok`, `cursor`, `composer`, and `opencode`. Keep five identity facts separate in every disclosure and receipt: target, harness/intermediary route, requested model, actual model, and receipt status.

**Fixed controller route tokens:** record exactly `codex`, `claude`, `grok-cli`, `cursor`, `composer`, `grok-cursor`, or `opencode` in the egress sanction. `grok-cli` maps target `grok` to its native harness; `grok-cursor` maps target `grok` through intermediary `cursor`. These controller tokens are not descriptive route labels.

- `cursor` means the Cursor harness with its configured default model.
- `composer` means a Composer-family model through Cursor.
- `grok` prefers its fixed native route; a Grok model through Cursor is a different intermediary and must be separately permitted and sanctioned.

For an ordered standing preference, preflight candidates in order without egress. Skip a candidate only when its harness and requested/default model are equivalent to the current host, or when observed evidence makes it unavailable. Attempt the documented adapter recipe first; local CLI help or version information may refine a compatible mapping only inside the same sanctioned harness/model family and only if the fixed adapter still enforces every restriction. An explicit model pin cannot become another model. The first qualified candidate becomes the one fixed recipient.

Before egress, list traversal may continue after a candidate is proven unavailable. After dispatch starts, the adapter receives one fixed recipient and must never switch recipients, providers, or intermediaries internally. A different recipient requires a separately resolved and sanctioned attempt after authoritative terminal or reaped state; it is never an in-flight fallback.

If a target asks for the same-host default with no distinct serving route or model, collapse to native execution. Record the target as requested and native as actual; do not create an external job merely to call the current host's default model through itself.

## Apply preference or requirement strength

Cross-model implementation routes are write- and shell-capable. Never request broader host permissions merely to make one reachable. If the current host boundary makes a route unavailable, apply the native fallback for its resolved mode instead of escaping that boundary.

**Preference-strength (`prefer`):** attempt the fixed route in both direct and automatic workflows. If preflight proves it unavailable, continue with native execution and prominently report requested versus actual route/model plus the observed `fallback_reason`.

**Requirement-strength (`require`):** keep the requested external identity fixed while the route is viable. If preflight proves it unavailable, disclose the reason once and continue on the current harness and session model without prompting, erroring, elevating the host boundary, or substituting another external recipient. A started attempt is not preflight-unavailable: no fallback begins until it reaches an authoritative terminal or reaped state.

## Sanction before egress

Before repository content or bounded mutation authority leaves the host, disclose and durably record:

- the binding `source` that authorizes external execution;
- the fixed recipient, provider, harness route, and every intermediary;
- the repository/unit material exposed, including the bounded source/unit packet and workspace content;
- the caller restrictions and which are adapter-enforced versus cooperative; and
- that linked worktree isolation contains accidental concurrent mutation but is not an OS security sandbox.

Treat a required restriction the adapter cannot enforce as route unavailable; apply `prefer` or `require` instead of weakening it silently. The sanction is route-specific. Any retry that changes a recipient or intermediary requires a newly resolved and sanctioned job.

## Bound worker authority

The worker receives one unit, one workspace, one fixed recipient, and only inherited authority. It may narrow scope or authority, never broaden either. Its packet grants no canonical commit, push, PR, shipping, recipient-switch, fallback, peer scheduling, or scope-expansion authority.

An external worker may edit only inside its controller-owned detached worktree. Do not instruct it to run `git add`, `git commit`, or another Git index write. Leave the completed working tree uncommitted; the host snapshots the tree. Codex `workspace-write` and Cursor `--sandbox enabled` cannot write the linked-worktree Git admin dir: the index lives in the shared Git common dir, outside the workspace. Worker commits are never required. Successful output terminalizes as an isolated transport commit from the complete working tree for the host to inspect. Only the host may apply output to the canonical checkout, run authoritative verification, create a host-only canonical commit, or proceed into the standalone or caller-owned tail.

A Codex `workspace-write` packet must also treat socket binds, OS permission checks, and peer-credential probes as host-owned. Preserve the host command and observed result; do not treat a sandbox `EPERM` as proof the host lacks the capability.

Ordinary synchronous native units stay in the active checkout. Ordinary native subagent isolation remains harness-owned. Only the external cross-model controller may create detached sibling worktrees under its private run root; this exception does not authorize `ce-work` to create worktrees for native execution. An active checkout that is already a linked worktree does not disable this route: the controller registers another detached **sibling** through the shared Git common directory under `/tmp/compound-engineering-<effective-uid>/ce-work/<run-id>/`, never a nested worktree beneath the active checkout.

## Preserve route and lifecycle receipts

Direct and return-to-caller runs expose the same receipt facts even though their prompting and shipping tails differ:

- `implementation_engine_binding`: resolved `mode`, `target`, `model`, and `source`;
- `requested_route` and `actual_route`, including every intermediary;
- `requested_model`, `actual_model`, and served-model receipt status;
- `fallback_reason` or `null`;
- `source_kind` (`plan` or `prompt`) and its controller-recorded digest;
- `run_id` and per-unit `unit_receipts` that distinguish process terminal state, integration, authoritative verification, host canonical commit, and cleanup;
- `plan_checkpoint`: a disclosed host commit only when the resolved selected plan was the sole canonical dirt;
- `blockers`; and
- `recovery_path` for preserved inspectable state.

Never infer success from detached-process completion alone. The run receipt is complete only after every required unit reaches its host-owned canonical state. A plan-only checkpoint is disclosed to the direct user or returned to an automatic caller; unrelated dirt never receives an implicit checkpoint and instead makes the external route unavailable.

## Build a source for bare-prompt work

A formal plan is not required when Phase 0 has already judged a bare prompt concrete enough to execute. Before controller initialization, write one bounded implementation brief directly to OS temp outside the repository. The brief is a distilled authority artifact, not a transcript: never include raw conversation history, unrelated session context, credentials, or speculative scope.

Populate these headings with concrete evidence from the request and Phase 0 discovery:

- `Request` — the current implementation request, paraphrased only to remove unrelated conversation;
- `Goal` — one observable outcome;
- `Scope` — expected files/surfaces and discovered patterns/tests;
- `Acceptance and verification` — behavior and authoritative checks that prove completion;
- `Constraints and exclusions` — inherited restrictions, non-goals, and unresolved boundaries; and
- `Units` — use one conservative `P1` unit by default; create more only when discovery establishes distinct goals, dependencies, expected files, and verification without guessing.

If any of Goal, bounded Scope, or Acceptance and verification cannot be populated, do not initialize or egress. Return to Phase 0 clarification or planning. Compute the brief's SHA-256, call controller `init` with `--prompt-brief <temp-path> --prompt-digest <sha256>`, and use the controller-owned private copy and digest thereafter. The caller's temp path is never authoritative after initialization. Invocation origin does not change this contract.

## Serial external-unit protocol

Run this protocol from the host checkout for one ready unit at a time. Resolve each bundled script from this skill's directory. The host remains the orchestrator throughout; the external CLI is only the bounded author. Use `git -C <canonical-checkout>` for every host Git operation; never rely on the shell's current directory after resolving or running a bundled script.

**Visibility and stop invariant:** use separate host tool calls and invoke only one state-changing controller transition per call until terminal scope inspection. Never generate or run a shell script that spans `start` through waiting or integration, spans multiple units, loops over the external runtime, or conditionally continues into another transition. `start` must return before supervision. Cap each runner wait at 60 seconds, then sync and publish progress in later host tool calls. After scope inspection, use the controller's single fail-stop `integrate` transaction; do not manually chain its internal Git/controller transitions. A nonzero controller, runner, verification, or Git exit ends that host tool call; inspect controller status and enter its prescribed restoration or recovery path before any later transition.

When set, `CE_WORK_RUNS_ROOT` is the parent CE Work directory containing all `<run-id>/` directories, not an individual run directory.

1. **Resolve, preflight, and sanction.** Apply the authority-and-scope resolution in `execution-engines.md`. For an ordered standing preference, skip an equivalent self-route and preflight each remaining candidate until the first qualifies; record every rejected candidate and reason. Verify the fixed adapter and caller restrictions, then disclose the sanction source, route/intermediaries, material exposed, and restriction posture. The controller `init` egress object uses the exact plural keys `route`, `intermediaries`, and `restrictions`: direct `codex`, `claude`, `grok-cli`, and `cursor` routes use `intermediaries: []`, while `composer` and `grok-cursor` use `intermediaries: ["cursor"]`. If all candidates are unavailable, both `prefer` and `require` disclose the attempted routes and reasons once, then continue natively on the current harness and session model without starting an external controller run.
2. **Establish the durable source and clean canonical state.** Initialize the private run with `unit-workspace.py` `init`; the controller owns creation of `/tmp/compound-engineering-<effective-uid>/ce-work/<run-id>` (under `$TMPDIR/compound-engineering-<effective-uid>/` instead when `/tmp` cannot host a writable private root, as in a sandbox that only allowlists `$TMPDIR`; `unit-workspace.py` resolves that itself). Do not pre-create the run directory. For a repository plan, pass `--plan <path> --plan-digest <sha256>`; the selected plan may be the only dirty path, and `unit-workspace.py` `checkpoint-plan` records that exact plan-only checkpoint. That exact plan-only state is checkpointable, not a route blocker. For a bounded prompt brief, pass `--prompt-brief <temp-path> --prompt-digest <sha256>`; the controller copies it into private run state, and the canonical checkout must already be clean because no repository source is eligible for checkpointing. Only checkpoint failure or unrelated dirt makes the route unavailable. Re-read canonical repository, branch, HEAD, source kind/digest, and cleanliness before preparing a unit.
   Once `init` returns `READY`, the engine decision is closed for that unit: the next implementation path is `prepare` then the fixed-author start, or a blocker that preserves the returned recovery path. Expected latency, orchestration cost, or a later judgment that native work would be simpler is not route unavailability and never authorizes canonical writes. Native implementation becomes eligible only through the later controller fallback gate.
3. **Prepare one bounded unit packet.** For a plan, include only the Goal Capsule, Definition of Done, active unit, relevant Verification Contract rows, cited R/F/AE/KTD excerpts, any Product Contract Key Decision whose exact `Governs R…` links name the active unit's cited R-IDs, inherited restrictions, expected files, evidence strategy, and explicit exclusions. For a prompt brief, include only the matching P-unit's Goal, Scope, Acceptance and verification, Constraints and exclusions, expected files, and evidence strategy. Do not expose the whole plan, prompt brief, or conversation by default. Write the packet source directly to OS temp outside the canonical checkout, such as a source file beneath the controller-returned recovery path; never draft it inside the repository and move or copy it later. Call `unit-workspace.py` `prepare` with that packet source path, unit id, recorded base, dependencies, wave fields, and the route-qualified activity posture. Use only the controller-returned `attempt_id`, controller-owned `packet_path`, and computed `packet_digest` for dispatch; never substitute the caller's source path, a caller-computed digest, or an assumed attempt name.
   `prepare` refuses only on unrelated dirt or a moved HEAD; the size and shape of the canonical checkout's git-ignored inventory (installed dependencies, virtualenvs, build caches, symlinks) never make the route unavailable.
4. **Start one fixed author.** The controller-issued authorization schema binds `run_id`, `unit_id`, and `attempt_id` to the fixed route/model/intermediary and packet contract. Call `peer-job-runner.py` `start --no-sweep --input-digest <controller-packet-digest>` with skill `ce-work`; the runner label must equal the unit id exactly, without the attempt id or another suffix, and `--result-path` must be `<controller-result-dir>/implementation-result.json`. Both `--input-digest` and the adapter's expected-packet argument must use the exact `packet_digest` returned by `prepare`; omitting the runner flag or recomputing/substituting either value makes the job ineligible for `record-job`. The runner and controller share the same root automatically: `CE_WORK_RUNS_ROOT` wins when set, otherwise both derive CE Work state from `CE_PEER_JOBS_ROOT`. Set `CE_PEER_HARD_SECS=7200` on every production CE Work runner start rather than relying on the shared runner default. Set `CE_PEER_IDLE_SECS=600` for route-qualified `incremental` activity and `CE_PEER_IDLE_SECS=0` for `hard-only` or otherwise untrustworthy activity. The 600-second window resets on progress and detects a stall; it is not a wall-clock maximum. Its production worker command is `cross-model-work.sh <authorization_path> <workspace> <unit-packet> <expected-packet-sha256> <result-dir>`; invoke the returned adapter path directly as the first worker argv, without a `bash`, `sh`, or `env` prefix, and pass only the controller-returned `authorization_path`, never a caller-supplied route string or ambient model override. The runner exports its controller-visible job id to the worker. Before prompt construction or external CLI start, the adapter must pass that runner-exported job id and obtain controller `authorize-dispatch` success by calling `unit-workspace.py` `authorize-dispatch`. That success reads the actual runner metadata and exact worker argv, rejecting a shell prefix or substituted adapter before egress, then atomically binds that job id to the exact attempt before egress along with the authorization digest, workspace, packet path and digest, and result directory while revalidating the controller-owned exact route, model, and intermediary contract. A second job for the attempt is refused. A missing or refused handshake, including hand-authored or cross-attempt authorization, refuses egress. Model pins come only from the authorized projection. The adapter's `--emit-adapter` mode remains introspection only and never authorizes production dispatch. Capture the returned job id immediately with `unit-workspace.py` `record-job`, using the controller-returned `attempt_id` verbatim; this idempotently confirms the same validated binding. Never hold one host tool call open for the external runtime and never let the adapter select another recipient.
5. **Observe without steering.** Interleave runner `status --skill ce-work` or `wait --skill ce-work --max-secs 60` calls with separate `unit-workspace.py` `sync-job` calls; every bare-job-id runner `status`, `wait`, `result`, or `reap` call must carry `--skill ce-work` under the same `CE_WORK_RUNS_ROOT` / `CE_PEER_JOBS_ROOT` selection used at start. Report unit, route, elapsed time, latest meaningful activity, activity posture, and terminal state after each cycle. For a qualified silent terminal-only route, `hard-only` is the normal posture: disable idle timeout, retain the universal hard cap, and never infer failure or fallback merely from absent incremental activity. A live or temporarily unreachable attempt is still authoritative: do not start fallback or duplicate work. Reap only through the explicit controller/runner path and retain its workspace.
6. **Terminalize complete Git output.** On authoritative `done`, call `unit-workspace.py` `terminalize`. Require its pinned synthetic transport commit to have the recorded base as sole parent and the complete final workspace tree. In a later host call, inspect the actual transport diff, changed paths, binary/mode/rename/delete evidence, adapter result, packet expected scope, and any scope-expansion request. A generated byproduct or any unexplained difference between actual transport paths, expected scope, and the worker's evidence is unexpected scope: preserve it for host resolution and do not acquire integration.
7. **Integrate through the fail-stop controller transaction.** After the separate scope-inspection call accepts the complete transport, invoke `unit-workspace.py integrate --run-id <run-id> --unit-id <unit-id> --commit-message <message> --verification-summary <summary> [--allowed-head <recorded-head>] -- <verification-command-and-argv>`. Pass a simple Verification Contract command as direct argv. If it contains shell syntax such as `$(...)`, a pipe, `&&`, a redirect, or a glob, invoke an explicit shell with pipe-failure handling on the first attempt, for example `-- bash -o pipefail -c 'test "$(cat delegated.txt)" = "expected"'`; quoting `$(...)` as a direct argument does not expand it. Do not manually chain or conditionally reproduce this transaction.
8. **Let `integrate` own canonical mutation.** It performs `unit-workspace.py` `integration-acquire`, `unit-workspace.py` `preflight`, `git cherry-pick --no-commit`, `unit-workspace.py` `mark-applied`, authoritative canonical verification with Python bytecode disabled, exact canonical-state reconciliation, `unit-workspace.py` `mark-verified`, one host-owned commit, and `unit-workspace.py` `mark-committed`. For a wave it also records `wave-advance`; after a reconciled commit it performs `unit-workspace.py` `cleanup` and `unit-workspace.py` `integration-release`.
9. **Treat its outcome as authoritative.** Capture the authoritative command's exit status directly; never infer a pass from stdout. Any change to tracked state — HEAD, branch, index, or status-visible paths — fails reconciliation. Git-ignored state is not part of that proof: the controller inventories ignored entries by metadata before and after verification and discloses `changed`, `removed`, and `created` counts with a bounded path sample as `ignored_state` in every verification receipt and return body; it never copies, restores, or deletes ignored files, so a verification command that mutates a cache or a local database changes only what native units already could. Ignored files the command creates stay in place. Empty ignored directories are not inventoried. `.env`-class files and local databases receive no protection from a misbehaving verification command. Verification and clean-state reconciliation happen before `mark-verified`; a rejected verification or pre-commit controller/Git step forbids the commit, calls `unit-workspace.py` `restore`, proves exact pre-fold equality, and releases the lock before fallback, retry, or another unit. If exact restoration cannot be proven, it retains the lock, workspace, transport ref, and recovery path. A crash after the canonical commit remains resumable evidence and must not be redispatched.
10. **Run the source-wide gates through the controller.** After every external unit is committed and cleaned, invoke `unit-workspace.py verify-run --run-id <run-id> --verification-summary <summary> -- <verification-command-and-argv>` for the plan-wide Verification Contract gates or the prompt brief's authoritative checks. Do not run those final commands directly in the canonical checkout or hide their status behind a later command. `verify-run` starts only from a clean canonical snapshot, captures the command's exit status directly with Python bytecode disabled, restores tracked state to the exact starting snapshot, discloses ignored-state divergence as `ignored_state` per step 9, and records a durable receipt. A failing gate retains its private log and blocks the return or shipping tail even after cleanup succeeds.

Project every transition into the direct commentary or return envelope: resolved source kind, route, plan checkpoint when applicable, dispatch, activity, terminal result, transport/scope inspection, integration, restoration if any, authoritative verification, canonical commit, cleanup, blocker, and recovery path. The serial protocol does not authorize parallel waves or automatic redispatch; those require their later gates.

## Parallel external-wave protocol

Use a wave only after the always-loaded Parallel Safety Check proves the ready units independent across dependencies, declared files, shared contracts/interfaces, migrations, lockfiles, generated/registry/config surfaces, environment singletons, and expected merge cost. Uncertainty selects serial execution. The scheduler remains a host responsibility: workers cannot add peers, broaden the wave, or change their own dependencies. Cap the wave at 3-5 workers.

1. Record one wave id, dependency order, and **one recorded wave base** from a clean canonical checkout. Prepare every member from that same base with distinct controller-owned workspaces. Synchronous native work remains in the active checkout; every concurrent worker is isolated.
2. Start the fixed-route jobs concurrently, observe them without steering, and terminalize every worker into its base-parented synthetic transport commit **before the first fold-in**. If any worker has not terminalized, integrate none. Inspect all actual transport inventories before mutation; same-path changes stop the wave, while shared-interface or other semantic contention returns to host resolution even when paths differ.
3. Integrate results sequentially in dependency order under the canonical integration lock. The first unit preflights against the wave base. A clean textual apply is not semantic compatibility: inspect scope, run that unit's authoritative verification against the advancing canonical tree, and create one host-owned canonical commit.
4. While that unit's lock is still held, call `unit-workspace.py` `wave-advance` with its exact recorded canonical commit. The controller accepts only the manifest-confirmed commit whose parent is the recorded pre-fold HEAD, then adds it to later siblings as one of the **exact earlier host-owned canonical commits**. Clean and release the completed unit before acquiring the next sibling's lock.
5. For each later result, pass only the current exact recorded wave commit to `preflight`. Unknown HEAD movement blocks. Even when three-way application is clean, repeat actual-scope inspection, semantic revalidation, authoritative verification, and a separate host commit. A transport commit is never merged or committed directly.
6. Conflict, scope expansion, semantic collision, test failure, or commit failure enters the serial protocol's exact restoration while the lock remains held. No sibling, retry, or fallback starts until restoration is proven and the lock is released; affected dependents remain queued. After readiness is recomputed, unaffected siblings may continue; the affected unit requires explicit host resolution, re-dispatch on the new base, or serial fallback; never blind-merge a colliding or stale result.
7. Repeated collision, broad unplanned edits, or inability to prove exact restoration disables more waves for the run. Preserve inspectable workspaces and refs; failure of one unit does not discard an unaffected sibling, but it never authorizes a dependent unit early.

## Resume and fallback exactly once

When the caller supplies a run id directly or through return-to-caller `implementation_run:<safe-id>`, use `unit-workspace.py` `resume --run-id <id>` and treat that id as authoritative. This recovery activation occurs before ordinary input classification and must not dispatch new work. Without one, a plan-backed run must use `resume --repo <canonical-checkout> --plan-digest <selected-plan-digest>` for discovery; never infer or select a run id by listing the shared run root. Prompt-backed runs require their disclosed run id; do not rediscover them from conversation text or a caller temp path. Resume only a single unfinished plan-backed run matching repository identity, branch, and plan digest; if several match, list the matching run ids and recovery paths and block selection rather than choosing one. Do not dispatch a new third run while matching unfinished state remains ambiguous. An unsafe or foreign-mode manifest is a blocker, not a candidate to skip.

When all units are terminal (`cleaned` or `native-completed`) and a successful plan-wide verification receipt already exists, the completed run is observation-only. Build the return envelope from the manifest and stored receipts; recovery must not rerun a Verification Contract gate, call `verify-run`, or execute test/build/format/install/generation commands merely to reconfirm completed evidence. A missing required receipt is a reported recovery blocker, not permission to improvise a loose verification tail.

Resume reconciles durable evidence, but must not redispatch, reapply, recommit, or run either owning tail. It may adopt exactly one matching unbound runner job, monitor a recorded live job, terminalize authoritative `done` output, continue an interrupted exact restoration, or record a verified canonical commit whose parent and tree match the pre-fold evidence. A transport ref written before its manifest row is reusable only when its sole parent and final tree match. Preflight records the expected post-apply tree and changed-path set: an exact match may reconcile an apply that landed before its manifest row, while unknown dirt blocks without destructive restoration.

Use the controller's `status`, `reap`, and `cleanup` operations for preserved work. Loss of contact leaves a recorded job live until runner evidence becomes terminal or an explicit reap records termination. Cleanup is idempotent after an interrupted worktree removal. Explicit abandonment requires the exact transport SHA, or the exact terminal job id when a failed/reaped attempt produced no transport.

After exact restoration, exact abandonment cleanup, and integration-lock release, re-dispatch a corrected unit under the same scalar `run_id` by calling `prepare` with a fresh `attempt_id`. Preserve the earlier attempt receipts; do not mint another run id or represent one logical run as a list of run ids. The controller refuses a retry that changes the unit's recorded dependency or wave/base contract.

Post-start fallback is a separate atomic gate. After authoritative failure, timeout, `died-without-result`, or exact restoration and lock release, disclose the unavailable route once and call `unit-workspace.py` `claim-fallback` before native implementation. The first `prefer` or `require` claim authorizes exactly one fallback on the current harness and session model; `FALLBACK_ALREADY_AUTHORIZED` means do not start it again. Requirement strength prevents external recipient substitution while the route is viable, but never turns an unavailable route into an error or user-choice gate. A live job or successful unreconciled output refuses the claim. If integration began, no retry, sibling, or fallback claim is eligible until exact restoration is recorded and the canonical checkout still equals that snapshot. After the claimed native implementation is committed and locally verified, call `unit-workspace.py` `complete-fallback` with the accepted HEAD, a SHA-256 digest of the local verification evidence, and a bounded summary; both `prefer` and `require` claims can complete. `FALLBACK_COMPLETED` closes that unit only. Once every unit is terminal, run the plan-wide Verification Contract through `verify-run`; do not report the run complete until the controller returns `RUN_VERIFIED` and stores its successful receipt.

## Preserve tail ownership

The engine changes only implementation authorship. A standalone invocation resumes `ce-work`'s quality and shipping workflow after local implementation. `mode:return-to-caller` returns implementation and local-verification receipts with `standalone_shipping_skipped: true`; it never runs simplify/review/PR/CI gates owned by the caller. External workers never inherit either tail.

**A successful controller `init` locks that unit to the selected cross-model engine.** From that point, advance it through this controller protocol or return blocked with its recovery path. Never reclassify it as trivial, abandon it for speed, or implement it natively unless the protocol later returns an explicit fallback authorization.

SHA-256: f3316717f4af8e314b97f95e987aa82023e99059a2327b791046a3ebbf920b82