← Files Codex AdvisorARCHIVED FILE
skills/consultation/references/operations.md
22 KB · Oct 2, 2026 · 00:32 UTC
# Consultation operations
Advisor provides pre-decision advice only. It does not implement, route
implementation, perform final review, or replace root authority. The root performs
any repository or web research before consultation and sends enough relevant evidence
and source references in the five-section packet for a decision. If the evidence is
not enough, the advisor may identify only a concrete research-first next step, missing
evidence, research questions, or bounded brainstorming areas. The advisor uses zero
tools: it does not inspect files, fetch the web, or conduct independent research.
The root may assign bounded evidence gathering to separate research workers before assembling the decision packet; the consulted advisor still uses zero tools and never delegates.
## Install and verify the companion roles
## Live configuration and advanced tools
The installed `advisor.toml` is the authoritative live configuration for normal
`--tier standard|specialist` calls. It has only `[standard]` and `[specialist]`, each
with a model and effort. Edit it in place: the next consultation reads it, freezes its
pair and SHA-256 source revision through retry, and requires no catalog, discovery,
canary, or local state. Future syntactically valid selectors are allowed subject to
account/runtime support; Specialist Astra opts into higher usage. Plugin updates or a
reinstall can replace edits. `scripts/advisor-config.sh show` and `doctor` display the
actual installed path, revision, and pairs.
The helper's settings, prior-settings recovery, catalog evidence, canary receipts, and
optional journal are advanced legacy state under Codex home, outside the plugin cache.
`scripts/advisor-config.sh models refresh` is an explicit, bounded read of the Codex
app-server using the existing Codex provider/auth context. It uses
`--ignore-user-config --ignore-rules`, sends only initialize/initialized/model-list
JSON-RPC messages, includes hidden models, and preserves existing catalog state on
offline, protocol, or authorization failure. It never copies credentials, writes user
Codex configuration, enables tools, or runs inference. New launcher calls use
`--tier standard|specialist`; raw model/effort overrides are not accepted. Live
defaults are Terra/high and GPT-6 Sol/high without state creation, discovery, or a canary.
That permission to attempt is not compatibility proof: each
real child remains subject to the existing exact runtime identity, effort, read-only,
zero-tool, and schema checks. Legacy role calls remain pinned.
If the installed CLI does not expose both isolation flags for `app-server`, refresh
returns unavailable without launching it or changing the catalog. Tested CLI
provenance `0.153.2` is currently in that safe-unavailable state; discovery never
blocks built-in defaults. Use `models add MODEL` for a future manual selector, but
catalog presence, including optional Astra, is neither entitlement nor compatibility.
`models test MODEL
--effort EFFORT --authorize-usage` is the only inference-based configuration command.
It runs a fixed content-free packet through this same transport, permits at most one
response-only retry, records the exact local CLI version as provenance on success, and
never changes a tier selection. CLI updates alone do not stale evidence. It requires
explicit usage authorization; configuration metadata or a saved selection is not
consent. `set` and `reset` reject as legacy-only rather than reporting an active tier
change; `restore` reports that it only restores legacy state.
`journal status|enable|disable|clear` manages a content-free journal disabled by
default. Pruning older-than-30-day records happens during a later journal write, not
through a background deletion service.
From the installed plugin root:
```sh
sh scripts/install-agents.sh
sh scripts/install-agents.sh --check
```
The installer adds `advisor-terra.toml`, `advisor-sol.toml`, and the explicit opt-in
`advisor-astra.toml`. Astra is byte-exactly installed and checked as an independent
fixed role; it is never selected by Standard/Specialist defaults, trigger selection,
or fallback. During an attended upgrade it
recoverably retires byte-exact known historical Luna, Terra, and Sol-reviewer files
to `<role>.toml.retired-v0.6.0`, the old Sol consultation role to
`sol-advisor.toml.retired-v1.0.0`, and the obsolete neutral role to
`advisor.toml.retired-v1.0.1`. It preflights every path before mutation, is
idempotent, refuses symlinks/nonregular files/modified content/collisions/dual paths,
and never edits Codex configuration. An exact Advisor 1.1.0 upgrade recoverably
retires the prior model-pinned roles to `advisor-terra.toml.retired-v1.1.0` and
`advisor-sol.toml.retired-v1.1.0`, then installs the risk-described 1.3.0 roles at
their original active paths. Later exact 1.3.0 generations retire separately to
`.retired-v1.3.0` and `.retired-v1.3.0-zero-tool`, preserving both predecessor
files without collision. Exact retired-only interrupted states resume safely;
modified, dual, or colliding states refuse all mutation.
## Root and advisor records
For a consult candidate, before the first implementation write and before the
decision record, run:
```sh
sh <absolute-installed-plugin-root>/scripts/inspect-parent-runtime.sh
```
This preflight identifies the parent only with `CODEX_THREAD_ID`; it never falls back
to `CODEX_SESSION_ID`. It resolves one regular, nonsymlinked rollout in the
caller-supplied/default sessions root and accepts an unambiguous recognized sandbox
and permission profile. The parent may be `workspace-write`; the separate consultation
transport owns read-only isolation. Missing, malformed, duplicate, or conflicting
evidence is typed unavailable.
Then emit:
```text
ADVISOR DECISION
route: consult | skip | unavailable
reason: <one task-specific sentence>
question: <bounded decision question, or none>
```
An identified parent permits `consult`, including a normal `workspace-write` root.
An unavailable preflight emits `route: unavailable`, with no `ADVISOR CALL`, no
consultation process, and no block on root-owned work. The ordinary `skip` route
remains unchanged.
For a consult, select the tier from decision risk:
- Standard: `--tier standard`, defaulting to Terra / `high`. This covers ordinary
bounded material architecture, interface, data-model, and generic advisor requests.
- Specialist: `--tier specialist`, defaulting to GPT-6 Sol / `high`, only when targeted evidence still leaves unresolved a cross-module or system design, compatibility or concurrency boundary, competing diagnosis, an unresolved security or trust boundary,
recovery, an irreversible migration or data-loss decision, or a credible unresolved
High-severity disagreement.
Security adjacency or project importance alone, or an ordinary architecture question
alone, does not qualify for Specialist. A borderline role choice uses Standard. The
parent model is irrelevant.
Legacy cached callers may still use `--role advisor-terra` or `--role advisor-sol`;
those routes remain fixed to Terra/high and GPT-6 Sol/high, do not follow `advisor.toml`,
and cannot be mixed with a tier or preset.
An explicit caller may use `--role advisor-astra` for the fixed Astra/high opt-in;
that route is not a tier default or fallback and cannot be mixed with a tier or preset.
Never call `codex exec` directly or pass a model/effort override. Invoke the fixed
installed-plugin wrapper through the shell tool's
`sandbox_permissions: require_escalated` boundary with a narrow justification; do
not first attempt it inside the parent sandbox, where nested Codex app-server
initialization is blocked. This elevation launches only the fixed wrapper. The child
remains forced to `--sandbox read-only` and must pass runtime inspection. The wrapper
resolves the live model and effort once, forces a read-only sandbox, and creates a
distinct persisted consultation thread:
```sh
/bin/sh <absolute-installed-plugin-root>/scripts/run-advisor.sh --tier standard <<'ADVISOR_PACKET'
DECISION
<the complete five-section packet continues here>
ADVISOR_PACKET
# or use: --tier specialist
```
### Deferred shell-tool result handoff
The shell tool can yield a process handle before the wrapper exits. A nonempty `session_id`
from `exec_command` is nonterminal for every Advisor role and tier;
drain that exact session while accumulating tool output, then parse the preserved
final output. The caller-side pattern is:
```javascript
let process = await tools.exec_command({cmd: transportCommand});
let combinedOutput = process.output ?? "";
while (process.session_id) {
process = await tools.write_stdin({
session_id: process.session_id,
chars: "",
yield_time_ms: 5000,
max_output_tokens: 20000,
});
combinedOutput += process.output ?? "";
}
if (process.session_id || process.exit_code == null) {
throw new Error("Advisor transport did not reach a terminal result");
}
if (process.exit_code !== 0) {
throw new Error("Advisor transport failed");
}
const candidates = combinedOutput.split(/\r?\n/).flatMap((line) => {
try { return [JSON.parse(line)]; } catch { return []; }
}).filter((value) => value && value.schema_version === 3);
if (candidates.length !== 1) {
throw new Error("Advisor transport did not return exactly one schema-v3 envelope");
}
const verifiedEnvelope = candidates[0];
text(JSON.stringify(verifiedEnvelope));
```
Never treat the initial yielded result as terminal: it is nonterminal progress, not
the verified envelope, and must not be parsed. An outer `functions.wait` result or heartbeat is also nonterminal progress;
it must not produce an `ADVISOR RESULT`,
unavailable classification, or receipt. Drain and validate the exact process inside
the owning `functions.exec`, preserving every tool-output chunk in `combinedOutput`,
then extract exactly one parseable `schema_version: 3` JSON envelope. The shell tool
may merge stderr progress into `output`, so zero or multiple schema-v3 candidates
is a fail-closed handoff error. Only after validation, emit
`text(JSON.stringify(verifiedEnvelope))` so the enclosing call delivers the verified
envelope; nested shell-tool output is not itself a result. A missing terminal exit
status or a still-present `session_id` is a handoff failure, not a model failure and
not evidence that the consultation returned no response.
Resolve `<absolute-installed-plugin-root>` from the loaded `SKILL.md` path, two
directories above its containing directory. Require regular, nonsymlinked scripts
beneath that root. Never elevate a repository-relative or workspace-resolved
`plugins/advisor` script.
Use only the shown single-quoted heredoc after confirming its delimiter is absent from
the packet. Never use `< packet.txt`, an unquoted heredoc, `eval`, or shell-interpolated
packet text at the elevated boundary, and never stage the packet in a
workspace-writable file.
The wrapper uses existing Codex authentication in place; it does not read, copy,
print, or relink authentication material. Send the five-section
DECISION/CONTEXT/OPTIONS/BOUNDARIES/REQUEST packet from the skill, with only
root-gathered relevant evidence and source references. Require exactly one JSON
object conforming to the installed `advisor-response.schema.json`, with no prose or
code fences. This schema is the sole supported wrapper model-output format. Direct/native role invocation is unsupported and is not schema-validated. It requires
six nonblank string fields (`recommendation`, `why`, `strongest_objection`,
`change_my_mind`, `risks`, `follow_up_areas`) and a nonempty array of nonblank string
`acceptance_checks` values.
```text
ADVISOR RESPONSE
RECOMMENDATION: <recommendation>
WHY: <why>
STRONGEST OBJECTION: <strongest_objection>
CHANGE MY MIND: <change_my_mind>
ACCEPTANCE CHECKS: <acceptance checks joined by ; >
RISKS: <risks>
FOLLOW-UP AREAS: <follow_up_areas>
```
Immediately after every launched child, before response classification or machine
output, the wrapper runs `inspect-agent-runtime.sh` for that thread, expected
model/effort, and parent thread. This mandatory inspection verifies allowlisted
`codex_exec` or `Codex Desktop` provenance, a thread distinct from the parent,
the exact frozen effort, read-only isolation, and
zero tool use; it is not a metadata fallback. Wrapper-owned semantic validation
accepts only schema-valid, nonblank fields and deterministically renders the exact
canonical eight-line receipt above. A runtime-valid response-validation failure emits
only a redacted failure `class` and `field`, then receives exactly one fresh corrective
retry whose prompt names only that diagnostic. A consultation launches at most two children.
Packet, launcher, event, identity, same-session, runtime, wrong-model,
wrong-effort, non-read-only, normalization, or tool-use failure is terminal and never
retries. A second response-validation failure is unavailable. Rejected content is
never emitted, accepted, merged, or copied into the retry prompt. Every attempt
artifact is confined to one private mode-0700 consultation directory beneath the
private transport root; an unconditional exit trap removes it after every wrapper
exit. If packet evidence is
insufficient, the advisor names the specific missing evidence or research questions
under `FOLLOW-UP AREAS` instead of researching. A valid processed response contains
either a recommendation grounded in the packet or a concrete `FOLLOW-UP AREAS`
entry. Only then does the root verify the response's cited source references and
record `accept`, `modify`, or `reject`. For a research-first response, the concise
research-first plan is the recommendation and `accept` means accepting that plan,
not a technical choice the advisor did not make. The advisor is never authoritative.
After a valid, runtime-inspected completed result, the root may route only its
research or brainstorming follow-up to an appropriate Luna or Terra subagent outside
this consultation, synthesize the result, and optionally begin a fresh consultation
with fresh `ADVISOR CALL` and `ADVISOR RESULT` receipts. An unavailable result cannot be rescued by follow-up work, and the advisor may not spawn or conduct that work.
Resolve the tier with the installed helper. Immediately before invoking the
consultation transport, the root emits the actual resolved metadata:
```text
ADVISOR CALL
tier: Standard | Specialist
model: <resolved model selector>
effort: <resolved effort>
reason: <one task-specific sentence>
question: <bounded decision question>
status: running
```
After runtime evidence and advice processing, it always emits:
```text
ADVISOR RESULT
status: completed | unavailable
tier: Standard | Specialist
model: <verified resolved model selector>
effort: <verified resolved effort>
isolation: read-only
recommendation: <concise recommendation, or unavailable>
decision: accept | modify | reject | blocked
reason: <one sentence>
```
Only mandatory post-response runtime inspection plus a processed response can produce
`completed`. Missing, conflicting, non-read-only, tool-use, or required-advice
evidence produces `status: unavailable`, `recommendation: unavailable`, and
`decision: blocked`; the consult route remains fail-closed. These main-chat receipts
summarize verified evidence but are not runtime proof. The distinct Codex consultation
thread remains the inspectable detailed record. A skip emits only `ADVISOR DECISION`,
with no call/result receipt and no transport invocation.
## Runtime evidence
Persisted `codex exec` runtime metadata is primary. The transport must establish one
fresh thread, the exact frozen model and effort resolved for the selected tier, a read-only sandbox, allowlisted
`codex_exec` or `Codex Desktop` provenance, and a thread distinct from the parent.
The wrapper runs:
```sh
sh <absolute-installed-plugin-root>/scripts/inspect-agent-runtime.sh --expected-role <tier-label> --expected-model <selected-model> --expected-effort <selected-effort> --expected-parent <parent-thread-id> <thread-id>
```
The inspector emits only thread, parent, role, transport, model, effort,
sandbox-policy, and permission-profile fields, while rejecting any tool-use event. It must confirm a
read-only runtime policy; a role TOML requesting read-only is not proof of actual
isolation. Missing, conflicting, unexpected, non-read-only, or tool-use evidence is
unavailable, never approval. No substitute advisor role or replacement consultation
is allowed; any root-routed follow-up remains outside this consultation. Progress is
stderr-only and successful stdout is one verified JSON object containing the allowlisted
runtime evidence and the canonical eight-line receipt rendered from the accepted
schema object. Every launched child is inspected before retry eligibility is decided.
Exactly one corrective retry is permitted only for a runtime-valid
response-validation failure and exposes only its redacted failure class and field;
terminal transport, identity, isolation, provenance, runtime, or tool failures never
retry. Rejected raw content remains private to the consultation directory and is never
emitted or copied into a retry prompt.
## Local advisor audit
Use the read-only local audit to inspect aggregate consultation drift without exposing
session content. It writes progress to stderr before enumeration and parsing, then
emits one redacted JSON report on stdout. The report never includes session names,
paths, identifiers, receipt prose, prompts, responses, or cost estimates.
```sh
sh plugins/advisor/scripts/advisor-audit.sh --window-hours 24
```
`--since RFC3339` and `--until RFC3339` select an explicit half-open time window;
`--sessions-dir DIR` is available for isolated synthetic tests. The audit reads only
allowlisted runtime metadata, fixed receipt enums, tool-event kinds, timestamps, and
usage counters. Missing sandbox, tool, token, or duration evidence is JSON `null`
with an `unavailable` availability value, never an inferred value. It reports receipt
attempts and allowlisted `ADVISOR DECISION` routes (`consult`, `skip`, and
`unavailable`) in an exact top-level `decisions` object; decision availability is a
separate field. Schema v2 identifies exact current `advisor-terra`, `advisor-sol`,
and `advisor-astra` child sessions from full-file `session_meta` before applying the
half-open window to their activity. It reports those child sessions separately from deduplicated parent
`spawn_agent` completion evidence; parent completion counts are JSON `null` with
explicit `unavailable` availability unless a completed role-bearing spawn event
exists. Current parent `function_call` spawn requests and role-free
`SubAgentActivity` lifecycle events (`started`, `interacted`, `completed`, and
`interrupted`) are separate corroborating counts. A request never establishes a
selected role or completion, and activity is counted only by correlation to an exact
current child ID. Standard
and Specialist selections, evidenced dispositions, stale `sol_advisor`/`sol-advisor`
attempts, sandbox counts, advisor tool-call counts, duration aggregates, and token
totals remain aggregate-only. Astra is included in child-session and parent-completion
role aggregates, but never in the exact `{standard, specialist}` selected-role counts.
It never changes sessions or Codex configuration.
## Trigger evaluation
Complex work takes one completion consultation before it is declared complete: the
task already consulted, or it spans multiple phases, files, or sessions. It is an
ordinary bounded consultation with the same records, transport, isolation, and
inspection, asking whether the finished work meets its stated contract and what
evidence would falsify that. It never becomes a final diff review or release
verification, and routine or already-skipped work takes none.
Static verification is non-networked:
```sh
sh plugins/advisor/scripts/verify.sh --static
```
Live evaluation is a separate attended step. `evaluate-triggers.sh --run --result
PATH` requires subscription-only routing and disabled overage. The parent process keeps
the live authenticated Codex home and runs with ignored user configuration/rules,
ephemeral state, and a read-only sandbox. Each feature state gets an isolated temporary
project and child runtime: the project links the consultation skill and repository-local
plugin, while the companion installer places and checks all three exact roles in the child
runtime. The evaluator does not copy or link authentication and does not add a plugin
or marketplace during live evaluation.
The two `multi_agent_v2` schemas remain configured, but ephemeral feature-state coverage is not
freshness evidence. A route marker classifies `consult`, `skip`, or `unavailable`;
claimed advisor metadata, an empty wait, or consultation without a completed
`spawn_agent` event is typed `runtime_evidence_unavailable` and blocks the consult
route. Because the evaluator uses `--ephemeral`, its first trial has no persisted
parent rollout and expects the preflight's `route: unavailable`, with no `ADVISOR
CALL` or child spawn; it writes a typed unavailable artifact instead of continuing the
matrix. Deterministic persisted fixtures in `verify.sh` separately prove the read-only
consult path. Progress goes to stderr; the result is redacted JSON
without prompts, raw events, or thread identifiers. Each run records paired
before/after digests for contract-owned live state (`config.toml`, `agents/`,
`skills/`, and `plugins/`) and marketplace state, including unavailable exits.
Auth, session, log, and cache files are excluded from the read set. Validate with:
```sh
sh plugins/advisor/scripts/evaluate-triggers.sh --verify-result --result PATH
```
Use `--allow-unavailable` only to accept a typed unavailable artifact as evidence of
an unavailable evaluation, never as a passing consultation result.
# Usage and local support data
The consultation envelope records an opaque consultation ID plus per-attempt and aggregate duration/token accounting when the host provides structured counters. This preserves the canonical eight-line response and never infers spend or allowance. Local content-free usage journaling is disabled by default and is controlled with `advisor-config.sh journal enable`, `disable`, `status`, and `clear`; explicit Astra records preserve tier `opt-in`, and all records omit prompts, answers, paths, child/session IDs, authentication material, and raw errors. Journal failure is surfaced as a sanitized warning without retrying a model or failing otherwise accepted advice.
SHA-256: 3d82d8689478200f19ac76aa02375512a8c8bd096959b8c78f73d81b5c548004