← Files Adaptive Codex OrchestratorARCHIVED FILE
docs/SECURITY.md
9.24 KB · Sep 30, 2026 · 23:14 UTC
# Security ## Security posture Adaptive Codex Orchestrator is a local, offline control and policy layer. Its security objective is to preserve normal Codex permission boundaries while adding deterministic mode state and bounded delegation guidance. It does not make delegated code safe by itself. The parent and user must review generated changes before relying on them. This is an independent community project and is not affiliated with or endorsed by OpenAI. ## Threat model The design considers: - Prompt text that resembles a control command accidentally or maliciously changing persistent state. - Commands embedded in quoted examples, fenced code, long pasted text, or conflicting instructions. - Malformed, stale, concurrently written, or future-version state. - Path disclosure through project preferences, logs, errors, or backup files. - Shell injection through prompt text, paths, Git discovery, or hook commands. - A hook receiving malformed or unsupported JSON. - A worker expanding scope, writing overlapping files, delegating recursively, or claiming validation without evidence. - A host omitting or misreporting parent or worker model identity. - Untrusted plugin code receiving the same local access as the Codex process. The model does not attempt to protect a machine from a malicious or already compromised Codex host, Python interpreter, operating system account, Git binary, or plugin package. It also cannot guarantee the correctness or safety of model-generated code. ## Trust boundaries 1. **User and host:** the user selects the parent model, reasoning setting, sandbox, and approval policy. The plugin does not change them. 2. **Host and hook:** trusted hooks receive event JSON and may read/write only plugin-owned state. Review `hooks/hooks.json` and `hooks/runtime.py` before granting trust. 3. **Control and execution planes:** deterministic code decides only control facts and context. The parent model owns architectural and product judgment. 4. **Parent and workers:** every worker gets a bounded contract. The parent reviews evidence and validation before integration. 5. **Plugin state and project:** the state key is a SHA-256 digest of a normalized root. The raw project path and source are not persisted. See [Architecture](ARCHITECTURE.md) for the detailed lifecycle. ## Hook risks and controls - Disabled ordinary requests emit no orchestration policy. - Invalid JSON, unknown events, unsupported schema, and unavailable state fail closed for plugin functionality while preserving normal Codex behavior. - Hooks do not preprocess or rewrite tool/subagent calls, auto-approve tools, alter permissions, elevate the sandbox, or request another model turn. - A compatibility notice is rate-limited to once per session. - Active hooks deliver a compact policy. A bounded numeric `active_policy_revision` session marker suppresses repeated delivery of the same revision on ordinary active turns; it contains no prompt or task text. - One-shot and session cleanup is scoped to the current session; project and global preferences and other sessions are preserved. - Hook trust is an explicit user/host decision. If hooks are not trusted, the one-shot skill is the supported fallback. ## Prompt command parsing risks and controls The parser is deterministic and uses no LLM. It normalizes Unicode with NFKC, case-folds, normalizes whitespace, masks fenced code, and attempts to exclude quoted examples before classifying intent. Status and explicit negation take priority over activation. Conflicting scopes, profiles, or enable/disable instructions preserve state and return an ambiguity fact. Prompt text exists in memory only for the current event. It must never be stored, logged, evaluated, or interpolated into a command. A long prompt needs both a recognized mode alias and an action phrase at a command-like boundary. ## State corruption and concurrency State is schema-versioned and deterministically serialized under `PLUGIN_DATA`. Writes use a same-directory temporary file followed by `os.replace` and a best-effort cross-platform lock. Readable malformed or unsupported state is preserved before clean recovery when storage permits. Failure to read or write state must not block ordinary Codex work. Recovery backups preserve raw prior bytes. Although this runtime never writes prompts or source into state, another process with write access could have put arbitrary content in a corrupt file. Treat corrupt backups as potentially sensitive local evidence; they are never interpreted or transmitted. Locks and user-only permissions are platform-specific best efforts, not a cross-platform security boundary. Do not share one writable `PLUGIN_DATA` directory between mutually untrusted operating-system users. ## Path and shell safety Project-root discovery invokes Git with an argument array, `shell=False`, captured output, and a short timeout. Missing Git, a timeout, and a non-Git directory fall back to the normalized current directory. Prompt text is never part of this command. Only a SHA-256 digest of the normalized root is retained. Hook commands must use the current host schema and platform-appropriate Python launcher. They must not assume Bash on Windows or concatenate user-controlled text into a shell command. ## Permission boundaries The plugin requires no API key, OAuth grant, external account, network access, or additional sandbox permission. It contains no path for changing `~/.codex/config.toml`, the model selector, reasoning level, tool approvals, or the sandbox. A mode toggle writes only to the plugin's `PLUGIN_DATA` location. ## Delegation safety - Delegation is optional and limited to bounded, reversible, independently verifiable work when it provides a material benefit. - Profile ceilings are two, four, and six workers for conservative, balanced, and fast, but task-aware safety caps reduce actual use to zero, one, or two where applicable. Trivial or clear single-file work and unsupported- completion checks use zero. A local reproducible bug defaults to zero and may use one read-only Explorer only for materially useful independent evidence. Four small or obvious module-test pairs also default to zero. At most two disjoint read-only Explorers are used only when substantial independent evidence makes the expected saving clearly exceed spawn and integration cost. Shared-state, authentication, authorization, permission, and tenant work may use at most one read-only Explorer and keeps the parent as the only writer. - Every profile permits at most one concurrent writer. There is no disjoint- file or fast-profile exception. - Nested delegation is forbidden without exception. - Detailed routing, worker-contract, and model references are loaded once only after the compact gate selects a real delegated task. - Each delegated subtask gets one Spark spawn attempt. Failure, a limit, or unsupported explicit model selection returns it to the parent without retry or host-default substitution. - Requested model settings stay separate from host-confirmed facts; a start- time model report is not completion or billing attestation. - Every worker returns exactly `conclusion`, `evidence`, `files_and_lines`, `tests_or_checks`, `risks`, and `recommended_parent_action`, with no other top-level fields. Missing concrete evidence or exact validation is incomplete. - Architecture, authentication, authorization, cryptography, destructive/data- loss-sensitive decisions, major public API/dependency changes, integration, and final validation judgment remain with the parent. - The parent may spot-check cited evidence and investigate gaps or conflicts, but must not repeat the same broad delegated exploration end to end. ## Unsupported and residual-risk scenarios - The plugin cannot verify the host's Ultra reasoning setting. - A missing model identifier prevents model-specific enforcement or reporting. - A host without reliable completion/session-end events may delay cleanup; a later event must remove stale one-shot state where supported. - Untrusted or unsupported hooks cannot provide persistent natural-language control. - A compromised dependency-free Python runtime, Git executable, host, or plugin directory is outside this control plane's protection. - Model output and delegated edits can still contain security defects and need human review appropriate to their impact. ## Reporting a vulnerability Do not include exploit details, credentials, prompts, proprietary source code, or private paths in a public issue. Until the repository owner publishes a dedicated security contact, report privately through the repository host's private security-advisory feature or the support channel designated in the release metadata. If neither exists, withhold sensitive details and ask the maintainer for a private channel first. Include the affected version, platform and Codex surface, minimal reproduction, security impact, and whether the issue requires trusted hooks. The maintainer should acknowledge, triage, coordinate remediation and disclosure, and credit the reporter when requested and appropriate. No response-time guarantee is made before a maintainer publishes one.
SHA-256: f3104659926c47bc0be8a07a91c856678a284d8856456eda360a850ac402b679