← Codex Engineering GuardrailsCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to Codex Engineering Guardrails
Snapshot Sep 30, 2026 · 23:13 UTC · version 1.1.1
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "code-verification",
"description": "Independently review, test, diagnose, and verify software against explicit requirements, repository standards, and material risks using traceable fresh evidence. Use when Codex is asked to review code or a diff, inspect a pull request, diagnose a failing check without changing code, run or assess tests, validate a bug fix or feature, audit code quality, identify coverage gaps, judge release readiness, or report whether an implementation is correct. Default to read-only source inspection and existing checks; create or change tests only when explicitly requested, and do not modify production code unless the user separately authorizes a fix.",
"included_files": [
{
"relative_path": "LICENSE",
"size_in_bytes": 1064
},
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 266
},
{
"relative_path": "references/verification-matrix.md",
"size_in_bytes": 6433
}
],
"skill_md_contents": "---\nname: code-verification\ndescription: Independently review, test, diagnose, and verify software against explicit requirements, repository standards, and material risks using traceable fresh evidence. Use when Codex is asked to review code or a diff, inspect a pull request, diagnose a failing check without changing code, run or assess tests, validate a bug fix or feature, audit code quality, identify coverage gaps, judge release readiness, or report whether an implementation is correct. Default to read-only source inspection and existing checks; create or change tests only when explicitly requested, and do not modify production code unless the user separately authorizes a fix.\n---\n\n# Code Verification\n\nDetermine what the evidence actually establishes. Keep verification independent from implementation assumptions, distinguish requirements from general quality, and do not repair findings under a review-only request.\n\n## Preserve verification independence\n\n- Treat the user's explicit requirements and final decisions as authoritative.\n- Default to read-only source inspection. Running existing checks is allowed, including normal temporary build or cache output, but do not edit source, tests, configuration, lockfiles, or infrastructure without explicit authorization.\n- Do not install dependencies, alter external systems, deploy, commit, push, or fix findings unless those actions are separately in scope.\n- Inspect applicable repository instructions, the current status, and the exact diff or target before drawing conclusions. Preserve all user-owned changes.\n- Evaluate the implementation against requirements and independent invariants, not against its own structure or comments.\n- Report adjacent findings without expanding the audit target.\n\nIf the user asks to add tests, enter a bounded test-writing mode: change only authorized tests, fixtures, and indispensable test infrastructure. A failing test that exposes a production defect is a valid result. Do not alter production code merely to make the new test pass without a separate command.\n\n## Define the verification contract\n\nEstablish:\n\n1. The target: files, diff, commit range, component, API, artifact, or running behavior.\n2. The baseline to compare against.\n3. Explicit requirements and acceptance criteria.\n4. Repository standards and supported environments.\n5. Material failure modes and risk boundaries.\n6. Allowed actions and available resources.\n7. The evidence necessary for a pass decision.\n\nWhen the specification is incomplete, separate confirmed requirements from assumptions. Ask only when an ambiguity would change the verdict or make a check unsafe.\n\n## Build traceability before execution\n\nCreate a compact working matrix:\n\n`requirement or risk -> observable behavior -> best evidence -> result`\n\nKeep two independent review axes:\n\n- **Specification:** Does the implementation do exactly what was requested, including edge and failure behavior?\n- **Engineering quality:** Is it correct, understandable, maintainable, secure, efficient enough, and consistent with repository conventions?\n\nPassing one axis must not conceal failure on the other. Read [references/verification-matrix.md](references/verification-matrix.md) when selecting checks, evaluating test quality, assigning severity, or coordinating parallel verification.\n\n## Review in evidence-producing passes\n\n### 1. Establish the changed surface\n\n- Resolve the exact baseline and final state.\n- Inspect changed callers, consumers, schemas, configuration, generated artifacts, and public interfaces.\n- Identify behavior that may change indirectly, including error paths and state transitions.\n\n### 2. Inspect tests before trusting them\n\n- Map tests to requirements and important risks.\n- Prefer tests through public behavior or stable seams.\n- Verify that expected values come from a specification, invariant, standard, known example, or independent calculation.\n- Look for tautological assertions, excessive mocking, missing negative cases, ignored failures, and tests that only mirror the implementation.\n- For regression tests, seek evidence that the test detects the prior defect when a safe baseline or deterministic reproduction exists.\n\n### 3. Review implementation quality\n\nCheck correctness first, then readability, maintainability, architecture, security, performance, operability, and repository consistency. Focus on observable consequences. Do not inflate personal style preferences into findings.\n\n### 4. Execute proportionate checks\n\nDiscover commands from repository documentation, manifests, CI configuration, and nearby conventions. Prefer this order:\n\n1. Focused reproduction or affected test.\n2. Affected unit, integration, and contract suites.\n3. Type checking, linting, formatting, and build validation.\n4. End-to-end or runtime checks at real boundaries.\n5. Security, performance, migration, concurrency, and recovery checks when those risks are present.\n\nCapture each command, environment, exit status, and material result. A green command proves only what it exercised.\n\nInvestigate failures enough to classify them as an implementation defect, test defect, environment limitation, flaky result, unrelated pre-existing failure, or unresolved. Do not silently retry until green or automatically fix the cause.\n\n## Use adaptive parallelism\n\nUse parallel execution only when it preserves independence and evidence quality.\n\n1. Detect available agents, processes, test workers, CPU, memory, and isolated environments. Keep a sequential fallback.\n2. Start with two to four workers and adapt to measured contention and project guidance.\n3. Parallelize orthogonal read-only review passes and checks that have independent state, such as linting, type analysis, isolated unit suites, or separate platform reviews.\n4. Give each reviewer a distinct question and raw target context. Do not leak another reviewer's conclusion as the expected answer.\n5. Serialize checks sharing a database, port, service, filesystem fixture, rate-limited API, device, account, or live environment unless the project provides proven isolation.\n6. Do not run concurrent commands when their build outputs, caches, snapshots, or generated files can race.\n7. Aggregate results centrally, deduplicate findings, resolve contradictions against primary evidence, and run any required integrated check on the final state.\n\nParallel agreement is not proof. One direct, reproducible result outweighs several unsupported summaries.\n\n## Grade evidence and findings\n\nUse these result states:\n\n- **Pass:** Fresh direct evidence covers every required criterion and no material unresolved finding remains.\n- **Partial:** Relevant evidence passes, but an important criterion, environment, or test layer is missing.\n- **Fail:** Direct evidence shows a material requirement or risk is not satisfied.\n- **Inconclusive:** Conflicting or insufficient evidence prevents a defensible decision.\n\nFor each finding, provide:\n\n- severity and confidence;\n- precise location or affected behavior;\n- violated requirement, invariant, or material risk;\n- reproduction or supporting evidence;\n- consequence and affected users or systems;\n- the smallest direction for remediation, without implementing it.\n\nOrder findings by severity. Use severity from impact, likelihood, reach, detectability, and recoverability rather than code style. Label inference as inference and distinguish confirmed, likely, historical, and unverified information.\n\n## Report the verdict\n\nLead with material findings. If none exist, say so plainly while still naming residual gaps. Then report:\n\n- specification verdict and engineering-quality verdict;\n- requirement-to-evidence summary;\n- commands executed with results;\n- coverage gaps, environment limitations, and flaky or pre-existing failures;\n- actions deliberately not performed because they were outside scope.\n\nDo not issue a pass based on code reading alone when executable verification was available, an old test run, coverage percentage without behavioral mapping, or another worker's unsupported conclusion.\n"
}SHA-256: 215260eedda86ad1c35a4b6465e121016fb604378833df71224266e2613b0ad0