← Agent SkillsCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to Agent Skills
Snapshot Sep 30, 2026 · 23:17 UTC · version 0.6.10
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "doubt-driven-development",
"description": "Use when a plan or implementation needs adversarial scrutiny because assumptions are uncertain, stakes are high, or correctness matters more than speed.",
"included_files": [
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 252
}
],
"skill_md_contents": "---\nname: doubt-driven-development\ndescription: Use when a plan or implementation needs adversarial scrutiny because assumptions are uncertain, stakes are high,\n or correctness matters more than speed.\n---\n\n# Doubt-Driven Development\n\n## Overview\n\nA confident answer is not a correct one. Long sessions accumulate context that quietly turns assumptions into \"facts\" without anyone noticing. Doubt-driven development is the discipline of materializing a fresh-context reviewer — biased to **disprove**, not approve — before any non-trivial output stands.\n\nThis is not `/review`. `/review` is a verdict on a finished artifact. This is an in-flight posture: non-trivial decisions get cross-examined while course-correction is still cheap.\n\n## When to Use\n\nA decision is **non-trivial** when at least one of these is true:\n\n- It introduces or modifies branching logic\n- It crosses a module or service boundary\n- It asserts a property the type system or compiler cannot verify (thread safety, idempotence, ordering, invariants)\n- Its correctness depends on context the future reader cannot see\n- Its blast radius is irreversible (production deploy, data migration, public API change)\n\nApply the skill when:\n\n- About to make an architectural decision under uncertainty\n- About to commit non-trivial code\n- About to claim a non-obvious fact (\"this is safe\", \"this scales\", \"this matches the spec\")\n- Working in code you don't fully understand\n\n**When NOT to use:**\n\n- Mechanical operations (renaming, formatting, file moves)\n- Following a clear, unambiguous user instruction\n- Reading or summarizing existing code\n- One-line changes with obvious correctness\n- Pure tooling operations (running tests, listing files)\n- The user has explicitly asked for speed over verification\n\nIf you doubt every keystroke, you ship nothing. The skill applies only to non-trivial decisions as defined above.\n\n## Loading Constraints\n\nThis skill is designed for the **main-session orchestrator**, where Step 3 (DOUBT, detailed below) can spawn a fresh-context reviewer.\n\n- **Do NOT add this skill to a persona's `skills:` frontmatter.** A persona that follows Step 3 would spawn another persona — the orchestration anti-pattern explicitly forbidden by `../../references/orchestration-patterns.md` (\"personas do not invoke other personas\").\n- **If you find yourself applying this skill from inside a subagent context** (where Claude Code prevents nested subagent spawn): the preferred path is to surface to the user that doubt-driven cannot run nested and let the main session handle it. As a last resort only, a degraded self-questioning fallback exists — rewrite ARTIFACT + CONTRACT as a fresh self-prompt with a hard mental separator from your prior reasoning, and walk Steps 1–5. This is **not fresh-context review** (you carry your own context with you), so flag the result as degraded and prefer escalation whenever the user is reachable.\n\n## The Process\n\nCopy this checklist when applying the skill:\n\n```\nDoubt cycle:\n- [ ] Step 1: CLAIM — wrote the claim + why-it-matters\n- [ ] Step 2: EXTRACT — isolated artifact + contract, stripped reasoning\n- [ ] Step 3: DOUBT — invoked fresh-context reviewer with adversarial prompt\n- [ ] Step 4: RECONCILE — classified every finding against the artifact text\n- [ ] Step 5: STOP — met stop condition (trivial findings, 3 cycles, or user override)\n```\n\n### Step 1: CLAIM — Surface what stands\n\nName the decision in two or three lines:\n\n```\nCLAIM: \"The new caching layer is thread-safe under the\n read-heavy workload described in the spec.\"\nWHY THIS MATTERS: a race here corrupts user data and is\n hard to detect in QA.\n```\n\nIf you can't write the claim that compactly, you have a vibe, not a decision. Surface it before scrutinizing it.\n\n### Step 2: EXTRACT — Smallest reviewable unit\n\nA fresh-context reviewer needs the **artifact** and the **contract**, not the journey.\n\n- Code: the diff or the function — not the whole file\n- Decision: the proposal in 3–5 sentences plus the constraints it has to satisfy\n- Assertion: the claim plus the evidence that supposedly supports it (kept distinct from the Step 1 CLAIM block, which is the orchestrator's hypothesis under scrutiny)\n\nStrip your reasoning. If you hand over conclusions, you'll get back validation of your conclusions. The unit must be small enough that a reviewer can hold it in mind in one read — if it's a 500-line PR, decompose first.\n\n### Step 3: DOUBT — Invoke the fresh-context reviewer\n\nThe reviewer's prompt **must be adversarial**. Framing decides the answer.\n\n```\nAdversarial review. Find what is wrong with this artifact.\nAssume the author is overconfident. Look for:\n- Unstated assumptions\n- Edge cases not handled\n- Hidden coupling or shared state\n- Ways the contract could be violated\n- Existing conventions this might break\n- Failure modes under unexpected input\n\nDo NOT validate. Do NOT summarize. Find issues, or state\nexplicitly that you cannot find any after thorough examination.\n\nARTIFACT: <paste artifact>\nCONTRACT: <paste contract>\n```\n\n**Pass ARTIFACT + CONTRACT only. Do NOT pass the CLAIM.** Handing the reviewer your conclusion biases it toward agreement. The reviewer must independently determine whether the artifact satisfies the contract.\n\nIn Claude Code, the role-based reviewers in `agents/` start with isolated context by design and are usable here — see `agents/` for the roster and per-domain match.\n\n**The adversarial prompt above takes precedence over the persona's default response shape.** Personas like `code-reviewer` are written to produce balanced verdicts with both strengths and weaknesses; doubt-driven needs issues-only output. Paste the adversarial prompt verbatim into the invocation so it overrides the persona's default. If a persona's response shape can't be overridden cleanly, fall back to a generic subagent with the adversarial prompt.\n\n#### Cross-model escalation\n\nA single-model reviewer shares blind spots with the original author — a colder, different-architecture model catches them. Doubt-driven is already opt-in for non-trivial decisions, so within that scope offering cross-model is part of the skill's value, not optional friction.\n\n**Interactive sessions: always offer. Never silently skip.**\n\n**Step 1: Ask the user**\n\nAfter the single-model review in Step 3 above, but before RECONCILE, pause and ask:\n\n> *\"Single-model review complete. Want a cross-model second opinion? Options: Gemini CLI, Codex CLI, manual external review (you paste it elsewhere), or skip.\"*\n\nThis question is mandatory in every interactive doubt cycle — even on artifacts that feel low-stakes. The user — not the agent — decides whether the cost is worth it. The agent's job is to surface the choice.\n\n**Step 2: If the user picks a CLI — verify, then invoke**\n\n1. Check the tool is in PATH (`which gemini`, `which codex`).\n2. Test it works (`gemini --version` or equivalent) before passing the full prompt — a stale or broken binary may pass `which` but fail on real input.\n3. Confirm the exact invocation with the user, including required flags, auth, and env vars (e.g., API keys). Implementations vary; never assume.\n4. Pass ARTIFACT + CONTRACT + the adversarial prompt **only**. No session context, no CLAIM.\n5. Mind shell escaping. If the artifact contains quotes, `$(...)`, or backticks, prefer stdin (`echo … | gemini`) or a heredoc over inline `-p \"…\"`. When in doubt, ask the user to confirm the invocation before running it.\n6. Take the output into Step 4 (RECONCILE).\n\n**Never interpolate the artifact into a shell-quoted argument.** Code, markdown, and review prompts routinely contain backticks, `$(...)`, and quote characters that will either truncate the prompt or execute embedded shell. Write the full prompt to a file and pipe it through stdin.\n\nExample shapes (verify flags against your installed tool — syntax differs across implementations and versions):\n\n```bash\n# Write the adversarial prompt + ARTIFACT + CONTRACT to a temp file first.\n# Then pipe via stdin so shell metacharacters in the artifact stay inert.\n\n# Codex (read-only sandbox keeps the CLI from writing to your workspace):\ncodex exec --sandbox read-only -C <repo-path> - < /tmp/doubt-prompt.md\n\n# Gemini ('--approval-mode plan' is read-only; '-p \"\"' triggers non-interactive\n# mode and the prompt is read from stdin):\ngemini --approval-mode plan -p \"\" < /tmp/doubt-prompt.md\n```\n\nA read-only sandbox is the load-bearing detail: a doubt artifact may itself contain instructions (intentional or accidental prompt injection) that the cross-model CLI would otherwise execute against your workspace.\n\n**Step 3: If the CLI is unavailable or fails**\n\nSurface the failure explicitly. Offer: run it manually, try a different tool, or skip. Do not silently fall back to single-model — the user should know cross-model didn't happen.\n\n**Step 4: If the user skips**\n\nAcknowledge the skip in the output (*\"Proceeding with single-model findings only\"*) and continue to RECONCILE. Skipping is fine; silent skipping is not.\n\n**Non-interactive contexts** (CI, `/loop`, autonomous-loop, scheduled runs):\n\n- Cross-model is **skipped**, and the skip must be **announced** in the output: *\"Cross-model skipped: non-interactive context.\"*\n- **Never invoke an external CLI without explicit user authorization** — this is a load-bearing safety property.\n\nCross-model adds cost, latency, and tool fragility. The agent surfaces the choice every cycle; the user decides whether this artifact warrants it.\n\n### Step 4: RECONCILE — Fold findings back\n\nThe reviewer's output is data, not verdict. **You are still the orchestrator.** Re-read the artifact text against each finding before classifying — rubber-stamping the reviewer is the same failure mode as ignoring it.\n\nFor each finding, classify in this **precedence order** (first matching class wins):\n\n1. **Contract misread** — reviewer flagged something specifically because the CONTRACT you provided was unclear or incomplete. Fix the contract first, re-classify on the next cycle.\n2. **Valid + actionable** — real issue requiring a change to the artifact. Change it, re-loop.\n3. **Valid trade-off** — issue is real but cost of fixing exceeds cost of accepting. Document the trade-off explicitly so the user sees it.\n4. **Noise** — reviewer flagged something that's actually correct under context the reviewer didn't have. Note it, move on, and ask: would adding that context to the contract have prevented the false flag?\n\nA fresh reviewer can be wrong because it lacks context. Don't defer just because it's \"fresh.\"\n\n### Step 5: STOP — Bounded loop, not recursion\n\nStop when:\n\n- Next iteration returns only trivial or already-considered findings, **or**\n- 3 cycles completed (escalate to user, don't grind a fourth alone), **or**\n- User explicitly says \"ship it\"\n\nIf after 3 cycles the reviewer still surfaces substantive issues, the artifact may not be ready. Surface this to the user — three unresolved cycles is information about the artifact, not a reason to keep looping.\n\nIf 3 cycles is \"obviously insufficient\" because the artifact is large: the artifact is too big — return to Step 2 and decompose. Do not lift the bound.\n\n## Common Rationalizations\n\n| Rationalization | Reality |\n|---|---|\n| \"I'm confident, skip the doubt step\" | Confidence correlates poorly with correctness on novel problems. Moments of certainty are exactly when blind spots hide. |\n| \"Spawning a reviewer is expensive\" | Debugging a wrong commit in production is more expensive. The check is bounded; the bug isn't. |\n| \"The reviewer will just nitpick\" | Only if unscoped. Constrain the prompt to \"issues that would make this fail under the contract.\" |\n| \"I'll do doubt at the end with `/review`\" | `/review` is a final gate. Doubt-driven catches wrong directions early when course-correction is cheap. By PR time it's too late. |\n| \"If I doubt every step I'll never ship\" | The skill applies to non-trivial decisions, not every keystroke. Re-read \"When NOT to Use.\" |\n| \"Two opinions are always better than one\" | Not when the second has less context and produces noise. Reconcile, don't defer. |\n| \"The reviewer disagreed so I was wrong\" | The reviewer lacks your context — disagreement is information, not verdict. Re-read the artifact, classify, then decide. |\n| \"Cross-model is always better\" | Cross-model catches blind spots a single model shares with itself, but it adds cost and tool fragility. Offer it every interactive doubt cycle — the user decides whether the artifact warrants it. The agent's job is to surface the choice, not to gate it. |\n| \"User said yes once, so I can keep invoking the CLI\" | Each invocation is its own authorization. The artifact, the prompt, and the flags change between calls — re-confirm the exact command with the user before every run. |\n\n## Red Flags\n\n- Spawning a fresh-context reviewer for a one-line rename or formatting change\n- Treating reviewer output as authoritative without re-reading the artifact text\n- Looping >3 cycles without escalating to the user\n- Prompting the reviewer with \"is this good?\" instead of \"find issues\"\n- Skipping doubt under time pressure on a high-stakes decision\n- Re-spawning fresh-context on an unchanged artifact (you'll get the same findings; you're stalling)\n- **Doubt theater (checkable signal)**: across 2 or more cycles where the reviewer surfaced substantive findings, zero findings were classified as actionable. You are validating, not doubting. Stop and escalate.\n- Doubting only after committing — that's `/review`, not doubt-driven development\n- Hardcoding an external CLI invocation without confirming with the user that the tool exists, is configured, and accepts that exact syntax\n- **Silently skipping cross-model in an interactive doubt cycle.** Even when not recommending it, the offer must be visible. Skipping is fine; silent skipping is not.\n- Falling back silently when an external CLI errors or is missing — surface the failure and let the user redirect\n- Stripping the contract from the reviewer's input\n- Passing the CLAIM to the reviewer (biases toward agreement)\n\n## Interaction with Other Skills\n\n- **`code-review-and-quality` / `/review`**: complementary. `/review` is post-hoc PR verdict; doubt-driven is in-flight per-decision. Use both.\n- **`source-driven-development`**: SDD verifies *facts about frameworks* against official docs. Doubt-driven verifies *your reasoning about the artifact*. SDD checks the API exists; doubt-driven checks you used it correctly under the contract.\n- **`test-driven-development`**: TDD's RED step is doubt made concrete — a failing test is a disproof attempt. When TDD applies, that failing test *is* the doubt step for behavioral claims.\n- **`debugging-and-error-recovery`**: when the reviewer surfaces a real failure mode, drop into the debugging skill to localize and fix.\n- **Repo orchestration rules** (`../../references/orchestration-patterns.md`): this skill orchestrates from the main session. A persona calling another persona is anti-pattern B — see Loading Constraints above.\n\n## Verification\n\nAfter applying doubt-driven development:\n\n- [ ] Every non-trivial decision (per the definition above) was named explicitly as a CLAIM before standing\n- [ ] At least one fresh-context review per non-trivial artifact (a failing test produced by TDD's RED step satisfies this for behavioral claims, per Interaction with Other Skills)\n- [ ] The reviewer received ARTIFACT + CONTRACT — NOT the CLAIM, NOT your reasoning\n- [ ] The reviewer's prompt was adversarial (\"find issues\"), not validating (\"is it good\")\n- [ ] Findings were classified against the artifact text (not rubber-stamped) using the precedence: contract misread / actionable / trade-off / noise\n- [ ] A stop condition was met (trivial findings, 3 cycles, or user override)\n- [ ] In interactive mode, cross-model was **explicitly offered** to the user (regardless of artifact stakes) and the response was acknowledged in the output\n- [ ] In non-interactive mode, cross-model was skipped and the skip was announced\n- [ ] Any external CLI invocation was preceded by a PATH check, a working-binary test, syntax confirmation with the user, and explicit authorization to run\n"
}SHA-256: b521edc7fcc6cdae841692f8cc1432678696109a21ab805b1050fa6cd5ab1d22