← Files Matt Skills CuratedARCHIVED FILE
skills/diagnosing-bugs/SKILL.md
5.21 KB · Oct 3, 2026 · 06:31 UTC
--- name: diagnosing-bugs description: "Diagnose hard bugs, intermittent flakes, and performance regressions using a tight feedback loop. Use when the user reports broken behavior, runtime exceptions, failing tests, flaky CI, or says \"debug this\" / \"diagnose this\" — even if no error message is provided. Do NOT use for routine test-driven feature development." --- # Diagnosing Bugs A disciplined feedback-loop methodology for diagnosing and resolving hard bugs, intermittent flakes, and performance regressions. ## Core Principle > **Never guess or hypothesize before establishing a fast, deterministic, automated feedback loop that reliably reproduces the red failure.** --- ## Core Invariants 1. **Mandatory Red Loop Before Theory**: Construct a fast, automated repro command and see it fail before formulating or testing any hypotheses. 2. **Credential Redaction First**: Redact all keys, authorization headers, tokens, and secrets with `<REDACTED>` before displaying outputs or logs. 3. **Single Variable Instrumentation**: Change only one variable at a time when probing hypothesis boundaries. 4. **Unique Debug Tagging**: Tag all temporary diagnostic logs with a unique searchable prefix (e.g. `[DEBUG-trace]`) for complete cleanup before commit. 5. **Regression Test at Public Seam**: Lock down the fix with a permanent test at a genuine architectural seam before declaring victory. --- ## Architecture & Map of Content (MOC) ``` [ Build Tight Loop ] ──► [ Reproduce & Minimise ] ──► [ 3–5 Ranked Hypotheses ] ──► [ Targeted Probe ] ──► [ Fix & Regression Test ] ──► [ Clean ] ``` | Phase | Core Objective | Key Deliverable | |---|---|---| | **Phase 1: Build Loop** | Construct automated pass/fail signal | Single executable command that goes red on this bug | | **Phase 2: Minimise** | Shrink failure to essential variables | Minimal load-bearing repro payload | | **Phase 3: Hypothesise** | Formulate 3–5 falsifiable predictions | Ranked hypothesis table with testable predictions | | **Phase 4: Instrument** | Test predictions with minimal probes | Tagged `[DEBUG-...]` logs or debugger inspection | | **Phase 5: Fix & Test** | Lock down behavior at public seam | Permanent automated regression test | | **Phase 6: Cleanup** | Remove diagnostic scaffolding | Verified clean git status & passing test suite | --- ## Step-by-Step Procedure (TWI) ### Step 1: Construct a Tight Feedback Loop (Phase 1) - **Action**: Build an automated runner (failing test, CLI fixture, curl script, or trace replay) that drives the bug path. - **Key Point**: The loop must be fast (< 5s), deterministic, and assert the user's exact symptom. - **Inline Checklist**: - [ ] Automated command exists and has been executed - [ ] Command goes red specifically on this bug symptom - [ ] Execution completes in seconds without manual intervention - **Why**: Staring at code without a feedback loop leads to guessing and confirmation bias. ### Step 2: Reproduce and Minimise (Phase 2) - **Action**: Run the loop to confirm reproduction, then systematically remove non-essential config, data, and steps. - **Key Point**: Every remaining line in the repro must be load-bearing (removing it turns the loop green). - **Why**: Minimal repros shrink the hypothesis space and convert cleanly into permanent regression tests. ### Step 3: Formulate Falsifiable Hypotheses (Phase 3) - **Action**: Generate 3–5 ranked hypotheses stating the explicit prediction each makes. - **Key Point**: Use format: *"If X is the cause, then changing Y will make the symptom disappear."* - **Why**: Single-hypothesis debugging anchors on first impressions and wastes turns. ### Step 4: Instrument and Isolate (Phase 4) - **Action**: Insert tagged probes (`[DEBUG-xxx]`) or inspect values at key boundaries. - **Key Point**: Change only one variable at a time; measure baselines before tuning performance bugs. - **Why**: Changing multiple variables simultaneously confounds cause and effect. ### Step 5: Fix, Verify, and Clean (Phases 5 & 6) - **Action**: Write the regression test, apply the minimal fix, verify the full suite, and purge all debug instrumentation. - **Inline Checklist**: - [ ] Regression test fails without fix and passes with fix - [ ] Original un-minimised repro confirmed green - [ ] All `[DEBUG-...]` tags purged from codebase - [ ] Root cause documented clearly in commit message --- ## Anti-Rationalization Guardrails | Tempting Rationalization | Binding Rule | Engineering Rationale | |---|---|---| | *"I think I see the bug in the code, let me fix it now."* | **No edits without an automated red feedback loop.** | Fixing code based on inspection often addresses symptoms while missing root causes. | | *"The bug is non-deterministic so we cannot automate a test."* | **Increase reproduction rate (stress, loops, pinned clocks).** | A 50% flake rate is debuggable; raise reproduction frequency until testable. | | *"I'll add logs everywhere and inspect all outputs."* | **Targeted, tagged probes only.** | Untargeted logging floods context and creates cleanup debt. | | *"The fix works in manual testing, skip the regression test."* | **Mandatory automated regression test at public seam.** | Without a regression test, the bug will silently return in future refactors. |
SHA-256: e7a947205389b94d28e958a4578af2fdcdc7e588a3af40886391e3d892e122cc