← Matt Skills CuratedCONTENT HISTORY

Update to Matt Skills Curated

Snapshot Sep 30, 2026 · 23:14 UTC · version 1.1.0

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "description": "Diagnose hard bugs, intermittent flakes, and performance regressions using a tight feedback loop. Use when the user reports broken behavior, runtime exceptions, failing tests, flaky CI, or says \"debug this\" / \"diagnose this\" — even if no error message is provided. Do NOT use for routine test-driven feature development.",
  "included_files": [
    {
      "relative_path": "agents/openai.yaml",
      "size_in_bytes": 103
    },
    {
      "relative_path": "scripts/hitl-loop.template.sh",
      "size_in_bytes": 1316
    }
  ],
  "name": "diagnosing-bugs",
  "skill_md_contents": "---\nname: diagnosing-bugs\ndescription: \"Diagnose hard bugs, intermittent flakes, and performance regressions using a tight feedback loop. Use when the user reports broken behavior, runtime exceptions, failing tests, flaky CI, or says \\\"debug this\\\" / \\\"diagnose this\\\" — even if no error message is provided. Do NOT use for routine test-driven feature development.\"\n---\n\n# Diagnosing Bugs\n\nA disciplined feedback-loop methodology for diagnosing and resolving hard bugs, intermittent flakes, and performance regressions.\n\n## Core Principle\n\n> **Never guess or hypothesize before establishing a fast, deterministic, automated feedback loop that reliably reproduces the red failure.**\n\n---\n\n## Core Invariants\n\n1. **Mandatory Red Loop Before Theory**: Construct a fast, automated repro command and see it fail before formulating or testing any hypotheses.\n2. **Credential Redaction First**: Redact all keys, authorization headers, tokens, and secrets with `<REDACTED>` before displaying outputs or logs.\n3. **Single Variable Instrumentation**: Change only one variable at a time when probing hypothesis boundaries.\n4. **Unique Debug Tagging**: Tag all temporary diagnostic logs with a unique searchable prefix (e.g. `[DEBUG-trace]`) for complete cleanup before commit.\n5. **Regression Test at Public Seam**: Lock down the fix with a permanent test at a genuine architectural seam before declaring victory.\n\n---\n\n## Architecture & Map of Content (MOC)\n\n```\n[ Build Tight Loop ] ──► [ Reproduce & Minimise ] ──► [ 3–5 Ranked Hypotheses ] ──► [ Targeted Probe ] ──► [ Fix & Regression Test ] ──► [ Clean ]\n```\n\n| Phase | Core Objective | Key Deliverable |\n|---|---|---|\n| **Phase 1: Build Loop** | Construct automated pass/fail signal | Single executable command that goes red on this bug |\n| **Phase 2: Minimise** | Shrink failure to essential variables | Minimal load-bearing repro payload |\n| **Phase 3: Hypothesise** | Formulate 3–5 falsifiable predictions | Ranked hypothesis table with testable predictions |\n| **Phase 4: Instrument** | Test predictions with minimal probes | Tagged `[DEBUG-...]` logs or debugger inspection |\n| **Phase 5: Fix & Test** | Lock down behavior at public seam | Permanent automated regression test |\n| **Phase 6: Cleanup** | Remove diagnostic scaffolding | Verified clean git status & passing test suite |\n\n---\n\n## Step-by-Step Procedure (TWI)\n\n### Step 1: Construct a Tight Feedback Loop (Phase 1)\n- **Action**: Build an automated runner (failing test, CLI fixture, curl script, or trace replay) that drives the bug path.\n- **Key Point**: The loop must be fast (< 5s), deterministic, and assert the user's exact symptom.\n- **Inline Checklist**:\n  - [ ] Automated command exists and has been executed\n  - [ ] Command goes red specifically on this bug symptom\n  - [ ] Execution completes in seconds without manual intervention\n- **Why**: Staring at code without a feedback loop leads to guessing and confirmation bias.\n\n### Step 2: Reproduce and Minimise (Phase 2)\n- **Action**: Run the loop to confirm reproduction, then systematically remove non-essential config, data, and steps.\n- **Key Point**: Every remaining line in the repro must be load-bearing (removing it turns the loop green).\n- **Why**: Minimal repros shrink the hypothesis space and convert cleanly into permanent regression tests.\n\n### Step 3: Formulate Falsifiable Hypotheses (Phase 3)\n- **Action**: Generate 3–5 ranked hypotheses stating the explicit prediction each makes.\n- **Key Point**: Use format: *\"If X is the cause, then changing Y will make the symptom disappear.\"*\n- **Why**: Single-hypothesis debugging anchors on first impressions and wastes turns.\n\n### Step 4: Instrument and Isolate (Phase 4)\n- **Action**: Insert tagged probes (`[DEBUG-xxx]`) or inspect values at key boundaries.\n- **Key Point**: Change only one variable at a time; measure baselines before tuning performance bugs.\n- **Why**: Changing multiple variables simultaneously confounds cause and effect.\n\n### Step 5: Fix, Verify, and Clean (Phases 5 & 6)\n- **Action**: Write the regression test, apply the minimal fix, verify the full suite, and purge all debug instrumentation.\n- **Inline Checklist**:\n  - [ ] Regression test fails without fix and passes with fix\n  - [ ] Original un-minimised repro confirmed green\n  - [ ] All `[DEBUG-...]` tags purged from codebase\n  - [ ] Root cause documented clearly in commit message\n\n---\n\n## Anti-Rationalization Guardrails\n\n| Tempting Rationalization | Binding Rule | Engineering Rationale |\n|---|---|---|\n| *\"I think I see the bug in the code, let me fix it now.\"* | **No edits without an automated red feedback loop.** | Fixing code based on inspection often addresses symptoms while missing root causes. |\n| *\"The bug is non-deterministic so we cannot automate a test.\"* | **Increase reproduction rate (stress, loops, pinned clocks).** | A 50% flake rate is debuggable; raise reproduction frequency until testable. |\n| *\"I'll add logs everywhere and inspect all outputs.\"* | **Targeted, tagged probes only.** | Untargeted logging floods context and creates cleanup debt. |\n| *\"The fix works in manual testing, skip the regression test.\"* | **Mandatory automated regression test at public seam.** | Without a regression test, the bug will silently return in future refactors. |\n"
}

SHA-256 of public snapshot: c3ea80b6fcbeb4b9cbb6b527a503865bd2e3b39153d5377d94e3b5a3423bdbe8