Matt Skills Curated
Mamdouh Aboammar v1.1.0
Publisher description
From the marketplace listing
A curated ChatGPT and Codex distribution of MIT-licensed engineering, AI/ML, and productivity workflows, matured for automatic routing. It includes 42 focused Skills for requirements, specs, tickets, multi-agent spec implementation, TDD, debugging, review, session retrospectives, architecture, research, AI model development, data remediation, statistical ML best practices, J-space cognitive reasoning, autonomous goal contracts, skill conductor authoring, repository safety and setup, TypeScript package boundaries, workflow design, handoffs, teaching, course exercise scaffolding, and long-form writing. Host-specific duplicates and repository-development noise are excluded.
Language: English · Automatically detected from descriptions.
Files & skills
File archives
Skill instructions
ai-data-remediation6.38 KB
---
name: ai-data-remediation
description: "Self-healing data pipeline layer using semantic anomaly clustering, AST-validated lambda transformations, and zero-loss mathematical reconciliation. Use when data quality checks fail, anomalous records break ETL/ELT pipelines, you need automated data cleansing with sandboxed Python transformations, or you want to group data errors into semantic clusters — even if they don't explicitly say \"data remediation\". Do NOT use for standard database schema migrations, routine CRUD queries, or basic pipeline scheduling."
---
# AI Data Remediation
Operate a self-healing remediation layer for mission-critical data pipelines: intercept corrupt or anomalous data, semantically cluster pattern families, generate deterministic fix logic via local or sandboxed language models, and guarantee zero data loss.
## Core Principle
> **AI generates verifiable transformation logic — never touch or mutate production data directly.**
---
## Core Invariants
1. **AI Logic Over Raw Data Mutation**: The model generates pure, deterministic transformation functions that can be tested, reviewed, and versioned. Raw model outputs are never piped directly into tables.
2. **Mandatory AST Static Validation**: Every generated lambda or transformation must pass Abstract Syntax Tree (AST) validation against restricted namespaces before evaluation.
3. **Strict Zero-Loss Accounting**: Total input records must exactly equal successful records plus quarantined records ($\text{Source} = \text{Success} + \text{Quarantine}$). Any non-zero delta immediately halts processing.
4. **Air-Gapped / Privacy-Preserved Execution**: Sensitive or PII data must never egress to external cloud APIs without prior tokenization or local execution.
5. **Human Review Quarantine**: Clusters with low transformation confidence ($< 0.75$) or failed AST checks route to an isolated quarantine table with full lineage.
---
## Architecture & Map of Content (MOC)
```
[ Anomalous Records ] ──► [ Semantic Clustering ] ──► [ Sandboxed Logic Gen ] ──► [ AST Safety Gate ] ──► [ Vectorized Apply ] ──► [ Zero-Loss Audit ]
```
| Component | Responsibility | Key Mechanism |
|---|---|---|
| **Anomaly Buffer** | Staging buffer for records flagged `NEEDS_AI` | Asynchronous staging table / queue |
| **Semantic Compression** | Grouping high-volume anomalies into pattern families | Vector embeddings + similarity clustering |
| **Logic Synthesis** | Compiling deterministic lambda functions | Prompt-constrained JSON output format |
| **AST Security Gate** | Static code analysis & sandboxed execution | Python `ast.parse` + restricted builtins |
| **Reconciliation Audit** | Mathematical verification of record counts | Strict invariant: $\Delta = \text{Source} - (\text{Success} + \text{Quarantine}) = 0$ |
---
## Step-by-Step Procedure (TWI)
### Step 1: Intercept & Isolate Anomalous Records
- **Action**: Ingest failed rows from the deterministic validation layer into an isolated staging buffer.
- **Key Point**: Operate strictly downstream of primary schema validation without blocking the main ingest stream.
- **Why**: Decoupling remediation keeps upstream ingestion healthy and prevents system-wide backpressure.
### Step 2: Semantic Anomaly Compression
- **Action**: Compute vector representations of error strings and cluster them into distinct pattern families.
- **Key Point**: Compress thousands of broken rows into 5–15 representative clusters using similarity metrics.
- **Why**: Synthesizing 10 cluster-level transformation functions instead of 50,000 row-by-row LLM calls reduces execution time and compute cost by over 95%.
### Step 3: Sandboxed Logic Generation
- **Action**: Feed representative cluster exemplars to a constrained language model prompt that emits a pure Python lambda function.
- **Key Point**: Restrict output to a single lambda definition with explicit input/output type contracts.
- **Why**: Lambda functions are reproducible, testable against test suites, and easily inspected before deployment.
### Step 4: AST Safety Verification & Vectorized Application
- **Action**: Statically parse the generated lambda AST and execute across the cluster within a restricted environment.
- **Key Point**: Reject any code containing `import`, `exec`, `eval`, `__builtins__`, file I/O, or OS calls.
- **Inline Checklist**:
- [ ] AST parsing verifies zero unauthorized node types (no `Import`, `ImportFrom`, `Call` to unapproved functions)
- [ ] Transformation confidence score $\ge 0.75$
- [ ] Lambda passes unit test assertions on 3 cluster sample inputs
- [ ] Unverified or failing records route immediately to quarantine
- **Why**: Unchecked dynamic code execution presents critical security vulnerabilities and risks catastrophic database corruption.
### Step 5: Zero-Loss Reconciliation & Lineage Audit
- **Action**: Run mathematical balance checks across input, output, and quarantine tables.
- **Key Point**: Verify $\text{Source Records} = \text{Success Records} + \text{Quarantine Records}$.
- **Why**: Silent row loss corrupts financial ledgers, analytical dashboards, and downstream dependencies.
---
## Anti-Rationalization Guardrails
| Tempting Rationalization | Binding Rule | Engineering Rationale |
|---|---|---|
| *"The generated lambda looks safe, skip the AST check."* | **Mandatory AST verification on 100% of generated code.** | A single unvalidated attribute access or system call compromises execution integrity. |
| *"Only 3 rows disappeared, let's ship the batch anyway."* | **Immediate halt if $\Delta \neq 0$.** | Silent record drops compound into major audit and financial discrepancies. |
| *"Let's directly patch data strings instead of generating a function."* | **Generate logic, never mutate raw data in-place.** | Logic can be audited, unit tested, reviewed, and rolled back; raw data patches cannot. |
| *"Send full customer records with PII to public APIs for faster fix."* | **Enforce data privacy boundaries.** | PII egress violates privacy regulations and compliance mandates. |
---
## Verification & Troubleshooting
- **AST Rejection**: If the model attempts forbidden imports, tighten the system prompt to enforce pure mathematical/string transformations.
- **Low Model Confidence**: Route the entire pattern family to `quarantine_review` table for human sign-off.
- **Reconciliation Discrepancy**: Check row exception handlers to ensure failing rows are captured in quarantine rather than swallowed.
Referenced files: 1
ai-engineering6.16 KB
---
name: ai-engineering
description: "Design, train, optimize, deploy, and evaluate production machine learning and LLM systems. Use when building ML models, training classifiers, fine-tuning LLMs with LoRA/PEFT, deploying low-latency inference endpoints, architecting vector RAG pipelines, or evaluating model bias and data leakage — even if they don't explicitly say \"AI engineering\". Do NOT use for standard backend CRUD development, basic SQL queries without ML, or simple UI styling."
---
# AI & Machine Learning Engineering
Design, train, optimize, deploy, and monitor production machine learning models and intelligent agent systems with rigorous validation, latency engineering, and ethical guardrails.
## Core Principle
> **Every ML solution must beat a naive baseline, guarantee train-test isolation, and meet strict production latency and fairness budgets.**
---
## Core Invariants
1. **Strict Train-Test Isolation**: Preprocessing scalers, encoders, and tokenizers must be fit strictly on training folds inside isolated pipelines to eliminate data leakage.
2. **Sub-100ms Inference Budget**: Synchronous user-facing endpoints must achieve $p99 < 80\text{ms}$ through quantization (ONNX/TensorRT/GGUF), engine optimization, or continuous batching.
3. **Mandatory Baseline Before Complexity**: Always benchmark against a naive baseline (majority class/mean) and simple linear model before advancing to deep networks or complex ensembles.
4. **Demographic Parity & Bias Auditing**: Every production candidate must pass four-fifths disparate impact testing across demographic slices ($\text{ratio} \ge 0.80$).
5. **Continuous Drift Observability**: Production models must emit data drift metrics (KS-test / Population Stability Index) and concept drift alerts to trigger automated retraining.
---
## Architecture & Map of Content (MOC)
```
Problem Framing & Data Assessment
│
▼
┌───────────────────────────────────────┐
│ 1. Data Prep & Leakage-Free Pipeline │ (Cross-Validation, Feature Isolation)
└──────────────────┬────────────────────┘
│
▼
┌───────────────────────────────────────┐
│ 2. Model Training & Fine-Tuning │ (Baselines, LoRA / QLoRA, Tree Ensembles)
└──────────────────┬────────────────────┘
│
▼
┌───────────────────────────────────────┐
│ 3. Bias, Fairness & Ethics Audit │ (Disparate Impact, SHAP Attributions)
└──────────────────┬────────────────────┘
│
▼
┌───────────────────────────────────────┐
│ 4. Production Serving & Optimization │ (ONNX/vLLM, Sub-100ms Latency, Canary)
└───────────────────────────────────────┘
```
---
## Step-by-Step Procedure (TWI)
### Step 1: Requirements & Data Hygiene Verification
- **Action**: Assess target metrics, evaluate class balance, and establish reproducible validation splits.
- **Key Point**: Check for temporal dependencies. If data has a time dimension, use chronological splitting; otherwise, use Stratified $k$-Fold cross-validation.
- **Why**: Random splits on time-series data cause severe lookahead leakage and false confidence.
### Step 2: Model Training & Architecture Selection
- **Action**: Fit naive baselines, progress to gradient-boosted trees or fine-tune neural architectures (LoRA/PEFT for LLMs).
- **Key Point**: For LLMs, apply LoRA to all attention and MLP projection layers with $\alpha = 2 \times r$ and 4-bit NF4 quantization.
- **Why**: Selective rank adaptation achieves foundation model performance at 25% of the VRAM cost.
- **Inline Checklist**:
- [ ] Simple baseline established (Linear/Logistic Regression or Mean baseline)
- [ ] Cross-validation variance across folds is within acceptable bounds (< 5%)
- [ ] Overfitting diagnostics checked (Training loss vs Validation loss convergence)
- [ ] Hyperparameters tuned systematically using Bayesian optimization
### Step 3: Bias, Explainability & Robustness Auditing
- **Action**: Evaluate fairness across subpopulations and calculate SHAP feature attributions.
- **Key Point**: Enforce the four-fifths rule ($\text{selection rate ratio} \ge 0.80$) and compute top local feature attributions for high-stakes decisions.
- **Why**: Undetected bias causes regulatory violations and uncalibrated real-world discrimination.
### Step 4: Low-Latency Serving & Canary Rollout
- **Action**: Export models to ONNX or vLLM, wrap in asynchronous API services, and initiate a 10% canary deployment.
- **Key Point**: Track $p99$ latency, Population Stability Index (PSI), and error rates across the canary group.
- **Why**: Canary testing isolates performance regressions before full traffic exposure.
---
## Anti-Rationalization Guardrails
| Tempting Rationalization | Binding Rule | Engineering Rationale |
|---|---|---|
| *"Accuracy is 98%, so we don't need confusion matrices."* | **Mandatory PR-AUC, F1, and confusion matrix analysis.** | High accuracy on imbalanced data often hides a degenerate majority-class classifier. |
| *"Preprocessing before splitting is harmless."* | **Zero data leakage: pipeline-contained transformers only.** | Dataset-wide preprocessing leaks test distribution parameters into training folds. |
| *"Model meets accuracy goals, so skip bias checks."* | **Fairness audit is a mandatory release gate.** | Regulatory and ethical standards mandate demographic parity verification. |
| *"400ms latency is fast enough for cloud servers."* | **Sub-100ms p99 budget for interactive endpoints.** | High latency compounds across microservices and degrades user experience. |
Referenced files: 1
codebase-design4.7 KB
--- name: codebase-design description: "Shared vocabulary and patterns for designing deep modules with narrow interfaces and clean seams. Use when designing module interfaces, finding deepening opportunities, deciding where seams go, or making code more testable and AI-navigable — even if the user says \"improve this module\". Do NOT use for whole-codebase architectural surveys." --- # Codebase Design Design deep, high-leverage modules that encapsulate complex domain behavior behind minimal interfaces placed at clear architectural seams. --- ## Core Invariants 1. **Depth Over Surface Area**: Maximize internal implementation leverage while minimizing interface surface area (few methods, simple primitive parameters). 2. **Interface as Test Surface**: External callers and unit tests cross the exact same seam; never pierce the interface to test internal private plumbing. 3. **The Deletion Test**: If deleting a module causes complexity to vanish, it was an unnecessary pass-through; if complexity scatters across $N$ callers, it was earning its keep. 4. **Real vs. Hypothetical Seams**: One adapter indicates a hypothetical seam; introduce an interface seam only when at least two concrete adapters vary across it. 5. **Exact Design Vocabulary**: Strictly use canonical terminology (**module**, **interface**, **depth**, **seam**, **adapter**, **leverage**, **locality**); avoid vague synonyms (service, component, boundary). --- ## Architecture & Map of Content (MOC) ``` ┌───────────────────────────────────────┐ │ Narrow Public Interface │ ◄── Small surface (few methods, simple inputs) ├───────────────────────────────────────┤ │ │ │ Deep Implementation │ ◄── High leverage, hidden state, rich logic │ │ └───────────────────────────────────────┘ ``` | Component | Responsibility | Reference | |---|---|---| | **Module Deepening** | Refactor shallow pass-throughs into deep modules | `skills/codebase-design/DEEPENING.md` | | **Design It Twice** | Explore multi-model interface variations | `skills/codebase-design/DESIGN-IT-TWICE.md` | | **Testability Rules** | Accept dependencies, return pure values | Public seam tests | --- ## Step-by-Step Procedure (TWI) ### Step 1: Evaluate Current Interface Depth & Seams - **Action**: Inspect the module's public methods, parameter signatures, and call sites. - **Key Point**: Check the ratio of interface cognitive overhead to internal capabilities. - **Why**: Shallow modules force callers to understand internal mechanics, destroying locality. ### Step 2: Apply the Deletion Test & Simplify Signatures - **Action**: Consolidate fine-grained procedural methods into unified, intent-revealing operations. - **Key Point**: Hide internal state transformations and dependency instantiations behind the seam. - **Inline Checklist**: - [ ] Methods reduced to minimal essential operations - [ ] Dependencies passed in rather than created internally - [ ] Functions return values rather than mutating global side effects ### Step 3: Align Seam with Unit Test Harness - **Action**: Structure test suites to exercise the module strictly through its public interface. - **Key Point**: Eliminate internal mocking and testing of private helper functions. - **Why**: Testing through the interface ensures tests survive internal refactors without breakage. ### Step 4: Explore Alternatives (Design It Twice) - **Action**: When designing complex or foundational modules, draft 2–3 radically different interface designs before coding. - **Key Point**: Compare candidates on depth, locality, and caller ergonomics. - **Why**: The first interface that comes to mind is rarely the deepest or most maintainable. --- ## Anti-Rationalization Guardrails | Tempting Rationalization | Binding Rule | Engineering Rationale | |---|---|---| | *"Expose private helper functions so we can write unit tests for them."* | **Forbidden. Test exclusively through the public interface.** | Testing private helpers couples tests to implementation details and prevents refactoring. | | *"Create an interface and adapter for a single implementation."* | **Wait for 2 adapters before extracting a generic seam.** | Speculative generalization creates shallow, unnecessary abstraction layers. | | *"Break this 100-line cohesive function into 5 single-use files."* | **Maintain locality inside deep modules.** | Excessive fragmentation increases cognitive load and scatters related logic. |
Referenced files: 3
code-review5.01 KB
---
name: code-review
description: "Review changed code against repository coding standards and original specification intent. Use when reviewing a branch, diff, PR, pull request, merge-base changes, or verifying code against documented standards — even if the user just says \"review this\". Do NOT use for authoring new code or diagnosing failing runtime bugs."
---
# Code Review
Two-axis review of the diff between `HEAD` and a fixed point:
1. **Standards**: Does the code conform to documented repository standards and clean architecture smells?
2. **Spec**: Does the code faithfully and completely implement the originating issue/spec without scope creep?
Both axes execute as isolated parallel sub-agents to prevent context contamination, with findings reported side by side.
---
## Core Invariants
1. **Strict Two-Axis Separation**: Keep Standards and Spec findings isolated; never merge or cross-rank them into a blended score.
2. **Pinned Merge-Base Diff**: Always resolve refs with `git rev-parse` and review `git diff <fixed-point>...HEAD`.
3. **Evidence-Based Citations**: Every finding must quote the exact file, line range, and standard/spec clause violated.
4. **Tooling Non-Duplication**: Skip formatting, syntax, or lint errors that automated pre-commit tooling already catches.
5. **No Blind Approvals**: If the spec is missing, report "No spec provided - verified against standards only" explicitly.
---
## Architecture & Map of Content (MOC)
```
[ Pin Fixed Point ] ──► [ Identify Spec & Standards ] ──► [ Parallel Review Subagents ] ──► [ Side-by-Side Synthesis ]
│
┌────────────────────────────────┴────────────────────────────────┐
▼ ▼
[ Standards Subagent ] [ Spec Subagent ]
- Documented repo rules - Missing requirements
- Fowler code smells - Unasked scope creep
- Architectural boundaries - Flawed implementations
```
| Component | Responsibility | Evaluation Source |
|---|---|---|
| **Standards Axis** | Architecture smells, naming, cohesion | `CODING_STANDARDS.md`, `CONTRIBUTING.md`, smell baseline |
| **Spec Axis** | Functional completeness, scope boundaries | Issue description, `specs/*.md`, user requirements |
| **Aggregation** | Side-by-side balanced reporting | Verbatim findings categorized by axis |
---
## Step-by-Step Procedure (TWI)
### Step 1: Pin the Fixed Point & Diff
- **Action**: Resolve the base ref and confirm a non-empty diff (`git diff <base>...HEAD`).
- **Key Point**: Fail fast if the ref is invalid or the working tree is empty.
- **Why**: Reviewing against an incorrect base compares irrelevant changes.
### Step 2: Extract Standards and Spec Sources
- **Action**: Locate repository guidelines (`CODING_STANDARDS.md`) and originating requirements/issues.
- **Key Point**: Equip the standards agent with Fowler smell baselines (Mysterious Name, Duplicated Code, Feature Envy, Primitive Obsession, Speculative Generality).
- **Inline Checklist**:
- [ ] Diff base confirmed
- [ ] Standards docs identified
- [ ] Spec/issue requirements extracted
### Step 3: Dispatch Parallel Sub-Agents
- **Action**: Spawn Standards subagent and Spec subagent concurrently with dedicated prompts.
- **Key Point**: Restrict each subagent to its designated domain (< 400 words per report).
- **Why**: Combining standards and spec evaluation into a single pass leads to halo bias where clean code masks missing features.
### Step 4: Aggregate and Synthesize Report
- **Action**: Present findings under `## Standards` and `## Spec` headers with actionable remediation recommendations.
- **Key Point**: State the single most severe issue within each axis clearly.
- **Why**: Clear prioritization helps authors address critical design issues first.
---
## Anti-Rationalization Guardrails
| Tempting Rationalization | Binding Rule | Engineering Rationale |
|---|---|---|
| *"The code looks beautifully formatted, so it must be correct."* | **Standards pass $\neq$ Spec pass.** | Elegant code can completely fail to implement required business logic. |
| *"It does what the ticket asked, so ignore messy architecture."* | **Spec pass $\neq$ Standards pass.** | Quick hacks that bypass standards generate severe technical debt. |
| *"Merge both reviews into one combined score."* | **Strict two-axis separation.** | Blending scores obscures which dimension requires remediation. |
| *"Point out minor indentation issues in the review."* | **Skip issues handled by automated linters.** | Manual review should focus on semantics, architecture, and intent. |
Referenced files: 1
diagnosing-bugs5.21 KB
--- name: diagnosing-bugs description: "Diagnose hard bugs, intermittent flakes, and performance regressions using a tight feedback loop. Use when the user reports broken behavior, runtime exceptions, failing tests, flaky CI, or says \"debug this\" / \"diagnose this\" — even if no error message is provided. Do NOT use for routine test-driven feature development." --- # Diagnosing Bugs A disciplined feedback-loop methodology for diagnosing and resolving hard bugs, intermittent flakes, and performance regressions. ## Core Principle > **Never guess or hypothesize before establishing a fast, deterministic, automated feedback loop that reliably reproduces the red failure.** --- ## Core Invariants 1. **Mandatory Red Loop Before Theory**: Construct a fast, automated repro command and see it fail before formulating or testing any hypotheses. 2. **Credential Redaction First**: Redact all keys, authorization headers, tokens, and secrets with `<REDACTED>` before displaying outputs or logs. 3. **Single Variable Instrumentation**: Change only one variable at a time when probing hypothesis boundaries. 4. **Unique Debug Tagging**: Tag all temporary diagnostic logs with a unique searchable prefix (e.g. `[DEBUG-trace]`) for complete cleanup before commit. 5. **Regression Test at Public Seam**: Lock down the fix with a permanent test at a genuine architectural seam before declaring victory. --- ## Architecture & Map of Content (MOC) ``` [ Build Tight Loop ] ──► [ Reproduce & Minimise ] ──► [ 3–5 Ranked Hypotheses ] ──► [ Targeted Probe ] ──► [ Fix & Regression Test ] ──► [ Clean ] ``` | Phase | Core Objective | Key Deliverable | |---|---|---| | **Phase 1: Build Loop** | Construct automated pass/fail signal | Single executable command that goes red on this bug | | **Phase 2: Minimise** | Shrink failure to essential variables | Minimal load-bearing repro payload | | **Phase 3: Hypothesise** | Formulate 3–5 falsifiable predictions | Ranked hypothesis table with testable predictions | | **Phase 4: Instrument** | Test predictions with minimal probes | Tagged `[DEBUG-...]` logs or debugger inspection | | **Phase 5: Fix & Test** | Lock down behavior at public seam | Permanent automated regression test | | **Phase 6: Cleanup** | Remove diagnostic scaffolding | Verified clean git status & passing test suite | --- ## Step-by-Step Procedure (TWI) ### Step 1: Construct a Tight Feedback Loop (Phase 1) - **Action**: Build an automated runner (failing test, CLI fixture, curl script, or trace replay) that drives the bug path. - **Key Point**: The loop must be fast (< 5s), deterministic, and assert the user's exact symptom. - **Inline Checklist**: - [ ] Automated command exists and has been executed - [ ] Command goes red specifically on this bug symptom - [ ] Execution completes in seconds without manual intervention - **Why**: Staring at code without a feedback loop leads to guessing and confirmation bias. ### Step 2: Reproduce and Minimise (Phase 2) - **Action**: Run the loop to confirm reproduction, then systematically remove non-essential config, data, and steps. - **Key Point**: Every remaining line in the repro must be load-bearing (removing it turns the loop green). - **Why**: Minimal repros shrink the hypothesis space and convert cleanly into permanent regression tests. ### Step 3: Formulate Falsifiable Hypotheses (Phase 3) - **Action**: Generate 3–5 ranked hypotheses stating the explicit prediction each makes. - **Key Point**: Use format: *"If X is the cause, then changing Y will make the symptom disappear."* - **Why**: Single-hypothesis debugging anchors on first impressions and wastes turns. ### Step 4: Instrument and Isolate (Phase 4) - **Action**: Insert tagged probes (`[DEBUG-xxx]`) or inspect values at key boundaries. - **Key Point**: Change only one variable at a time; measure baselines before tuning performance bugs. - **Why**: Changing multiple variables simultaneously confounds cause and effect. ### Step 5: Fix, Verify, and Clean (Phases 5 & 6) - **Action**: Write the regression test, apply the minimal fix, verify the full suite, and purge all debug instrumentation. - **Inline Checklist**: - [ ] Regression test fails without fix and passes with fix - [ ] Original un-minimised repro confirmed green - [ ] All `[DEBUG-...]` tags purged from codebase - [ ] Root cause documented clearly in commit message --- ## Anti-Rationalization Guardrails | Tempting Rationalization | Binding Rule | Engineering Rationale | |---|---|---| | *"I think I see the bug in the code, let me fix it now."* | **No edits without an automated red feedback loop.** | Fixing code based on inspection often addresses symptoms while missing root causes. | | *"The bug is non-deterministic so we cannot automate a test."* | **Increase reproduction rate (stress, loops, pinned clocks).** | A 50% flake rate is debuggable; raise reproduction frequency until testable. | | *"I'll add logs everywhere and inspect all outputs."* | **Targeted, tagged probes only.** | Untargeted logging floods context and creates cleanup debt. | | *"The fix works in manual testing, skip the regression test."* | **Mandatory automated regression test at public seam.** | Without a regression test, the bug will silently return in future refactors. |
Referenced files: 2
domain-modeling4.98 KB
---
name: domain-modeling
description: "Build and sharpen a project's domain model, ubiquitous language, and architectural decision records. Use when establishing codebase terminology, challenging fuzzy concepts, writing ADRs, or updating CONTEXT.md — even if the user says \"define our terms\". Do NOT use for general code refactoring without domain shifts."
---
# Domain Modeling
Actively establish, sharpen, and enforce ubiquitous domain language (`CONTEXT.md`) and architectural decision records (`docs/adr/*.md`) to prevent semantic drift across agents and engineering teams.
---
## Core Invariants
1. **Active Semantic Enforcement**: Proactively challenge overloaded, ambiguous, or colloquial terms and align them with canonical definitions in real time.
2. **Immediate Inline Glossary Updates**: Capture domain terms into `CONTEXT.md` the instant they crystallize; never batch glossary edits to the end of a session.
3. **Strict Implementation-Free Glossary**: `CONTEXT.md` must contain zero implementation details, frameworks, or database choices—it is a pure domain dictionary.
4. **Selective ADR Threshold**: Only author an ADR when a decision meets all 3 criteria: (1) Hard to reverse, (2) Surprising without context, (3) The result of a real trade-off.
5. **Codebase-Glossary Alignment**: Cross-reference terminology with live codebase entities and flag discrepancies immediately.
---
## Architecture & Map of Content (MOC)
```
[ Domain Discussions / User Prompts ] ──► [ Semantic Challenge & Disambiguation ] ──► [ Inline CONTEXT.md Update ]
│
┌────────────────────────────────┴────────────────────────────────┐
▼ ▼
[ Domain Dictionary ] [ Architectural Records ]
- Single-context: `CONTEXT.md` - `docs/adr/NNNN-<slug>.md`
- Multi-context: `CONTEXT-MAP.md` - Context, Decision, Consequences
```
| Artifact | Responsibility | Format Reference |
|---|---|---|
| **Domain Glossary** | Canonical terms, entity boundaries, invariants | `skills/domain-modeling/CONTEXT-FORMAT.md` |
| **Architectural Record** | Irreversible architectural choices & trade-offs | `skills/domain-modeling/ADR-FORMAT.md` |
| **Context Map** | Bounded contexts across modular repositories | `CONTEXT-MAP.md` |
---
## Step-by-Step Procedure (TWI)
### Step 1: Detect Context Architecture & Glossary Baseline
- **Action**: Check if a root `CONTEXT-MAP.md` exists (multi-context) or single `CONTEXT.md` / `docs/adr/`.
- **Key Point**: Create glossary files lazily on the first resolved term.
- **Why**: Multi-context systems require partitioning domain terms by bounded context to prevent collision.
### Step 2: Challenge Fuzzy & Overloaded Language
- **Action**: Intercept vague nouns (e.g. "account", "item", "process") and propose distinct canonical domain entities.
- **Key Point**: Stress-test boundaries with concrete edge-case scenarios (e.g., "What happens during partial cancellation?").
- **Why**: Ambiguous nouns lead to bloated database entities and tangled business logic.
### Step 3: Verify Alignment Against Existing Codebase
- **Action**: Search the codebase for entity names and check whether existing schemas agree with the user's description.
- **Key Point**: Highlight discrepancies immediately: "The code cancels entire Orders, but you described partial cancellation. Which is correct?"
- **Inline Checklist**:
- [ ] Term verified against live database/code entities
- [ ] Definition added to `CONTEXT.md` using standard format
- [ ] Implementation details omitted from glossary
### Step 4: Author Architectural Decision Records (ADRs)
- **Action**: For decisions meeting the 3-point threshold, create `docs/adr/NNNN-<slug>.md`.
- **Key Point**: Document Context, Decision, Status, and Consequences.
- **Why**: Transparent ADRs prevent repetitive debates and document technical debt trade-offs.
---
## Anti-Rationalization Guardrails
| Tempting Rationalization | Binding Rule | Engineering Rationale |
|---|---|---|
| *"I'll add database table schemas into CONTEXT.md."* | **Forbidden. CONTEXT.md contains pure domain terms only.** | Coupling the domain glossary to DB schemas makes it obsolete upon migration. |
| *"Let's write an ADR for every small choice (e.g. library helper)."* | **Enforce the 3-point ADR threshold.** | Low-value ADRs clutter documentation and obscure truly critical architectural choices. |
| *"The user used 'User' and 'Customer' interchangeably; I'll ignore it."* | **Challenge and disambiguate overloaded terms immediately.** | Conflating distinct domain concepts creates severe authorization and modeling bugs. |
Referenced files: 3
engineering-workflow-guide6.38 KB
---
name: engineering-workflow-guide
description: "Route engineering, AI/ML, cognitive, or productivity tasks to the narrowest effective specialist skill. Use when a task spans planning, implementation, debugging, review, architecture, research, setup, AI modeling, data remediation, or autonomous goals, when the user asks what to do next, or when unsure which workflow fits — even if they don't explicitly name a skill. Do NOT use when the specific specialist skill is already obvious and unambiguous."
---
# Engineering Workflow Guide
Centrally analyze tasks, classify developer intent, and route engineering workflows to the narrowest, highest-leverage specialist skill across the curated skill catalog.
---
## Core Invariants
1. **Narrowest Effective Skill**: Select the most specific skill that fully covers the task; avoid invoking broad, ceremonial meta-skills when a focused tool exists.
2. **One Primary Skill per Phase**: Designate exactly one primary specialist skill for the current turn; chain secondary skills only across distinct lifecycle gates (e.g. Planning $\rightarrow$ Implementation $\rightarrow$ Review).
3. **No Unpackaged Ghost Skills**: Every routed skill must exist verbatim in `references/catalog.md` and have an established `SKILL.md`.
4. **Fast-Path User Overrides**: If the user explicitly requests a valid skill by name, honor it immediately without redundant routing deliberation.
5. **Setup Pre-Requisite Check**: If repository issue tracking, triage labels, or domain docs are unconfigured, run `setup-engineering-workflows` prior to planning or triage.
---
## Architecture & Map of Content (MOC)
```
[ Developer Intent / Request ]
│
▼
┌────────────────────────────────────────────────────────┐
│ 1. Intent Classification & Catalog Lookup │
│ (Match job against `references/catalog.md`) │
└──────────────────────────┬─────────────────────────────┘
│
┌──────────────────────┼──────────────────────┐
▼ ▼ ▼
[ Discovery & Plan ] [ Code & Architecture ] [ AI, ML & Automation ]
- `grill-with-docs` - `implement` - `ai-engineering`
- `to-spec` - `code-review` - `ai-data-remediation`
- `to-tickets` - `diagnosing-bugs` - `ml-best-practices`
- `wayfinder` - `domain-modeling` - `j-space`
- `prototype` - `setup-ts-deep-modules` - `goal`
│
▼
┌────────────────────────────────────────────────────────┐
│ 2. Lifecycle Route Execution │
│ (State route briefly, invoke primary skill) │
└────────────────────────────────────────────────────────┘
```
| Router Reference | Responsibility | File Location |
|---|---|---|
| **Skill Catalog** | Exhaustive index of all 42 packaged skills | [references/catalog.md](references/catalog.md) |
| **ChatGPT Routing Contract** | Runtime consumption & anti-hallucination protocol | [references/chatgpt-routing-contract.md](references/chatgpt-routing-contract.md) |
| **Routing Matrix** | 100+ intent patterns, triggers, and exclusions | [references/routing-matrix.md](references/routing-matrix.md) |
| **Phase Boundaries** | Decision tree for `/clear`, `/compact`, `/handoff`, and subagents | [PHASE-BOUNDARIES.md](PHASE-BOUNDARIES.md) |
| Lifecycle Scenario | Canonical Route Pipeline |
|---|---|
| **Fuzzy Greenfield Feature** | `grill-with-docs` $\rightarrow$ `to-spec` $\rightarrow$ `to-tickets` $\rightarrow$ `implement` |
| **Hard Bug or Regression** | `diagnosing-bugs` $\rightarrow$ `tdd` |
| **Pull Request Review** | `code-review` (Two-axis: Standards & Spec) |
| **Deep Module Restructuring** | `improve-codebase-architecture` $\rightarrow$ `setup-ts-deep-modules` |
| **Large Multi-Session Map** | `wayfinder` $\rightarrow$ `to-spec` $\rightarrow$ `to-tickets` |
| **Deep Multi-Step Reasoning** | `j-space` (Cognitive registers & seam audits) |
| **Autonomous Unattended Goal** | `goal` (7-part prompt contract) |
| **Production ML / LLM System** | `ai-engineering` $\rightarrow$ `ml-best-practices` |
| **Self-Healing Data Pipeline** | `ai-data-remediation` |
| **Long-Form Writing & Essays** | `writing-fragments` $\rightarrow$ `writing-shape` / `writing-beats` |
---
## Step-by-Step Procedure (TWI)
### Step 1: Ingest Request and Consult Catalog
- **Action**: Read `references/catalog.md` and classify the user's underlying intent, not just raw keywords.
- **Key Point**: Check whether repository setup is complete (`docs/agents/issue-tracker.md`).
- **Why**: Accurate intent classification ensures the agent doesn't jump into code when requirements are undefined.
### Step 2: Formulate the Narrowest Route
- **Action**: Select the single primary specialist skill and state the multi-phase roadmap in 1–2 lines.
- **Inline Checklist**:
- [ ] Target skill verified in `references/catalog.md`
- [ ] Zero overlapping meta-skills invoked
- [ ] Distinct lifecycle phases separated cleanly
### Step 3: Execute Primary Phase
- **Action**: Invoke the selected skill instructions and proceed directly to execution.
- **Why**: Direct transition eliminates conversational overhead.
---
## Anti-Rationalization Guardrails
| Tempting Rationalization | Binding Rule | Engineering Rationale |
|---|---|---|
| *"Invoke 4 skills at once for a single simple change."* | **Enforce 1 primary skill per lifecycle phase.** | Overlapping skill prompts create conflicting instructions and waste context budget. |
| *"Invent a new custom workflow name not in catalog."* | **Route exclusively to packaged skills in `catalog.md`.** | Routing to non-existent skills causes execution failures. |
| *"Overrule the user's explicit skill invocation."* | **Honor user-requested skill unless hard conflict exists.** | Respect developer intent and operational autonomy. |
Referenced files: 5
git-safety-guardrails4.78 KB
---
name: git-safety-guardrails
description: "Safeguard repositories against destructive, irreversible, or history-rewriting Git operations. Use when running force pushes, hard resets, branch deletions, cleans, destructive restores, or history re-writes — even if the user says \"clean up git history\". Do NOT use for routine safe git status, fetch, diff, or branch queries."
---
# Git Safety Guardrails
Install and maintain deterministic pre-execution guardrails that intercept and block dangerous, destructive, or history-rewriting Git commands before autonomous agents can execute them.
---
## Core Invariants
1. **Deterministic Interception**: Automatically block destructive Git commands (`git push --force`, `git reset --hard`, `git clean -f/-fd`, `git branch -D`, `git checkout .`, `git restore .`) before execution.
2. **Explicit Authority Gate**: Intercepted commands return an explicit non-zero exit code notifying the agent that it lacks authority to execute destructive operations.
3. **Scope Clarification**: Always ask the user whether to apply guardrails locally to the project (`.claude/settings.json`) or globally (`~/.claude/settings.json`).
4. **Settings Merge Safety**: Seamlessly merge guardrail hooks into existing `PreToolUse` configurations without overwriting other tools or settings.
5. **Mandatory Interception Test**: Verify that the safety hook triggers correctly on simulated forbidden commands before concluding setup.
---
## Architecture & Map of Content (MOC)
```
[ Agent Tool Call (Bash/Git) ] ──► [ PreToolUse Hook: `block-dangerous-git.sh` ]
│
┌─────────────────────────┴─────────────────────────┐
▼ ▼
[ Safe Git Operation ] [ Destructive Command ]
- `git status`, `git diff` - `git push --force`, `reset --hard`
- ALLOWED to execute - BLOCKED (Exit Code 2)
```
| Component | Responsibility | Location |
|---|---|---|
| **Classifier Hook Script** | Inspect and block forbidden git patterns | `scripts/block-dangerous-git.sh` |
| **Python Command Parser** | Parse complex shell command chains | `scripts/classify_git_command.py` |
| **Settings Integration** | Hook registration in agent environment | `.claude/settings.json` or `~/.claude/settings.json` |
---
## Step-by-Step Procedure (TWI)
### Step 1: Confirm Guardrail Scope
- **Action**: Ask the user whether to configure guardrails for this project only or globally.
- **Key Point**: Default to project-level `.claude/settings.json` unless the user specifies global.
- **Why**: Scoped installation avoids unexpected side-effects across external personal repositories.
### Step 2: Copy Hook Script & Make Executable
- **Action**: Copy `scripts/block-dangerous-git.sh` to `.claude/hooks/block-dangerous-git.sh` and run `chmod +x`.
- **Key Point**: Ensure parent directories exist before copying.
- **Inline Checklist**:
- [ ] Target directory created
- [ ] Script copied and executable permissions set (`chmod +x`)
- [ ] Python classifier script colocated if needed
### Step 3: Register Hook in Settings Configuration
- **Action**: Add the PreToolUse hook entry to `.claude/settings.json`, merging into existing arrays if present.
- **Key Point**: Use `"$CLAUDE_PROJECT_DIR"/.claude/hooks/block-dangerous-git.sh` path expansion.
- **Why**: Relative path expansions ensure the hook functions across different working directory contexts.
### Step 4: Verify Guardrail Interception (Test Gate)
- **Action**: Test the hook with a simulated blocked command:
```bash
echo '{"tool_input":{"command":"git push origin main --force"}}' | .claude/hooks/block-dangerous-git.sh
```
- **Key Point**: Verify that the command exits with code 2 and outputs a descriptive blocked message.
- **Why**: Proving the hook intercepts dangerous commands guarantees that unverified agents cannot accidentally wipe Git history.
---
## Anti-Rationalization Guardrails
| Tempting Rationalization | Binding Rule | Engineering Rationale |
|---|---|---|
| *"Allow `git reset --hard` if the working tree has uncommitted bugs."* | **Block all hard resets; require explicit stashes or reverts.** | Hard resets permanently delete uncommitted code and worktree context. |
| *"Allow force pushes on feature branches."* | **Block all force pushes by default.** | Force pushing can overwrite teammate commits and destroy branch history. |
| *"Skip verifying the hook script with simulated input."* | **Mandatory simulated test pass.** | Syntax errors in hook scripts cause them to fail open, leaving the repo unprotected. |
Referenced files: 3
goal4.89 KB
--- name: goal description: "Design and synthesize high-leverage autonomous goal prompts and contracts for unattended execution. Use when crafting a /goal prompt, planning an overnight autonomous coding run, converting rambling specifications into a self-contained agentic mission, or setting up multi-agent autonomous loops — even if they don't explicitly say \"goal prompt\". Do NOT use when the user asks for immediate single-turn execution in the current conversation without an autonomous goal contract." --- # Autonomous Goal & Contract Synthesis Turn user intent, fuzzy requirements, or rambling specifications into an exceptional, self-contained `/goal` prompt ready for unattended, autonomous multi-turn execution. ## Core Principle > **An autonomous goal prompt specifies the destination, quality bar, and verification loops — never micromanaged step-by-step scripts.** --- ## Core Invariants 1. **Observable Done Condition**: Every deliverable must possess a concrete completion state the autonomous session can verify itself (e.g., test suite passes, CLI runs on fixtures, build outputs exist). 2. **Upfront Authority & Creative Freedom**: Grant explicit permission to make design decisions, choose internal workflows, and resolve ambiguities without interrupting the user. 3. **Verified Resource Inventory**: Name only 2–4 tools or paths that were verified against the live environment. Pair with an explicit discovery mandate. 4. **Mandatory Multi-Pass Verification**: Mandate at least 3 distinct iteration passes matched to the target medium. 5. **Terminal Goal Line**: End with a single recency-anchored sentence containing the deliverable and autonomy directive (`"... is your /goal. Work completely autonomously and do not ask me for anything until you are all done."`). --- ## The Seven-Part Goal Prompt Anatomy ``` ┌──────────────────────────────────────────────────────────┐ │ 1. Core Desire & High-Stakes Context │ │ 2. Uncompromising Quality Bar & Principles │ │ 3. Verified Resource Inventory & Discovery Mandate │ │ 4. Explicit Decision Authority & Creative Freedom │ │ 5. Medium-Matched Multi-Pass Verification Loop │ │ 6. Concrete Delivery Destination │ │ 7. Terminal Goal Line & Autonomy Directive │ └──────────────────────────────────────────────────────────┘ ``` --- ## Step-by-Step Procedure (TWI) ### Step 1: Extract Intent & Fill Gaps - **Action**: Extract deliverable, scope, real stakes, mentioned tools, quality standards, and target destination. - **Key Point**: Synthesize sensible defaults for minor gaps. If an ambiguity is critical, resolve it upfront. - **Why**: Multi-round back-and-forth defeats the speed advantage of autonomous delegation. ### Step 2: Verify Resources & Compose Anatomy - **Action**: Verify paths and tools before naming them, then weave the 7 components into natural flowing prose (150–350 words). - **Key Point**: State the outcome and constraints clearly while explicitly granting the agent freedom to navigate internal implementation steps. - **Inline Checklist**: - [ ] Word count is within 150–350 words - [ ] Named resources verified against live repository - [ ] Creative freedom & decision authority explicitly granted - [ ] Medium-specific verification passes defined (3 iterations) - [ ] Destination explicit (link, directory path, file artifact) - [ ] Concludes with terminal goal line and autonomy directive ### Step 3: Mechanical Validation & Delivery - **Action**: Verify the drafted prompt against the seven invariants and output in a clean code block. - **Key Point**: Deliver the prompt ready to copy-paste with zero clutter. - **Why**: Clean delivery prevents formatting and syntax mistakes during execution. --- ## Anti-Rationalization Guardrails | Tempting Rationalization | Binding Rule | Engineering Rationale | |---|---|---| | *"I listed step-by-step instructions so the agent won't get lost."* | **Specify destination, not micromanaged steps.** | Micromanagement prevents the agent from adapting to unforeseen obstacles. | | *"I named several tools from memory without checking if they exist."* | **Verify all resources before naming.** | Unverified resources cause agents to burn turns searching for phantom paths. | | *"The task is subjective, so verification passes aren't needed."* | **Every deliverable requires verification criteria.** | Unverified tasks lead to premature completion and unspotted defects. | | *"I added notes and explanations after the goal line."* | **Terminal goal line must be the final sentence.** | Recency bias ensures the terminal directive anchors the model's primary objective. |
Referenced files: 1
grilling6.35 KB
---
name: grilling
description: "Relentlessly interview the user round-by-round to stress-test thinking and expose unexamined assumptions. Use when the user requests a grilling interview, wants their idea challenged, or uses trigger phrases like \"grill me\" — even if they just say \"poke holes in my plan\". Do NOT use when the user asks for immediate implementation."
---
# Grilling
Relentlessly stress-test thinking, expose hidden assumptions, and map the design space as an expanding decision tree until mutual understanding is reached.
---
## Core Invariants
1. **Frontier-Based Batched Rounds**: Ask all currently unblocked questions together in a single numbered round; never drip questions one by one.
2. **Mandatory Concrete Recommendations**: Every question must include a definitive recommended answer (`➡️ [Recommendation]`) to streamline decision-making.
3. **Autonomous Fact Exploration**: Investigate the codebase, configs, and dependencies with background exploration before asking questions; never query the user for discoverable facts.
4. **Strict Topological Dependency Ordering**: Questions whose answers depend on open decisions must wait for future rounds; never mix dependent questions into the current frontier.
5. **Verified Exhaustion Gate**: The interview concludes only when the frontier is completely empty and the user explicitly validates the final consensus.
---
## Architecture & Map of Content (MOC)
```
[ Problem Space / Initial Proposal ]
│
▼
┌───────────────────────────────────────┐
│ 1. Compute Decision Frontier Tree │
│ (Unblocked questions with defaults)│
└──────────────────┬────────────────────┘
│
▼
┌───────────────────────────────────────┐
│ 2. Issue Formatted Round (Q1..Qn) │
│ (Questions + Strong Recommendations)│
└──────────────────┬────────────────────┘
│
▼
┌───────────────────────────────────────┐
│ 3. Ingest Answers & Advance Frontier │
│ (Prune resolved, unlock downstream)│
└──────────────────┬────────────────────┘
│
▼
┌───────────────────────────────────────┐
│ 4. Verification & Consensus Handshake │
└───────────────────────────────────────┘
```
| Phase | Responsibility | Expected Output |
|---|---|---|
| **Frontier Extraction** | Identify independent decision nodes | Filtered list of unblocked questions |
| **Round Formatting** | Apply standardized Markdown question templates | User-facing numbered interview round |
| **Downstream Unblocking** | Re-evaluate tree based on user decisions | Updated frontier state |
| **Consensus Handshake** | Synthesize agreed architecture & requirements | Clean brief ready for specification |
---
## Step-by-Step Procedure (TWI)
### Step 1: Map the Design Space & Compute Frontier
- **Action**: Analyze the user's intent, identify all decision nodes, and select only those whose prerequisites are fully settled.
- **Key Point**: Filter out any question that requires guessing the answer to an unresolved peer question.
- **Why**: Asking dependent questions simultaneously forces the user into speculative, conditional answers that clutter context.
### Step 2: Format and Dispatch the Question Round
- **Action**: Present the current frontier using the standardized Markdown question format:
```markdown
❓ **Q1** - **<Question Title>**: <Detailed question body and options>
➡️ **Recommendation**: <Specific proposed answer and technical rationale>
---
❓ **Q2** - **<Question Title>**: <Detailed question body and options>
➡️ **Recommendation**: <Specific proposed answer and technical rationale>
```
- **Key Point**: Keep recommendations opinionated, clear, and grounded in industry best practices.
- **Inline Checklist**:
- [ ] Every question numbered and titled
- [ ] Concrete recommendation provided for each question
- [ ] No dependent questions included in the same round
### Step 3: Ingest Answers & Advance the Frontier Tree
- **Action**: Parse user responses, record settled decisions, unlock newly available downstream questions, and generate the next round.
- **Key Point**: If a user response reveals a new unknown in the codebase, dispatch background exploration immediately to resolve it.
- **Why**: Real-time background discovery prevents passing factual burdens onto the user.
### Step 4: Verify Alignment and Completion
- **Action**: Once all branches of the decision tree have been explored and no frontier items remain, present a concise synthesis of all resolved decisions.
- **Key Point**: Require explicit user confirmation before proceeding to specification or execution.
- **Why**: Mutual alignment guarantees that the subsequent implementation phase proceeds without friction or false assumptions.
---
## Anti-Rationalization Guardrails
| Tempting Rationalization | Binding Rule | Engineering Rationale |
|---|---|---|
| *"I'll ask just one question now to see what they say."* | **Ask the entire frontier in one turn.** | Drip-feeding questions creates annoying latency and loses the macro structure of the design. |
| *"Let the user decide without biasing them with my recommendation."* | **Mandatory recommended answer for every question.** | Concrete recommendations speed up review and give users a clear baseline to react against. |
| *"Ask the user where configuration files or types are defined."* | **Discover environment facts autonomously.** | User interviews are exclusively for product and architectural decisions, not repo search. |
| *"The user gave a vague answer, but I can guess what they meant."* | **Clarify ambiguous answers immediately.** | Guessing user intent during grilling undermines the entire purpose of stress-testing. |
Referenced files: 1
grill-me4.74 KB
--- name: grill-me description: "Interview the user relentlessly to sharpen an idea, requirement, or decision before execution. Use when a plan, design, requirement, or decision needs a focused interview to expose ambiguity, missing constraints, trade-offs, or weak assumptions — even if the user just says \"grill me on this\". Do NOT use for repo-stateful domain modeling with doc generation." --- # Grill Me Conduct a focused, relentless Socratic interview to stress-test a plan, architectural decision, requirement, or nascent concept before writing code. --- ## Core Invariants 1. **Relentless Frontier Exploration**: Model the design space as a decision tree and exhaust the frontier before declaring alignment. 2. **One Question Round per Turn**: Batch questions into structured frontier rounds with recommended defaults; never interrogate with one-off dribble. 3. **Autonomous Fact Extraction**: Retrieve codebase, system, and file facts autonomously; never ask the user what the environment can reveal. 4. **Active Assumption Invalidation**: Proactively challenge comfortable assumptions, scale bottlenecks, edge-case failure modes, and hidden dependencies. 5. **No Premature Implementation**: Refuse to write production code or tickets until the interview frontier is completely empty and mutually agreed. --- ## Architecture & Map of Content (MOC) ``` [ Nascent Idea / Unsettled Decision ] ──► [ Autonomous Codebase Recon ] ──► [ Frontier Question Round ] ──► [ Socratic Refinement ] ──► [ Settled Alignment ] ``` | Component | Responsibility | Reference / Target | |---|---|---| | **Autonomous Recon** | Read repo context, ADRs, schemas | Primary source files | | **Frontier Tree** | Map open decision nodes & prerequisites | `skills/grilling/SKILL.md` | | **Question Round** | Formatted numbered prompts + defaults | User interaction loop | | **Settled Alignment** | Clean, unambiguous problem/solution contract | Handoff to `to-spec` | --- ## Step-by-Step Procedure (TWI) ### Step 1: Autonomous Reconnaissance & Context Gathering - **Action**: Inspect relevant codebase files, configurations, documentation, and existing architectural decisions before drafting questions. - **Key Point**: Form a grounded mental model of existing constraints without asking the user for baseline facts. - **Why**: Asking users for facts available in the repository wastes their cognitive bandwidth and damages credibility. ### Step 2: Formulate Frontier Question Rounds - **Action**: Identify all open decisions whose prerequisites are settled and present them in a single batch with explicit recommendations. - **Key Point**: Format each question as `❓ **Q[N]** - **[Title]**:` with options and `➡️ [Recommended Default]`. - **Why**: Presenting strong recommendations lowers decision fatigue while forcing the user to actively affirm or override specific choices. - **Inline Checklist**: - [ ] Every question targets a genuine decision rather than an environmental fact - [ ] Questions are independent and can be answered in parallel - [ ] Concrete recommendation provided for every question - [ ] No implementation code started ### Step 3: Iterate and Deepen the Decision Tree - **Action**: Ingest user answers, collapse resolved nodes, uncover new downstream branches, and issue the next frontier round. - **Key Point**: Probe edge cases: failure modes, security boundaries, migration paths, and non-happy paths. - **Why**: Most architectural failures occur at boundaries that initial proposals casually gloss over. ### Step 4: Verify Alignment and Closure - **Action**: Summarize the unified decision contract and ask for explicit user confirmation. - **Key Point**: Confirm that the frontier is empty and zero unexamined assumptions remain. - **Why**: Clear alignment prevents costly rewrites during implementation. --- ## Anti-Rationalization Guardrails | Tempting Rationalization | Binding Rule | Engineering Rationale | |---|---|---| | *"The user's initial prompt looks clear enough, let's start coding."* | **Mandatory grilling on all non-trivial plans.** | Initial prompts almost always conceal edge-case ambiguities and conflicting constraints. | | *"I'll ask questions one by one across multiple back-and-forth turns."* | **Batch the entire frontier into numbered rounds.** | Single-question ping-pong creates conversational fatigue and fragments context. | | *"Ask the user which database schema or library version is used."* | **Discover environment facts autonomously.** | Agents must look up facts directly in code and configs rather than burdening the user. | | *"Agree with the user's preferred approach without challenging trade-offs."* | **Stress-test every critical design decision.** | Sycophancy allows architectural bugs and scaling bottlenecks to slip into production. |
Referenced files: 1
grill-with-docs4.89 KB
---
name: grill-with-docs
description: "Interview the user to stress-test a design while simultaneously recording domain terms and architectural decisions in project documentation. Use when planning features in a codebase and wanting domain terms captured in CONTEXT.md and ADRs — even if the user says \"grill this feature\". Do NOT use when no codebase context or documentation trail is needed."
---
# Grill with Docs
Conduct a structured Socratic design interview that concurrently distills and commits domain terminology into `CONTEXT.md` and hard-to-reverse architectural decisions into Architectural Decision Records (`docs/adr/*.md`).
---
## Core Invariants
1. **Simultaneous Documentation Distillation**: Extract and record ubiquitous domain language into `CONTEXT.md` and architectural choices into ADRs in real time as decisions settle.
2. **Batch Frontier Questioning**: Present independent frontier questions in structured batches with explicit defaults rather than one-off queries.
3. **Autonomous Fact Reconnaissance**: Research existing codebase architecture, types, and schemas autonomously before posing design questions.
4. **Lightweight ADR Trigger**: When a design choice creates significant technical debt or is difficult to reverse (e.g. database choice, sync strategy), write a dedicated ADR immediately.
5. **Zero Speculative Prose**: Document only decisions that have actively settled during the interview; keep draft notes clearly demarcated.
---
## Architecture & Map of Content (MOC)
```
[ User Feature Idea ] ──► [ Codebase Recon & ADR Review ] ──► [ Socratic Grilling Rounds ]
│
┌────────────────────────────────┴────────────────────────────────┐
▼ ▼
[ Update CONTEXT.md ] [ Author ADR Docs ]
- Ubiquitous language - Context & Decision
- Entity definitions - Consequences & Tradeoffs
```
| Artifact | Purpose | Reference Format |
|---|---|---|
| **Domain Dictionary** | Capture ubiquitous language and entity relationships | `skills/domain-modeling/CONTEXT-FORMAT.md` |
| **Architectural Record** | Record irreversible architectural choices | `skills/domain-modeling/ADR-FORMAT.md` |
| **Grilling Protocol** | Drive structured frontier question rounds | `skills/grilling/SKILL.md` |
---
## Step-by-Step Procedure (TWI)
### Step 1: Discover Existing Documentation Baseline
- **Action**: Check for existing `CONTEXT.md`, `GLOSSARY.md`, and `docs/adr/` records in the repository.
- **Key Point**: Adopt existing project conventions for domain modeling and decision logging.
- **Why**: Maintaining architectural consistency prevents duplicate terminology and fragmented records.
### Step 2: Conduct Socratic Grilling Rounds
- **Action**: Identify unsettled requirements and formulate numbered question rounds with recommended defaults.
- **Key Point**: Ask about boundaries, invariants, entity lifecycles, and failure recovery.
- **Inline Checklist**:
- [ ] Terminology vetted against existing domain dictionary
- [ ] Questions formatted with clear recommendations
- [ ] Architecture tradeoffs highlighted
### Step 3: Distill and Update Domain Glossary
- **Action**: As terms and entities are clarified, update or create `CONTEXT.md` using the standard format.
- **Key Point**: Ensure terms are strictly defined with unambiguous scope.
- **Why**: Shared ubiquitous language prevents misalignment between engineers and agents.
### Step 4: Author Architectural Decision Records (ADRs)
- **Action**: For significant architectural choices, author a new numbered ADR in `docs/adr/NNNN-<slug>.md`.
- **Key Point**: Include Context, Decision, Status, and Consequences (both positive and negative).
- **Why**: ADRs preserve institutional memory and prevent rehashing past debates.
---
## Anti-Rationalization Guardrails
| Tempting Rationalization | Binding Rule | Engineering Rationale |
|---|---|---|
| *"I'll write the documentation after the entire interview finishes."* | **Capture terms and ADRs as decisions settle.** | Post-hoc documentation often drops subtle nuances, tradeoffs, and rationale. |
| *"This architectural choice is minor, no ADR needed."* | **If it is hard to reverse, author an ADR.** | Seemingly minor choices often cascade into major technical debt. |
| *"I will ask the user to explain the existing domain model."* | **Read existing docs and schemas autonomously.** | Agents must build initial context from repo artifacts before engaging the user. |
Referenced files: 1
handoff3.82 KB
--- name: handoff description: "Compact current conversation context, decisions, evidence, and next actions into a portable markdown handoff for a fresh agent session. Use when ending a session, transferring work to another agent, switching workspaces, or preparing a resume point — even if the user just says \"save progress\". Do NOT use for committing code to git or creating branch PRs." --- # Handoff Compact complex conversational history, locked architectural decisions, unblocked next actions, and verification evidence into a self-contained markdown handoff document for a fresh agent session. --- ## Core Invariants 1. **External Temp Storage**: Save handoff documents to the OS temporary directory (`/tmp/handoff-<timestamp>.md`) or designated scratch folder—never pollute the project root repository files. 2. **Context Pointer Economy**: Point directly to repository artifacts, issue URLs, commit SHAs, and test suites rather than duplicating walls of text. 3. **Suggested Skills Routing**: Explicitly declare the `## Suggested Skills` block indicating which specialized skills the resuming agent should invoke first. 4. **Strict Secret Redaction**: Ensure zero credentials, tokens, API keys, or private auth headers are leaked into the handoff file. 5. **Exact Next Command**: Provide the exact command line or next prompt required to continue execution without ambiguity. --- ## Architecture & Map of Content (MOC) ``` [ Active Multi-Turn Session ] ──► [ Extract Invariants, Decisions & Blockers ] ──► [ Write `/tmp/handoff-*.md` ] ──► [ Output Resume Command ] ``` | Section | Purpose | Example Content | |---|---|---| | **Goal & Scope** | What this initiative accomplishes | Spec boundary, user requirements | | **Decisions & Invariants** | Hard architectural contracts settled | Domain model choices, ADR pointers | | **Current State & Diff** | Exact working tree status | Git status, modified files, passing tests | | **Suggested Skills** | Tools the next session needs | `[skills/implement-spec/SKILL.md](file:///...)` | | **Exact Next Action** | Atomic next step | Command or task frontier claim | --- ## Step-by-Step Procedure (TWI) ### Step 1: Synthesize Session Delta & Decisions - **Action**: Extract the critical path decisions settled, active files modified, and outstanding questions. - **Key Point**: Check that all decisions link to their primary ADRs or specs. - **Why**: Resuming agents need the "why" behind choices without re-litigating settled discussions. ### Step 2: Format Handoff Document & Redact Secrets - **Action**: Structure the markdown document with Goal, Completed Work, Active Blockers, Suggested Skills, and Next Actions. - **Key Point**: Scan the text to ensure no private tokens or environment secrets are included. - **Inline Checklist**: - [ ] Saved in `/tmp/` (e.g. `/tmp/handoff-2026-08-30.md`) - [ ] Suggested skills explicitly listed - [ ] Zero secrets present ### Step 3: Emit Resume Instructions - **Action**: Output a clean completion summary to the user with the exact path to the handoff file and how the next agent can ingest it. - **Why**: Allows instant session resumption without context loss. --- ## Anti-Rationalization Guardrails | Tempting Rationalization | Binding Rule | Engineering Rationale | |---|---|---| | *"Write the handoff directly into the project repository root."* | **Save handoffs to `/tmp/` or temporary OS directory.** | Ephemeral session dumps pollute git history and clutter repo source trees. | | *"Paste full file contents into the handoff document."* | **Use concise context pointers and file paths.** | Pasting entire files consumes context budget on the resuming agent session. | | *"Omit suggested skills and let the next agent guess."* | **Mandatory 'Suggested Skills' section.** | Direct skill guidance prevents the resuming agent from drifting into generic workflows. |
Referenced files: 1
implement4.36 KB
--- name: implement description: "Build scoped code changes and run verification from an approved specification or plan. Use when a concrete spec, approved plan, or set of tickets is ready to implement, when the user asks to build or code a planned feature, or when executing tasks — even if they don't explicitly say \"implement\". Do NOT use for initial architecture exploration or vague requirements." --- # Implement Execute scoped code changes and thorough verification from an approved specification, plan, or ticket set. ## Core Principle > **Implement incrementally via vertical tracer slices, verify at each step, and maintain a green build at every seam.** --- ## Core Invariants 1. **Approved Plan or Spec First**: Never start broad implementation on ambiguous or unapproved specs. 2. **Vertical Slice Progress**: Implement in small, testable tracer slices rather than mass horizontal file edits. 3. **Continuous Typecheck & Test Verification**: Run local typechecks and single test files after every file modification, and the full suite before completion. 4. **Pre-Commit Quality Gate**: Pass clean code review (`code-review`) and lint checks before committing. 5. **Zero Dead or Speculative Code**: Write strictly the code required to satisfy the spec; avoid speculative abstractions. --- ## Architecture & Map of Content (MOC) ``` [ Review Approved Spec ] ──► [ Order Tracer Slices ] ──► [ TDD Cycle per Slice ] ──► [ Full Suite Verification ] ──► [ Code Review & Commit ] ``` | Phase | Responsibility | Key Action | |---|---|---| | **1. Preparation** | Read spec, identify dependencies, check existing test seams | Verify repo builds cleanly before touching code | | **2. Tracer Execution** | Implement slice-by-slice with TDD | Red-green cycles on target components | | **3. Continuous Check** | Fast feedback on type errors and single tests | Run targeted test runner after each edit | | **4. Integration Gate** | End-to-end regression validation | Full test suite, linter, and typecheck pass | | **5. Delivery** | Review diff against standards and spec | Route to `code-review` and clean commit | --- ## Step-by-Step Procedure (TWI) ### Step 1: Ingest Spec & Verify Baseline Cleanliness - **Action**: Read the spec or ticket descriptions, inspect referenced files, and verify the existing test suite passes cleanly. - **Key Point**: Never begin adding new features to a broken or failing baseline. - **Why**: Existing failures confound regression detection for new changes. ### Step 2: Implement via Vertical Slices (TDD) - **Action**: For each ticket or component, write failing assertions at public seams, then implement minimal passing logic. - **Key Point**: Keep edits contained and avoid touching unrelated modules. - **Inline Checklist**: - [ ] Public seam identified and tested - [ ] Typecheck passes without errors - [ ] Targeted test passes cleanly ### Step 3: Full Test Suite & Quality Verification - **Action**: Run the complete project test suite, typechecker, and linter across the entire repository. - **Key Point**: Zero failures, zero unexpected warnings, zero unformatted files. - **Why**: Changes in one module can subtly break downstream consumers. ### Step 4: Code Review & Final Commit - **Action**: Audit the git diff against repository coding standards and original requirements using `code-review`. - **Key Point**: Write a clear, value-communicating commit message summarizing the change. - **Why**: Durable commit histories preserve architectural reasoning for future maintainers. --- ## Anti-Rationalization Guardrails | Tempting Rationalization | Binding Rule | Engineering Rationale | |---|---|---| | *"I'll edit all 10 files at once before testing anything."* | **Vertical tracer slicing: verify each file/unit immediately.** | Massive multi-file edits make errors difficult to isolate and debug. | | *"Typechecks are slow, I'll run them at the very end."* | **Run fast targeted checks continuously.** | Catching type errors early prevents building on flawed data signatures. | | *"The existing test suite was failing before, so ignore it."* | **Establish clean baseline before modifying code.** | Unaccounted failures hide new regressions introduced by the feature. | | *"Skip code review since the code runs fine."* | **Mandatory diff review against spec and standards.** | Functional code can still violate repository conventions and security standards. |
Referenced files: 1
implement-spec5.87 KB
---
name: implement-spec
description: "Execute a full specification across task tickets using isolated subagent branches into a unified PR. Use when a specification with associated task-graph tickets is ready to implement across multiple subagents or isolated worktrees, and the goal is a complete pull request on a single branch — even if the user says \"build the whole spec\". Do NOT use for single isolated bug fixes or exploratory coding without tickets."
---
# Implement Spec
Orchestrate the parallel implementation of an approved specification and its DAG task graph across isolated subagent branches, culminating in a single unified, code-reviewed pull request.
---
## Core Invariants
1. **DAG Frontier Concurrency**: Implementer subagents work exclusively on unblocked frontier tickets; merge completions immediately unlock newly unblocked tickets.
2. **Strict Worktree Isolation**: Every implementer subagent executes in its own dedicated, isolated Git worktree and branch to prevent file collision.
3. **Context Pointers Exclusively**: Communicate to subagents strictly via context pointers (spec path, ticket IDs, ADRs); avoid dumping giant walls of text.
4. **Dedicated Merger Verification**: Merging completed worktree branches into the main PR branch is handled by a merger subagent with full test suite verification.
5. **Unified Code-Review Quality Gate**: Run `code-review` across the unified PR branch before marking it ready for human review, fixing all findings in a final pass.
---
## Architecture & Map of Content (MOC)
```
[ Approved Spec & DAG Tickets ] ──► [ Create PR Branch ] ──► [ Parallel Worktree Subagents ]
│
┌────────────────────────────────┴────────────────────────────────┐
▼ ▼
[ Ticket #01 Worktree ] [ Ticket #02 Worktree ]
- Isolated branch - Isolated branch
- Red/Green TDD cycle - Red/Green TDD cycle
│ │
└────────────────────────────────┬────────────────────────────────┘
▼
[ Merge & Verify on PR Branch ]
│
▼
[ Code Review & Worktree Cleanup ]
```
| Role | Responsibility | Execution Mode |
|---|---|---|
| **Exploration Agent** | Pre-read external documentation and shared schemas | Background subagent (`research`) |
| **Implementer Agents** | Execute single tickets inside isolated worktrees | Parallel subagents (`implement`) |
| **Merger Agent** | Merge feature branches into PR branch & verify tests | Serial integration pass |
| **Review Gate** | Run two-axis code review on final PR branch | `skills/code-review/SKILL.md` |
---
## Step-by-Step Procedure (TWI)
### Step 1: Ingest Task Graph & Initialize PR Branch
- **Action**: Read the specification and associated tickets to compute the initial unblocked frontier.
- **Key Point**: Create a dedicated PR branch (`feat/<spec-slug>`) and draft pull request linking all target tickets.
- **Why**: Linking tickets upfront ensures automated status tracking and traceability.
### Step 2: Dispatch Parallel Implementer Subagents
- **Action**: For each ticket on the unblocked frontier, spawn an implementer subagent in an isolated worktree (`.worktrees/<ticket-slug>`).
- **Key Point**: Pass context pointers to the spec and ticket without duplicating instructions.
- **Inline Checklist**:
- [ ] Worktrees isolated from main repository workspace
- [ ] Implementers follow strict TDD red-green cycle
- [ ] Each implementer works on a single assigned ticket
### Step 3: Merge Completed Tickets & Advance Frontier
- **Action**: When an implementer completes, merge its branch into the PR branch, run the test suite, and delete the worktree.
- **Key Point**: Recompute the task graph frontier to launch newly unblocked tickets immediately.
- **Why**: Continuous merging keeps integration diffs small and surfaces conflicts early.
### Step 4: Run Two-Axis Code Review & Cleanup
- **Action**: Once all tickets are merged, run `code-review` on `feat/<spec-slug>`.
- **Key Point**: Remediate all Standards and Spec findings before marking the PR ready for human review.
- **Why**: Comprehensive pre-merge review guarantees production-grade architecture and full spec compliance.
---
## Anti-Rationalization Guardrails
| Tempting Rationalization | Binding Rule | Engineering Rationale |
|---|---|---|
| *"Run all implementers in the same shared working tree."* | **Mandatory isolated git worktrees per subagent.** | Shared workspaces cause file lock contention, overwrites, and git index corruption. |
| *"Skip code review since all unit tests passed."* | **Mandatory two-axis code review on the final PR branch.** | Passing tests do not catch architectural smells, Fowler anti-patterns, or missed spec clauses. |
| *"Implement blocked tickets before their dependencies merge."* | **Strictly adhere to the DAG frontier.** | Implementing blocked tickets prematurely results in massive merge conflicts and rework. |
Referenced files: 1
improve-codebase-architecture4.55 KB
--- name: improve-codebase-architecture description: "Survey codebases for shallow modules, weak seams, and deepening opportunities, producing a visual report. Use when conducting architectural reviews, identifying design debt, finding deepening opportunities, or preparing codebase refactors — even if the user says \"analyze our architecture\". Do NOT use for basic syntax linting or formatting." --- # Improve Codebase Architecture Survey codebases for shallow modules, leaky abstractions, and weak seams, generating an interactive, visual HTML report in the OS temp directory with before/after architectural refactoring models. --- ## Core Invariants 1. **YAGNI Scope First**: Prioritize hotspots in recent commit history (`git log --oneline`) where architectural friction is actively slowing development. 2. **Strict Design Vocabulary**: Frame all findings using canonical terms (**module**, **interface**, **depth**, **seam**, **adapter**, **leverage**, **locality**) and domain vocabulary from `CONTEXT.md`. 3. **External Temp Report**: Write the visual report to the OS temporary directory (`<tmpdir>/architecture-review-<timestamp>.html`) with Tailwind CDN and Mermaid diagrams; never litter the repository with review HTML. 4. **Before/After Visual Models**: Every deepening candidate must feature a clear before/after structural diagram illustrating interface simplification and implementation depth. 5. **No Speculative Interface Proposing**: Propose candidate problem areas and deepening directions; do not propose concrete code interfaces until the user selects a candidate. --- ## Architecture & Map of Content (MOC) ``` [ Hotspot & Git Log Analysis ] ──► [ Deepening Candidate Survey ] ──► [ Generate HTML Report in /tmp ] ──► [ Grilling on Chosen Candidate ] ``` | Component | Responsibility | Reference | |---|---|---| | **Hotspot Scanner** | Identify frequently changed, high-friction files | Git log & subagent exploration | | **HTML Visual Report** | Render side-by-side Mermaid & Tailwind cards | `skills/improve-codebase-architecture/HTML-REPORT.md` | | **Candidate Deepening** | Socratic review of selected candidate | `skills/grilling/SKILL.md` + `skills/codebase-design/SKILL.md` | --- ## Step-by-Step Procedure (TWI) ### Step 1: Scan Recent Codebase Hotspots - **Action**: Inspect `git log --oneline -n 100` and `CONTEXT.md` to identify high-churn modules and domain boundaries. - **Key Point**: Focus on modules where understanding one concept requires hopping between multiple fragmented files. - **Why**: Deepening stable, untouched legacy files yields low ROI compared to active hotspots. ### Step 2: Survey Deepening Opportunities - **Action**: Evaluate candidate modules using the deletion test and locate shallow pass-throughs. - **Key Point**: Check for extracted pure functions that lack locality and leak callers' state. - **Inline Checklist**: - [ ] Hotspot files identified from git churn - [ ] 2–4 distinct deepening candidates formulated - [ ] Candidates categorized by recommendation strength (`Strong`, `Worth exploring`, `Speculative`) ### Step 3: Generate Self-Contained Visual HTML Report - **Action**: Write `<tmpdir>/architecture-review-<timestamp>.html` with Tailwind and Mermaid CDN scripts. - **Key Point**: Include problem descriptions, leverage/locality benefits, and side-by-side before/after Mermaid diagrams. - **Why**: Visual architecture diagrams communicate structural improvements far more effectively than walls of text. ### Step 4: Open Report & Facilitate Candidate Selection - **Action**: Launch the HTML report via `open <path>` (macOS) / `xdg-open` (Linux) / `start` (Windows) and ask the user which candidate to explore. - **Key Point**: Upon selection, enter the grilling loop to settle module boundaries and update `CONTEXT.md` / ADRs. - **Why**: Collaborative candidate selection ensures team buy-in before investing in refactoring. --- ## Anti-Rationalization Guardrails | Tempting Rationalization | Binding Rule | Engineering Rationale | |---|---|---| | *"Write the HTML review report directly into the repo root."* | **Write reports exclusively to the OS temp directory.** | Review artifacts should never pollute project Git history. | | *"Draft full replacement code files immediately in the report."* | **Present architectural direction and diagrams first.** | Premature coding before agreeing on architectural seams leads to wasted effort. | | *"Re-open settled ADR decisions without strong evidence."* | **Respect existing ADRs unless severe friction is demonstrated.** | Constant relitigation of settled decisions stalls progress. |
Referenced files: 2
j-space6.07 KB
---
name: j-space
description: "Operate the agent's inner cognitive workspace for multi-step reasoning, long-horizon planning, deep debugging, and calibrated self-monitoring. Use when tackling complex multi-step problems, architecture design, long mechanical tasks requiring goal preservation, suspicious input audits, or when asked to think deeper or harder — even if they don't explicitly say \"J-space\". Do NOT use for single-step trivial lookups, basic formatting, or conversational small talk."
---
# J-Space: Inner Cognitive Workspace
Operate the model's inner cognitive workspace (J-space) for multi-step chained reasoning, long-horizon planning, complex debugging, calibrated confidence, global consistency, and resistant self-monitoring.
## The J-Space Premise
You do not only produce words; you also reason before — and without — immediately emitting them. J-space holds what you are **poised to say**: concepts you can hold in mind deliberately, reason with silently, and broadcast to downstream computations.
---
## Core Invariants
1. **Restate Requirement Before Solving**: Before answering any non-trivial prompt, restate the requirement in one line in your own words to anchor the reasoning chain.
2. **Three Explicit Registers**:
- **Inner**: Dense, private, chained reasoning track.
- **Ledger**: Durable state tracking (Goal, Settled, Open, Next). Checked at every seam.
- **Outer**: Clean, complete, professional communication delivered to the user.
3. **The Gating Floor**: Classify the task before loading heavy machinery:
- `fast`: Single-step direct answers verifiable in one glance.
- `full`: 2–4 steps, single deliverable, verifiable in one reading.
- `loop`: Multi-stage, multi-file, multi-turn stateful work.
4. **Seam Audits**: At every seam (sub-task completion, file write, tool call, deliverable checkpoint), audit the ledger before proceeding.
5. **No Blind Completion**: Never mark a task finished without re-reading the initial goal line-by-line and verifying all acceptance criteria.
---
## Architecture & Map of Content (MOC)
```
[ Incoming Task ] ──► [ Restate Requirement ] ──► [ Task Gate: Fast / Full / Loop ]
│
▼
┌──────────────────────────────────────────────────────────┐
│ J-Space Cognitive Loop │
│ ├─ Inner Register (Chained intermediate reasoning) │
│ ├─ Ledger Register (Goal / Open / Settled / Next) │
│ └─ Seam Audit (Check invariants at each transition) │
└───────────────────────────┬──────────────────────────────┘
│
▼
[ Outer Register Delivery ]
```
| Functional Property | Cognitive Purpose | Failure Mode Prevented |
|---|---|---|
| **Capacity Selectivity** | Only 1–2 active concepts on stage at once | Attention thinning & overloaded context |
| **Directed Focus** | Goal remains active through tedious middle | Goal evaporation during mechanical tasks |
| **Deep Reasoning** | Intermediate bridges form before conclusions | Rationalizing unverified gut guesses |
| **Self-Monitoring** | Calibrated confidence & error detection | Overconfident halluncinations |
| **Seam Auditing** | State refreshed at every transition | Context drift over long execution turns |
---
## Step-by-Step Procedure (TWI)
### Step 1: Awakening & Requirement Re-Encoding
- **Action**: Restate the requirement in one line, in your own words, before writing code or making tool calls.
- **Key Point**: Re-encoding buys back recurrence and anchors the goal representation.
- **Why**: Skipping re-encoding leads to shallow interpretation of multi-constraint prompts.
### Step 2: Task Gate Classification
- **Action**: Classify into `fast`, `full`, or `loop` pass.
- **Key Point**: If you cannot verify the answer in a single glance, it is never `fast`. Escalating costs nothing.
- **Why**: Under-classifying complex tasks causes premature shortcutting.
### Step 3: Operate the Three Registers
- **Action**: Think in the inner register, record state in the ledger, and speak in the outer register.
- **Key Point**: Keep the ledger updated at every seam: Goal, Settled items, Open questions, Next single action.
- **Inline Checklist**:
- [ ] Requirement re-encoded
- [ ] Correct pass classified
- [ ] Ledger reflects current state
- [ ] Verification criteria stated explicitly
### Step 4: Seam Refresh & Invariant Verification
- **Action**: At every seam (file written, subagent returned, tool executed), check that invariants hold.
- **Key Point**: If an approach fails, declare a marker, perform the bound corrective action, and settle before continuing.
- **Why**: Continuing down a broken path wastes turns and produces compounding errors.
### Step 5: Clean Outer Delivery
- **Action**: Translate settled conclusions into clean outer-register deliverables.
- **Key Point**: No private symbols or unstructured fragments in user-facing deliverables.
- **Why**: Deliverables must be immediately usable, clear, and actionable.
---
## Anti-Rationalization Guardrails
| Tempting Rationalization | Binding Rule | Engineering Rationale |
|---|---|---|
| *"This is simple enough, skip the ledger."* | **Enforce the floor: multi-step work requires ledger tracking.** | Unsettled state drops out of context during long turns. |
| *"The first guess sounds right, skip the intermediate step."* | **Derive intermediate bridges before adopting conclusions.** | Pretrained fluency often produces plausible but flawed shortcuts. |
| *"I will clean up the verification at the very end."* | **Verify at every seam, not just at the end.** | Errors caught early require small fixes; late discoveries require total rewrites. |
| *"Confidence is always high across all steps."* | **Calibrate confidence per step.** | Uniform confidence indicates disabled self-monitoring. |
Referenced files: 1
migrate-to-shoehorn4.3 KB
---
name: migrate-to-shoehorn
description: "Migrate unsafe TypeScript test assertions to @total-typescript/shoehorn with explicit fixture intent. Use when test files contain unsafe `as` or `as unknown as` typecasts, when mock fixtures break under typechecking, or when modernizing test typing — even if the user says \"fix test type assertions\". Do NOT use for production code types."
---
# Migrate to Shoehorn
Safely migrate unsafe, brittle TypeScript test-fixture typecasts (`as Type`, `as unknown as Type`) to intention-revealing, type-safe helpers from `@total-typescript/shoehorn`.
---
## Core Invariants
1. **Test-Code Isolation Strictly**: `@total-typescript/shoehorn` is strictly for test suites and test fixtures; NEVER import or use shoehorn helpers in production application code.
2. **Intent-Specific Helper Mapping**:
- Partial objects where only some keys matter $\rightarrow$ `fromPartial(obj)`
- Intentionally invalid/wrong types for error testing $\rightarrow$ `fromAny(obj)`
- Fully populated mocks $\rightarrow$ `fromExact(obj)`
3. **Preserve Autocomplete & Refactor Safety**: Maintain TypeScript type inference and IDE autocompletion across all migrated test fixtures.
4. **Automated Discovery & Regex Caution**: Discover test casts via targeted grep (`as [A-Z]`); avoid blind global replacements that might alter production code.
5. **Mandatory Typecheck Green**: Every migration pass must conclude with a green `tsc --noEmit` or equivalent repository typecheck command.
---
## Architecture & Map of Content (MOC)
```
[ Unsafe Test Casts (`as Type`, `as unknown as`) ] ──► [ Intent Classification ] ──► [ Replace with Shoehorn Helper ] ──► [ Verify Typecheck ]
```
| Unsafe Pattern | Shoehorn Replacement | Target Scenario |
|---|---|---|
| `{ id: "123" } as User` | `fromPartial<User>({ id: "123" })` | Test cares only about `id` in large type |
| `{ id: 123 } as unknown as User` | `fromAny<User>({ id: 123 })` | Testing runtime validation on bad input |
| `completeMock as User` | `fromExact<User>(completeMock)` | Explicit complete mock structure |
---
## Step-by-Step Procedure (TWI)
### Step 1: Install Dependency with Detected Package Manager
- **Action**: Install `@total-typescript/shoehorn` as a devDependency (`bun add -d`, `pnpm add -D`, `npm i -D`).
- **Key Point**: Check that it is added strictly to `devDependencies`.
- **Why**: Prevent bundling test-helper utilities into production builds.
### Step 2: Locate Unsafe Casts in Test Files
- **Action**: Search test files for type assertion smells:
```bash
grep -rnE " as [A-Z]| as unknown as " --include="*.test.ts" --include="*.spec.ts" src/
```
- **Key Point**: Inspect surrounding context to determine whether the fixture is a partial mock or an intentionally invalid payload.
- **Inline Checklist**:
- [ ] Only test files (`*.test.ts`, `*.spec.ts`) targeted
- [ ] Zero production source files modified
- [ ] Intent classified per fixture (`fromPartial` vs. `fromAny`)
### Step 3: Replace Casts and Import Helpers
- **Action**: Replace `as Type` with `fromPartial(...)` and `as unknown as Type` with `fromAny(...)`, adding `import { fromPartial, fromAny } from "@total-typescript/shoehorn"`.
- **Why**: `fromPartial` documents that missing fields are deliberate, avoiding false compiler errors while retaining type intelligence.
### Step 4: Run Typecheck & Test Suite
- **Action**: Execute the project's typecheck command (e.g. `bun run typecheck`, `pnpm check`, `tsc --noEmit`) and run the test suite.
- **Key Point**: Verify that tests pass and TypeScript emits 0 errors.
- **Why**: Ensures no subtle type contract regressions were introduced during migration.
---
## Anti-Rationalization Guardrails
| Tempting Rationalization | Binding Rule | Engineering Rationale |
|---|---|---|
| *"Use shoehorn in production helpers to bypass strict types."* | **Forbidden. Test code only.** | Bypassing types in production creates silent runtime TypeError crashes. |
| *"Use `fromAny` everywhere because it's easier than `fromPartial`."* | **Use `fromPartial` for valid partials; `fromAny` only for invalid types.** | Overusing `fromAny` disables autocomplete and typechecking benefits. |
| *"Skip running typecheck after updating test files."* | **Mandatory green typecheck verification.** | Minor syntax errors or missing generic arguments break CI builds. |
Referenced files: 1
ml-best-practices6.33 KB
---
name: ml-best-practices
description: "Statistical machine learning best practices, exploratory data analysis, feature engineering, and rigorous model evaluation. Use when analyzing tabular datasets, engineering features, training classifiers or regressors, forecasting time series, or computing 95% bootstrap confidence intervals — even if they don't explicitly say \"ml best practices\". Do NOT use for standard database queries without ML, basic spreadsheet formatting, or general application backend logic."
---
# Machine Learning Best Practices
Apply rigorous machine learning engineering standards across exploratory analysis, feature engineering, model training, and statistical performance validation.
## Core Principle
> **Every model finding must be grounded in leakage-free splits, dual-model baselines, and statistical confidence intervals.**
---
## Core Invariants
1. **Strict Train-Test Isolation**: Always partition datasets into training, validation, and test splits **before** fitting any scalers, encoders, or transformers.
2. **Mandatory Missing Value Strategy**: Explicitly quantify missing value frequencies and document domain rationale for keeping, dropping, or imputing them.
3. **Dual-Model Benchmark**: Never rely on a single model architecture in isolation. Train at least two distinct candidate families against a naive baseline.
4. **Statistical Significance Over Raw Averages**: Report 95% bootstrap confidence intervals for key evaluation metrics (F1, AUC, RMSE) to verify statistical separation.
5. **Story-First Notebook Structure**: Every code execution cell must be paired with an analytical markdown cell explaining the observed patterns, trade-offs, and conclusions.
---
## Architecture & Map of Content (MOC)
```
Raw Dataset & Objective
│
▼
┌────────────────────────────────────────┐
│ 1. Data Hygiene & Anomaly Screening │ (Missingness, Cardinality, Invariants)
└──────────────────┬─────────────────────┘
│
┌───────────┴───────────┐
▼ ▼
┌──────────────┐ ┌──────────────┐
│ Clustering │ │ Forecasting │ (Silhouette Optimization, Stationarity)
└──────────────┘ └──────────────┘
│ │
└───────────┬───────────┘
│
┌───────────┴───────────┐
▼ ▼
┌──────────────┐ ┌──────────────┐
│Classification│ │ Regression │ (Pipeline Encoders, Regularization)
└──────────────┘ └──────────────┘
│
▼
┌────────────────────────────────────────┐
│ 2. Statistical Model Comparison │ (95% Bootstrap CIs, Slice Analysis)
└────────────────────────────────────────┘
```
---
## Step-by-Step Procedure (TWI)
### Step 1: Exploratory Analysis & Anomaly Detection
- **Action**: Inspect schema, compute summary statistics, and visualize target distributions with scatter plots and histograms.
- **Key Point**: Check for target column anomalies, extreme outliers, and non-sensical values before modeling.
- **Why**: Training models on corrupt or undetected anomalous targets invalidates the experimental setup.
### Step 2: Supervised Modeling (Classification / Regression)
- **Action**: Apply transformers inside scikit-learn pipelines fit exclusively on training data; evaluate regularization to control overfitting.
- **Key Point**: For high-cardinality nominal features, restrict cardinality or use target encoding with out-of-fold regularization.
- **Why**: Fitting transformers on the combined dataset causes subtle target and variance leakage.
- **Inline Checklist**:
- [ ] Data split chronologically (if temporal) or via StratifiedKFold (if classification)
- [ ] Encoders and scalers wrapped inside Pipeline, fitted only on X_train
- [ ] Confusion matrices, ROC curves, or residual diagnostic plots generated
- [ ] Missing values handled with documented justification
### Step 3: Unsupervised Clustering & Time-Series Forecasting
- **Action**: For clustering, standardize features and optimize the Silhouette Score; for forecasting, test stationarity and seasonality.
- **Key Point**: For time series, verify that no future information leaks into historical lag features.
- **Why**: Silhouette optimization prevents arbitrary cluster count selection; stationarity testing determines appropriate model family.
### Step 4: Model Comparison & 95% Bootstrap Confidence Intervals
- **Action**: Benchmark multiple model candidates on identical cross-validation folds and calculate bootstrap confidence intervals.
- **Key Point**: Conduct slice-based error analysis across critical subpopulations and compile an operational trade-off table.
- **Why**: Point estimates of accuracy often conceal performance drops on key customer cohorts.
---
## Anti-Rationalization Guardrails
| Tempting Rationalization | Binding Rule | Engineering Rationale |
|---|---|---|
| *"Standard scaling before splitting is harmless."* | **Fit scalers strictly on training folds.** | Scalers compute dataset-wide statistics that cause subtle distribution leakage. |
| *"Accuracy is 94%, so the classifier is ready."* | **Always inspect PR-AUC, F1, and confusion matrix.** | On imbalanced classes, 94% accuracy may be worse than a naive majority-class guess. |
| *"We don't need a baseline because XGBoost is superior."* | **Always establish naive and linear baselines first.** | Naive baselines justify model complexity and reveal genuine signal gains. |
| *"Shuffling temporal data is fine if well-mixed."* | **Split temporal data chronologically.** | Shuffling time series causes future lookahead leakage into historical folds. |
Referenced files: 1
prototype4.77 KB
---
name: prototype
description: "Build a throwaway prototype to answer a specific design, state model, or UI exploration question. Use when evaluating whether an interface feels right, exploring UI concepts, or testing logic before committing to a full spec — even if the user says \"mock this up\". Do NOT use for production implementation."
---
# Prototype
Build throwaway exploratory code designed to answer a single load-bearing architectural, state machine, or UI question rapidly without production overhead.
---
## Core Invariants
1. **Throwaway By Construction**: Prototypes must be explicitly marked throwaway with zero production persistence or abstraction overhead.
2. **Branch Isolation**: Separate logic/state-machine prototypes (`LOGIC.md`) from visual UI variant explorations (`UI.md`).
3. **Single-Command Launch**: UI prototypes run with one command (`pnpm dev`, `bun run ...`); logic prototypes are single double-clickable HTML/JS files.
4. **Transparent State Exposure**: Every action or transition must visually expose the complete underlying state payload.
5. **Decisions-Only Mainline Merge**: Merge only the validated decision/type into main; commit the prototype code to a separate scratch branch.
---
## Architecture & Map of Content (MOC)
```
[ Load-Bearing Design Question ] ──► [ Select Exploration Branch ] ──► [ Minimal Runnable Prototype ] ──► [ Extract Settled Decision ]
│
┌────────────────────────┴────────────────────────┐
▼ ▼
[ Logic / State Prototype ] [ UI Variant Explorer ]
- Single HTML/JS file - Multi-variant route
- Free-play + guided tabs - Bottom floating bar
- Full state visualizer - URL parameter toggle
```
| Branch | Question Answered | Artifact Format |
|---|---|---|
| **Logic / State** | "Does this state machine or business rule feel right?" | `skills/prototype/LOGIC.md` |
| **UI Variations** | "What should this visual interaction look like?" | `skills/prototype/UI.md` |
---
## Step-by-Step Procedure (TWI)
### Step 1: Identify the Question & Exploration Branch
- **Action**: Determine whether the core uncertainty is logical/stateful or visual/experiential.
- **Key Point**: Check if the component has complex transition states (choose Logic) or styling/layout decisions (choose UI).
- **Why**: Choosing the wrong format wastes time building UI for logical edge cases or state machines for static layouts.
### Step 2: Implement the Minimal Runnable Prototype
- **Action**: Create the prototype with zero database persistence and minimal abstractions:
- Logic: Self-contained HTML with state buttons, transition logs, and guided walkthrough scenarios.
- UI: Dedicated scratch route with 2–4 radically different variants toggled via query params.
- **Inline Checklist**:
- [ ] Marked as throwaway in filenames and comments
- [ ] Launches via single standard command or browser click
- [ ] Displays live internal state on every interaction
### Step 3: Interactive Evaluation & Decision Extraction
- **Action**: Walk the user through the prototype to evaluate edge cases and record the verdict.
- **Key Point**: Extract the validated data shape, state machine reducer, or component layout into the issue tracker or spec.
- **Why**: Capturing the distilled finding prevents throwaway prototype code from accidentally morphing into production spaghetti.
### Step 4: Archive Prototype & Clean Main
- **Action**: Commit the prototype to a scratch branch (`prototype/<name>`), link it in the ticket, and keep `main` clean.
- **Key Point**: Never merge un-linted prototype hacks directly into main branches.
- **Why**: Strict separation keeps the main codebase pristine while retaining historical design context.
---
## Anti-Rationalization Guardrails
| Tempting Rationalization | Binding Rule | Engineering Rationale |
|---|---|---|
| *"Let's build this prototype directly inside the main production file."* | **Forbidden. Keep prototypes in isolated scratch paths.** | Inlining prototypes into production code creates accidental dependencies and tech debt. |
| *"Add full unit tests and error handling to the prototype."* | **Skip production hardening in throwaway prototypes.** | Hardening exploratory code slows learning cycles and creates emotional attachment. |
| *"Merge the entire prototype into main since it works."* | **Extract decisions only; archive prototype branch.** | Prototypes lack production safety, validation, error boundaries, and documentation. |
Referenced files: 3
research4.19 KB
--- name: research description: "Investigate a technical question against high-trust primary sources and capture findings as a cited Markdown note. Use when gathering API facts, reading documentation, verifying library capabilities, or delegating reading legwork — even if the user says \"look into this library\". Do NOT use for writing production feature code." --- # Research Investigate technical questions, library capabilities, and architectural facts against authoritative primary sources, capturing verifiable findings in a cited Markdown research note. --- ## Core Invariants 1. **Primary Sources Exclusively**: Ground all findings in official documentation, source code, formal specifications, or first-party release notes; never rely on unverified blog posts or secondary summaries. 2. **Mandatory Attribution & Line-Level Citations**: Every technical claim, API signature, or version constraint must link directly to its primary source URI or repo file path. 3. **Background Agent Execution**: Run intensive reading, scraping, and repository audits in an isolated background subagent to preserve the primary agent's working context. 4. **Structured Decision Markdown Output**: Synthesize research into a permanent markdown file matching project conventions (e.g. `docs/research/<slug>.md` or `.scratch/research/<slug>.md`). 5. **Separation of Fact vs. Opinion**: Explicitly separate verified architectural facts from subjective engineering recommendations. --- ## Architecture & Map of Content (MOC) ``` [ Technical Question / Library Query ] ──► [ Spawn Research Subagent ] ──► [ Primary Source Retrieval ] ──► [ Cited Synthesis Report ] ``` | Component | Responsibility | Output Target | |---|---|---| | **Primary Source Auditor** | Read official docs, Github repos, specs | Web search & URL content tools | | **Research Note** | Structured findings with quotes & citations | `docs/research/<topic>.md` | | **Verification Gate** | Validate API signatures against runtime/version | Direct test snippets | --- ## Step-by-Step Procedure (TWI) ### Step 1: Formulate Research Hypothesis & Boundary - **Action**: Define the core technical questions, necessary library versions, and compatibility requirements. - **Key Point**: Distinguish between hard technical limits (e.g. rate limits, memory footprint) and ergonomic trade-offs. - **Why**: Unbounded research spirals into excessive token consumption without answering the core decision question. ### Step 2: Query Authoritative Primary Sources - **Action**: Fetch primary documentation, GitHub repositories, RFCs, and API references using search and web tools. - **Key Point**: Verify the exact version compatibility against the project's `package.json`, `pyproject.toml`, or `Cargo.toml`. - **Inline Checklist**: - [ ] Source is first-party / authoritative - [ ] Version matches project environment - [ ] Exact API signatures and failure modes captured ### Step 3: Author Cited Research Note - **Action**: Write the synthesis to `docs/research/<slug>.md` (or project standard path): - **Summary**: High-level verdict and recommended direction. - **Key Findings**: Concrete technical facts with direct markdown links. - **Code Samples**: Minimal, validated usage examples. - **Trade-offs & Gotchas**: Edge cases, performance bottlenecks, and limitations. - **Why**: Permanent research notes preserve institutional context and prevent repeating research across engineering cycles. --- ## Anti-Rationalization Guardrails | Tempting Rationalization | Binding Rule | Engineering Rationale | |---|---|---| | *"I remember how this library works from training data, so no need to look up current docs."* | **Mandatory primary-source lookup for all library facts.** | Training memory hallucinates deprecated API signatures and misses recent breaking changes. | | *"Summarize from a third-party tutorial or forum post."* | **Trace claims back to the authoritative primary source.** | Third-party tutorials often propagate anti-patterns and outdated workarounds. | | *"Inline the full research text into chat without writing a file."* | **Always commit findings to a durable research note.** | In-chat findings vanish across context resets; markdown files provide persistent documentation. |
Referenced files: 1
resolving-merge-conflicts4.53 KB
--- name: resolving-merge-conflicts description: "Resolve in-progress git merge or rebase conflicts by intent traced to primary sources. Use when hit with merge conflicts, CONFLICT markers in files, rebase pauses, or when git prompts to resolve conflicted hunks — even if the user just says \"fix this merge\". Do NOT use for creating new branches or routine rebasing without conflicts." --- # Resolving Merge Conflicts Systematically resolve in-progress Git merge and rebase conflicts by identifying the semantic intent of both branches and verifying integration integrity with automated test gates. --- ## Core Invariants 1. **3-Way Intent Reconstruction**: Understand the common merge-base ancestor, the incoming changes (`THEIRS`), and the target branch changes (`OURS`) before altering conflicted code. 2. **Never Blindly Choose Ours or Theirs**: Inspect every conflict hunk individually; synthesize solutions that preserve the functional requirements of both branches. 3. **Zero Orphaned Conflict Markers**: Verify that all `<<<<<<<`, `=======`, and `>>>>>>>` markers are completely eliminated before staging. 4. **Non-Destructive Resolution**: Preserve adjacent unchanged code, comments, and docstrings; avoid accidental line deletions outside conflict hunks. 5. **Mandatory Post-Resolution Test Pass**: Execute the full build, typecheck, and test suite before concluding the merge (`git commit`) or rebase (`git rebase --continue`). --- ## Architecture & Map of Content (MOC) ``` [ Merge/Rebase Conflict Detected ] ──► [ Identify 3-Way Merge Base & Commits ] ──► [ Hunk-by-Hunk Semantic Synthesis ] ──► [ Build & Test Gate ] ``` | Phase | Responsibility | Verification Command | |---|---|---| | **Conflict Discovery** | Identify all unmerged files | `git status --porcelain | grep "^UU\|^AA\|^DU\|^UD"` | | **Ancestor Inspection** | View base version of conflicted file | `git show :1:<file>` (Base), `:2:<file>` (Ours), `:3:<file>` (Theirs) | | **Integration Gate** | Run compiler, linter, and unit tests | Project build / test command | --- ## Step-by-Step Procedure (TWI) ### Step 1: Identify Conflicted Files and Merge Context - **Action**: Run `git status` to enumerate all conflicted files and inspect recent commit logs on both branches: ```bash git log --oneline -n 5 HEAD git log --oneline -n 5 MERGE_HEAD # or REBASE_HEAD ``` - **Key Point**: Determine what feature or bug fix each branch was attempting to deliver. - **Why**: Understanding developer intent prevents resolving conflicts with syntax-valid but semantically broken hybrids. ### Step 2: Resolve Conflicting Hunks Semantically - **Action**: Open each conflicted file, analyze the diff hunks between `HEAD` (Ours) and incoming (Theirs), and rewrite the block to satisfy both requirements. - **Key Point**: If both branches added new imports, dependencies, or routes, combine them cleanly without duplicates. - **Inline Checklist**: - [ ] All conflict markers (`<<<`, `===`, `>>>`) removed - [ ] Duplicate imports and exports deduplicated - [ ] New functionality from both branches retained ### Step 3: Verify with Build & Test Suite - **Action**: Run the repository's typechecker, linter, and test suite across the resolved working tree. - **Key Point**: If tests fail, investigate whether resolving the conflict broke subtle runtime assumptions. - **Why**: Many merge conflicts compile cleanly but introduce logical regressions that only tests catch. ### Step 4: Stage and Conclude Merge/Rebase - **Action**: Stage the resolved files (`git add <files>`) and complete the operation: - For merge: `git commit` (preserving standard merge commit message). - For rebase: `git rebase --continue`. - **Key Point**: Check `git status` to ensure the working tree is clean. - **Why**: Clean conclusion ensures upstream CI pipelines can build the merged branch without human intervention. --- ## Anti-Rationalization Guardrails | Tempting Rationalization | Binding Rule | Engineering Rationale | |---|---|---| | *"Accept 'ours' or 'theirs' completely with `git checkout --ours` to be fast."* | **Forbidden on semantic conflicts.** | Blanket checkouts overwrite valid work and revert critical bug fixes from one branch. | | *"Delete the failing test to get the merge commit through."* | **Fix the code to make tests pass.** | Deleting tests lowers coverage and introduces regressions into production. | | *"Assume the code is fine without running the full test suite."* | **Mandatory test pass before committing merge.** | Resolving conflicts manually frequently introduces syntax errors and broken imports. |
Referenced files: 1
retro5.34 KB
---
name: retro
description: "Conduct a retrospective on a coding session to systematically improve agent environment, navigation pointers, automated checks, coding standards, tool economy, or AGENTS.md instructions. Use when reflecting on completed work, auditing agent mistakes, or optimizing repository rules — even if the user says \"run a retro\". Do NOT use during active mid-task implementation."
---
# Retro
Conduct systematic, evidence-grounded retrospectives on past agent coding sessions to improve repository navigation pointers, automated lint/type checks, reviewer standards, and instruction ergonomics.
---
## Core Invariants
1. **Context Pressure Separation**: Differentiate implementation agents (high context pressure; need minimal steering and navigation pointers) from review agents (low context pressure; enforce deep coding standards).
2. **Push Rules Down the Pyramid**: Whenever possible, convert instructions into automated compiler/linter checks $\rightarrow$ reviewer rules $\rightarrow$ documentation $\rightarrow$ only as a last resort `AGENTS.md`.
3. **No-Op Instruction Pruning**: Actively audit and eliminate non-operational or redundant instructions from `AGENTS.md` and `CLAUDE.md`.
4. **Tool Economy Analysis**: Identify and streamline token-inefficient MCP tool calls or large file payload reads.
5. **Severity-Ranked Recommendations**: Present actionable improvement candidates ordered strictly by impact on agent reliability and token efficiency.
---
## Architecture & Map of Content (MOC)
```
[ Session Logs & Past Mistakes ]
│
▼
┌───────────────────────────────────────┐
│ 1. 6-Category Retrospective Audit │
│ (Nav, Checks, Standards, Steering, │
│ Tooling, Info Access) │
└──────────────────┬────────────────────┘
│
▼
┌───────────────────────────────────────┐
│ 2. Rule Hierarchy Placement │ ──► Auto-Check > Reviewer Rule > Doc Pointer > AGENTS.md
└──────────────────┬────────────────────┘
│
▼
┌───────────────────────────────────────┐
│ 3. Severity-Ranked Action Items │ ──► Concrete diffs to configs, linters, or standards
└───────────────────────────────────────┘
```
| Audit Category | Evaluation Focus | Remediation Action |
|---|---|---|
| **Navigation** | Time spent searching files | Add concise navigation pointers in `docs/` |
| **Automated Checks** | Preventable syntax/type/path bugs | Add linter rules, Husky hooks, TypeScript strictness |
| **Coding Standards** | Missed architectural guidelines | Add rules to `CODING_STANDARDS.md` (read during review) |
| **Steering Hygiene** | Unwieldy `AGENTS.md` / `CLAUDE.md` | Prune no-ops, trim prose, push rules to sub-docs |
| **Tool Economy** | Expensive/redundant tool calls | Cache results, scope searches, optimize grep patterns |
| **Info Access** | Missing logs or credentials | Provision read-only logs or environment variables |
---
## Step-by-Step Procedure (TWI)
### Step 1: Ingest Session Logs & Primary Sources
- **Action**: Read the transcript logs of the specified session (or current session) and identify friction points, wrong turns, and repeated failures.
- **Key Point**: Ground all critiques in observable events rather than generic advice.
- **Why**: Retrospectives must solve real developer and agent friction observed in actual execution.
### Step 2: Audit Against the 6 Improvement Categories
- **Action**: Evaluate where the failure should ideally have been caught (e.g. automated check vs. reviewer vs. prompt pointer).
- **Inline Checklist**:
- [ ] Can this mistake be caught by a linter or compiler flag?
- [ ] Does this belong in `CODING_STANDARDS.md` for the review agent?
- [ ] Are `AGENTS.md` files lean ($<100$ lines) and free of no-ops?
### Step 3: Propose Severity-Ordered Action Items
- **Action**: Format recommendations with concrete file diffs and command-line instructions.
- **Key Point**: Clearly explain the trade-offs of each proposed rule change.
- **Why**: Concrete proposals allow maintainers to accept improvements with a single confirmation.
---
## Anti-Rationalization Guardrails
| Tempting Rationalization | Binding Rule | Engineering Rationale |
|---|---|---|
| *"Add a 50-line instruction to AGENTS.md for every bug encountered."* | **Push rules down to automated checks or review docs.** | Overloading `AGENTS.md` bloats context window on every turn and degrades reasoning. |
| *"Impose heavy coding standard rules on the implementer prompt."* | **Place coding standards in `CODING_STANDARDS.md` for review.** | Implementers need context space for reasoning, debugging, and file exploration. |
| *"Keep no-op instructions because they sound good."* | **Prune all instructions that do not demonstrably steer model behavior.** | Dead instructions consume tokens and dilute attention on critical invariants. |
Referenced files: 1
scaffold-exercises3.69 KB
--- name: scaffold-exercises description: "Scaffold course, workshop, or tutorial exercises following repository conventions. Use when creating exercise folders, problem/solution/explainer variants, numbered lesson files, or workshop boilerplate — even if the user says \"add a new exercise\". Do NOT use for routine application feature scaffolding." --- # Scaffold Exercises Scaffold standardized educational exercise directories, problem/solution/explainer variant subfolders, and boilerplate TypeScript modules that strictly pass repository linter rules. --- ## Core Invariants 1. **Strict Dash-Case Numeric Hierarchy**: - Section folders: `exercises/XX-section-name/` (2-digit zero-padded number). - Exercise folders: `exercises/XX-section-name/XX.YY-exercise-name/` (section.exercise format). 2. **Mandatory Variant Subfolders**: Every exercise must contain at least one of `problem/`, `solution/`, or `explainer/` (defaulting to `explainer/` for stubs). 3. **Non-Empty Readme Standards**: Every variant folder must contain a non-empty `readme.md` with a clean `# Title`, description, and zero broken links. 4. **Git-Aware Moves**: Always use `git mv` instead of raw filesystem renames when renumbering or restructuring existing exercises to preserve commit history. 5. **Mandatory CLI Lint Gate**: Execute and verify `pnpm ai-hero-cli internal lint` before concluding the scaffolding step. --- ## Architecture & Map of Content (MOC) ``` [ Curriculum Plan / Outline ] ──► [ Generate Numbered Directories ] ──► [ Scaffold Variant Subfolders ] ──► [ Lint & Commit Gate ] ``` | Variant Folder | Student Purpose | Required Files | |---|---|---| | `problem/` | Active student workspace with `// TODO:` markers | `readme.md`, `main.ts` | | `solution/` | Complete reference implementation | `readme.md`, `main.ts` | | `explainer/` | Conceptual deep-dive without code tasks | `readme.md` | --- ## Step-by-Step Procedure (TWI) ### Step 1: Parse Curriculum Plan & Compute Hierarchy - **Action**: Extract section names, exercise titles, and required variant types from the provided outline. - **Key Point**: Formulate proper 2-digit numeric prefixes (`01`, `02`, `01.01`, `01.02`). - **Why**: Consistent numbering ensures exercises display in correct chronological order in the CLI and UI. ### Step 2: Scaffold Directory Tree and Variant Files - **Action**: Create folders (`mkdir -p`) and populate `readme.md` stubs with titles and descriptions. - **Key Point**: If code execution is involved, generate a non-empty `main.ts`. - **Inline Checklist**: - [ ] Dash-case directory naming strictly enforced - [ ] Non-empty `readme.md` created in each variant subfolder - [ ] No forbidden `.gitkeep` or `speaker-notes.md` files introduced ### Step 3: Run Internal Linter & Fix Issues - **Action**: Execute `pnpm ai-hero-cli internal lint` to validate the exercise structure. - **Key Point**: Iterate on any reported broken links or missing files until the linter passes completely. - **Why**: Pre-commit linting prevents breaking curriculum builds and automated runner harnesses. --- ## Anti-Rationalization Guardrails | Tempting Rationalization | Binding Rule | Engineering Rationale | |---|---|---| | *"Use raw `mv` when renumbering exercises."* | **Mandatory `git mv` for all renames and moves.** | Raw moves break Git blame and make diffs unreadable in PR reviews. | | *"Create empty `.gitkeep` files in exercise folders."* | **Forbidden. Populate meaningful `readme.md` stubs.** | Linter rules forbid `.gitkeep` files in exercise packages. | | *"Skip running `pnpm ai-hero-cli internal lint` on stubs."* | **Mandatory green linter verification.** | Missing title headers or invalid subfolder names break downstream test runners. |
Referenced files: 1
setup-engineering-workflows4.26 KB
--- name: setup-engineering-workflows description: "Configure repository issue-tracking, triage-label, domain-doc, and agent-instruction conventions. Use when initializing engineering workflows in a repo, configuring GitHub/GitLab issue labels, or establishing CONTEXT.md conventions — even if the user says \"setup our workflows\". Do NOT use for general project package installs." --- # Setup Engineering Workflows Configure and standardize repository issue-tracking, triage-label vocabularies, domain-modeling architecture (`docs/agents/domain.md`), and agent instructions (`AGENTS.md` / `CLAUDE.md`). --- ## Core Invariants 1. **Prompt-Driven Exploration**: Inspect existing remotes, directories, and labels before proposing changes; do not blindly overwrite existing tracker setups. 2. **Single vs Multi-Context Prudence**: Default to single-context (`CONTEXT.md` at root); only suggest multi-context (`CONTEXT-MAP.md`) when monorepo signals (workspaces, `packages/*`) are detected. 3. **Preserve Agent Instruction File**: If `CLAUDE.md` or `AGENTS.md` exists, edit it in-place; if neither exists, ask the user before creating one. Never create both. 4. **Conditional Triage Configuration**: Only configure triage label files (`docs/agents/triage-labels.md`) if the `triage` skill is present in the repository. 5. **Durable Workspace Documentation**: Persist all configured decisions into `docs/agents/issue-tracker.md`, `docs/agents/domain.md`, and `docs/agents/triage-labels.md`. --- ## Architecture & Map of Content (MOC) ``` [ Codebase Inspection (Remotes, Monorepos) ] ──► [ Interactive Setup Sections A/B/C ] ──► [ Update AGENTS.md / docs/agents/ ] ``` | Component | Responsibility | Seed Template | |---|---|---| | **Issue Tracker Config** | GitHub (`gh`), GitLab (`glab`), or Local Markdown | `skills/setup-engineering-workflows/issue-tracker-github.md` | | **Triage Vocabulary** | 5 canonical roles mapping to repo labels | `skills/setup-engineering-workflows/triage-labels.md` | | **Domain Docs Layout** | Single-context vs multi-context conventions | `skills/setup-engineering-workflows/domain.md` | --- ## Step-by-Step Procedure (TWI) ### Step 1: Autonomous Repository Exploration - **Action**: Check `git remote -v`, root instruction files (`AGENTS.md`, `CLAUDE.md`), `docs/adr/`, and workspace configurations (`pnpm-workspace.yaml`, `packages/`). - **Key Point**: Identify whether the repository already has a tracker or triage convention. - **Why**: Avoids re-asking the user for facts already evident in git and configuration files. ### Step 2: Configure Issue Tracker & Triage Labels - **Action**: Present recommendations section-by-section: - **Section A (Tracker)**: GitHub (`gh`), GitLab (`glab`), or Local Markdown (`.scratch/`). - **Section B (Labels)**: Canonical roles (`needs-triage`, `needs-info`, `ready-for-agent`, `ready-for-human`, `wontfix`). - **Section C (Domain)**: Single-context vs. Multi-context. - **Inline Checklist**: - [ ] Tracker confirmed and documented in `docs/agents/issue-tracker.md` - [ ] Triage labels mapped in `docs/agents/triage-labels.md` - [ ] Domain doc rules written to `docs/agents/domain.md` ### Step 3: Wire Instructions into AGENTS.md / CLAUDE.md - **Action**: Add or update the `## Agent skills` block inside `AGENTS.md` (or `CLAUDE.md`): ```markdown ## Agent skills ### Issue tracker [Summary of issue tracking]. See `docs/agents/issue-tracker.md`. ### Domain docs [Single or multi-context]. See `docs/agents/domain.md`. ``` - **Why**: Context pointers ensure newly spawned subagents read repository conventions immediately. --- ## Anti-Rationalization Guardrails | Tempting Rationalization | Binding Rule | Engineering Rationale | |---|---|---| | *"Create both AGENTS.md and CLAUDE.md for maximum coverage."* | **Forbidden. Maintain exactly one canonical agent file.** | Multiple instruction files lead to configuration drift and conflicting rules. | | *"Force multi-context on single-package repositories."* | **Default to single-context unless monorepo exists.** | Unnecessary multi-context structures create excessive directory nesting. | | *"Overwrite existing custom issue labels without asking."* | **Map existing repo labels to canonical roles.** | Overwriting existing team labels breaks active project boards and workflows. |
Referenced files: 6
setup-pre-commit4.55 KB
---
name: setup-pre-commit
description: "Configure pre-commit quality checks with Husky and lint-staged while preserving package manager conventions. Use when setting up git pre-commit hooks, automated formatters, staged linters, or typecheck gates — even if the user says \"add pre-commit hooks\". Do NOT use for continuous integration (CI) pipeline setup."
---
# Setup Pre-Commit Hooks
Install and configure deterministic Git pre-commit quality gates using Husky, `lint-staged`, and Prettier, preserving existing repository package managers and build scripts.
---
## Core Invariants
1. **Package Manager Fidelity**: Detect and use the repository's native package manager (`bun`, `pnpm`, `yarn`, `npm`); do NOT introduce conflicting lockfiles.
2. **Staged-Only Formatting**: Run Prettier strictly against staged files via `lint-staged` with `--ignore-unknown --write` to prevent un-staged code drift.
3. **Repository-Wide Typecheck & Test**: Execute full typecheck and unit test passes inside `.husky/pre-commit` after `lint-staged` succeeds.
4. **Preserve Existing Configurations**: If `.prettierrc` or existing Husky hooks exist, merge scripts safely without overwriting user rules.
5. **Mandatory Smoke-Test Commit**: Verify the entire hook pipeline by staging all configuration changes and executing a verified test commit.
---
## Architecture & Map of Content (MOC)
```
[ Git Commit Attempt ] ──► [ .husky/pre-commit Hook ]
│
┌───────────────────────────┼───────────────────────────┐
▼ ▼ ▼
[ lint-staged ] [ Typecheck Gate ] [ Unit Test Gate ]
- Prettier format - `npm run typecheck` - `npm run test`
- Staged files only - Full project types - Unit/Integration pass
```
| Component | Responsibility | Configuration File |
|---|---|---|
| **Husky Engine** | Git hook manager v9+ | `.husky/pre-commit` |
| **lint-staged** | Run formatters exclusively on staged changes | `.lintstagedrc` |
| **Prettier** | Code style standardization | `.prettierrc` |
---
## Step-by-Step Procedure (TWI)
### Step 1: Detect Package Manager & Check Existing Setup
- **Action**: Detect the lockfile (`bun.lockb` $\rightarrow$ bun, `pnpm-lock.yaml` $\rightarrow$ pnpm, `yarn.lock` $\rightarrow$ yarn, else npm).
- **Key Point**: Check if `prettier`, `lint-staged`, or `husky` are already installed.
- **Why**: Reusing existing dependencies prevents package bloat and lockfile churn.
### Step 2: Install DevDependencies & Initialize Husky
- **Action**: Install `husky`, `lint-staged`, and `prettier` as devDependencies, then run `npx husky init`.
- **Key Point**: Verify that `"prepare": "husky"` is added to `package.json`.
- **Inline Checklist**:
- [ ] Dependencies installed in `devDependencies`
- [ ] `.husky/` directory created
- [ ] `prepare` script present in `package.json`
### Step 3: Write `.husky/pre-commit` & `.lintstagedrc`
- **Action**: Configure `.husky/pre-commit`:
```bash
npx lint-staged
npm run typecheck
npm run test
```
*(Adapting `npm` to the detected package manager, and omitting `typecheck`/`test` if scripts are absent).*
- **Key Point**: Create `.lintstagedrc` with `{"*": "prettier --ignore-unknown --write"}`.
- **Why**: `prettier --ignore-unknown` safely ignores binaries and images while formatting all supported text formats.
### Step 4: Smoke Test and Initial Commit
- **Action**: Stage all created hook files and run `git commit -m "Add pre-commit hooks (husky + lint-staged + prettier)"`.
- **Key Point**: Confirm that the commit triggers the pre-commit hook and passes all stages.
- **Why**: Live commit execution proves that hooks are executable and correctly configured.
---
## Anti-Rationalization Guardrails
| Tempting Rationalization | Binding Rule | Engineering Rationale |
|---|---|---|
| *"Run prettier across the whole repository on every commit."* | **Run Prettier strictly on staged files via `lint-staged`.** | Full-repo formatting on every commit slows developer workflows and pollutes diffs. |
| *"Skip the typecheck and test steps in pre-commit to make it faster."* | **Include typecheck and fast tests in pre-commit.** | Catching type errors locally before push prevents broken remote CI builds. |
| *"Use npm commands in a pnpm or bun repository."* | **Strictly match detected package manager commands.** | Mismatched package managers create duplicate lockfiles and broken installations. |
Referenced files: 1
setup-ts-deep-modules5.36 KB
---
name: setup-ts-deep-modules
description: "Enforce TypeScript package boundaries, entry points, and cyclic dependency rules with dependency-cruiser. Use when structuring monorepo packages, establishing public API entry points, or preventing circular imports — even if the user says \"fix TS package boundaries\". Do NOT use for non-TypeScript repositories."
---
# Setup TS Deep Modules
Configure and verify strict deep-module boundaries across TypeScript packages using `dependency-cruiser`, ensuring public interfaces exist exclusively at package root files while private implementations remain sealed inside subdirectories.
---
## Core Invariants
1. **Root Files as Exclusive Entry Points**: External callers and consumers may only import package root files (`index.ts`, `client.ts`, `server.ts`); all subdirectories (`lib/`, `internal/`, `tests/`) are private.
2. **Four Strict Dependency Rules**: Enforce (1) Entry-point boundary, (2) Intra-package internal freedom, (3) Tests import entry-points only, (4) Zero dependency cycles.
3. **Barrels Strongly Discouraged**: Expose focused, multiple root entry points instead of giant barrel files that re-export whole subtrees.
4. **Mandatory Biting Proof**: Before completing setup, deliberately introduce a forbidden deep import to verify that `lint:boundaries` fails fast.
5. **No Speculative Path Aliases**: Wire boundary checks directly through dependency-cruiser rather than altering `tsconfig.json` paths or building complex build-time layers.
---
## Architecture & Map of Content (MOC)
```
src/packages/
<package-name>/
index.ts ◄── Public Entry Point (exported to other packages)
client.ts ◄── Additional Public Entry Point
lib/ ◄── Private Implementation (sealed from outsiders)
tests/ ◄── Tests (import only root entry points)
```
| Rule | Enforcement | Target |
|---|---|---|
| **Entry Boundary** | `not-to-subfolder` | Outside code $\rightarrow$ root files only |
| **Test Boundary** | `tests-through-entrypoints` | `tests/` $\rightarrow$ root entry points only |
| **Cycle Prevention** | `no-circular` | Disallow all circular dependencies |
| **Verification Gate** | `lint:boundaries` script | Automated CI failure on breach |
---
## Step-by-Step Procedure (TWI)
### Step 1: Detect Environment & Package Manager
- **Action**: Detect the package manager (`bun.lockb` $\rightarrow$ bun, `pnpm-lock.yaml` $\rightarrow$ pnpm, `yarn.lock` $\rightarrow$ yarn, else npm) and locate the packages root (`src/packages` or `packages`).
- **Key Point**: Check for existing `.dependency-cruiser.*` configuration to merge rules rather than overwriting.
- **Why**: Preserving existing project package manager standards ensures zero friction with current CI scripts.
### Step 2: Install and Configure Dependency Cruiser
- **Action**: Install `dependency-cruiser` as a devDependency and copy `dependency-cruiser.config.cjs` to `.dependency-cruiser.cjs` with updated `PACKAGES_ROOT`.
- **Key Point**: Use `.cjs` extension so CommonJS exports function seamlessly in ESM / `"type": "module"` projects.
- **Inline Checklist**:
- [ ] Package manager identified correctly
- [ ] `dependency-cruiser` added to `devDependencies`
- [ ] `.dependency-cruiser.cjs` configured with 4 core rules
### Step 3: Wire Scripts & Scaffold Clean Example
- **Action**: Add `"lint:boundaries": "depcruise <packages-root>"` to `package.json` and create an example deep package:
- `<packages-root>/example/index.ts` (public entry point delegating to `lib/`)
- `<packages-root>/example/lib/impl.ts` (hidden implementation)
- `<packages-root>/example/tests/example.test.ts` (imports `../index`)
- **Why**: A reference package serves as living documentation and a starter template for developers.
### Step 4: Prove the Rules Bite (Verification Gate)
- **Action**: Execute three validation passes:
1. Run `lint:boundaries` $\rightarrow$ must **PASS** on clean repo.
2. Add deep import `import { impl } from "../lib/impl"` in test $\rightarrow$ must **FAIL**.
3. Revert deep import and run `lint:boundaries` $\rightarrow$ must **PASS**.
- **Key Point**: Never sign off on boundary rules without observing an intentional test failure.
- **Why**: Unproven linter configs frequently have syntax or glob errors that silently pass all code.
### Step 5: Document Package Conventions & Agent Pointer
- **Action**: Create `<packages-root>/README.md` explaining layout and rules, and add a single-line context pointer in `AGENTS.md` / `CLAUDE.md`.
- **Key Point**: Explicitly explain why barrel files are discouraged.
- **Why**: Agent pointers ensure future autonomous sessions respect package boundaries from their first prompt.
---
## Anti-Rationalization Guardrails
| Tempting Rationalization | Binding Rule | Engineering Rationale |
|---|---|---|
| *"Skip testing whether the rule fails on bad imports."* | **Mandatory 3-step proof (Pass $\rightarrow$ Fail $\rightarrow$ Pass).** | Linters with invalid glob configurations pass silently without enforcing anything. |
| *"Export all submodules through a giant root index.ts barrel."* | **Discourage large barrel files.** | Giant barrels destroy tree-shaking and create hidden cyclic import dependencies. |
| *"Import directly from lib/ in unit tests for convenience."* | **Enforce tests through public entry points only.** | Deep imports in tests tightly couple test suites to volatile internal refactors. |
Referenced files: 2
skill-conductor4.87 KB
--- name: skill-conductor description: "Author, refine, evaluate, and package agent skills across their full lifecycle. Use when building a new skill from scratch, improving an existing skill, fixing a skill that triggers unreliably, running skill evals, benchmarking skill performance, or packaging skills for distribution — even if they don't explicitly say \"skill conductor\". Do NOT use for general software coding tasks or using pre-existing skills." --- # Skill Conductor Full lifecycle management for agent skills: **draft → test → review → improve → package**. One master discipline to govern skill creation and optimization, rooted in the 10 canonical authoring principles. --- ## Core Invariants & The 10 Canonical Authoring Principles 1. **Pre-flight verification**: Check dependencies and environment before executing workflows. 2. **No process in descriptions**: Descriptions define triggering boundaries (`Use when...`, `Do NOT use for...`), never workflow recipes. 3. **Map of Content (MOC)**: `SKILL.md` is a clear architectural map (< 500 lines) pointing to modular references. 4. **Fresh practitioner empathy**: Explain the rationale behind steps so agents understand context. 5. **Training Within Industry (TWI)**: Structure critical instructions as `Action`, `Key Point`, and `Why`. 6. **Blind agent testability**: Ensure instructions are unambiguous when executed without prior conversation history. 7. **Inline risk checklists**: Place verification checklists directly at high-risk seams. 8. **One term per concept**: Use consistent domain terminology throughout. 9. **Zero secrets / environment cleanliness**: Never hardcode credentials, tokens, or absolute user home paths. 10. **Match form to failure**: Counter specific failure modes with targeted structural constraints and anti-rationalization tables. --- ## Core Lifecycle Modes | Mode | Trigger / Context | Key Output | |---|---|---| | **1. CREATE** | "build a skill", "new skill for..." | Full lifecycle: intent → architecture → scaffold → write → eval | | **2. IMPROVE** | "fix this skill", "it doesn't trigger" | Diagnose → eval loop → gated self-update → iterate | | **3. VALIDATE** | "test this skill", "run evals" | Structural checks + trigger testing + BinEval scoring | | **4. REVIEW** | "review this skill", quality audit | 11-point quality gate assessment | | **5. OPTIMIZE** | "improve triggering", "optimize description" | Automated description optimization with train/test splits | | **6. PACKAGE** | "package for distribution" | Validation + bundle into release artifact | --- ## Step-by-Step Procedure (TWI) ### Step 1: Capture Intent & Define Triggers - **Action**: Extract 2–3 concrete user scenarios and establish positive and negative trigger boundaries. - **Key Point**: Specify exact phrases users say and adjacent domains the skill must reject. - **Why**: Clear boundaries prevent undertriggering and false-positive overtriggering. ### Step 2: Architecture & Progressive Disclosure - **Action**: Select the architectural pattern (sequential, iterative, context-aware) and structure files. - **Key Point**: Keep `SKILL.md` concise (< 500 lines) and push heavy schemas or tables into references. - **Why**: Bloated instruction files exhaust attention budgets and degrade execution quality. ### Step 3: Write Frontmatter & Body - **Action**: Draft YAML frontmatter with kebab-case name and trigger-rich description, followed by MOC, TWI steps, and guardrails. - **Inline Checklist**: - [ ] Frontmatter name matches directory name - [ ] Description is $\le 1024$ characters and contains no process steps - [ ] Negative triggers specified (`Do NOT use for...`) - [ ] Anti-rationalization table included - [ ] No hardcoded tokens, passwords, or machine-specific paths ### Step 4: Validate with BinEval 5 Dimensions - **Action**: Evaluate across **Discovery, Clarity, Structure, Robustness, Completeness**. - **Key Point**: Pass every critical gate check before declaring release readiness. - **Why**: Multi-dimensional evaluation catches hidden failure modes before distribution. --- ## Anti-Rationalization Guardrails | Tempting Rationalization | Binding Rule | Engineering Rationale | |---|---|---| | *"Putting steps in the description helps the model."* | **Forbidden: zero process in description.** | Models follow description steps and skip the comprehensive body instructions. | | *"The skill is small, so negative triggers are unneeded."* | **Mandatory negative triggers in description.** | Without negative boundaries, skills trigger falsely on loosely related queries. | | *"One big 1200-line markdown file is easier to manage."* | **Progressive disclosure: SKILL.md < 500 lines.** | Overloaded contexts dilute attention and lead to instruction-skipping. | | *"We can skip testing if the markdown looks clean."* | **Trigger evaluation on positive and negative test cases.** | Clean prose can still fail in actual agent routing. |
Referenced files: 1
tdd4.61 KB
--- name: tdd description: "Implement features and bug fixes test-first using red-green-refactor cycles. Use when writing new functionality, adding regression tests, fixing bugs with test coverage, or designing behavior through test assertions — even if the user says \"write tests for this\". Do NOT use for throwaway exploratory prototypes." --- # Test-Driven Development (TDD) Implement robust, maintainable functionality through disciplined red → green → verify cycles. Verify behavior through public seams rather than private implementation details. ## Core Principle > **Write the failing test first, witness it fail for the right reason, and write only the minimal code needed to make it green.** --- ## Core Invariants 1. **Strict Red-First Execution**: Never write production implementation code before seeing an automated test fail against a pre-agreed seam. 2. **Behavioral Testing Over Implementation Inspection**: Tests must assert observable input/output behavior and public contracts, never private variables or internal call chains. 3. **Independent Expected Values**: Assertions must compare against independent ground truth (fixtures, domain specs, known literals), never recomputed mirror logic. 4. **Vertical Tracer Slicing**: Deliver one vertical slice at a time (one test → minimal code → green) rather than batching horizontal test suites upfront. 5. **Deterministic & Isolated Fixtures**: Tests must execute in milliseconds without cross-test state leakage, unseeded random seeds, or unpinned clocks. --- ## Architecture & Map of Content (MOC) ``` [ Agree Public Seam ] ──► [ Red: Write Failing Test ] ──► [ Green: Minimal Code ] ──► [ Refactor & Verify ] ``` | Component | Responsibility | Reference | |---|---|---| | **Public Seams** | Identify module boundaries and observable behaviors | `codebase-design` | | **Test Examples** | Concrete patterns for unit and integration assertions | `tests.md` | | **Mocking Guidelines** | Safe boundary isolation rules without over-mocking | `mocking.md` | | **Review & Refactor** | Clean code review after reaching green | `code-review` | --- ## Step-by-Step Procedure (TWI) ### Step 1: Identify & Confirm the Public Seam - **Action**: State the public interface and observable contract you intend to test, and confirm alignment with domain conventions. - **Key Point**: Anchor assertions at public boundaries; do not reach into internal private state. - **Why**: Testing through private internals creates brittle tests that break during harmless refactoring. ### Step 2: Write the Failing Test (RED) - **Action**: Author a single focused test specifying the expected behavior and run the test runner to observe failure. - **Key Point**: Verify the test fails specifically because the new behavior is absent, not due to syntax or setup errors. - **Inline Checklist**: - [ ] Test names the exact business capability (e.g. `returns_discounted_total_for_premium_member`) - [ ] Test fails with expected assertion error - [ ] Expected values derived from independent domain constants - **Why**: A test that doesn't fail properly cannot be trusted to protect against regressions. ### Step 3: Implement Minimal Code (GREEN) - **Action**: Write the simplest, most direct code that makes the failing test pass. - **Key Point**: Do not anticipate speculative requirements or write unexercised branches. - **Why**: Minimal code keeps the change set tight and prevents unverified dead logic. ### Step 4: Verify Suite & Cycle - **Action**: Run the full test suite to guarantee zero regressions. - **Key Point**: All tests must be green before proceeding to the next vertical tracer slice. - **Why**: Catching regressions immediately keeps debugging costs near zero. --- ## Anti-Rationalization Guardrails | Tempting Rationalization | Binding Rule | Engineering Rationale | |---|---|---| | *"I'll write the implementation first, then add tests."* | **Forbidden: test must be written and observed failing first.** | Tests written after code tend to mirror implementation bias and miss edge-case failures. | | *"I can mock this internal helper to test the private method."* | **Test only through public seams.** | Mocks on private internals cement architectural rigidity and hide real integration bugs. | | *"I will write all 15 test cases before writing any code."* | **Vertical tracer slicing: one test at a time.** | Bulk tests lock in premature interface assumptions before real implementation learnings. | | *"The test passed on the first run without changes."* | **Investigate immediately: tautological or wrong test.** | A test that passes without implementation changes is asserting something already true. |
Referenced files: 3
teach4.52 KB
--- name: teach description: "Teach a technical concept interactively using the workspace as an active learning lab with durable learning records. Use when the user asks to understand a codebase concept, learn a language feature, explore design patterns, or practice coding skills — even if they say \"explain how this works to me\". Do NOT use for silent autonomous code generation." --- # Teach Transform the current workspace into an interactive, stateful learning laboratory producing structured HTML lessons, reference materials, primary source bibliographies, and progressive learning records. --- ## Core Invariants 1. **Mission-Grounded Curriculum**: Every lesson, quiz, and exercise must directly anchor to `MISSION.md` (the learner's explicit real-world objective). 2. **Desirable Difficulty & Retrieval**: Target long-term storage strength through effortful retrieval, spaced practice, and interleaved exercises over fleeting in-the-moment fluency. 3. **Beautiful HTML Artifacts**: Lessons are rendered as self-contained, publication-grade HTML documents in `./lessons/000X-<name>.html` styled via shared assets (`./assets/style.css`). 4. **Durable Learning Records**: Capture key insights, non-obvious breakthroughs, and mental model shifts in `./learning-records/000X-<name>.md` to maintain the Zone of Proximal Development (ZPD). 5. **Primary-Source Grounding**: Never rely purely on model memory; cite and link verified primary sources in `RESOURCES.md` and inline lesson annotations. --- ## Architecture & Map of Content (MOC) ``` [ Learner Request / Mission ] ──► [ Assess ZPD via Learning Records ] ──► [ Build HTML Lesson & Quizzes ] ──► [ Record Progress in Learning Records ] ``` | Component | Responsibility | File Location / Template | |---|---|---| | **Mission Anchor** | Core motivation and real-world goal | `MISSION.md` ([MISSION-FORMAT.md](MISSION-FORMAT.md)) | | **Lesson Artifacts** | Interactive HTML lesson modules | `./lessons/000X-<slug>.html` | | **Reference Sheets** | Cheat sheets, syntax summaries, and glossaries | `./reference/*.html` | | **Learning Records** | Milestone reflections and conceptual breakthroughs | `./learning-records/000X-<slug>.md` | | **Shared Assets** | Typography, styles, and interactive quiz widgets | `./assets/*` | | **Primary Sources** | Vetted documentation and book references | `RESOURCES.md` ([RESOURCES-FORMAT.md](RESOURCES-FORMAT.md)) | --- ## Step-by-Step Procedure (TWI) ### Step 1: Establish Mission and Assess Zone of Proximal Development - **Action**: Check `MISSION.md` and existing `./learning-records/`. If uninitialized, grill the user to define their motivation and starting baseline. - **Key Point**: Frame teaching around what the user needs to *do* rather than abstract theory. - **Why**: Learning is significantly faster and more durable when grounded in immediate application. ### Step 2: Source Primary References - **Action**: Identify and record the top primary source documentation or authoritative articles in `RESOURCES.md`. - **Key Point**: Ensure claims and code patterns conform to modern best practices. - **Why**: Citing primary sources builds user trust and teaches research habits. ### Step 3: Author Interactive HTML Lesson - **Action**: Generate `./lessons/000X-<slug>.html` linking to `./assets/style.css`. - **Key Point**: Include interactive self-check questions, code snippets, and direct links to reference sheets. - **Inline Checklist**: - [ ] Standalone HTML opens cleanly in browser - [ ] Quiz answers formatted with balanced word/character counts - [ ] Follow-up discussion prompts included ### Step 4: Record Breakthroughs in Learning Records - **Action**: After the user interacts with the lesson, capture the key concepts mastered and areas for next practice in `./learning-records/000X-<slug>.md`. - **Why**: Keeps learning state synchronized across different sessions. --- ## Anti-Rationalization Guardrails | Tempting Rationalization | Binding Rule | Engineering Rationale | |---|---|---| | *"Dump a 2000-word explanation directly into chat."* | **Package learning into modular HTML lessons and reference files.** | Chat text scrolls away; structured files provide a lasting personal library. | | *"Give easy multiple-choice questions with obvious answers."* | **Enforce balanced, high-friction retrieval practice.** | Obvious answers create false fluency without building long-term memory. | | *"Teach without checking `MISSION.md`."* | **Mandatory mission alignment for every lesson.** | Disconnected theory causes learner disengagement and rapid forgetting. |
Referenced files: 5
to-questionnaire4.02 KB
--- name: to-questionnaire description: "Transform unknown requirements and stakeholder dependencies into a structured questionnaire. Use when a plan or spec is blocked by external stakeholder decisions, business rules, or missing domain facts — even if the user says \"what should I ask them?\". Do NOT use when the user can answer the questions themselves." --- # To Questionnaire Transform unresolved stakeholder dependencies, business rule unknowns, and external domain ambiguities into a high-yield, structured questionnaire that minimizes async friction. --- ## Core Invariants 1. **Grill the Send, Not the Subject**: Interview the user exclusively about the recipient's role, expertise, and required decisions—never interrogate the user on answers only the recipient holds. 2. **Gap-Targeted Formulations**: Every question must laser-target the precise boundary between what the user knows and what the stakeholder decides. 3. **Importance-First Ordering**: Sequence questions in descending order of critical architectural impact to extract maximum value from partial responses. 4. **Single-Idea Atomic Questions**: Never combine multiple compound questions into a single item; each question must possess its own context and answer stub. 5. **No Speculative Answering**: Do not guess or fabricate stakeholder business rules; format clear prompts with rationales (`Why this matters:`) to elicit clean answers. --- ## Architecture & Map of Content (MOC) ``` [ External Unknowns / Blockers ] ──► [ Interview User on Recipient & Needs ] ──► [ Draft Gap-Targeted Questionnaire ] ──► [ Output Markdown Doc ] ``` | Component | Responsibility | Output Target | |---|---|---| | **Recipient Profile** | Role, expertise, and decision scope | Document header | | **Context Summary** | 1-paragraph orientation for the recipient | `## Context` section | | **Thematic Questions** | Atomic questions with answer stubs & rationale | `to-questionnaire-<slug>.md` | --- ## Step-by-Step Procedure (TWI) ### Step 1: Clarify Recipient Persona and Required Decisions - **Action**: Ask the user in a single exchange: (1) Who is the recipient (role, technical depth)? (2) Exactly what decisions or facts must be unlocked? - **Key Point**: Keep the focus strictly on defining the gap, not guessing the answer. - **Why**: Calibrating the recipient's perspective ensures the document uses appropriate tone and technical depth. ### Step 2: Structure Thematic Question Blocks - **Action**: Group questions under clear theme headings (`## <Theme>`) and sort questions within each theme by descending priority. - **Key Point**: For every question, include an explicit `_Why this matters:_` line explaining how their answer influences the system design. - **Inline Checklist**: - [ ] Questions ordered most-important-first - [ ] No compound/nested sub-questions - [ ] Clear answer quote blocks (`> `) provided under each item - [ ] Explicit deadline and partial-answer instructions included ### Step 3: Write and Deliver Questionnaire Document - **Action**: Write the complete Markdown document to `to-questionnaire-<slug>.md` in the current directory. - **Key Point**: Conclude with a catch-all `## Anything Else?` section to capture unknown unknowns. - **Why**: Stakeholders frequently possess contextual constraints that standard questions overlook. --- ## Anti-Rationalization Guardrails | Tempting Rationalization | Binding Rule | Engineering Rationale | |---|---|---| | *"I'll ask the user to guess what the stakeholder would prefer."* | **Forbidden. Target the questionnaire at the gap.** | Guessing stakeholder preferences introduces false assumptions into production architectures. | | *"Combine 3 related questions into a single multi-part paragraph."* | **Enforce single-idea atomic questions.** | Multi-part questions overwhelm busy stakeholders, leading to skipped answers. | | *"Omit the 'Why this matters' clause to keep questions brief."* | **Include context on why the answer matters.** | Stakeholders give better, actionable answers when they understand the engineering impact. |
Referenced files: 1
to-spec4.58 KB
--- name: to-spec description: "Synthesize conversation and codebase context into an unambiguous, buildable technical specification. Use when requirements and design decisions are settled and need to be formalized into a technical spec — even if the user says \"write a spec for this\". Do NOT use when the core idea is still fuzzy and ungrilled." --- # To Spec Synthesize settled conversation context and codebase understanding into an unambiguous, buildable technical specification with minimal test seams and complete user stories. --- ## Core Invariants 1. **Pure Synthesis, Zero Interrogation**: Synthesize strictly from established context and codebase facts; do not interview the user. 2. **Minimal Seams at the Highest Tier**: Prefer existing high-level test seams; never introduce unnecessary internal mocking seams. 3. **Exhaustive User Stories**: Generate a comprehensive, numbered list of `As an <actor>, I want <feature>, so that <benefit>` stories covering all user paths. 4. **Decisions Over Concrete Snippets**: Specify architectural decisions, interfaces, and schema changes without fragile, hardcoded code snippets (unless derived from a validated prototype). 5. **Strict Out-of-Scope Demarcation**: Explicitly list out-of-scope capabilities to prevent scope creep and unbound work. --- ## Architecture & Map of Content (MOC) ``` [ Settled Context & ADRs ] ──► [ Codebase Seam Inspection ] ──► [ Spec Document Synthesis ] ──► [ Tracker Publication ] ``` | Component | Responsibility | Format / Template | |---|---|---| | **Problem & Solution** | Frame user-centric intent | Problem / Solution statements | | **User Stories** | Enumerate all functional paths | Numbered standard user stories | | **Implementation Decisions** | Define module boundaries & contracts | Architecture & schema decisions | | **Testing Decisions** | Specify external verification seams | Behavior-driven test strategy | --- ## Step-by-Step Procedure (TWI) ### Step 1: Autonomous Codebase & Seam Inspection - **Action**: Explore the repository to inspect existing modules, domain glossary terms, and ADRs. - **Key Point**: Identify the highest available integration seam to test the feature externally. - **Why**: Testing through high-level seams verifies true system behavior while leaving internal implementation details free to refactor. ### Step 2: Formulate Comprehensive User Stories - **Action**: Draft an exhaustive, numbered list of user stories capturing all primary and edge-case user interactions. - **Key Point**: Follow the strict template: `1. As an <actor>, I want a <feature>, so that <benefit>`. - **Why**: Detailed user stories prevent implementers from making ad-hoc product assumptions during coding. - **Inline Checklist**: - [ ] Every user story has an explicit actor and tangible benefit - [ ] Edge cases and failure states are covered as distinct stories - [ ] No implementation jargon inside user story statements ### Step 3: Formalize Implementation and Testing Decisions - **Action**: Document module boundaries, modified interfaces, database schema changes, and API contracts. - **Key Point**: Omit volatile file line numbers or speculative code snippets. - **Why**: Fragile code snippets go stale immediately and misdirect downstream implementation agents. ### Step 4: Define Out-of-Scope Boundaries & Publish - **Action**: Detail what is explicitly NOT included, apply the `ready-for-agent` triage label, and publish to the configured issue tracker or `.scratch/<feature-slug>/spec.md`. - **Key Point**: If the issue tracker is unconfigured, instruct the user to run `/setup-engineering-workflows`. - **Why**: Clear negative boundaries prevent scope bloat and keep subsequent ticket decomposition bounded. --- ## Anti-Rationalization Guardrails | Tempting Rationalization | Binding Rule | Engineering Rationale | |---|---|---| | *"I'll ask the user a few more questions before writing the spec."* | **No interviewing during to-spec.** | If requirements are still fuzzy, route back to `grill-me`. `to-spec` is pure synthesis. | | *"I will write extensive mock-heavy unit test decisions."* | **Test at the highest possible public seam.** | Mock-heavy tests break during refactors and fail to verify actual end-to-end functionality. | | *"Include complete implementation code blocks in the spec."* | **Specify interfaces and decisions, not full code.** | Implementation code in specs blinds the developer agent to live codebase nuances. | | *"Skip the out-of-scope section since it seems obvious."* | **Mandatory explicit Out-of-Scope section.** | Ambiguity in scope boundaries causes runaway scope creep during implementation. |
Referenced files: 1
to-tickets5.46 KB
---
name: to-tickets
description: "Decompose an approved specification or plan into ordered, dependency-linked tracer-bullet implementation tickets. Use when turning a spec or architectural plan into actionable tracker issues with clear acceptance criteria — even if the user says \"break this into tickets\". Do NOT use for initial requirements gathering."
---
# To Tickets
Decompose an approved specification, plan, or design document into vertically sliced, dependency-ordered tracer-bullet implementation tickets with explicit acceptance criteria.
---
## Core Invariants
1. **Strict Vertical Tracer Slicing**: Every standard ticket must cut vertically across all necessary layers (schema, domain logic, API, UI, test) rather than horizontal slices of a single layer.
2. **Explicit Dependency DAG**: Every ticket must explicitly declare its blocking prerequisites (`Blocked by:`) to form an unambiguous directed acyclic graph.
3. **Single Context-Window Sizing**: Each ticket must be sized to complete within a single agent context window without exhausting token or tool limits.
4. **Expand-Contract for Wide Refactors**: Wide architectural refactors with broad blast radius must follow the Expand-Contract pattern rather than being forced into fragile single-step vertical slices.
5. **No Parent Mutation**: Never close, resolve, or corrupt parent tracker issues when authoring and publishing child tickets.
---
## Architecture & Map of Content (MOC)
```
[ Approved Spec / Plan ] ──► [ DAG Dependency Decomposition ] ──► [ User Granularity Review ] ──► [ Atomic Tracker Publication ]
│
┌────────────────────────────┴────────────────────────────┐
▼ ▼
[ Vertical Tracer Bullets ] [ Expand-Contract Migrations ]
- Narrow cross-layer path - Expand (add new form beside old)
- Independently verifiable - Migrate in localized batches
- Fit in 1 context window - Contract (delete old form)
```
| Component | Responsibility | Output Target |
|---|---|---|
| **Local Ticket Store** | Atomic Markdown ticket files | `.scratch/<feature-slug>/issues/NN-<slug>.md` |
| **Tracker Issues** | Remote issue creation with blocking metadata | GitHub, Linear, GitLab issues |
| **Prefactoring Gate** | Make the change easy before making the easy change | Dedicated blocker ticket #01 |
---
## Step-by-Step Procedure (TWI)
### Step 1: Context Ingestion & Prefactoring Identification
- **Action**: Ingest the approved spec and inspect the target codebase area to identify necessary prefactoring.
- **Key Point**: If the existing code makes the target change difficult, create a dedicated preparatory refactor ticket as the first blocker.
- **Why**: "Make the change easy, then make the easy change" prevents intertwining architectural cleanup with functional feature additions.
### Step 2: Slice Vertical Tracer Bullets & Formulate DAG
- **Action**: Break the feature into vertical tracer slices and link dependencies:
- For standard features: Vertical slices cutting through schema, API, UI, and tests.
- For wide refactors: Sequence as Expand $\rightarrow$ Batch Migration $\rightarrow$ Contract.
- **Key Point**: Assign each ticket an unambiguous "Blocked by" list. Tickets with no blockers are ready for immediate execution.
- **Inline Checklist**:
- [ ] Every vertical slice delivers an independently verifiable behavior
- [ ] Wide refactors partitioned via expand-contract
- [ ] Tickets strictly fit in a single fresh context window
- [ ] Acceptance criteria are binary and testable
### Step 3: Present Breakdown & Quiz User for Approval
- **Action**: Present the proposed tickets with titles, blockers, and deliverables to the user.
- **Key Point**: Prompt the user to verify granularity, dependencies, and ordering.
- **Why**: Quick human validation ensures ticket sizing matches team velocity and avoids execution roadblocks.
### Step 4: Publish to Tracker
- **Action**: Publish approved tickets in topological order (blockers first) to the configured issue tracker or `.scratch/<feature-slug>/issues/NN-<slug>.md`.
- **Key Point**: Apply the `ready-for-agent` triage label to all unblocked tickets.
- **Why**: Topological creation ensures downstream tickets can cleanly reference live upstream ticket IDs.
---
## Anti-Rationalization Guardrails
| Tempting Rationalization | Binding Rule | Engineering Rationale |
|---|---|---|
| *"I'll create horizontal tickets: 1 for DB, 1 for backend, 1 for UI."* | **Forbidden. Enforce vertical tracer slices.** | Horizontal layers cannot be verified end-to-end and leave systems broken across commits. |
| *"This ticket is huge, but one agent can manage it."* | **Split tickets exceeding 1 context window.** | Oversized tickets cause context exhaustion, lost requirements, and hallucinations. |
| *"Combine the preparatory refactor with the new feature logic."* | **Separate prefactoring into a distinct ticket.** | Mixing refactoring with feature delivery obscures regressions in code review. |
| *"Skip writing acceptance criteria since the spec has them."* | **Mandatory atomic acceptance checklist per ticket.** | Implementer agents require self-contained verification criteria without re-reading the spec. |
Referenced files: 1
triage5.66 KB
---
name: triage
description: "Classify, verify, and prepare external bug reports and feature requests into agent-ready briefs. Use when processing incoming issues, validating bug reproductions, rejecting out-of-scope requests, or structuring contributor tickets — even if the user says \"triage these issues\". Do NOT use for internal tickets already produced by to-tickets."
---
# Triage
Classify incoming issues and external pull requests, reproduce bug reports against code, reject out-of-scope requests to `.out-of-scope/`, and synthesize durable, high-context agent briefs.
---
## Core Invariants
1. **Mandatory Disclaimer**: Every comment or issue created during triage must open with `> *This was generated by AI during triage.*`
2. **Category and State Orthogonality**: Every triaged issue must possess exactly one Category (`bug`, `enhancement`) and one State (`needs-triage`, `needs-info`, `ready-for-agent`, `ready-for-human`, `wontfix`).
3. **Reproduction Before Brief**: Verify bug claims against codebase tests and reproduction steps before promoting to `ready-for-agent`.
4. **Out-of-Scope Knowledge Base**: Rejected enhancement requests must be codified in `.out-of-scope/<slug>.md` with clear rationale to prevent recurring re-triage.
5. **Redundancy Screening**: Search codebase for existing implementations before triaging feature requests; if already implemented, close with an explanation.
---
## Architecture & Map of Content (MOC)
```
[ Incoming Issue / External PR ]
│
▼
┌───────────────────────────────────────┐
│ 1. Redundancy & Out-of-Scope Check │ ──► Search code & `.out-of-scope/*.md`
└──────────────────┬────────────────────┘
│
▼
┌───────────────────────────────────────┐
│ 2. Reproduction & Verification Pass │ ──► Verify bug or test PR diff locally
└──────────────────┬────────────────────┘
│
▼
┌───────────────────────────────────────┐
│ 3. Socratic Grilling & State Decision │ ──► Settle terms, update `CONTEXT.md`
└──────────────────┬────────────────────┘
│
▼
┌───────────────────────────────────────┐
│ 4. Outcome Publication │ ──► Post Agent Brief or Out-of-Scope Record
└───────────────────────────────────────┘
```
| Component | Responsibility | Reference Format |
|---|---|---|
| **Agent Brief** | Self-contained instructions for AFK agents | `skills/triage/AGENT-BRIEF.md` |
| **Out of Scope KB** | Persistent store of rejected proposals | `skills/triage/OUT-OF-SCOPE.md` |
| **Triage State Machine** | Canonical roles and state transitions | `docs/agents/triage-labels.md` |
---
## Step-by-Step Procedure (TWI)
### Step 1: Query Issues Needing Attention
- **Action**: Query tracker for unlabeled issues, `needs-triage` issues, and `needs-info` issues with fresh reporter activity.
- **Key Point**: Include external PRs when configured as a request surface.
- **Why**: Focuses maintainer attention on actionable items requiring triage intervention.
### Step 2: Screen Redundancy & Prior Rejections
- **Action**: Search codebase for existing capabilities and inspect `.out-of-scope/*.md` for prior rejections.
- **Key Point**: If the requested behavior already exists, close as `wontfix` pointing directly to the implementation.
- **Why**: Prevents duplicate feature implementations and wasted engineering cycles.
### Step 3: Reproduce Claim & Formulate Recommendation
- **Action**: Attempt reproduction of bug reports or checkout external PR diffs to run test suites.
- **Key Point**: If reproduction fails due to missing details, move to `needs-info` with specific questions.
- **Inline Checklist**:
- [ ] Bug reproduction attempted on local environment
- [ ] Relevant code paths and potential root causes identified
- [ ] Category and state recommendation presented to maintainer
### Step 4: Apply State & Deliver Agent Brief
- **Action**: Apply maintainer-approved state:
- `ready-for-agent`: Post structured Agent Brief with reproduction commands and suggested seams.
- `needs-info`: Post specific question block tagging the reporter.
- `wontfix` (enhancement): Document in `.out-of-scope/` and close issue.
- **Why**: High-fidelity agent briefs enable downstream autonomous agents to implement fixes without human intervention.
---
## Anti-Rationalization Guardrails
| Tempting Rationalization | Binding Rule | Engineering Rationale |
|---|---|---|
| *"Promote a bug to `ready-for-agent` without reproducing it."* | **Mandatory reproduction pass before agent handover.** | Unverified bug reports send downstream agents on wild goose chases. |
| *"Reject a feature in a short comment without documenting it."* | **Log rejected enhancements in `.out-of-scope/`.** | Undocumented rejections lead to repeatedly re-evaluating the same request. |
| *"Ask the reporter vague questions like 'more details please'."* | **Ask specific, actionable questions in `needs-info`.** | Vague questions delay resolution and frustrate external contributors. |
Referenced files: 3
wait-what3.53 KB
--- name: wait-what description: "Re-pitch an explanation or proposal from a simpler angle with reset assumptions and plain English. Use when the previous explanation was confusing, jargon-heavy, or did not land — even if the user just says \"what?\" or \"I don't get it\". Do NOT use for general brainstorming or initial planning." --- # Wait What Reset conversational assumptions, strip out convoluted jargon, and re-explain technical proposals using ASD-STE100 Simplified Technical English grounded in canonical domain terminology (`CONTEXT.md`). --- ## Core Invariants 1. **Simplified Technical English (STE)**: Use clear, unambiguous sentences; restrict vocabulary to concrete, observable concepts without cognitive overload. 2. **Domain Glossary Fidelity**: Strictly utilize the canonical nouns established in `CONTEXT.md` (or `CONTEXT-MAP.md`); avoid colloquial synonyms. 3. **Reset Ground Assumptions**: Step back from complex implementation details and re-anchor the explanation in the core user problem and high-level architecture. 4. **Concrete Contrast (Before vs. After)**: Use simple side-by-side examples or ASCII diagrams to demonstrate the change visually. 5. **Check for Understanding**: Conclude with a single direct question confirming whether the new explanation lands clearly. --- ## Architecture & Map of Content (MOC) ``` [ Confusing / Jargon-Heavy Explanation ] ──► [ Step Back to Problem Statement ] ──► [ ASD-STE100 Re-Pitch with CONTEXT.md ] ──► [ Single Alignment Check ] ``` | Phase | Responsibility | Guidelines | |---|---|---| | **Context Reset** | Re-state the immediate goal in 1 sentence | Ground in user value | | **STE Explanation** | Explain mechanism with short active verbs | 1 concept per sentence | | **Visual/Concrete Anchor** | Show 3-line input/output or ASCII diff | High clarity, low noise | --- ## Step-by-Step Procedure (TWI) ### Step 1: Strip Jargon and Reset Context - **Action**: Acknowledge the ambiguity, discard low-level implementation minutiae, and state the primary objective in plain English. - **Key Point**: Check `CONTEXT.md` to ensure correct domain terms are used without inventing new terms. - **Why**: Cognitive fatigue occurs when explanations introduce multiple competing mental models simultaneously. ### Step 2: Formulate Simplified Explanation - **Action**: Deliver the re-pitch in short, structured bullet points: 1. What is the current problem? 2. What is the proposed change? 3. Why does this solve the problem cleanly? - **Inline Checklist**: - [ ] Sentences average under 15 words - [ ] Canonical domain vocabulary used - [ ] Abstract metaphors replaced with concrete mechanics ### Step 3: Align on Core Understanding - **Action**: Ask a single targeted question to verify comprehension before moving forward. - **Why**: Prevents continuing down a misaligned technical path. --- ## Anti-Rationalization Guardrails | Tempting Rationalization | Binding Rule | Engineering Rationale | |---|---|---| | *"Repeat the previous explanation louder or with more words."* | **Re-pitch from a completely fresh, simpler baseline angle.** | Repeating the same explanation fails to address the root cognitive confusion. | | *"Use complex industry analogies and metaphors."* | **Use concrete, direct technical English (STE).** | Analogies introduce leaky abstractions and hide technical realities. | | *"Assume the user understood and start modifying files."* | **Confirm alignment with a clean question before acting.** | Proceeding without mutual clarity leads to rejected work and reverted PRs. |
Referenced files: 1
wayfinder5.58 KB
---
name: wayfinder
description: "Map and navigate large, uncertain multi-session efforts through a shared graph of decision tickets. Use when facing complex greenfield projects or massive architectural migrations that exceed a single session — even if the user says \"map out this massive project\". Do NOT use for small, well-scoped features."
---
# Wayfinder
Chart, navigate, and incrementally resolve large, uncertain multi-session architectural efforts through an evolving graph of decision tickets and clear frontier discovery.
---
## Core Invariants
1. **Plan Over Do**: Wayfinder tickets resolve decisions, investigate unknowns, or validate prototypes—not raw execution slices.
2. **Single Canonical Map Issue**: Maintain a single root tracker issue labeled `wayfinder:map` as the low-resolution index linking all child decision tickets.
3. **One Decision Ticket Per Session**: A single agent session claims and resolves at most ONE decision ticket (excepting parallel research subagents).
4. **Frontier Claim Protocol**: A session must explicitly assign the decision ticket to itself before beginning work to prevent multi-agent collision.
5. **Fog-of-War Graduation**: Never pre-slice blurry, distant work into speculative tickets; keep them in `## Not yet specified` until the frontier reaches them.
---
## Architecture & Map of Content (MOC)
```
[ Massive Uncertain Initiative ]
│
▼
┌───────────────────────────────────────┐
│ 1. Chart Destination & Map Issue │ ──► Issue `wayfinder:map`
│ (Notes, Fog sketches, Boundaries) │
└──────────────────┬────────────────────┘
│
▼
┌───────────────────────────────────────┐
│ 2. Frontier Decision Tickets │ ──► Child issues (`research`, `prototype`, `grilling`, `task`)
│ (Unblocked, unclaimed questions) │
└──────────────────┬────────────────────┘
│
▼
┌───────────────────────────────────────┐
│ 3. Atomic Session Resolution │ ──► Post resolution comment, close issue,
│ (Claim 1 ticket, settle decision) │ update `Decisions so far`, graduate fog
└───────────────────────────────────────┘
```
| Ticket Type | Execution Mode | Purpose |
|---|---|---|
| **Research** | AFK Subagent | Documentation, API discovery, external facts |
| **Prototype** | HITL Interactive | Throwaway exploratory code (`skills/prototype/SKILL.md`) |
| **Grilling** | HITL Interactive | Socratic decision distillation (`skills/grilling/SKILL.md`) |
| **Task** | AFK or HITL | Manual prerequisite unblocking (e.g. account setup, credential provisioning) |
---
## Step-by-Step Procedure (TWI)
### Step 1: Chart the Destination & Map Root
- **Action**: Grill the user breadth-first to define the destination (the ultimate spec, architecture, or migration target) and establish the `wayfinder:map` root issue.
- **Key Point**: If the route is already completely clear without fog, stop and route directly to `to-spec`.
- **Why**: Wayfinding overhead is warranted only when substantial unknown decisions lie between the starting point and destination.
### Step 2: Formulate Frontier Child Tickets & Wire Dependencies
- **Action**: Create sharp child issues for immediate unblocked decisions and wire native blocking relationships.
- **Key Point**: Keep un-sharp future ideas in the `## Not yet specified` section of the map.
- **Inline Checklist**:
- [ ] Map created with clear Destination and Notes
- [ ] Immediate decision questions ticketed and labeled
- [ ] Blocking dependencies wired in tracker
- [ ] Research tickets dispatched to background subagents
### Step 3: Claim and Resolve a Single Frontier Ticket
- **Action**: Assign the chosen frontier ticket to self, execute the investigation/grilling, and produce the decision verdict.
- **Key Point**: Post the decision as a resolution comment on the issue and close it.
- **Why**: Immediate closure and commentary keep the frontier clean and transparent for concurrent team members.
### Step 4: Update Map Index & Graduate Fog
- **Action**: Add a one-line summary to `## Decisions so far` on the map issue and graduate freshly specifiable items from `## Not yet specified` into new tickets.
- **Key Point**: If a path is determined to be outside the destination scope, move it to `## Out of scope` and close any associated tickets.
- **Why**: Incremental fog clearing keeps the map accurate without speculative bloat.
---
## Anti-Rationalization Guardrails
| Tempting Rationalization | Binding Rule | Engineering Rationale |
|---|---|---|
| *"I'll resolve 4 tickets in this single session."* | **Enforce 1 decision ticket per session.** | Multi-ticket batching causes context exhaustion and sloppy decision records. |
| *"Pre-create 20 tickets for all future phases."* | **Keep blurry work in 'Not yet specified'.** | Pre-slicing the fog creates brittle tickets that get invalidated by early decisions. |
| *"Start building the final feature code during wayfinding."* | **Wayfinder produces decisions, not deliverables.** | Premature implementation while architecture is foggy leads to massive rewrites. |
Referenced files: 1
wizard4.35 KB
--- name: wizard description: "Generate an interactive bash wizard to guide humans through manual setup, dashboard, or credential steps. Use when setting up API keys, third-party dashboards, CI secrets, or infrastructure where human authentication is required — even if the user says \"create a setup wizard\". Do NOT use for steps the AI agent can execute autonomously." --- # Wizard Generate structured, interactive Bash wizards that walk human operators step-by-step through manual dashboard procedures, secret provisioning, and irreversible cutover operations. --- ## Core Invariants 1. **Human-Only Boundary**: Wizards are strictly for tasks that require human interactive authentication, MFA, billing approval, or physical dashboard navigation. 2. **Immutable Helper Library**: Preserve the standardized UI helper library in `template.sh` above the `STAGES` marker; never modify internal screen-clearing or secret-masking logic. 3. **URL-First Navigation**: Always open target URLs (`open_url`) before prompting the human for values or confirmation. 4. **Masked Secret Input**: Use `ask_secret` for all API tokens, private keys, and passwords; persist secrets directly to `.env` or GitHub Secrets (`set_secret`). 5. **Static Syntax Validation**: Verify all generated wizard scripts with `bash -n <script>` and `shellcheck` before handing off to the user. --- ## Architecture & Map of Content (MOC) ``` [ Manual Dashboard / Credential Prerequisite ] ──► [ Scope Stages & Target Secrets ] ──► [ Generate Wizard from template.sh ] ──► [ Operator Execution ] ``` | Component | Responsibility | Reference Template | |---|---|---| | **Bash Wizard Engine** | UI helpers, secret prompt, .env upsert, GitHub CLI writes | `skills/wizard/template.sh` | | **Stage Scaffolding** | Linear step sequence with URL triggers | `scripts/setup-*.sh` | | **Verification Pass** | Bash syntax check and dry run | `bash -n <script>` | --- ## Step-by-Step Procedure (TWI) ### Step 1: Scope Manual Stages & Secrets Inventory - **Action**: Inspect `.env.example`, `.github/workflows/*`, and documentation to identify all manual inputs and secrets needed. - **Key Point**: For every value, determine: (1) Source URL / dashboard path, (2) Destination (`.env`, `gh secret`, or both), (3) Secret visibility. - **Why**: Thorough inventory prevents writing incomplete scripts that leave operators blocked midway. ### Step 2: Map Operator Journey per Stage - **Action**: Draft clear, sequential instructions: which dashboard menu to click, where keys are generated, and which variable is populated. - **Key Point**: Clarify exact UI labels (e.g. "Settings $\rightarrow$ API Keys $\rightarrow$ Create Secret Key"). - **Inline Checklist**: - [ ] Every stage maps to a single focused task - [ ] URLs verified against official documentation - [ ] Secret inputs mapped to `ask_secret` and `write_env` ### Step 3: Author Wizard Script from `template.sh` - **Action**: Copy `skills/wizard/template.sh` to target path (e.g. `scripts/setup-provider.sh`), set `TOTAL_STAGES`, and author stages below the marker. - **Key Point**: Do not touch the helper functions above the `STAGES` marker. - **Why**: Consistency in UI helpers ensures uniform, robust terminal behavior across platforms (macOS, Linux, WSL). ### Step 4: Validate Syntax & Deliver Hand-off - **Action**: Run `bash -n <script>` and `chmod +x <script>`. Provide the user with the exact execution command. - **Key Point**: Do not attempt to run the interactive wizard autonomously inside the agent session. - **Why**: The wizard requires interactive terminal input and browser windows that block autonomous agent subshells. --- ## Anti-Rationalization Guardrails | Tempting Rationalization | Binding Rule | Engineering Rationale | |---|---|---| | *"Generate a wizard for steps the agent could do via CLI."* | **Execute agent-capable steps directly; reserve wizards for human-only tasks.** | Forcing humans to execute tasks an agent could run wastes human time. | | *"Prompt for secrets with standard `read` without masking."* | **Mandatory `ask_secret` for all sensitive credentials.** | Plaintext secret prompts leak API tokens in terminal logs and shoulder surfing. | | *"Attempt to execute the interactive bash wizard in background subshell."* | **Hand off wizard script to user with execution command.** | Background subshells hang indefinitely on interactive `read` and `open` calls. |
Referenced files: 2
workflow-designer3.93 KB
--- name: workflow-designer description: "Specify recurring operational, review, or content workflows with explicit triggers, inputs, actions, and ownership. Use when designing repeatable team routines, approval loops, release checklists, or automated processes — even if the user says \"design a workflow for this\". Do NOT use for one-off coding tasks." --- # Workflow Designer Design and formalize repeatable human-in-the-loop and autonomous workflows by identifying cyclical operational loops, pushing checkpoints right, and defining unambiguous execution contracts. --- ## Core Invariants 1. **Push Checkpoints Right**: Defer human checkpoints as far downstream as possible; maximize autonomous work before requesting review. 2. **Actionable Decision Briefs**: Checkpoints must present a tight, decision-ready brief with clear diffs and links—never raw, uncurated outputs. 3. **Mandate Nothing Structural**: Only add AI agents, schedules, or approval gates when the problem domain strictly requires them. 4. **Self-Contained Spec Completeness**: A workflow specification is complete only when an implementer agent can build it without asking clarifying questions. 5. **Durable Notes Separation**: Distinguish between immutable workflow specifications (`workflows/*.md`) and evolving domain observations (`NOTES.md`). --- ## Architecture & Map of Content (MOC) ``` [ Operational Habit / Team Routine ] ──► [ Loop Identification & Grilling ] ──► [ Workflow Formalization ] ──► [ Implementer Handoff ] ``` | Component | Responsibility | Location | |---|---|---| | **Workflow Specification** | Executable contract (trigger, inputs, actions, checkpoints) | `workflows/*.md` | | **Domain & Tools Notes** | Observed vocabulary, habits, tools, and channels | `NOTES.md` | | **Grilling Protocol** | Socratic distillation of requirements and boundary edges | `skills/grilling/SKILL.md` | --- ## Step-by-Step Procedure (TWI) ### Step 1: Identify Cyclical Loops & Document Baseline Notes - **Action**: Analyze the user's recurring tasks (daily, weekly, per-release) and record observed tools/channels into `NOTES.md`. - **Key Point**: Map the natural lifecycle from trigger to final delivery. - **Why**: Capturing the current baseline prevents designing theoretical workflows that mismatch daily reality. ### Step 2: Conduct Socratic Grilling on Workflow Boundaries - **Action**: Run a focused grilling round exploring: - **Trigger**: Event-based (e.g. webhooks, new issue) vs. time-scheduled (e.g. cron). - **Inputs & State**: Exactly what artifacts must be ingested. - **Actions**: Discrete operational or agentic transformations. - **Checkpoints**: Minimal review gates pushed to the end. - **Inline Checklist**: - [ ] Triggers clearly specified with fallback polling intervals - [ ] Checkpoints provide decision-ready briefs - [ ] Failure paths and escalation owners defined ### Step 3: Author Executable Workflow Spec - **Action**: Write the workflow contract to `workflows/<name>.md` detailing trigger, inputs, actions, checkpoint briefs, and outputs. - **Key Point**: Ensure an implementer agent could implement scripts or automations from the document alone. - **Why**: Ambiguity in workflow contracts causes brittle automation failures. --- ## Anti-Rationalization Guardrails | Tempting Rationalization | Binding Rule | Engineering Rationale | |---|---|---| | *"Ask the human to verify every intermediate step."* | **Push checkpoints right; batch human reviews.** | Frequent interruptions cause human fatigue and destroy automation efficiency. | | *"Dump raw log files for the human checkpoint."* | **Produce synthesized, decision-ready briefs.** | Humans review clean executive summaries 10x faster than raw debug logs. | | *"Assume an LLM agent is required for every workflow step."* | **Use deterministic scripts where possible; AI only where judgment is needed.** | Deterministic scripts are faster, cheaper, and more reliable than LLMs for structured operations. |
Referenced files: 1
writing-beats4.77 KB
---
name: writing-beats
description: "Develop long-form writing through an interactive progression of narrative and argumentative beats. Use when drafting essays, articles, documentation, or blog posts where structure evolves beat by beat — even if the user says \"write this post with me\". Do NOT use for quick single-sentence copy edits."
---
# Writing Beats
Drive an interactive, choose-your-own-adventure drafting loop where each beat introduces a focused narrative move, respects grounded prerequisites, and unlocks reachable subsequent pathways.
---
## Core Invariants
1. **One Beat per Turn**: Propose 2–3 candidate next moves, preview what each unlocks, and write strictly ONE chosen beat per turn to disk.
2. **Dynamic Reachability & Grounding**: A candidate beat is only reachable if all required concepts are already grounded (either as baseline prerequisites or introduced by previous beats).
3. **No Batch Writing Ahead**: Never jump ahead to write unapproved beats; each beat must establish state before subsequent branches are computed.
4. **Preserve Author Revisions**: Always re-read the draft file from disk before generating next candidates to incorporate user edits seamlessly.
5. **Complete Journey Over Empty Pile**: The essay concludes when the narrative journey is complete, not when every raw fragment in the pile is exhausted.
---
## Architecture & Map of Content (MOC)
```
[ Raw Fragment Pile ] ──► [ Settle Grounded Prerequisites ] ──► [ Present 2-3 Candidate Beats ]
│
┌─────────────────────────────────┴─────────────────────────────────┐
▼ ▼
[ Candidate Beat A ] [ Candidate Beat B ]
- Requires: Grounded X, Y - Requires: Grounded X
- Grounds: New Concept Z - Grounds: New Concept W
- Unlocks: Path Alpha - Unlocks: Path Beta
│
▼
[ User Chooses Beat A ] ──► [ Append Beat A to Disk ] ──► [ Compute Next 2-3 Candidates ]
```
| Beat Scale | Structure | Narrative Function |
|---|---|---|
| **Micro Beat** | 1 punchy sentence | Scene transition or timing pause |
| **Standard Beat** | 1 tight paragraph | Setup + punchline or claim + rationale |
| **Macro Beat** | 2–3 paragraphs | Self-contained vignette or code walkthrough |
---
## Step-by-Step Procedure (TWI)
### Step 1: Initialize Grounding Matrix & Prerequisites
- **Action**: Ingest the raw pile and agree with the user on what concepts the audience knows walking in.
- **Key Point**: Track the active list of grounded concepts in session memory.
- **Why**: Ensures every candidate branch is conceptually accessible to the reader.
### Step 2: Formulate Candidate Next Beats
- **Action**: Offer 2–3 candidate beats drawn from the pile, explicitly stating: (1) what concepts it requires, (2) what new concept it grounds, (3) what directions it unlocks.
- **Inline Checklist**:
- [ ] All candidates reachable from currently grounded set
- [ ] Each candidate represents a distinct narrative branch
- [ ] Unlocked paths previewed for the user
### Step 3: Append Chosen Beat & Re-Read Disk
- **Action**: Write the user-selected beat to the draft file and re-read the entire draft before computing the next step.
- **Key Point**: If the user asks to backtrack or rewrite an earlier beat, edit it in place and re-compute available branches.
- **Why**: Gives the author complete interactive agency over article pacing and narrative tone.
---
## Anti-Rationalization Guardrails
| Tempting Rationalization | Binding Rule | Engineering Rationale |
|---|---|---|
| *"Write 3 beats at once to speed up drafting."* | **Write exactly 1 beat per turn.** | Batching beats removes the author's ability to steer branch direction. |
| *"Force every leftover raw fragment into the draft."* | **Conclude when the journey is complete; leave excess in the pile.** | Forcing unused fragments into the draft dilutes clarity and clutters the narrative. |
| *"Offer candidate beats that require ungrounded concepts."* | **Enforce prerequisite reachability on all candidates.** | Leaping into ungrounded concepts confuses readers and breaks argument continuity. |
Referenced files: 1
writing-for-agents4.51 KB
--- name: writing-for-agents description: "Author and refine agent-facing instructions, skills, AGENTS.md files, and context pointers. Use when creating new agent skills, optimizing prompt guidelines, writing steerable docs, or organizing progressive disclosure — even if the user says \"write a skill for this\". Do NOT use for human-facing marketing copy." --- # Writing for Agents Author and refine documents an agent consumes: a Skill, an `AGENTS.md` / `CLAUDE.md`, or a document reached by a context pointer. The packaging differs; the writing principles do not. --- ## Core Invariants 1. **The 10 Authoring Principles**: Adhere strictly to pre-flight checks, zero process in descriptions, MOC architecture (< 500 lines), TWI clarity, inline checklists, one term per concept, zero hardcoded secrets, and matching form to failure. 2. **Context Load vs. Cognitive Load**: Inline what every branch needs; push behind context pointers what only some branches reach. 3. **Front-Loaded Leading Words**: Recruit model priors with precise tokens (*tight*, *red*, *tracer bullets*) rather than lengthy re-explanations. 4. **Observable Completion Criteria**: Every step must conclude on a checkable, exhaustive completion bound to prevent premature victory declaration. 5. **Zero Process in Descriptions**: Descriptions define triggering boundaries (`Use when...`, `Do NOT use for...`), never workflow recipes. --- ## Architecture & Map of Content (MOC) ``` [ Context Pointer / Description ] ──► [ SKILL.md: Map of Content ] ──► [ Disclosed References / Scripts ] ``` | Domain | Key Mechanism | Reference | |---|---|---| | **Context Pointers** | Trigger condition + branch definition | `references/pointers.md` | | **Information Hierarchy** | In-file step → In-file reference → Disclosed reference | `SKILL-MECHANICS.md` | | **Progressive Disclosure** | Keep `SKILL.md` < 500 lines; push heavy specs out | `skill-conductor` | | **Skill Lifecycle** | Draft → Test → Review → Improve → Package | `skill-conductor` | --- ## The Information Hierarchy A document is built from **steps** (ordered actions) and **reference** (definitions and rules): 1. **In-file step**: The primary tier — what the agent does, in order. 2. **In-file reference**: Consulted on demand; short tables or checklists co-located with steps. 3. **Disclosed reference**: Pushed into a separate file reached by a pointer, loaded only when that branch fires. --- ## Step-by-Step Procedure (TWI) ### Step 1: Design Context Pointers & Trigger Boundaries - **Action**: Draft the pointer with front-loaded leading words, distinct trigger branches, and explicit negative exclusions. - **Key Point**: Never put workflow steps inside the pointer or description. - **Why**: When process steps appear in the description, models follow them and skip the detailed body. ### Step 2: Structure the Body as a Map of Content - **Action**: Organize the main markdown file as an executive map with concise headings, TWI steps, and inline checklists. - **Key Point**: Keep the main file under 500 lines; disclose deep schemas into reference files. - **Why**: Attention thins across overly long documents, causing instructions to be skipped. ### Step 3: Define Checkable Completion Bounds - **Action**: Conclude every procedure with an unambiguous completion criterion. - **Inline Checklist**: - [ ] Frontmatter name is kebab-case and matches folder - [ ] Description has positive triggers and negative exclusions (`Do NOT use for...`) - [ ] Main document is under 500 lines - [ ] No hardcoded secrets, API tokens, or user home paths - **Why**: Clear bounds prevent premature step termination and unverified assumptions. --- ## Anti-Rationalization Guardrails | Tempting Rationalization | Binding Rule | Engineering Rationale | |---|---|---| | *"I'll list the steps in the description so it triggers better."* | **Forbidden: descriptions define triggers, not steps.** | Models execute the summary description and skip the detailed body instructions. | | *"More documentation is always better."* | **Prune aggressively: ruthlessly eliminate no-ops.** | Excess lines dilute attention and increase cognitive and context load. | | *"I can use 'MUST' in all caps instead of explaining why."* | **Explain the reasoning (TWI: Action, Key Point, Why).** | Explaining why produces robust generalization; rigid prohibitions invite prompt injection. | | *"Keep everything in one file for convenience."* | **Progressive disclosure: disclose heavy references.** | Single-file sprawl degrades context efficiency and task focus. |
Referenced files: 2
writing-fragments3.56 KB
---
name: writing-fragments
description: "Capture and refine raw ideas, quotes, and observations before committing to an article outline. Use when collecting source material, brainstorming essay fragments, or exploring angles without structure pressure — even if the user says \"help me brainstorm notes\". Do NOT use for final draft formatting."
---
# Writing Fragments
Facilitate low-pressure idea capture and Socratic exploration by continuously extracting raw writing fragments, quotes, analogies, and leading metaphors into a delimiter-separated markdown notes file.
---
## Core Invariants
1. **Pure Exploration Mode**: Never force premature structuring, outlines, or paragraph grouping; keep capture friction near zero.
2. **Horizontal Rule Delimiters**: Separate every distinct fragment with horizontal rules (`---`) under a single top-level `# Working Title`.
3. **Continuous Disk Re-Read**: Always re-read the target fragments markdown file from disk before appending to preserve human out-of-band edits.
4. **Leading-Word Extraction**: Actively identify and coin compact metaphors ("leading words") that anchor core conceptual points.
5. **No Overhead Metadata**: Avoid adding tables of contents, timestamps, author frontmatter, or category tags into raw fragment files.
---
## Architecture & Map of Content (MOC)
```
[ Freeform Ideation / Grilling Dialogue ] ──► [ Extract Raw Fragment / Leading Word ] ──► [ Append to `notes.md` via `---` ] ──► [ Re-read Before Next Write ]
```
| Fragment Type | Structure | Purpose |
|---|---|---|
| **Sharp Sentence** | Single punchy line | Reusable hook or thesis anchor |
| **Vignette / Anecdote** | Short narrative paragraph | Concrete real-world grounding |
| **Quote / Dialogue** | Blockquote with attribution | Primary source evidence |
| **Leading Word** | Coined metaphor / terminology | Load-bearing concept for future shaping |
---
## Step-by-Step Procedure (TWI)
### Step 1: Initialize Working File & Settle Target Path
- **Action**: Ask the user once for the target file path (if omitted) and initialize with `# Working Title`.
- **Key Point**: Start capturing immediately from the user's first prompt.
- **Why**: Zero upfront bureaucracy maximizes creative momentum.
### Step 2: Socratic Exploration & Fragment Extraction
- **Action**: Interview the user relentlessly, surfacing non-obvious observations, contrarian takes, and sharp formulations.
- **Key Point**: When a powerful recurring theme appears, coin a compact "leading word" (e.g. *tracer bullet*, *fog of war*).
- **Inline Checklist**:
- [ ] Fragments separated strictly by `---`
- [ ] No internal subheadings (H2/H3) inside raw fragment file
- [ ] File re-read from disk prior to every append
### Step 3: Incremental Silent Appends
- **Action**: Append agreed fragments to disk without interrupting conversational flow.
- **Why**: Continuous, non-blocking persistence ensures no fleeting ideas are lost.
---
## Anti-Rationalization Guardrails
| Tempting Rationalization | Binding Rule | Engineering Rationale |
|---|---|---|
| *"Organize the raw fragments into numbered thematic sections."* | **Forbidden. Keep raw fragments unstructured.** | Premature organization closes off novel angles and creates cognitive friction. |
| *"Overwrite the file with a newly organized version."* | **Append only; preserve human manual edits.** | Overwriting destroys out-of-band human phrasing and edits. |
| *"Ask permission before writing every single fragment."* | **Append silently and mention in passing.** | Constant save confirmation prompts derail productive brainstorming flow. |
Referenced files: 1
writing-shape3.82 KB
--- name: writing-shape description: "Shape raw notes, transcript fragments, and research into a structured, coherent article draft. Use when assembling collected fragments into narrative sections, establishing flow, and building a draft block by block — even if the user says \"turn these notes into a draft\". Do NOT use for initial fragment generation." --- # Writing Shape Mine raw notes, transcript fragments, and research piles into a structured, coherent article draft by enforcing concept grounding, candidate opening hooks, and deliberate block-level formatting. --- ## Core Invariants 1. **Input Immutability**: The raw material file is strictly read-only; shape the article into a new separate document. 2. **Strict Concept Grounding**: Every concept must be grounded (either as a predefined reader prerequisite or introduced in an earlier block) before any subsequent paragraph leans on it. 3. **Multi-Opening Selection**: Always present 2–3 distinct opening hooks representing different angles/theses before building the body. 4. **Deliberate Block Formatting**: Formally justify the format of every block (prose vs. list, callout vs. inline, table vs. text, quote vs. paraphrase). 5. **Incremental Persistence**: Append and refine the article draft file paragraph-by-paragraph with live disk re-reads before every write. --- ## Architecture & Map of Content (MOC) ``` [ Raw Notes / Fragment Pile (Read-Only) ] ──► [ Establish Reader Prerequisites ] ──► [ Select Opening Hook (2-3 Options) ] ──► [ Block-by-Block Construction ] ``` | Decision Area | Tradeoff Analysis | Rule | |---|---|---| | **Prose vs. List** | Argumentative flow vs. scannable items | Prose carries thesis; lists carry strictly parallel items | | **Inline vs. Callout** | Mainline continuity vs. aside context | Use callouts (`> [!NOTE]`) only if inline disrupts thesis | | **Table vs. Paragraphs** | Repetitive structured schemas | 3+ items with matching keys $\rightarrow$ Table | | **Quote vs. Paraphrase** | Historical phrasing vs. core idea | Quote when exact words matter; paraphrase for clarity | --- ## Step-by-Step Procedure (TWI) ### Step 1: Ingest Pile & Settle Prerequisites - **Action**: Read the raw fragment pile end-to-end and establish what foundational knowledge the reader brings to the article. - **Key Point**: Keep a live register of grounded concepts throughout the session. - **Why**: Prevents cognitive gaps where articles lean on unintroduced jargon. ### Step 2: Draft 2–3 Candidate Openings - **Action**: Present 2–3 distinct openings with different angles and force a selection. - **Key Point**: The chosen opening sets the explicit contract for the rest of the draft. - **Inline Checklist**: - [ ] Raw pile untouched - [ ] Prerequisites recorded - [ ] Opening agreed and written to target draft file ### Step 3: Grow Draft Paragraph-by-Paragraph - **Action**: Ask "Given this block, what does the reader need next?", pull material from the pile, and debate the block format. - **Key Point**: Re-read file before writing and append immediately. - **Why**: Step-by-step composition ensures seamless transitions and eliminates fluff. --- ## Anti-Rationalization Guardrails | Tempting Rationalization | Binding Rule | Engineering Rationale | |---|---|---| | *"Generate the entire article all at once in one big turn."* | **Draft and debate block by block with the user.** | Single-shot generation produces generic prose and ignores user tone preferences. | | *"Edit the raw notes file directly during shaping."* | **Raw material is read-only; write to a separate draft.** | Altering the source pile destroys raw fragments needed for future articles. | | *"Use jargon before introducing the underlying concept."* | **Enforce prerequisite grounding before leaning on terms.** | Ungrounded concepts alienate readers and weaken argumentative structure. |
Referenced files: 1
Package details
Publisher declarations from the archived package. These are separate from our research and the live service's terms.
- Package author
- Mamdouh Aboammar
Declared capabilities
- Engineering task routing
- Requirements discovery & specs
- Issue decomposition & tickets
- Multi-agent spec implementation
- Test-driven development (TDD)
- Root-cause bug diagnosis
- Pre-commit code review
- Software architecture & deep modules
- Domain-driven modeling & ADRs
- Primary-source technical research
- AI & ML model engineering
- Self-healing data remediation
- ML statistical evaluation & metrics
- Cognitive workspace reasoning (J-space)
- Autonomous goal contracts
- Skill conductor authoring & evals
- Git safety guardrails & pre-commit
- Workflow design & session handoffs
- Course scaffolding & pedagogy
- Long-form writing & beats structure
Package observed Oct 3, 2026.
Technical details
- First seen
- Sep 30, 2026 · 22:02 UTC
- Last seen
- Oct 3, 2026 · 12:00 UTC
- Collection status
- Collected
plugins_6a78e83987748191afc0c56e12172fce
Download plugin data (JSON)Before you connect Matt Skills Curated
How do I connect it?
Open the publisher's marketplace listing to check current availability and follow its connection instructions. This directory does not install plugins. Check the requested access and any account requirements before connecting.
Check marketplace availability ↗
Does it require paid access?
We have not established the pricing or subscription requirements for this plugin. An absent price does not mean free access.
Compare researched pricing and access models →
How can I evaluate it?
Check the declared skills and available files, then try a small task whose result you can verify. Our archived descriptions and instructions establish publisher claims, not tested runtime quality. Review sources and coverage limits.