← Matt Skills CuratedCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to Matt Skills Curated
Snapshot Sep 30, 2026 · 23:14 UTC · version 1.1.0
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"description": "Design, train, optimize, deploy, and evaluate production machine learning and LLM systems. Use when building ML models, training classifiers, fine-tuning LLMs with LoRA/PEFT, deploying low-latency inference endpoints, architecting vector RAG pipelines, or evaluating model bias and data leakage — even if they don't explicitly say \"AI engineering\". Do NOT use for standard backend CRUD development, basic SQL queries without ML, or simple UI styling.",
"included_files": [
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 111
}
],
"name": "ai-engineering",
"skill_md_contents": "---\nname: ai-engineering\ndescription: \"Design, train, optimize, deploy, and evaluate production machine learning and LLM systems. Use when building ML models, training classifiers, fine-tuning LLMs with LoRA/PEFT, deploying low-latency inference endpoints, architecting vector RAG pipelines, or evaluating model bias and data leakage — even if they don't explicitly say \\\"AI engineering\\\". Do NOT use for standard backend CRUD development, basic SQL queries without ML, or simple UI styling.\"\n---\n\n# AI & Machine Learning Engineering\n\nDesign, train, optimize, deploy, and monitor production machine learning models and intelligent agent systems with rigorous validation, latency engineering, and ethical guardrails.\n\n## Core Principle\n\n> **Every ML solution must beat a naive baseline, guarantee train-test isolation, and meet strict production latency and fairness budgets.**\n\n---\n\n## Core Invariants\n\n1. **Strict Train-Test Isolation**: Preprocessing scalers, encoders, and tokenizers must be fit strictly on training folds inside isolated pipelines to eliminate data leakage.\n2. **Sub-100ms Inference Budget**: Synchronous user-facing endpoints must achieve $p99 < 80\\text{ms}$ through quantization (ONNX/TensorRT/GGUF), engine optimization, or continuous batching.\n3. **Mandatory Baseline Before Complexity**: Always benchmark against a naive baseline (majority class/mean) and simple linear model before advancing to deep networks or complex ensembles.\n4. **Demographic Parity & Bias Auditing**: Every production candidate must pass four-fifths disparate impact testing across demographic slices ($\\text{ratio} \\ge 0.80$).\n5. **Continuous Drift Observability**: Production models must emit data drift metrics (KS-test / Population Stability Index) and concept drift alerts to trigger automated retraining.\n\n---\n\n## Architecture & Map of Content (MOC)\n\n```\nProblem Framing & Data Assessment\n │\n ▼\n┌───────────────────────────────────────┐\n│ 1. Data Prep & Leakage-Free Pipeline │ (Cross-Validation, Feature Isolation)\n└──────────────────┬────────────────────┘\n │\n ▼\n┌───────────────────────────────────────┐\n│ 2. Model Training & Fine-Tuning │ (Baselines, LoRA / QLoRA, Tree Ensembles)\n└──────────────────┬────────────────────┘\n │\n ▼\n┌───────────────────────────────────────┐\n│ 3. Bias, Fairness & Ethics Audit │ (Disparate Impact, SHAP Attributions)\n└──────────────────┬────────────────────┘\n │\n ▼\n┌───────────────────────────────────────┐\n│ 4. Production Serving & Optimization │ (ONNX/vLLM, Sub-100ms Latency, Canary)\n└───────────────────────────────────────┘\n```\n\n---\n\n## Step-by-Step Procedure (TWI)\n\n### Step 1: Requirements & Data Hygiene Verification\n- **Action**: Assess target metrics, evaluate class balance, and establish reproducible validation splits.\n- **Key Point**: Check for temporal dependencies. If data has a time dimension, use chronological splitting; otherwise, use Stratified $k$-Fold cross-validation.\n- **Why**: Random splits on time-series data cause severe lookahead leakage and false confidence.\n\n### Step 2: Model Training & Architecture Selection\n- **Action**: Fit naive baselines, progress to gradient-boosted trees or fine-tune neural architectures (LoRA/PEFT for LLMs).\n- **Key Point**: For LLMs, apply LoRA to all attention and MLP projection layers with $\\alpha = 2 \\times r$ and 4-bit NF4 quantization.\n- **Why**: Selective rank adaptation achieves foundation model performance at 25% of the VRAM cost.\n- **Inline Checklist**:\n - [ ] Simple baseline established (Linear/Logistic Regression or Mean baseline)\n - [ ] Cross-validation variance across folds is within acceptable bounds (< 5%)\n - [ ] Overfitting diagnostics checked (Training loss vs Validation loss convergence)\n - [ ] Hyperparameters tuned systematically using Bayesian optimization\n\n### Step 3: Bias, Explainability & Robustness Auditing\n- **Action**: Evaluate fairness across subpopulations and calculate SHAP feature attributions.\n- **Key Point**: Enforce the four-fifths rule ($\\text{selection rate ratio} \\ge 0.80$) and compute top local feature attributions for high-stakes decisions.\n- **Why**: Undetected bias causes regulatory violations and uncalibrated real-world discrimination.\n\n### Step 4: Low-Latency Serving & Canary Rollout\n- **Action**: Export models to ONNX or vLLM, wrap in asynchronous API services, and initiate a 10% canary deployment.\n- **Key Point**: Track $p99$ latency, Population Stability Index (PSI), and error rates across the canary group.\n- **Why**: Canary testing isolates performance regressions before full traffic exposure.\n\n---\n\n## Anti-Rationalization Guardrails\n\n| Tempting Rationalization | Binding Rule | Engineering Rationale |\n|---|---|---|\n| *\"Accuracy is 98%, so we don't need confusion matrices.\"* | **Mandatory PR-AUC, F1, and confusion matrix analysis.** | High accuracy on imbalanced data often hides a degenerate majority-class classifier. |\n| *\"Preprocessing before splitting is harmless.\"* | **Zero data leakage: pipeline-contained transformers only.** | Dataset-wide preprocessing leaks test distribution parameters into training folds. |\n| *\"Model meets accuracy goals, so skip bias checks.\"* | **Fairness audit is a mandatory release gate.** | Regulatory and ethical standards mandate demographic parity verification. |\n| *\"400ms latency is fast enough for cloud servers.\"* | **Sub-100ms p99 budget for interactive endpoints.** | High latency compounds across microservices and degrades user experience. |\n"
}SHA-256 of public snapshot: 18d68287c0b8cf2b719b303c0e007654c2e031fffe92d5e6a8f4b5e96607f2e5