← Matt Skills CuratedCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to Matt Skills Curated
Snapshot Sep 30, 2026 · 23:14 UTC · version 1.1.0
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"description": "Self-healing data pipeline layer using semantic anomaly clustering, AST-validated lambda transformations, and zero-loss mathematical reconciliation. Use when data quality checks fail, anomalous records break ETL/ELT pipelines, you need automated data cleansing with sandboxed Python transformations, or you want to group data errors into semantic clusters — even if they don't explicitly say \"data remediation\". Do NOT use for standard database schema migrations, routine CRUD queries, or basic pipeline scheduling.",
"included_files": [
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 115
}
],
"name": "ai-data-remediation",
"skill_md_contents": "---\nname: ai-data-remediation\ndescription: \"Self-healing data pipeline layer using semantic anomaly clustering, AST-validated lambda transformations, and zero-loss mathematical reconciliation. Use when data quality checks fail, anomalous records break ETL/ELT pipelines, you need automated data cleansing with sandboxed Python transformations, or you want to group data errors into semantic clusters — even if they don't explicitly say \\\"data remediation\\\". Do NOT use for standard database schema migrations, routine CRUD queries, or basic pipeline scheduling.\"\n---\n\n# AI Data Remediation\n\nOperate a self-healing remediation layer for mission-critical data pipelines: intercept corrupt or anomalous data, semantically cluster pattern families, generate deterministic fix logic via local or sandboxed language models, and guarantee zero data loss.\n\n## Core Principle\n\n> **AI generates verifiable transformation logic — never touch or mutate production data directly.**\n\n---\n\n## Core Invariants\n\n1. **AI Logic Over Raw Data Mutation**: The model generates pure, deterministic transformation functions that can be tested, reviewed, and versioned. Raw model outputs are never piped directly into tables.\n2. **Mandatory AST Static Validation**: Every generated lambda or transformation must pass Abstract Syntax Tree (AST) validation against restricted namespaces before evaluation.\n3. **Strict Zero-Loss Accounting**: Total input records must exactly equal successful records plus quarantined records ($\\text{Source} = \\text{Success} + \\text{Quarantine}$). Any non-zero delta immediately halts processing.\n4. **Air-Gapped / Privacy-Preserved Execution**: Sensitive or PII data must never egress to external cloud APIs without prior tokenization or local execution.\n5. **Human Review Quarantine**: Clusters with low transformation confidence ($< 0.75$) or failed AST checks route to an isolated quarantine table with full lineage.\n\n---\n\n## Architecture & Map of Content (MOC)\n\n```\n[ Anomalous Records ] ──► [ Semantic Clustering ] ──► [ Sandboxed Logic Gen ] ──► [ AST Safety Gate ] ──► [ Vectorized Apply ] ──► [ Zero-Loss Audit ]\n```\n\n| Component | Responsibility | Key Mechanism |\n|---|---|---|\n| **Anomaly Buffer** | Staging buffer for records flagged `NEEDS_AI` | Asynchronous staging table / queue |\n| **Semantic Compression** | Grouping high-volume anomalies into pattern families | Vector embeddings + similarity clustering |\n| **Logic Synthesis** | Compiling deterministic lambda functions | Prompt-constrained JSON output format |\n| **AST Security Gate** | Static code analysis & sandboxed execution | Python `ast.parse` + restricted builtins |\n| **Reconciliation Audit** | Mathematical verification of record counts | Strict invariant: $\\Delta = \\text{Source} - (\\text{Success} + \\text{Quarantine}) = 0$ |\n\n---\n\n## Step-by-Step Procedure (TWI)\n\n### Step 1: Intercept & Isolate Anomalous Records\n- **Action**: Ingest failed rows from the deterministic validation layer into an isolated staging buffer.\n- **Key Point**: Operate strictly downstream of primary schema validation without blocking the main ingest stream.\n- **Why**: Decoupling remediation keeps upstream ingestion healthy and prevents system-wide backpressure.\n\n### Step 2: Semantic Anomaly Compression\n- **Action**: Compute vector representations of error strings and cluster them into distinct pattern families.\n- **Key Point**: Compress thousands of broken rows into 5–15 representative clusters using similarity metrics.\n- **Why**: Synthesizing 10 cluster-level transformation functions instead of 50,000 row-by-row LLM calls reduces execution time and compute cost by over 95%.\n\n### Step 3: Sandboxed Logic Generation\n- **Action**: Feed representative cluster exemplars to a constrained language model prompt that emits a pure Python lambda function.\n- **Key Point**: Restrict output to a single lambda definition with explicit input/output type contracts.\n- **Why**: Lambda functions are reproducible, testable against test suites, and easily inspected before deployment.\n\n### Step 4: AST Safety Verification & Vectorized Application\n- **Action**: Statically parse the generated lambda AST and execute across the cluster within a restricted environment.\n- **Key Point**: Reject any code containing `import`, `exec`, `eval`, `__builtins__`, file I/O, or OS calls.\n- **Inline Checklist**:\n - [ ] AST parsing verifies zero unauthorized node types (no `Import`, `ImportFrom`, `Call` to unapproved functions)\n - [ ] Transformation confidence score $\\ge 0.75$\n - [ ] Lambda passes unit test assertions on 3 cluster sample inputs\n - [ ] Unverified or failing records route immediately to quarantine\n- **Why**: Unchecked dynamic code execution presents critical security vulnerabilities and risks catastrophic database corruption.\n\n### Step 5: Zero-Loss Reconciliation & Lineage Audit\n- **Action**: Run mathematical balance checks across input, output, and quarantine tables.\n- **Key Point**: Verify $\\text{Source Records} = \\text{Success Records} + \\text{Quarantine Records}$.\n- **Why**: Silent row loss corrupts financial ledgers, analytical dashboards, and downstream dependencies.\n\n---\n\n## Anti-Rationalization Guardrails\n\n| Tempting Rationalization | Binding Rule | Engineering Rationale |\n|---|---|---|\n| *\"The generated lambda looks safe, skip the AST check.\"* | **Mandatory AST verification on 100% of generated code.** | A single unvalidated attribute access or system call compromises execution integrity. |\n| *\"Only 3 rows disappeared, let's ship the batch anyway.\"* | **Immediate halt if $\\Delta \\neq 0$.** | Silent record drops compound into major audit and financial discrepancies. |\n| *\"Let's directly patch data strings instead of generating a function.\"* | **Generate logic, never mutate raw data in-place.** | Logic can be audited, unit tested, reviewed, and rolled back; raw data patches cannot. |\n| *\"Send full customer records with PII to public APIs for faster fix.\"* | **Enforce data privacy boundaries.** | PII egress violates privacy regulations and compliance mandates. |\n\n---\n\n## Verification & Troubleshooting\n\n- **AST Rejection**: If the model attempts forbidden imports, tighten the system prompt to enforce pure mathematical/string transformations.\n- **Low Model Confidence**: Route the entire pattern family to `quarantine_review` table for human sign-off.\n- **Reconciliation Discrepancy**: Check row exception handlers to ensure failing rows are captured in quarantine rather than swallowed.\n"
}SHA-256 of public snapshot: 1f27397032319e4c05af51e062cc75005266acfd9f9aaf6a1a23024887b76ae2