{"id":16779,"plugin_id":"plugins_6a68c8b958b88191b2bfeae31847c8da","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:13:46.708Z","digest":"737e8ac367d770a9a32f948c99100595c5e4f5cfbb7ad271a713bd21f2749c0f","against":null,"payload":{"name":"ml-experiment-standards","description":"Always invoke for training, validating, tuning, benchmarking, or claiming readiness of a predictive model. Covers leakage audits, spatial and grouped splits, metrics, reproducibility, and honest reporting. Invoke especially when spatial dependence, split design, or deployment geography is unknown; uncertainty is a reason to use this skill. Do not trigger for descriptive EDA or non-predictive statistical inference.","included_files":[{"relative_path":"agents/openai.yaml","size_in_bytes":232},{"relative_path":"references/authoritative-sources.md","size_in_bytes":767},{"relative_path":"references/spatial-cv-protocol.md","size_in_bytes":3293}],"skill_md_contents":"---\nname: ml-experiment-standards\ndescription: >-\n  Always invoke for training, validating, tuning, benchmarking, or claiming\n  readiness of a predictive model. Covers leakage audits, spatial and grouped\n  splits, metrics, reproducibility, and honest reporting. Invoke especially\n  when spatial dependence, split design, or deployment geography is unknown;\n  uncertainty is a reason to use this skill. Do not trigger for descriptive\n  EDA or non-predictive statistical inference.\nlicense: MIT\nmetadata:\n  author: Muhammed Enes Duran\n---\n\n# ML Experiment Standards\n\nPurpose: every ML job (quick prototypes included) is reproducible,\nleakage-free, and metric-justified. These are not optional polish; every\nskipped item typically returns as \"the model collapsed in production\" or\n\"the result didn't replicate\".\n\n## 1. EDA comes first\n\nBefore any model, produce and show: distributions, missingness rates,\noutliers, target balance, salient correlations. Metric and loss choice\ndepend on this information; a model recommendation without EDA is a guess.\n\n## 2. Leakage audit\n\nAt every split decision, answer explicitly (and write the answer as a code\ncomment): \"Does the training set contain indirect information about any\ntest sample?\"\n\n| Data type | Correct split | Why |\n|---|---|---|\n| Independent samples | Stratified k-fold | Preserves class ratios |\n| Time series | TimeSeriesSplit / walk-forward | Future must not leak into past |\n| **Spatial data** | Spatial block CV — see `references/spatial-cv-protocol.md` | Neighbors are near-duplicates |\n| Grouped data (patient, parcel, scene) | GroupKFold | A group must not straddle the split |\n\n- Scalers/encoders/imputers are **fit on train only**; the clean path is\n  `sklearn.pipeline.Pipeline` — CV then fits correctly by construction.\n- Target-derived features (target encoding etc.) must be computed\n  out-of-fold, and shown to be.\n\nThe spatial protocol in `references/spatial-cv-protocol.md` is the single\ncanonical source for this repo — other skills link here; do not restate it.\n\n## 3. Metric selection — justified\n\nNever choose a metric by default; write a one-sentence rationale:\n\n- Imbalanced classes → **F1 / AUC-PR**, not accuracy (accuracy rewards\n  majority-class memorization).\n- Segmentation → **IoU/Dice** (pixel accuracy is inflated by background).\n- Regression → RMSE (sensitive to large errors) vs MAE (robust) vs R²\n  (variance explained) — justify from the use case.\n- Every point estimate gets uncertainty: bootstrap CI or mean ± std across\n  CV folds. A single number hides whether a difference is signal or noise.\n\n## 4. Reproducibility skeleton\n\nEvery training script follows this shape (script-first; no notebook magic):\n\n```python\n\"\"\"Experiment: <name>. Goal and success criterion: <one sentence>.\"\"\"\nfrom dataclasses import dataclass, asdict\nimport json, random\nimport numpy as np\n\n@dataclass\nclass Config:\n    seed: int = 42\n    lr: float = 1e-3\n    batch_size: int = 32\n    epochs: int = 100\n    patience: int = 10  # early stopping\n\ndef set_seed(seed: int) -> None:\n    random.seed(seed)\n    np.random.seed(seed)\n    # if torch: torch.manual_seed(seed); torch.cuda.manual_seed_all(seed)\n\ndef main(cfg: Config) -> None:\n    set_seed(cfg.seed)\n    ...  # data -> split -> pipeline -> train -> evaluate\n    with open(\"runs/run_meta.json\", \"w\", encoding=\"utf-8\") as f:\n        json.dump({\"config\": asdict(cfg), \"metrics\": metrics}, f, indent=2)\n\nif __name__ == \"__main__\":\n    main(Config())\n```\n\n- Config lives in a dataclass/YAML, never hardcoded — sweeps and run\n  comparison depend on it.\n- Pin library versions (`pip freeze > requirements.txt`).\n- Use MLflow/W&B when available; the JSON log above is the minimum.\n\n## 5. Deep learning extras\n\n- **Loss rationale**: Dice/Dice+CE for imbalanced segmentation; write why.\n  Focal only after comparison — not a free win.\n- **Augmentation rationale**: state which transforms respect the physics\n  of the problem (orientation-dependent tasks forbid some rotations;\n  multispectral forbids naive color jitter).\n- **Overfitting control**: early stopping with patience + a train/val\n  curve in the report; no curve, no \"the model is good\".\n- **Capacity order**: small model + simple baseline first (logistic\n  regression, RF); a deep model that can't beat the baseline is a data\n  problem, not an architecture problem.\n- EO-specific chipping/inference details → `geo-deep-learning`.\n\n## 6. System context (MLOps)\n\nPosition every model in its chain in one paragraph: data source →\ncleaning → features/versioning → training → evaluation → deployment\n(batch/real-time) → monitoring (data/model drift). Even for a prototype,\nnote \"what this step becomes in production\".\n\n## 7. Report format\n\n```\n## Experiment: <name>\n- Data: n=<>, split: <strategy + rationale>\n- Baseline: <model> → <metric ± CI>\n- Model: <model> → <metric ± CI>\n- Leakage audit: <what was checked>\n- Next step: <single recommendation>\n```\n\nWhen reporting differences, respect statistical honesty: if the gap\ndoesn't exceed the across-fold std, say \"no clear difference\" — no\np-hacking, no selective reporting.\n\n## Execution contract\n\n- **Workflow:** define prediction target and decision use; establish a baseline; audit leakage; create spatially valid splits; train reproducibly; quantify uncertainty; inspect errors and deployment fit.\n- **Decision rules:** apply this skill only to predictive model experiments; use spatial statistics for inference, geostatistics for sampled-surface estimation, and descriptive analysis without forcing a model.\n- **Verification protocol:** reproduce from a clean environment, compare against baseline across folds or seeds, inspect spatial residuals, verify split independence, and test the final decision threshold.\n- **Failure modes:** invalidate uplift claims for leakage, post-split preprocessing, inappropriate metrics, non-independent test units, selective runs, or train-serving skew.\n- **Deliverables:** experiment configuration, split and seed manifest, baseline and model metrics with uncertainty, leakage audit, error analysis, artifacts, and deployment caveats.\n- **Source freshness:** consult [the authoritative source registry](references/authoritative-sources.md) before using version-sensitive split, metric, or reproducibility APIs.\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}