← GeoAI SkillsCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to GeoAI Skills
Snapshot Sep 30, 2026 · 23:13 UTC · version 0.4.0
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "ml-experiment-standards",
"description": "Always invoke for training, validating, tuning, benchmarking, or claiming readiness of a predictive model. Covers leakage audits, spatial and grouped splits, metrics, reproducibility, and honest reporting. Invoke especially when spatial dependence, split design, or deployment geography is unknown; uncertainty is a reason to use this skill. Do not trigger for descriptive EDA or non-predictive statistical inference.",
"included_files": [
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 232
},
{
"relative_path": "references/authoritative-sources.md",
"size_in_bytes": 767
},
{
"relative_path": "references/spatial-cv-protocol.md",
"size_in_bytes": 3293
}
],
"skill_md_contents": "---\nname: ml-experiment-standards\ndescription: >-\n Always invoke for training, validating, tuning, benchmarking, or claiming\n readiness of a predictive model. Covers leakage audits, spatial and grouped\n splits, metrics, reproducibility, and honest reporting. Invoke especially\n when spatial dependence, split design, or deployment geography is unknown;\n uncertainty is a reason to use this skill. Do not trigger for descriptive\n EDA or non-predictive statistical inference.\nlicense: MIT\nmetadata:\n author: Muhammed Enes Duran\n---\n\n# ML Experiment Standards\n\nPurpose: every ML job (quick prototypes included) is reproducible,\nleakage-free, and metric-justified. These are not optional polish; every\nskipped item typically returns as \"the model collapsed in production\" or\n\"the result didn't replicate\".\n\n## 1. EDA comes first\n\nBefore any model, produce and show: distributions, missingness rates,\noutliers, target balance, salient correlations. Metric and loss choice\ndepend on this information; a model recommendation without EDA is a guess.\n\n## 2. Leakage audit\n\nAt every split decision, answer explicitly (and write the answer as a code\ncomment): \"Does the training set contain indirect information about any\ntest sample?\"\n\n| Data type | Correct split | Why |\n|---|---|---|\n| Independent samples | Stratified k-fold | Preserves class ratios |\n| Time series | TimeSeriesSplit / walk-forward | Future must not leak into past |\n| **Spatial data** | Spatial block CV — see `references/spatial-cv-protocol.md` | Neighbors are near-duplicates |\n| Grouped data (patient, parcel, scene) | GroupKFold | A group must not straddle the split |\n\n- Scalers/encoders/imputers are **fit on train only**; the clean path is\n `sklearn.pipeline.Pipeline` — CV then fits correctly by construction.\n- Target-derived features (target encoding etc.) must be computed\n out-of-fold, and shown to be.\n\nThe spatial protocol in `references/spatial-cv-protocol.md` is the single\ncanonical source for this repo — other skills link here; do not restate it.\n\n## 3. Metric selection — justified\n\nNever choose a metric by default; write a one-sentence rationale:\n\n- Imbalanced classes → **F1 / AUC-PR**, not accuracy (accuracy rewards\n majority-class memorization).\n- Segmentation → **IoU/Dice** (pixel accuracy is inflated by background).\n- Regression → RMSE (sensitive to large errors) vs MAE (robust) vs R²\n (variance explained) — justify from the use case.\n- Every point estimate gets uncertainty: bootstrap CI or mean ± std across\n CV folds. A single number hides whether a difference is signal or noise.\n\n## 4. Reproducibility skeleton\n\nEvery training script follows this shape (script-first; no notebook magic):\n\n```python\n\"\"\"Experiment: <name>. Goal and success criterion: <one sentence>.\"\"\"\nfrom dataclasses import dataclass, asdict\nimport json, random\nimport numpy as np\n\n@dataclass\nclass Config:\n seed: int = 42\n lr: float = 1e-3\n batch_size: int = 32\n epochs: int = 100\n patience: int = 10 # early stopping\n\ndef set_seed(seed: int) -> None:\n random.seed(seed)\n np.random.seed(seed)\n # if torch: torch.manual_seed(seed); torch.cuda.manual_seed_all(seed)\n\ndef main(cfg: Config) -> None:\n set_seed(cfg.seed)\n ... # data -> split -> pipeline -> train -> evaluate\n with open(\"runs/run_meta.json\", \"w\", encoding=\"utf-8\") as f:\n json.dump({\"config\": asdict(cfg), \"metrics\": metrics}, f, indent=2)\n\nif __name__ == \"__main__\":\n main(Config())\n```\n\n- Config lives in a dataclass/YAML, never hardcoded — sweeps and run\n comparison depend on it.\n- Pin library versions (`pip freeze > requirements.txt`).\n- Use MLflow/W&B when available; the JSON log above is the minimum.\n\n## 5. Deep learning extras\n\n- **Loss rationale**: Dice/Dice+CE for imbalanced segmentation; write why.\n Focal only after comparison — not a free win.\n- **Augmentation rationale**: state which transforms respect the physics\n of the problem (orientation-dependent tasks forbid some rotations;\n multispectral forbids naive color jitter).\n- **Overfitting control**: early stopping with patience + a train/val\n curve in the report; no curve, no \"the model is good\".\n- **Capacity order**: small model + simple baseline first (logistic\n regression, RF); a deep model that can't beat the baseline is a data\n problem, not an architecture problem.\n- EO-specific chipping/inference details → `geo-deep-learning`.\n\n## 6. System context (MLOps)\n\nPosition every model in its chain in one paragraph: data source →\ncleaning → features/versioning → training → evaluation → deployment\n(batch/real-time) → monitoring (data/model drift). Even for a prototype,\nnote \"what this step becomes in production\".\n\n## 7. Report format\n\n```\n## Experiment: <name>\n- Data: n=<>, split: <strategy + rationale>\n- Baseline: <model> → <metric ± CI>\n- Model: <model> → <metric ± CI>\n- Leakage audit: <what was checked>\n- Next step: <single recommendation>\n```\n\nWhen reporting differences, respect statistical honesty: if the gap\ndoesn't exceed the across-fold std, say \"no clear difference\" — no\np-hacking, no selective reporting.\n\n## Execution contract\n\n- **Workflow:** define prediction target and decision use; establish a baseline; audit leakage; create spatially valid splits; train reproducibly; quantify uncertainty; inspect errors and deployment fit.\n- **Decision rules:** apply this skill only to predictive model experiments; use spatial statistics for inference, geostatistics for sampled-surface estimation, and descriptive analysis without forcing a model.\n- **Verification protocol:** reproduce from a clean environment, compare against baseline across folds or seeds, inspect spatial residuals, verify split independence, and test the final decision threshold.\n- **Failure modes:** invalidate uplift claims for leakage, post-split preprocessing, inappropriate metrics, non-independent test units, selective runs, or train-serving skew.\n- **Deliverables:** experiment configuration, split and seed manifest, baseline and model metrics with uncertainty, leakage audit, error analysis, artifacts, and deployment caveats.\n- **Source freshness:** consult [the authoritative source registry](references/authoritative-sources.md) before using version-sensitive split, metric, or reproducibility APIs.\n"
}SHA-256: 737e8ac367d770a9a32f948c99100595c5e4f5cfbb7ad271a713bd21f2749c0f