← GeoAI SkillsCONTENT HISTORY

Update to GeoAI Skills

Snapshot Sep 30, 2026 · 23:13 UTC · version 0.4.0

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "geostatistics-interpolation",
  "description": "Turn scattered point measurements into continuous surfaces with quantified uncertainty: variogram modeling, ordinary/universal/regression kriging, IDW, and spatially honest cross-validation. Use when unobserved values must be estimated from sparse samples such as stations, wells, or soundings. Trigger on \"interpolate\", \"kriging\", \"variogram\", or \"IDW\"; a named interpolation method is sufficient even when the requested surface is described informally as a heatmap. Do not use for point-density heatmaps, zonal aggregation, or raster resampling without value interpolation.",
  "included_files": [
    {
      "relative_path": "agents/openai.yaml",
      "size_in_bytes": 231
    },
    {
      "relative_path": "references/authoritative-sources.md",
      "size_in_bytes": 800
    }
  ],
  "skill_md_contents": "---\nname: geostatistics-interpolation\ndescription: >-\n  Turn scattered point measurements into continuous surfaces with quantified\n  uncertainty: variogram modeling, ordinary/universal/regression kriging,\n  IDW, and spatially honest cross-validation. Use when unobserved values must\n  be estimated from sparse samples such as stations, wells, or soundings.\n  Trigger on \"interpolate\", \"kriging\", \"variogram\", or \"IDW\"; a named\n  interpolation method is sufficient even when the requested surface is\n  described informally as a heatmap. Do not use for point-density heatmaps,\n  zonal aggregation, or raster resampling without value interpolation.\nlicense: MIT\nmetadata:\n  author: Muhammed Enes Duran\n---\n\n# Geostatistics & Interpolation\n\nPurpose: interpolation that reports what it doesn't know. The difference\nbetween a professional product and a pretty raster is the uncertainty\nsurface and an honest cross-validation — both are non-optional here.\n\n## Method selection\n\n| Situation | Method |\n|---|---|\n| Dense, smooth phenomenon, quick look | IDW (report power parameter; test 1-3) |\n| Physical phenomenon with spatial structure, need uncertainty | **Ordinary kriging** (default professional choice) |\n| Clear trend (elevation gradient, coastal effect) | Universal kriging or regression kriging on covariates |\n| Strong covariates available (DEM, land cover, distances) | Regression kriging / random-forest residual kriging |\n| Categorical target | Indicator kriging |\n| Honeycomb-free tessellation, no extrapolation wanted | Natural neighbor |\n\nIDW is a reasonable baseline but has no error model and bullseyes around\nextremes; say so when delivering IDW-only products.\n\n## Exploratory phase (before any interpolation)\n\n- Map the points with values; look for duplicates at identical coordinates\n  (average or offset them — kriging matrices go singular otherwise).\n- Histogram + skew: strongly skewed variables (rainfall, concentrations)\n  usually want a log/normal-score transform; back-transform predictions\n  properly (bias correction for lognormal kriging).\n- Trend check: regress value on x, y, and candidate covariates; visible\n  trend → universal/regression kriging path.\n- Declustering if sampling is preferential (dense where values are high).\n\n## Variogram discipline\n\nThe variogram is a MODELING decision, not an auto-fit output:\n\n```python\nimport gstools as gs\n\nbin_center, gamma = gs.vario_estimate((x, y), values, max_dist=dmax)  # dmax ≈ half extent\nmodel = gs.Exponential(dim=2)\nmodel.fit_variogram(bin_center, gamma, nugget=True)\nprint(model)   # report: nugget, sill, range — plus the fitted plot\n```\n\n- Max lag ≈ half the domain diameter; ≥30 pairs per bin.\n- Check anisotropy with directional variograms (0/45/90/135°); geological\n  and meteorological fields are often anisotropic — fit an anisotropic\n  model rather than ignoring it.\n- Interpret and report the parameters in words: nugget (measurement error +\n  micro-scale variance), range (correlation distance), sill. A nugget near\n  the sill means the data barely support interpolation — say that honestly.\n- Never interpolate meaningfully beyond the variogram range from the\n  nearest sample; mask or flag those cells.\n\n## Kriging execution\n\n```python\nfrom pykrige.ok import OrdinaryKriging\n\nok = OrdinaryKriging(x, y, values, variogram_model=\"exponential\",\n                     variogram_parameters={\"sill\": s, \"range\": r, \"nugget\": n},\n                     coordinates_type=\"euclidean\")\nz, ss = ok.execute(\"grid\", gridx, gridy)   # ss = kriging VARIANCE — keep it!\n```\n\n- Work in a projected CRS (kriging distances in degrees are wrong except\n  with explicitly geographic-aware models).\n- Deliver TWO rasters: prediction AND kriging standard deviation\n  (`np.sqrt(ss)`). The uncertainty map drives where to sample next and\n  where not to trust the map.\n- Search neighborhood: 16-32 nearest points typical; document it.\n\n## Validation — leave-one-out and beyond\n\n- **LOOCV** for small n: report ME (bias ≈ 0?), RMSE, and standardized\n  RMSE (should be ≈ 1 if kriging variances are honest).\n- For clustered samples, LOOCV flatters; also run spatial block CV\n  (canonical recipe: `ml-experiment-standards` →\n  `references/spatial-cv-protocol.md`) and report both.\n- Scatter plot observed vs predicted with 1:1 line; map the CV residuals —\n  spatially clustered residuals mean missing trend/covariate.\n- Compare against the dumb baseline (global mean, IDW): kriging must earn\n  its complexity.\n\n## Reporting template\n\n```\n## Surface: <variable>\n- n = <>, extent, CRS, transform applied: <log/none>\n- Variogram: <model>, nugget/sill/range = <...>, anisotropy: <...>\n- Method: <OK/UK/RK + covariates>\n- LOOCV: ME <>, RMSE <>, RMSE_std <>; block-CV RMSE <>\n- Uncertainty: kriging SD raster delivered; area beyond reliable range masked\n```\n\n## Pitfalls checklist\n\n- Auto-fitted variogram accepted without looking at the plot.\n- Kriging in EPSG:4326 over large extents.\n- Prediction raster delivered without its variance/SD companion.\n- Log-kriging back-transformed by naive `exp()` (bias!).\n- Extrapolation far beyond the data hull presented at equal confidence.\n- Duplicated station coordinates crashing or silently distorting the fit.\n\n## Execution contract\n\n- **Workflow:** inspect sampling and transform needs; model spatial dependence; fit candidate surfaces; validate against simple baselines; quantify uncertainty; mask unsupported extrapolation.\n- **Decision rules:** choose IDW or trend methods for transparent baselines, kriging when a defensible variogram exists, and regression kriging when covariates improve blocked validation.\n- **Verification protocol:** examine variogram diagnostics, LOOCV and spatial block-CV residuals, baseline uplift, uncertainty calibration, and behavior beyond the sample hull.\n- **Failure modes:** withhold a confident surface when sampling is too sparse, dependence is absent, duplicates dominate, distance units are invalid, or uncertainty is uncalibrated.\n- **Deliverables:** prediction surface, uncertainty surface, variogram and parameters, validation table, sampling-support mask, and method limitations.\n- **Source freshness:** consult [the authoritative source registry](references/authoritative-sources.md) before relying on package APIs or defaults and record the checked date.\n"
}

SHA-256: 832a71118c0334cde51e67a404b879ca4583b56954ae5e48e51e69489746e172