← GeoAI SkillsCONTENT HISTORY

Update to GeoAI Skills

Snapshot Sep 30, 2026 · 23:13 UTC · version 0.4.0

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "spatial-statistics",
  "description": "Always invoke before testing a geographic pattern for clustering, hotspots, dependence, or explanatory regression, even when aggregation or ordinary OLS is proposed as routine. Covers Moran's I, LISA, Getis-Ord Gi*, weights, MAUP and scale sensitivity for areas/grids, residual dependence, and spatial lag/error/GWR/MGWR models. Use ML standards for predictive evaluation and geostatistics for continuous surfaces from sparse samples.",
  "included_files": [
    {
      "relative_path": "agents/openai.yaml",
      "size_in_bytes": 214
    },
    {
      "relative_path": "references/authoritative-sources.md",
      "size_in_bytes": 750
    }
  ],
  "skill_md_contents": "---\nname: spatial-statistics\ndescription: >-\n  Always invoke before testing a geographic pattern for clustering, hotspots,\n  dependence, or explanatory regression, even when aggregation or ordinary\n  OLS is proposed as routine. Covers Moran's I, LISA, Getis-Ord Gi*, weights,\n  MAUP and scale sensitivity for areas/grids, residual dependence, and\n  spatial lag/error/GWR/MGWR models. Use ML standards for predictive\n  evaluation and geostatistics for continuous surfaces from sparse samples.\nlicense: MIT\nmetadata:\n  author: Muhammed Enes Duran\n---\n\n# Spatial Statistics\n\nPurpose: answer \"is it clustered, where, and why\" with defensible inference.\nThe core discipline: spatial data violates independence assumptions, so\nstandard statistics silently overstate significance — every analysis here\nstarts with weights design and ends with residual diagnostics.\n\n## Spatial weights (W) — the analysis IS the weights\n\nEvery result downstream depends on W; choose it for substantive reasons and\nrun a sensitivity check with one alternative:\n\n| Weights | Use when |\n|---|---|\n| Queen/Rook contiguity | Irregular polygons (admin units, parcels) |\n| K-nearest neighbors | Points; islands present (contiguity leaves them unconnected) |\n| Distance band | Physical process with known range |\n| Kernel (distance-decayed) | Smooth influence, GWR-style local models |\n\n```python\nfrom libpysal.weights import Queen\n\nw = Queen.from_dataframe(gdf, use_index=True)\nprint(f\"islands: {w.islands}\")   # unconnected units break stats — fix or document\nw.transform = \"r\"                # row-standardize (default for Moran/lag models)\n```\n\nAlways report: weights type, parameters, number of islands, and whether\nresults survive an alternative W.\n\n## Global → local workflow\n\n1. **Global Moran's I** (`esda.Moran`, permutation inference ≥999) —\n   answers \"any clustering at all?\" Report I, p_sim, and the permutation\n   distribution, not the analytical p.\n2. **LISA / local Moran** (`esda.Moran_Local`) — maps WHERE: High-High,\n   Low-Low clusters, High-Low/Low-High outliers. Correct for multiple\n   testing (FDR at minimum) before coloring a map — uncorrected LISA maps\n   overstate clusters and this is the field's most common abuse.\n3. **Getis-Ord Gi\\*** (`esda.G_Local`, star=True) — hot/cold spots of\n   intensity (a distinct question from Moran clusters — Gi* finds\n   concentrations of high values, LISA finds similarity structure).\n4. Rates, not counts, for population-based phenomena; use Empirical Bayes\n   smoothing (`esda.smoothing`) for small-population units before any of\n   the above — raw rates in sparse units are noise.\n\n## Point patterns\n\n- Separate first-order intensity (density varies) from second-order\n  interaction (points attract/repel) — KDE describes the former, Ripley's\n  K/L (`pointpats`) tests the latter.\n- Always test against an inhomogeneous null when the study area has obvious\n  density gradients (population, roads); CSR against a city is a strawman.\n- KDE bandwidth drives the story: report it, justify it (Silverman/CV), and\n  show one alternative.\n\n## Spatial regression decision path\n\nRun OLS first, then diagnose — never start with a spatial model:\n\n```python\nfrom spreg import OLS\nols = OLS(y, X, w=w, spat_diag=True, moran=True, name_y=\"price\", name_x=xnames)\n```\n\nDecision (Anselin's rule via LM tests): LM-Lag significant & LM-Error not →\n**spatial lag (SAR)**; reverse → **spatial error (SEM)**; both → compare\nrobust LM versions; neither → OLS stands (report that as a finding).\nInterpretation caveats: in SAR, coefficients are NOT marginal effects —\nreport direct/indirect (spillover) effects. In SEM, spatial structure is\nnuisance correlation, no spillover story allowed.\n\n**GWR/MGWR** (`mgwr`): when relationships plausibly vary over space.\nBandwidth by AICc search; map local coefficients WITH local t-values masked\nfor insignificance; MGWR when predictors operate at different scales.\nGWR is exploratory — resist causal language on local coefficients.\n\n## Inference honesty\n\n- Permutation p-values over analytical ones wherever available.\n- Multiple testing: n local tests = n units; FDR-correct.\n- MAUP (modifiable areal unit problem): results can flip with unit\n  aggregation — if the aggregation level is a choice, test one alternative\n  and disclose.\n- Spatial autocorrelation in residuals after modeling = model still wrong;\n  report residual Moran's I for every final model.\n- Correlation ≠ causation applies doubly here: spatially confounded\n  variables (everything correlates with \"distance to coast\") demand\n  explicit identification strategies before causal claims.\n\n## Reporting template\n\n```\n## Spatial analysis: <question>\n- Units & n, variable(s), rate smoothing: <...>\n- W: <type/params>, islands: <n>, sensitivity W: <type>\n- Global: Moran's I = <> (p_perm = <>)\n- Local: <k> significant clusters after FDR; map attached\n- Model: <OLS/SAR/SEM/GWR> chosen because <LM diagnostics>\n- Residual Moran's I: <> — <interpretation>\n- Caveats: MAUP, W-sensitivity, causal limits\n```\n\n## Execution contract\n\n- **Workflow:** define inferential question and unit; inspect distributions and rates; construct and justify spatial weights; run global before local tests; fit models if needed; diagnose residual dependence; report uncertainty.\n- **Decision rules:** use spatial statistics for dependence and inference, geostatistics for interpolating sampled continuous surfaces, and predictive ML when out-of-sample prediction is the primary goal.\n- **Verification protocol:** test alternative weights and aggregation, use valid permutation or model inference, correct local multiplicity, inspect residual Moran's I, and distinguish association from causation.\n- **Failure modes:** withhold inferential claims for arbitrary weights, islands ignored, unstable MAUP results, uncorrected multiple tests, residual autocorrelation, or unsupported causal language.\n- **Deliverables:** analysis-ready variables, weights specification, global and local results, corrected significance, diagnostic maps, model and residual checks, sensitivity analysis, and caveats.\n- **Source freshness:** consult [the authoritative source registry](references/authoritative-sources.md) before applying version-sensitive statistical APIs or defaults.\n"
}

SHA-256: 62d8184cc28a72aaaeb2948749b6be815b7186cf62f874c72434673fb5a37cb2