← GeoAI SkillsCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to GeoAI Skills
Snapshot Sep 30, 2026 · 23:13 UTC · version 0.4.0
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "postgis-spatial-sql",
"description": "Invoke whenever spatial SQL or its execution backend is the decision: PostGIS, DuckDB Spatial, SpatiaLite, ST_* functions, recurring spatial joins, concurrent/growing workloads, or large GeoParquet queries. Covers backend selection, schemas, GiST/BRIN indexes, KNN, geometry versus geography, correctness benchmarks, and EXPLAIN optimization. Use PostGIS for managed concurrent services and embedded engines for bounded local analytics when evidence supports that choice. Use geo-data-engineering for acquisition, conversion, and file-based ETL without spatial SQL.",
"included_files": [
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 216
},
{
"relative_path": "references/authoritative-sources.md",
"size_in_bytes": 781
}
],
"skill_md_contents": "---\nname: postgis-spatial-sql\ndescription: >-\n Invoke whenever spatial SQL or its execution backend is the decision:\n PostGIS, DuckDB Spatial, SpatiaLite, ST_* functions, recurring spatial\n joins, concurrent/growing workloads, or large GeoParquet queries. Covers\n backend selection, schemas, GiST/BRIN indexes, KNN, geometry versus\n geography, correctness benchmarks, and EXPLAIN optimization. Use PostGIS\n for managed concurrent services and embedded engines for bounded local\n analytics when evidence supports that choice. Use geo-data-engineering for\n acquisition, conversion, and file-based ETL without spatial SQL.\nlicense: MIT\nmetadata:\n author: Muhammed Enes Duran\n---\n\n# PostGIS & Spatial SQL\n\nPurpose: correct-and-fast spatial SQL. The two recurring failure modes are\nsemantic (geometry vs geography, SRID mismatches → wrong answers) and\nperformance (missing index usage → hour-long joins); this skill guards\nboth.\n\n## When the database is the right tool\n\nMove from files/GeoPandas to PostGIS when any of: features > a few\nmillion, concurrent readers/writers, repeated ad-hoc querying, a serving\nAPI on top, or transactional integrity needs. For single-shot analytical\nscans over GeoParquet, **DuckDB Spatial** is often the fastest\nzero-install path — same SQL mindset, no server.\n\nWhen requirements are incomplete, do not turn this heuristic into a final\nrecommendation. First obtain current and forecast data volume, concurrency,\ndelivery and mutation pattern, latency/SLA, serving needs, and operational\nownership (including backup and recovery). Define representative ingestion,\njoin, and read queries for both viable backends; compare runtime and resource\nuse only after row counts, join cardinality, SRID, geometry validity, and sample\noutputs agree. Include this benchmark and correctness plan in the current\nresponse; do not merely offer to draft it later.\n\n## Schema fundamentals\n\nThis runnable example assumes the data is contained in UTM zone 33N. Replace\nEPSG:32633 with a projected CRS verified for the actual area of interest.\n\n```sql\nCREATE TABLE parcels (\n id bigint GENERATED ALWAYS AS IDENTITY PRIMARY KEY,\n parcel_no text NOT NULL,\n landuse text,\n area_m2 double precision, -- unit in the name, always\n geom geometry(MultiPolygon, 32633) NOT NULL\n);\nCREATE INDEX parcels_geom_gix ON parcels USING gist (geom);\nANALYZE parcels;\n```\n\n- **Type the geometry column fully**: `geometry(MultiPolygon, SRID)` — an\n untyped `geometry` column happily accepts mixed garbage.\n- Promote to Multi* on load (`ST_Multi`) so Polygon/MultiPolygon mixing\n never bites.\n- **geometry vs geography**: geometry in a projected SRID for regional\n analysis (fast, full function set); geography (SRID 4326) when the\n extent is global/cross-zone and you want meters without picking a\n projection (slower, smaller function set). Never store in 4326 geometry\n and call `ST_Area` expecting m² — that's square degrees.\n- Never use EPSG:3857/Web Mercator for area or length measurement. When the\n analysis CRS is not yet known, either use 4326 geography for a geodesic\n result or stop and select a verified local/equal-area CRS; do not present a\n known-distorting CRS as a runnable measurement alternative.\n- **Any stored geometry column you recommend must be typed with its SRID.**\n Advising a \"second projected geometry column\" for repeated measurement is\n incomplete until it is written as `geometry(<Type>, <SRID>)` with the index\n and the populating `ST_Transform`. An untyped column recommended as a fix\n reintroduces the mixed-SRID problem it was meant to solve:\n\n ```sql\n ALTER TABLE parcels ADD COLUMN geom_32633 geometry(MultiPolygon, 32633);\n UPDATE parcels SET geom_32633 = ST_Transform(geom, 32633);\n CREATE INDEX parcels_geom_32633_gix ON parcels USING gist (geom_32633);\n ```\n- GiST index on every geometry column, `ANALYZE` after bulk loads; BRIN\n only for huge, spatially-ordered, append-only tables.\n- Load paths: `ogr2ogr -f PostgreSQL`, `shp2pgsql`, or GeoPandas\n `to_postgis` (small/medium). `COPY` beats INSERT by orders of magnitude.\n\n## Correct spatial predicates\n\n- `ST_Intersects` for \"touches at all\", `ST_Contains`/`ST_Within` for\n containment, `ST_DWithin(a, b, dist)` for proximity — **never**\n `ST_Distance(a,b) < dist` (that form can't use the index).\n- The classic point-in-polygon join:\n\n```sql\nSELECT p.id, a.district\nFROM points p\nJOIN admin a ON ST_Intersects(a.geom, p.geom); -- GiST on both sides\n```\n\n- KNN nearest-neighbor with the distance operator (index-assisted):\n\n```sql\nSELECT h.id, h.name\nFROM hospitals h\nORDER BY h.geom <-> (SELECT geom FROM incident WHERE id = 42)\nLIMIT 3;\n```\n\n`<->` gives true-distance ordering on modern PostGIS for geometry; wrap\nwith `ST_DWithin` to bound the search when tables are huge.\n\n## Performance playbook\n\n1. `EXPLAIN (ANALYZE, BUFFERS)` first — confirm the GiST index is used\n (look for \"Index Scan ... _gix\"); a Seq Scan on a big spatial join\n means a rewrite, not a bigger server.\n2. Same SRID on both sides of every predicate — `ST_Transform` inside a\n join predicate kills index use; store a transformed, indexed copy\n instead.\n3. Big-polygon problem: country/basin-sized geometries make index bboxes\n useless → `ST_Subdivide` into a work table (typical 10-100× speedup on\n joins against them).\n\nThe following example assumes `countries(country_id, geom)`.\n\n```sql\nCREATE TABLE country_parts AS\nSELECT c.country_id, part.geom\nFROM countries AS c\nCROSS JOIN LATERAL ST_Subdivide(c.geom, 256) AS part(geom);\n\nCREATE INDEX country_parts_geom_gix ON country_parts USING gist (geom);\nANALYZE country_parts;\n```\n\n`ST_Subdivide` is a set-returning function; do not access its result as\n`(ST_Subdivide(...)).geom`.\n\n4. Validity in-database: `ST_IsValid` audit, `ST_MakeValid` repair, add a\n `CHECK (ST_IsValid(geom))` if writers are untrusted.\n5. Simplify for serving, not for analysis: keep full-resolution geometry;\n generate `ST_SimplifyPreserveTopology` copies or vector tiles\n (`ST_AsMVT`) for the web tier.\n6. Batch updates in transactions; `VACUUM ANALYZE` after churn.\n\n## Common analytical patterns\n\n```sql\n-- Area-weighted aggregation (e.g., population into custom zones)\nSELECT z.zone_id,\n SUM(b.pop * ST_Area(ST_Intersection(z.geom, b.geom)) / ST_Area(b.geom)) AS pop_est\nFROM zones z JOIN blocks b ON ST_Intersects(z.geom, b.geom)\nGROUP BY z.zone_id;\n\n-- Dissolve with attribute\nSELECT landuse, ST_Multi(ST_Union(geom))::geometry(MultiPolygon, 32633) AS geom\nFROM parcels GROUP BY landuse;\n```\n\nArea-weighted interpolation assumes uniform density within source units —\nstate that assumption when reporting. Validity repair is `ST_MakeValid`,\nnever `ST_Buffer(geom, 0)`.\n\n## DuckDB Spatial quick path\n\n```sql\nINSTALL spatial; LOAD spatial;\nSELECT a.name, count(*)\nFROM 'admin.parquet' a, 'points.parquet' p\nWHERE ST_Intersects(a.geom, p.geom)\nGROUP BY a.name;\n```\n\nReads GeoParquet/Shapefile/GPKG directly, parallel by default — ideal for\none-off large joins and pipeline steps without a server. No GiST; it plans\nits own joins — benchmark, don't assume.\n\n## Verification protocol\n\n1. Row-count accounting query after each join/overlay CTE.\n2. `SELECT DISTINCT ST_SRID(geom), GeometryType(geom)` on every table\n touched — one query kills two classic bug families.\n3. Sample 5 output features rendered over a basemap (QGIS connects\n directly) — numbers can pass while geometries are garbage.\n4. Treat every `sql` fence presented as runnable as a syntax and alias\n boundary: it must execute top-to-bottom after stated schema assumptions.\n Never put angle-bracket placeholders, ellipses, pseudocode, abandoned joins,\n or incomplete aliases inside it. If a schema value such as an SRID is\n unknown, ask for it or keep the template in a labeled `text` block.\n\n## Pitfalls checklist\n\n- `ST_Area`/`ST_Length` on 4326 geometry (square degrees).\n- EPSG:3857/Web Mercator for area or length measurement (systematic distortion).\n- `ST_Distance < x` instead of `ST_DWithin` (no index).\n- `ST_Transform` in join predicates.\n- Untyped geometry columns with mixed SRIDs.\n- Country-sized polygons joined without `ST_Subdivide`.\n- `buffer(0)` as validity repair (silent part loss) — `ST_MakeValid`.\n- Serving full-resolution geometries to web clients.\n\n## Execution contract\n\n- **Workflow:** inspect schema, SRID, geometry type, size, and query goal; choose predicates and indexes; write auditable CTEs; inspect the plan; reconcile results; operationalize safely.\n- **Decision rules:** use PostGIS for concurrent, repeated, or transactional spatial workloads; use file pipelines or DuckDB Spatial for bounded one-off transformations when a server adds no value.\n- **Verification protocol:** assert SRID and geometry invariants, account for rows at each join, compare indexed plans and timings, sample geometries on a map, and test boundary semantics.\n- **Failure modes:** block release for mixed SRIDs, accidental many-to-many explosion, invalid geometries, non-indexable predicates, geography/geometry unit confusion, or unexplained plan regressions.\n- **Deliverables:** self-contained parameterized SQL or migration with consistent CTE/table aliases, indexes and rationale, query plan evidence, row accounting, sample validation, expected schema, performance notes, and rollback guidance.\n- **Source freshness:** consult [the authoritative source registry](references/authoritative-sources.md) for the deployed database and extension versions before selecting functions or plans.\n"
}SHA-256: fad6f68bc315c8ef28492ef9fd83cc1aaa4a22459ea86d7ef10e6b534a397753