{"id":26453,"plugin_id":"plugin_asdk_app_69e0086d87088191a3edc052fa50c29f","kind":"skill","collection_source":"plugin_package","comparison_source":null,"observed_at":"2026-10-04T18:02:39.415Z","digest":"496f2ec10701bb33d8368db3292dbd987855529e11d32697d21fc61c15c6bd67","against":null,"payload":{"description":"Diagnose and fix excessive Postgres egress (network data transfer) in a codebase. Use when a user mentions high database bills, unexpected data transfer costs, network transfer charges, egress spikes, \"why is my Neon bill so high\", \"database costs jumped\", SELECT * optimization, query overfetching, reduce Neon costs, optimize database usage, or wants to reduce data sent from their database to their application. Also use when reviewing query patterns for cost efficiency, even if the user doesn't explicitly mention egress or data transfer.","included_files":[],"name":"neon-postgres-egress-optimizer","skill_md_contents":"---\nname: neon-postgres-egress-optimizer\ndescription: >-\n  Diagnose and fix excessive Postgres egress (network data transfer) in a codebase.\n  Use when a user mentions high database bills, unexpected data transfer costs,\n  network transfer charges, egress spikes, \"why is my Neon bill so high\",\n  \"database costs jumped\", SELECT * optimization, query overfetching,\n  reduce Neon costs, optimize database usage, or wants to reduce data sent\n  from their database to their application. Also use when reviewing query\n  patterns for cost efficiency, even if the user doesn't explicitly mention\n  egress or data transfer.\nmetadata:\n  parent: neon\n  source: https://github.com/neondatabase/agent-skills/tree/main/skills/neon-postgres-egress-optimizer\n---\n\n**FIRST**: Use the parent `neon` skill for a Neon overview, getting started with Neon, Neon development best practices, and more.\n\nIf the `neon` skill is not installed, fetch it from https://neon.com/docs/ai/skills/neon/SKILL.md or install it with:\n\n```bash\nneon skills -s neon -y\n```\n\n# Postgres Egress Optimizer\n\nGuide the user through diagnosing and fixing application-side query patterns that cause excessive data transfer (egress) from their Postgres database. Most high egress bills come from the application fetching more data than it uses.\n\nWork the four steps in order: **diagnose** which queries transfer the most data, **analyze** the codebase behind them, **fix** the anti-patterns, then **verify** nothing broke and the transfer actually dropped.\n\n## Step 1: Diagnose\n\nIdentify which queries transfer the most data. The primary tool is the `pg_stat_statements` extension.\n\n### Check if pg_stat_statements is available\n\n```sql\nSELECT 1 FROM pg_stat_statements LIMIT 1;\n```\n\nIf this errors, the extension needs to be created:\n\n```sql\nCREATE EXTENSION IF NOT EXISTS pg_stat_statements;\n```\n\nOn Neon the extension is available by default, but it may still need this CREATE EXTENSION step.\n\n### Handle empty stats\n\nStats are cleared when a Neon compute scales to zero and restarts. If the stats are empty or the compute recently woke up:\n\n1. Reset the stats to start a clean measurement window: `SELECT pg_stat_statements_reset();`\n2. Let the application run under representative traffic for at least an hour.\n3. Return and run the diagnostic queries below.\n\nIf the user has stats from a production database, use those. If they have no access to production stats, proceed to Step 2 and analyze the codebase directly — code-level patterns are often sufficient to identify the worst offenders.\n\n### Diagnostic queries\n\nRun these to identify the top egress contributors. Focus on queries that return many rows, return wide rows (JSONB, TEXT, BYTEA columns), or are called very frequently.\n\n**Queries returning the most total rows:**\n\n```sql\nSELECT query, calls, rows AS total_rows, rows / calls AS avg_rows_per_call\nFROM pg_stat_statements\nWHERE calls > 0\nORDER BY rows DESC\nLIMIT 10;\n```\n\n**Queries returning the most rows per execution** (poorly scoped SELECTs, missing pagination):\n\n```sql\nSELECT query, calls, rows AS total_rows, rows / calls AS avg_rows_per_call\nFROM pg_stat_statements\nWHERE calls > 0\nORDER BY avg_rows_per_call DESC\nLIMIT 10;\n```\n\n**Most frequently called queries** (candidates for caching):\n\n```sql\nSELECT query, calls, rows AS total_rows, rows / calls AS avg_rows_per_call\nFROM pg_stat_statements\nWHERE calls > 0\nORDER BY calls DESC\nLIMIT 10;\n```\n\n**Longest running queries** (not a direct egress measure, but helps identify problem queries during a spike):\n\n```sql\nSELECT query, calls, rows AS total_rows,\n  round(total_exec_time::numeric, 2) AS total_exec_time_ms\nFROM pg_stat_statements\nWHERE calls > 0\nORDER BY total_exec_time DESC\nLIMIT 10;\n```\n\n### Interpret the results\n\nRank findings by estimated egress impact:\n\n- **High row count + wide rows** = biggest egress. A query returning 1,000 rows where each row includes a 50KB JSONB column transfers ~50MB per call.\n- **Extreme call frequency** on even small queries adds up. A query called 50,000 times/day returning 10 rows each = 500,000 rows/day.\n- **Cross-reference with the schema** to identify which columns are wide. Look for JSONB, TEXT, BYTEA, and large VARCHAR columns.\n\n## Step 2: Analyze the Codebase\n\nFor each query identified in Step 1, or for each database query in the codebase if no stats are available, check:\n\n- Does it select only the columns the response needs?\n- Does it return a bounded number of rows (LIMIT/pagination)?\n- Is it called frequently enough to benefit from caching?\n- Does it fetch raw data that gets aggregated in application code?\n- Does it use a JOIN that duplicates parent data across child rows?\n\n## Step 3: Fix\n\nApply the appropriate fix for each problem found. Below are the most common egress anti-patterns and how to fix them.\n\n### Unused columns (SELECT \\*)\n\n**Problem:** The query fetches all columns but the application only uses a few. Large columns (JSONB blobs, TEXT fields) get transferred over the wire and discarded.\n\n**Fix:** Name only the columns the response needs.\n\n**Before:**\n\n```sql\nSELECT * FROM products;\n```\n\n**After:**\n\n```sql\nSELECT id, name, price, image_urls FROM products;\n```\n\n### Missing pagination\n\n**Problem:** A list endpoint returns all rows with no LIMIT. This is an unbounded egress risk — every new row in the table increases data transfer on every request. Flag this regardless of current table size.\n\nThis is easy to miss because the application may work fine with small datasets. But at scale, an unpaginated endpoint returning 10,000 rows with even moderate column widths can transfer hundreds of megabytes per day.\n\n**Fix:** Bound the result set with `ORDER BY` plus `LIMIT`/`OFFSET`.\n\n**Before:**\n\n```sql\nSELECT id, name, price FROM products;\n```\n\n**After:**\n\n```sql\nSELECT id, name, price FROM products\nORDER BY id\nLIMIT 50 OFFSET 0;\n```\n\nWhen adding pagination, check whether the consuming client already supports paginated responses. If not, pick sensible defaults and document the pagination parameters in the API.\n\n### High-frequency queries on static data\n\n**Problem:** A query is called thousands of times per day but returns data that rarely changes. Every call transfers the same rows from the database. This pattern is only visible from `pg_stat_statements` — the code itself looks normal.\n\nLook for queries with extremely high call counts relative to other queries. Common examples: configuration tables, category lists, feature flags, user role definitions.\n\n**Fix:** Add a caching layer between the application and the database so it avoids hitting the database on every request.\n\n### Application-side aggregation\n\n**Problem:** The application fetches all rows from a table and then computes aggregates (averages, counts, sums, groupings) in application code. The full dataset transfers over the wire even though the result is a small summary.\n\n**Fix:** Push the aggregation into SQL.\n\n**Before:** The application fetches entire tables and aggregates in code with loops or `.reduce()`.\n\n**After:**\n\n```sql\nSELECT p.category_id,\n       AVG(r.rating) AS avg_rating,\n       COUNT(r.id) AS review_count\nFROM reviews r\nINNER JOIN products p ON r.product_id = p.id\nGROUP BY p.category_id;\n```\n\n### JOIN duplication\n\n**Problem:** A JOIN between a wide parent table and a child table duplicates all parent columns across every child row. If a product has 200 reviews and the product row includes a 50KB JSONB column, the join sends that 50KB × 200 = ~10MB for a single request.\n\nThis is distinct from the SELECT \\* problem. Even if you select only needed columns, a JOIN still repeats the parent data for every child row. The fix is structural: avoid the join entirely.\n\n**Fix:** Split the join into two queries, one per table.\n\n**Before:**\n\n```sql\nSELECT * FROM products\nLEFT JOIN reviews ON reviews.product_id = products.id\nWHERE products.id = 1;\n```\n\n**After (two separate queries):**\n\n```sql\nSELECT id, name, price, description, image_urls FROM products WHERE id = 1;\nSELECT id, user_name, rating, body FROM reviews WHERE product_id = 1;\n```\n\nTwo queries instead of one JOIN. The product data is fetched once. The reviews are fetched once. No duplication.\n\n## Step 4: Verify\n\nAfter applying fixes:\n\n1. **Run existing tests** to confirm nothing broke.\n2. **Check the responses** — make sure the API still returns the same data shape. Column selection and pagination changes can break clients that depend on specific fields or full result sets.\n3. **Measure the improvement** — if pg_stat_statements data is available, reset it (`SELECT pg_stat_statements_reset();`), let traffic run, then re-run the diagnostic queries to compare before and after.\n\n## Neon Infrastructure as Code (`neon.ts`)\n\nThe fixes above cut **egress** (data transferred out of Postgres). The other big non-prod cost lever is **compute**, and you can codify it durably in `neon.ts` — Neon's infrastructure-as-code file (see the `neon` skill for the full reference) — so dev, preview, and CI branches stay cheap by default instead of relying on per-branch flags:\n\n```bash\nnpm i @neon/config\n```\n\n```typescript\n// neon.ts\nimport { defineConfig } from \"@neon/config/v1\";\n\nexport default defineConfig({\n  branch: (branch) => {\n    if (branch.exists || branch.isDefault) return {}; // don't touch prod\n    return {\n      ttl: \"7d\", // ephemeral branches auto-expire instead of accruing storage\n      postgres: {\n        computeSettings: {\n          autoscalingLimitMinCu: 0.25, // scale to zero when idle\n          autoscalingLimitMaxCu: 1, // cap autoscaling on throwaway branches\n          suspendTimeout: \"5m\",\n        },\n      },\n    };\n  },\n});\n```\n\n```bash\nneon config apply   # apply to the current branch (neon deploy is an alias)\n```\n\nThis is complementary, not a substitute: query-pattern fixes are what actually reduce egress charges, while these settings keep non-production compute and storage from quietly inflating the same bill. Because `neon checkout` applies the policy when it creates a branch, new dev/preview branches inherit the cheap profile automatically.\n\n## Further Reading\n\n- https://neon.com/docs/introduction/network-transfer.md\n- https://neon.com/docs/introduction/cost-optimization.md\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}