{"id":17798,"plugin_id":"plugins_6a7b1e3e30948191aea92f131b0f6ca9","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:14:21.357Z","digest":"a843ca43301d2eec47a7f0790dbad2e28020a6b7e44d88100186dd9d76c6e9ce","against":null,"payload":{"description":"Use when profiling, benchmarking, or optimizing Go code — includes the measure-first methodology, the pprof-driven decision tree (which symptom maps to which fix), allocation reduction, capacity hints, hot-path patterns (strconv vs fmt, repeated string→byte conversions, strings.Builder), and runtime tuning. Apply proactively whenever a user mentions slowness, allocations, GC pressure, or asks for benchmarks, even if no specific pattern is named.","included_files":[{"relative_path":"agents/openai.yaml","size_in_bytes":228},{"relative_path":"references/allocation-and-memory.md","size_in_bytes":5011},{"relative_path":"references/benchmarking-and-pprof.md","size_in_bytes":4690},{"relative_path":"references/concrete-patterns.md","size_in_bytes":5374}],"name":"go-performance","skill_md_contents":"---\nname: go-performance\ndescription: \"Use when profiling, benchmarking, or optimizing Go code — includes the measure-first methodology, the pprof-driven decision tree (which symptom maps to which fix), allocation reduction, capacity hints, hot-path patterns (strconv vs fmt, repeated string→byte conversions, strings.Builder), and runtime tuning. Apply proactively whenever a user mentions slowness, allocations, GC pressure, or asks for benchmarks, even if no specific pattern is named.\"\nlicense: MIT\ncompatibility: \"Designed for Claude Code or similar AI coding agents. Methodology is Go-version-neutral; `b.Loop()` and PGO require Go 1.21+/1.24+.\"\nallowed-tools: Read Edit Write Glob Grep Bash(go:*) Bash(golangci-lint:*)\n---\n\n# Go Performance\n\nPerformance work in Go follows one rule: **measure first**. Intuition about bottlenecks is wrong roughly 80% of the time. Profile, hypothesise, change *one thing*, re-measure. The patterns in this skill apply only on hot paths — premature optimisation makes code worse without making it faster.\n\n## Core Rules\n\n1. **Profile before optimising.** `go test -bench`, `pprof`, `fgprof` — never guess.\n2. **One change at a time.** Multi-change \"optimisation\" passes are unreviewable.\n3. **Compare with `benchstat`.** Single runs lie; you need ≥6 runs to see signal.\n4. **Allocation reduction usually beats CPU micro-optimisation** — the GC is fast but not free.\n5. **Rule out external bottlenecks first.** If 90% of latency is the DB, faster Go code is irrelevant.\n6. **Document optimisations in comments.** Future readers will revert \"ugly\" code without context.\n\n## Iterative Methodology\n\nThe cycle is: **define goal → write benchmark → measure baseline → diagnose → improve one thing → re-measure → commit with the diff.**\n\n```bash\n# baseline\ngo test -bench=BenchmarkHotPath -benchmem -count=6 ./pkg/... | tee /tmp/report-1.txt\n\n# (apply ONE change)\n\n# compare\ngo test -bench=BenchmarkHotPath -benchmem -count=6 ./pkg/... | tee /tmp/report-2.txt\nbenchstat /tmp/report-1.txt /tmp/report-2.txt\n```\n\nIf `benchstat` shows no statistically significant change, the optimisation didn't work — revert it. Keep the `/tmp/report-*.txt` files as an audit trail; paste the `benchstat` output in the commit body.\n\n> Read [references/benchmarking-and-pprof.md](references/benchmarking-and-pprof.md) for benchmark writing, pprof workflow, and `b.Loop()` (Go 1.24+).\n\n## Rule Out External Bottlenecks First\n\nBefore optimising any Go code, check that the bottleneck is actually in your process:\n\n- **`fgprof`** — captures on-CPU and off-CPU (I/O wait) time. If off-CPU dominates, the issue is elsewhere.\n- **Goroutine profile** — many goroutines blocked in `net.(*conn).Read` or `database/sql` means external I/O is the limit.\n- **Distributed tracing** — span breakdown shows which upstream is slow.\n\nIf the bottleneck is external (DB, downstream API, disk), fix that — query tuning, indexes, connection pools, caching. No Go-level change will help.\n\n## Decision Tree: Where Is Time Spent?\n\n| Symptom (from pprof) | Action |\n|---|---|\n| High `alloc_objects` / `alloc_space` | reduce allocations (preallocate, pool, struct fields) |\n| One function dominates CPU profile | inline-friendly rewrite, avoid reflection, simpler algorithm |\n| High GC% / OOM kills | tune `GOMEMLIMIT`, `GOGC`; reduce live heap |\n| Goroutines blocked on I/O | concurrency, batching, connection pool tuning |\n| Same computation many times | memoise / `singleflight` / cache |\n| Wrong algorithm (O(n²) where O(n) exists) | fix algorithm before anything else |\n| Mutex profile hot | reduce critical section, sharded locks, `sync.Pool` |\n\n> Read [references/allocation-and-memory.md](references/allocation-and-memory.md) for allocation patterns, `sync.Pool`, struct alignment, and escape analysis.\n\n## Concrete High-ROI Patterns\n\nThese are the small changes that consistently show up in profiles. Apply them when the symptom matches — not preemptively.\n\n### 1. `strconv` over `fmt` for primitives\n\n```go\n// Bad — fmt parses a format string\ns := fmt.Sprint(n)\n\n// Good — direct conversion, ~2x faster, half the allocations\ns := strconv.Itoa(n)\n```\n\n| | ns/op | allocs |\n|---|---|---|\n| `fmt.Sprint(n)` | ~143 | 2 |\n| `strconv.Itoa(n)` | ~64 | 1 |\n\n### 2. Move constant `[]byte` conversions out of loops\n\n```go\n// Bad — allocates on every iteration\nfor i := 0; i < n; i++ {\n    w.Write([]byte(\"hello\"))\n}\n\n// Good — convert once\nhello := []byte(\"hello\")\nfor i := 0; i < n; i++ {\n    w.Write(hello)\n}\n```\n\nAbout 7x faster in a tight loop.\n\n### 3. Preallocate slice and map capacity\n\n```go\n// Bad — repeated growth, O(n) copies per growth\nout := []Result{}\nfor _, x := range input {\n    out = append(out, transform(x))\n}\n\n// Good — zero reallocations\nout := make([]Result, 0, len(input))\nfor _, x := range input {\n    out = append(out, transform(x))\n}\n```\n\nSlice capacity is **exact**: `make([]T, 0, n)` allocates exactly `n` slots. Map capacity is a **hint** about bucket count, but still avoids the worst rehashes.\n\n| | Time |\n|---|---|\n| no capacity | ~2.48s |\n| with capacity | ~0.21s |\n\nAbout 12x faster on the synthetic benchmark.\n\n### 4. `strings.Builder` for loop-built strings\n\n`s += w` in a loop is O(n²). Use `strings.Builder`, with `Grow(n)` when the final size is estimable.\n\n### 5. Pass small fixed-size values\n\n`*string`, `*int`, `*time.Time` add indirection without saving anything — strings and time.Time are already small headers. Use pointers only for mutation, types ~128B+, types embedding sync primitives, or where `nil` is meaningful.\n\n> Read [references/concrete-patterns.md](references/concrete-patterns.md) for the full pattern catalogue with benchmark numbers.\n\n## Anti-Patterns\n\n| Anti-pattern | Why it hurts | Do this instead |\n|---|---|---|\n| Optimising without `pprof` | wrong target, wasted effort | profile first |\n| Default `http.Client` for high-throughput callers | `MaxIdleConnsPerHost: 2` bottleneck | configure `Transport` |\n| Logging inside hot loops | prevents inlining, allocates even when disabled | `slog.LogAttrs`, gate by level |\n| `panic`/`recover` as control flow | stack trace allocation | error returns |\n| `reflect.DeepEqual` in production | 50-200x slower than typed comparison | `slices.Equal`, `maps.Equal`, `bytes.Equal` |\n| `unsafe` without a benchmark | rarely justified | benchmark + comment with numbers |\n| No `GOMEMLIMIT` in containers | OOM kills under load | set to ~80% of container limit |\n\n## Verification Checklist\n\n- [ ] A benchmark exists for the function being optimised.\n- [ ] Baseline `/tmp/report-1.txt` was captured before any change.\n- [ ] Each change is a single commit with `benchstat` output in the body.\n- [ ] `benchstat` shows the change is statistically significant (`p < 0.05`).\n- [ ] Profile (`pprof`) confirms the targeted hotspot actually moved.\n- [ ] Optimisations on production paths have an explanatory comment.\n- [ ] `GOMEMLIMIT` is configured for any containerised long-running process.\n\n## Enforce With Linters\n\nMechanical anti-patterns belong to CI:\n\n- `gocritic` — flags `fmt.Sprint(x)` for primitives, repeated allocations.\n- `prealloc` — slices that could be preallocated.\n- `gocyclo` / `funlen` — proxies for code that is hard to optimise.\n- `fieldalignment` (go vet) — struct layout for memory reduction.\n\n## References\n\n- [references/benchmarking-and-pprof.md](references/benchmarking-and-pprof.md) — writing benchmarks, `benchstat`, pprof workflow, `b.Loop()`\n- [references/allocation-and-memory.md](references/allocation-and-memory.md) — escape analysis, `sync.Pool`, struct alignment, backing-array leaks\n- [references/concrete-patterns.md](references/concrete-patterns.md) — full pattern catalogue with numbers\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}