{"id":17794,"plugin_id":"plugins_6a7b1e3e30948191aea92f131b0f6ca9","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:14:21.294Z","digest":"ee237caf58d4b5e7bac58813c75f25c16e1a27c66a275d364dfa3390a35da66f","against":null,"payload":{"description":"Use when instrumenting Go services with metrics and distributed traces, or wiring exemplars and request-id propagation. Covers Prometheus patterns (Counter/Gauge/Histogram, low-cardinality labels), OpenTelemetry tracing (TracerProvider, span attributes, errors, context propagation), and metric ↔ trace correlation so a P99 spike jumps to the offending trace. Logging: see go-logging.","included_files":[{"relative_path":"agents/openai.yaml","size_in_bytes":236},{"relative_path":"references/anti-patterns.md","size_in_bytes":4324},{"relative_path":"references/correlation.md","size_in_bytes":3909},{"relative_path":"references/metrics.md","size_in_bytes":4213},{"relative_path":"references/tracing.md","size_in_bytes":4279}],"name":"go-observability","skill_md_contents":"---\nname: go-observability\ndescription: \"Use when instrumenting Go services with metrics and distributed traces, or wiring exemplars and request-id propagation. Covers Prometheus patterns (Counter/Gauge/Histogram, low-cardinality labels), OpenTelemetry tracing (TracerProvider, span attributes, errors, context propagation), and metric ↔ trace correlation so a P99 spike jumps to the offending trace. Logging: see go-logging.\"\nlicense: MIT\ncompatibility: \"Designed for Claude Code or similar AI coding agents. Requires Go 1.21+ (for log/slog context variants used in correlation snippets). Prometheus client_golang and OpenTelemetry Go SDK.\"\nallowed-tools: Read Edit Write Glob Grep Bash(go:*) Bash(golangci-lint:*)\n---\n\n# Go Observability — Metrics, Traces, and Correlation\n\nProduction Go services need at least two always-on signals to be debuggable: **metrics** (aggregated measurements for alerting and SLOs) and **traces** (per-request flow showing where time went). The third deliverable is **correlation**: a P99 metric spike must lead, in one click, to the trace that caused it.\n\n> **Logging is not in this skill.** Structured logging with `log/slog`, log levels, and zap/logrus/zerolog migration belong to the **go-logging** skill. This skill only references logs in the context of correlating them with traces.\n\n## Core Rules\n\n1. **A feature is not done until it is observable.** New code MUST export at least: one Counter for operations, one Counter for errors, one Histogram for latency.\n2. **Histograms, not Summaries, for latency.** Summaries cannot be aggregated across instances; Histograms support `histogram_quantile()` server-side.\n3. **Label cardinality is bounded.** Never put unbounded values (user IDs, full URLs, request IDs) in Prometheus labels. Use route patterns, status classes, method.\n4. **Context flows everywhere.** A function that does I/O takes `ctx context.Context` as its first argument. No context = no trace propagation = no correlation.\n5. **Record errors on the span.** When a span ends in failure, call `span.RecordError(err)` and `span.SetStatus(codes.Error, ...)`. A green span hides a real failure.\n6. **Correlate or it didn't happen.** Inject `trace_id` into logs, attach exemplars to histograms. Otherwise the three signals are three different products.\n\n## Signal Decision\n\nPick the signal that matches the question. Do not log what should be a metric.\n\n| Question | Signal | Tool |\n|---|---|---|\n| How often does X happen? Error rate? Rate-per-second? | Metric (Counter) | Prometheus |\n| What is the P99 latency of endpoint /orders? | Metric (Histogram) | Prometheus + `histogram_quantile` |\n| Where did this one slow request spend its time? | Trace | OpenTelemetry |\n| Why does latency spike at 14:32? | Metric → exemplar → trace | Prometheus + OTel + exemplars |\n| What concrete error message did this request hit? | Log (see go-logging) | `log/slog` |\n\n> Read [references/metrics.md](references/metrics.md) for Counter/Gauge/Histogram patterns, naming, and PromQL-as-comments.\n> Read [references/tracing.md](references/tracing.md) for TracerProvider setup, span attributes, and `otelhttp` middleware.\n\n## Metrics — the 60-Second Setup\n\n```go\nimport \"github.com/prometheus/client_golang/prometheus\"\n\n// rate(http_requests_total{code=~\"5..\"}[5m]) / rate(http_requests_total[5m])\nvar httpRequests = prometheus.NewCounterVec(\n    prometheus.CounterOpts{\n        Name: \"http_requests_total\",\n        Help: \"Total HTTP requests by method, route, status class.\",\n    },\n    []string{\"method\", \"route\", \"code\"}, // ALL bounded\n)\n\n// histogram_quantile(0.99, sum by (le, route) (rate(http_request_duration_seconds_bucket[5m])))\nvar httpLatency = prometheus.NewHistogramVec(\n    prometheus.HistogramOpts{\n        Name:    \"http_request_duration_seconds\",\n        Help:    \"HTTP request latency in seconds.\",\n        Buckets: prometheus.DefBuckets,\n    },\n    []string{\"method\", \"route\"},\n)\n```\n\nThe comment above each metric is the PromQL it is designed to answer. This makes the metric discoverable and grep-able from a dashboard.\n\n## Traces — the 60-Second Setup\n\n```go\nimport (\n    \"go.opentelemetry.io/otel\"\n    \"go.opentelemetry.io/otel/codes\"\n)\n\nfunc (s *OrderService) Create(ctx context.Context, in CreateOrderInput) (*Order, error) {\n    ctx, span := otel.Tracer(\"order-service\").Start(ctx, \"OrderService.Create\")\n    defer span.End()\n    span.SetAttributes(attribute.String(\"order.user_id\", in.UserID))\n\n    order, err := s.repo.Insert(ctx, in)\n    if err != nil {\n        span.RecordError(err)\n        span.SetStatus(codes.Error, \"insert failed\")\n        return nil, fmt.Errorf(\"creating order: %w\", err)\n    }\n    return order, nil\n}\n```\n\nEvery service method, every DB query, every external HTTP call gets a span. Context **must** flow into `s.repo.Insert(ctx, ...)` so the DB span is a child of the service span.\n\n## Correlation\n\n### Metrics → Traces with Exemplars\n\nAn exemplar attaches a single trace_id to a histogram observation. In Grafana, click the dot on a P99 spike and you land on the trace.\n\n```go\nobs := httpLatency.WithLabelValues(r.Method, routePattern)\nsc := trace.SpanContextFromContext(ctx)\nif eo, ok := obs.(prometheus.ExemplarObserver); ok && sc.IsValid() {\n    eo.ObserveWithExemplar(elapsed.Seconds(),\n        prometheus.Labels{\"trace_id\": sc.TraceID().String()})\n} else {\n    obs.Observe(elapsed.Seconds())\n}\n```\n\n### Logs → Traces\n\nUse the `otelslog` bridge (see go-logging skill) so every `slog.InfoContext(ctx, ...)` call automatically emits `trace_id` and `span_id`. You can then grep logs by trace_id when starting from a trace, or jump from a log line to the trace.\n\n> Read [references/correlation.md](references/correlation.md) for exemplar wiring details, request-id propagation, and the end-to-end \"metric spike → trace → log\" workflow.\n\n## Context Propagation\n\n```go\n// Bad — breaks trace propagation; the DB call starts a new root trace.\nresult, err := db.Query(\"SELECT ...\")\n\n// Good — the DB span is a child of the HTTP request span.\nresult, err := db.QueryContext(ctx, \"SELECT ...\")\n```\n\nEvery I/O boundary takes ctx: HTTP client, database, gRPC client, message queue. Use `otelhttp.NewTransport` to instrument outbound HTTP and `otelhttp.NewHandler` for inbound.\n\n## Anti-Patterns\n\n| Anti-pattern | Why it hurts | Do this instead |\n|---|---|---|\n| `Summary` for latency | Cannot aggregate across replicas; no quantile flexibility | `Histogram` + `histogram_quantile()` |\n| `userID` as a Prometheus label | Unbounded cardinality → Prometheus OOM | Hash to a bucketed `user_tier` or drop |\n| Logging the same error you return | Duplicate lines, no single source of truth | Return wrapped, log once at the boundary |\n| No `RecordError` on failed spans | Trace shows green, alert fires red | Always pair error returns with `span.RecordError` + `SetStatus` |\n| Trace without context propagation | Each layer starts a new root trace | First argument of every I/O method is `ctx context.Context` |\n| Global metric defined inside a handler | Re-registers on every call → panic | Declare metrics at package level, register once |\n| Histogram with 200 buckets | High storage cost, slow queries | Start with `DefBuckets`, tune from real data |\n\n## Verification Checklist\n\nBefore marking the change done:\n\n- [ ] Each new code path emits at least one counter (operations) and one histogram (latency)\n- [ ] All metric labels are bounded (method, route pattern, status class — not IDs)\n- [ ] Each PromQL query the metric is meant to answer appears as a comment above the metric declaration\n- [ ] Every I/O method takes `ctx` first; downstream calls pass it through\n- [ ] Every service method, DB call, and outbound HTTP call has a span\n- [ ] Failure paths call `span.RecordError(err)` and `span.SetStatus(codes.Error, ...)`\n- [ ] At least one histogram uses `ObserveWithExemplar` for trace correlation\n- [ ] Logs and traces are correlated (see go-logging skill: `otelslog` bridge)\n\n## References\n\n- [references/metrics.md](references/metrics.md) — Counter/Gauge/Histogram, naming, cardinality, PromQL examples\n- [references/tracing.md](references/tracing.md) — TracerProvider, spans, attributes, `otelhttp`, sampling\n- [references/correlation.md](references/correlation.md) — Exemplars, trace_id in logs, end-to-end workflow\n- [references/anti-patterns.md](references/anti-patterns.md) — Detailed walkthrough of each anti-pattern with code\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}