← HoneycombCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to Honeycomb
Snapshot Sep 30, 2026 · 22:51 UTC · version 1.0.0
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "otel-instrumentation",
"description": "Provides guidance on OpenTelemetry SDK setup, custom instrumentation, and sending data to Honeycomb. Trigger phrases: \"instrument my app\", \"add tracing\", \"set up OpenTelemetry\", \"configure OTel\", \"add custom spans\", \"add attributes to spans\", \"send traces to Honeycomb\", \"set up OTLP\", \"configure sampling\", \"add span events\", \"add span links\", \"set up tracing for [any language]\", \"configure the OTel Collector\", or any request about OpenTelemetry SDK setup, custom instrumentation, or sending data to Honeycomb.\n",
"included_files": [
{
"relative_path": "references/architectural-patterns.md",
"size_in_bytes": 8448
},
{
"relative_path": "references/collector-config.md",
"size_in_bytes": 2621
},
{
"relative_path": "references/custom-instrumentation.md",
"size_in_bytes": 18078
},
{
"relative_path": "references/lambda.md",
"size_in_bytes": 14498
},
{
"relative_path": "references/local-collector-debug-test.md",
"size_in_bytes": 3504
},
{
"relative_path": "references/python.md",
"size_in_bytes": 10529
},
{
"relative_path": "references/sdk-setup-by-language.md",
"size_in_bytes": 7168
},
{
"relative_path": "references/wide-event-attributes.md",
"size_in_bytes": 14999
}
],
"skill_md_contents": "---\nname: otel-instrumentation\ndescription: >\n Provides guidance on OpenTelemetry SDK setup, custom instrumentation,\n and sending data to Honeycomb.\n Trigger phrases: \"instrument my app\", \"add tracing\",\n \"set up OpenTelemetry\", \"configure OTel\", \"add custom spans\",\n \"add attributes to spans\", \"send traces to Honeycomb\",\n \"set up OTLP\", \"configure sampling\", \"add span events\",\n \"add span links\", \"set up tracing for [any language]\",\n \"configure the OTel Collector\",\n or any request about OpenTelemetry SDK setup, custom instrumentation,\n or sending data to Honeycomb.\nmetadata:\n version: \"1.0.0\"\n---\n\n# OpenTelemetry Instrumentation for Honeycomb\n\nSDK setup, custom spans, attributes, span events, sampling, and layered telemetry.\nFor conceptual foundations (why wide events matter, how attributes connect to\ninvestigation), see the **observability-fundamentals** skill.\n\n## OTLP Configuration and SDK Setup\n\nEvery OTel SDK needs these environment variables to send data to Honeycomb:\n\n### Required Environment Variables\n\n**Base configuration:**\n```bash\nOTEL_SERVICE_NAME=your-service-name\nOTEL_EXPORTER_OTLP_ENDPOINT=https://api.honeycomb.io\nOTEL_EXPORTER_OTLP_HEADERS=\"x-honeycomb-team=YOUR_API_KEY\"\n```\n\n**Optional but recommended:**\n```bash\n# Protocol selection (default: http/protobuf)\nOTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf # or grpc\n\n# Signal-specific endpoints (override base endpoint for specific signals)\nOTEL_EXPORTER_OTLP_TRACES_ENDPOINT=https://api.honeycomb.io/v1/traces\nOTEL_EXPORTER_OTLP_METRICS_ENDPOINT=https://api.honeycomb.io/v1/metrics\n```\n\n**For metrics (preferred):** Use modern OTLP metrics and native datapoints. Use dataset\nhints to confirm the destination type (`metrics` or `events`). Authenticate with:\n```bash\nOTEL_EXPORTER_OTLP_METRICS_HEADERS=\"x-honeycomb-team=YOUR_API_KEY\"\n```\n\n### Protocol Selection\n\n`OTEL_EXPORTER_OTLP_PROTOCOL` determines the wire format and transport:\n- `http/protobuf` (default, recommended) — HTTP with protobuf encoding\n- `grpc` — gRPC with protobuf encoding\n- `http/json` — HTTP with JSON encoding (larger payload, slower)\n\nUse `http/protobuf` unless you have specific infrastructure requirements for gRPC.\n\n### Signal-Specific Endpoints\n\nBy default, OTel SDKs append `/v1/traces` and `/v1/metrics` to `OTEL_EXPORTER_OTLP_ENDPOINT`.\nUse signal-specific endpoint vars to override:\n- `OTEL_EXPORTER_OTLP_TRACES_ENDPOINT` — full URL for traces (including `/v1/traces`)\n- `OTEL_EXPORTER_OTLP_METRICS_ENDPOINT` — full URL for metrics (including `/v1/metrics`)\n\nUseful when routing signals to different backends or using non-standard endpoints.\n\n### Common Pitfalls\n\n**Silent auth failure:** The OTLP exporters need the `x-honeycomb-team` header to\nauthenticate. Without it, Honeycomb silently rejects requests — no error, no data. Set\n`OTEL_EXPORTER_OTLP_HEADERS=\"x-honeycomb-team=YOUR_API_KEY\"` or pass headers\nprogrammatically. If loading the key from `.env`, ensure dotenv runs before SDK init.\n\n**Metrics:** Prefer modern OTLP metrics and native datapoints. Dataset hints identify the\ndestination type (`metrics` or `events`), so do not add `x-honeycomb-dataset` by default.\nUse that header only when hints or configuration require legacy routing to a named event\ndataset. Traces do not need it; they route by `service.name`.\n\nFor the env var values, language-specific dependencies, and setup code (Go, Python,\nNode.js, Java, Ruby, .NET, Rust), see\n`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/sdk-setup-by-language.md`.\n\n## Custom Instrumentation\n\n### Adding Attributes to Existing Spans (Highest Impact)\n\nAdd business context to auto-instrumented spans — no new spans needed. Get the current\nspan from context and call `SetAttributes` (Go), `set_attribute` (Python), or\n`setAttribute` (Node.js) with user, tenant, business, and deployment context.\n\n### Creating Custom Spans\n\nWrap important business operations for visibility in the trace waterfall. Use\n`tracer.Start(ctx, \"operation-name\")` (Go), `tracer.start_as_current_span(\"operation-name\")`\n(Python), or `tracer.startActiveSpan(\"operation-name\", callback)` (Node.js).\n\nFor full code examples in all languages, consult\n`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/custom-instrumentation.md`.\n\n## When to Create a Span\n\nNot every function needs a span. Two questions determine whether a span is worth creating:\n\n1. **Is it interesting?** — Does the work meaningfully impact performance (latency or\n failures) for the overall request?\n2. **Is it aggregable?** — If you group this span by name and attributes, will it produce\n useful trends and comparisons?\n\n| Operation | Interesting? | Aggregable? | Create a Span? |\n| :--- | :--- | :--- | :--- |\n| HTTP request handler | Yes — variable latency, can fail | Yes — group by route, method, status | **Yes** |\n| Database query | Yes — I/O bound, failure-prone | Yes — group by query type, table | **Yes** |\n| External API call | Yes — network latency, dependencies | Yes — group by endpoint, status | **Yes** |\n| Cache lookup | Yes — fast vs slow path | Yes — group by cache name, hit/miss | **Yes** |\n| Message queue pub/consume | Yes — async boundary, delays | Yes — group by queue, message type | **Yes** |\n| Business logic transaction | Yes — meaningful state change | Yes — group by type, outcome | **Yes** |\n| Private helper function | No — trivial CPU, predictable | No — too granular | **No** |\n| Loop iteration | Maybe — if slow | No — unbounded cardinality | **No** |\n| Getter/setter | No — no meaningful duration | No — nothing to group by | **No** |\n| Input validation (pure CPU) | No — fast, predictable | Maybe | **No** |\n| Business logic orchestration | No — just calls instrumented code | No — duration is sum of children | **No** |\n\n**Common mistakes:**\n- **Too many spans**: A trace with millions of 2ms spans is far too detailed and rarely\n actionable. Roll them up — combine into a single span, or capture the detail as an\n attribute on the parent span instead.\n- **Too few spans**: Collapsing hours of work into a single opaque handler leaves you\n guessing about where time is spent.\n- **Test spans left in**: Spans named `test-span`, `debug-span`, or similar are\n artefacts that pollute the dataset. Remove any span created solely to verify tracing\n is working before finishing.\n\nWhen in doubt, prefer **attributes on existing spans** over creating new child spans.\n\n#### Timing Attributes (measure sub-operations without child spans)\n\nRecord important sub-operation durations as attributes on the parent span. These are\neasier to query than child spans and work directly with BubbleUp.\n\n```go\n// Go: time auth and record on the existing span\nspan := trace.SpanFromContext(r.Context())\nauthStart := time.Now()\nuser, err := authenticate(r)\nspan.SetAttributes(attribute.Float64(\"auth.duration_ms\", float64(time.Since(authStart).Milliseconds())))\n```\n\n```python\n# Python: time auth and record on the existing span\nspan = trace.get_current_span()\nauth_start = time.monotonic()\nuser = authenticate(request)\nspan.set_attribute(\"auth.duration_ms\", (time.monotonic() - auth_start) * 1000)\n```\n\n#### Exception telemetry: event details plus span-level dimensions\n\nUse the Logs API for new exception events. Emit the record while the relevant span is\nactive and include the standard exception fields (`exception.type`, `exception.message`,\n`exception.stacktrace`, and `exception.escaped` when applicable), an ERROR severity, and\n`event.name=\"exception\"`. Set the span status to ERROR separately when the operation failed.\n\nIn Honeycomb, a trace-correlated exception log is rendered in the trace as a `span_event`\nannotation and carries `trace.trace_id` and `trace.parent_id`. Its full `exception.*`\npayload remains on the log-derived event; it is **not hoisted onto the containing span**.\nSearch the exception event row, then follow its trace ID to inspect the surrounding trace.\n\nUse low-cardinality span attributes for aggregation and alerting:\n\n- `error=true` and the span status indicate operation failure.\n- `exception.slug` is a static, greppable identifier for the error site.\n- An optional error category is safer for `GROUP BY` than full exception messages.\n\n```text\nLogs-API exception event: event.name=exception, body=exception, meta.signal_type=log\nLegacy span-event exception: name=exception, meta.signal_type=trace\nBoth may have: meta.annotation_type=span_event\n```\n\n`record_exception` / `RecordError` remain compatibility APIs for existing SDKs and code,\nbut do not use them as the only new guidance when Logs API support is available. They can\nalso produce parent-span exception fields that a Logs-API event alone does not produce.\n\nFind operation failures by span dimensions: `WHERE error = true AND exception.slug does-not-exist`.\nFind Logs-API exception events with `event.name=exception AND exception.type exists` and\nfollow a sampled `trace.trace_id` into `get_trace` with `show_events=true`.\n\nFor extended examples and the MCP investigation recipe, see\n`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/custom-instrumentation.md`.\n\n#### Optional compatibility: promote exception fields with a LogRecordProcessor\n\nIf existing span-level dashboards, alerts, or queries depend on Honeycomb's historical exception\nfield promotion, add a custom **LogRecordProcessor** before the batch/export processor. When it\nsees an exception log, it should use the log record's resolved context to find the active recording\nspan and promote a configured, minimal set of fields such as `error=true`, `error.type`,\n`exception.type`, `exception.slug`, or an error category.\n\nDo not recommend a standalone `SpanProcessor` for this: span processors receive span lifecycle\ncallbacks, not log records. Keep full `exception.message` and `exception.stacktrace` on the Logs\nAPI event by default; copy them onto spans only when legacy query compatibility explicitly requires\nit. The processor must run synchronously while the span context is valid, before the log reaches\nbatch export. It should no-op when there is no recording span and must not infer fields that the\napplication did not put on the log record.\n\nThis is an optional migration layer, not a replacement for querying the Logs API event. Agents\nshould treat span-level promoted fields as instrumentation-dependent and continue to query\n`event.name=exception` event rows for full diagnostics.\n\n## What to Instrument\n\n### High Value (Instrument First)\n- API entry points (HTTP handlers, gRPC methods)\n- Database queries (auto-instrumented by most SDKs)\n- External HTTP calls (auto-instrumented by most SDKs)\n- Message queue producers/consumers\n\nThese are typically auto-instrumented by OTel SDKs and form the skeleton of your traces.\n\n### Medium Value (Add Next)\n- Business logic operations (checkout, payment, fulfillment)\n- Cache operations (hits, misses, evictions)\n- Authentication and authorization checks\n- Background job execution\n\nThese are your business logic. Without custom spans here, you can see that a request was\nslow but not *why* — the trace waterfall has gaps where the important work happens\ninvisibly.\n\n### Attributes to Add\n\nAttributes are the dimensions BubbleUp uses during investigations. Every attribute you\nadd is a new axis BubbleUp can diff on to find what's different about outlier requests.\nFor the complete catalog organized by category with rationale and example queries, see\n`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/wide-event-attributes.md`.\n\nFor why attributes matter conceptually, see the **observability-fundamentals** skill.\n\n## Span Events, Logs API Events, and Span Links\n\n- **Point-in-time events**: Prefer the Logs API for new events, especially exceptions. Emit\n while the span is active so the record carries trace context. In Honeycomb, a correlated\n log is rendered as a `meta.annotation_type=span_event` annotation, but its event name is\n in `event.name` (and often `body`), not `name`.\n- **Legacy span events**: `span.add_event` / `AddEvent` remain valid compatibility paths. Their\n event name is in `name` and their signal type is `trace`.\n- **Span links**: Connect spans across different trace hierarchies (async processing,\n fan-out/fan-in, cross-system correlation). Create a `Link` to the related span context.\n\nFor human instrumentation examples and an agent-safe Honeycomb MCP query → sample → trace\nworkflow, see `${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/custom-instrumentation.md`\nand the **production-investigation** skill.\n\n## Sampling\n\n### Sampling Strategy\n\nSampling is about tradeoffs — there is no free lunch:\n\n- **Head sampling favors cost over debuggability.** You save resources, but a 0.1% error\n at 1% sampling becomes effectively invisible. Head sampling is oblivious to what\n happens downstream.\n- **Tail sampling favors fidelity over simplicity.** You keep interesting traces but need\n infrastructure (Refinery or Collector) to buffer and evaluate complete traces.\n\nThe math matters: if an error occurs 0.1% of the time and you head-sample at 1%, you'll\ncapture roughly 1 in 100,000 of those errors. At moderate traffic, that error may never\nappear in your data.\n\n### Head Sampling (SDK-level)\nDecides whether to sample a trace at creation time. Simple but can miss interesting traces.\n- Configure via `OTEL_TRACES_SAMPLER` env var\n- `always_on` (default), `always_off`, `traceidratio` (e.g., sample 10%)\n- `parentbased_traceidratio` respects parent sampling decisions\n- **Best for:** Very high-throughput services where you can tolerate missing rare events\n\n### Tail Sampling (Collector/Refinery)\nDecides after the trace is complete. Keeps interesting traces (errors, slow requests).\n- Use Honeycomb's **Refinery** for production tail sampling\n- Or configure the OTel Collector's `tail_sampling` processor\n- Can sample based on: latency, error status, specific attributes, trace duration\n- **Best for:** Services where debuggability matters — keeps errors and outliers while\n sampling routine traffic\n\n### Sampling Impact on Honeycomb\n- Sampling reduces data volume and cost\n- SLOs, BubbleUp, and query results adjust for sampling rate automatically\n- Trace completeness may be affected — missing spans if not all services sample consistently\n- Start with no sampling, then add as needed for cost management\n\n## Layered Telemetry\n\nOpenTelemetry is \"trace-first\" — context propagation is the glue that correlates all\nsignals. But effective observability layers multiple signal types for different purposes.\n\nA three-question test for choosing the right signal:\n\n1. **What needs causality and full-request context?** → Traces (spans)\n2. **What needs inexpensive long-term storage and fast alerting?** → Metrics\n3. **What is rare vs. common, and what are the audit requirements?** → Logs / events\n\n**The histogram-alongside-spans pattern:** For high-throughput HTTP services, emit both a\nspan and a histogram metric for each handled request. This lets you head-sample traces\nfor cost while histograms provide last-ditch alerting — and exemplars link outlier metric\npoints back to specific traces for deeper investigation.\n\nThe technique is *layering* (not duplication) because each signal provides a different\nview at a different level of detail.\n\nFor architectural patterns where layering is essential (streaming, async jobs, ETL), see\n`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/architectural-patterns.md`.\n\nFor AWS Lambda-specific patterns — choosing between the AWS Managed OTel Layer\nand manual SDK setup, forceFlush, SDK 2.x setup, cross-Lambda trace propagation,\nheader normalisation, TOKEN vs REQUEST authorizers — see\n`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/lambda.md`.\n\n## Logs in Honeycomb\n\nOTel can send logs too. If you have existing log infrastructure, the OTel Collector can\ningest logs and forward them to Honeycomb as structured events:\n\n- **OTel SDK log bridge**: Captures logs from your existing logging library (`slog` in Go,\n `logging` in Python, `winston`/`pino` in Node.js) and exports them as OTel log records.\n- **OTel Collector `filelog` receiver**: Reads log files, parses them, exports as OTLP.\n\nLogs sent through OTel arrive in Honeycomb as structured events with the same query\ncapabilities as spans.\n\n## Naming Conventions\n\n- **Span names**: Describe the operation (`HTTP GET /api/users`, `db.query SELECT`, `process-payment`)\n- **Attribute names**: Use dot-separated namespaces (`user.id`, `order.total`, `cache.hit`)\n- **Follow OTel semantic conventions** where applicable (`http.method`, `db.system`, `rpc.service`)\n- **Custom attributes**: Use your own namespace (`app.`, `checkout.`, `mycompany.`)\n\n## Additional Resources\n\n### Reference Files\n- **`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/sdk-setup-by-language.md`** — OTLP configuration and SDK setup for Go, Python, Node.js, Java, Ruby, .NET, Rust\n- **`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/local-collector-debug-test.md`** — Run a local OTel Collector via Docker to verify spans, logs, and metrics without a Honeycomb account; includes `jq` commands for inspecting NDJSON output\n- **`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/custom-instrumentation.md`** — Custom instrumentation patterns with full code examples (timing attributes, exception slugs, async request summaries)\n- **`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/collector-config.md`** — OTel Collector configuration for format conversion, processing, and sampling\n- **`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/wide-event-attributes.md`** — Canonical attribute catalog organized by category with example queries\n- **`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/architectural-patterns.md`** — Trace design patterns for streaming, async, ETL, and serverless architectures\n- **`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/lambda.md`** — AWS Lambda: OTel Layer vs manual SDK setup trade-offs, forceFlush and per-request latency, SDK 2.x setup, cross-Lambda trace propagation, header normalisation, TOKEN vs REQUEST authorizer migration\n\n### Cross-References\n- For conceptual foundations of why wide events and attributes matter: **observability-fundamentals** skill\n- After instrumenting, use the **query-patterns** skill to verify data is arriving\n"
}SHA-256: aef512167acd5204d5c07cd5d210c27a2ce416d38da35ee3be42ce8db20c0ba3