← Plugin catalog
Developer Tools
Honeycomb
honeycomb.io v1.0.0
Publisher description
From the marketplace listing
Honeycomb helps users investigate observability data, inspect traces and AI conversations, run queries and BubbleUp analyses, and manage boards, SLOs, triggers, markers, and notification recipients through ChatGPT.
Language: English · Automatically detected from descriptions.
Files & skills
File archives
Plugin package2 files · 685 BytesBrowse files →
beeline-migration3 files · 3.72 KBBrowse files →
create-honeycomb-board3 files · 6.97 KBBrowse files →
metrics-queries4 files · 10.1 KBBrowse files →
observability-fundamentals2 files · 5.63 KBBrowse files →
otel-genai-instrumentation8 files · 25 KBBrowse files →
otel-instrumentation9 files · 32.2 KBBrowse files →
otel-migration6 files · 15.7 KBBrowse files →
production-investigation4 files · 8.79 KBBrowse files →
query-patterns6 files · 13.7 KBBrowse files →
slos-and-triggers4 files · 6.93 KBBrowse files →
verify-recent-trace1 files · 1.39 KBBrowse files →
Skill instructions
beeline-migration4.77 KB
---
name: beeline-migration
description: >
Step-by-step guide for migrating from Honeycomb Beelines (End of Life) to
OpenTelemetry instrumentation.
Trigger phrases: "migrate from Beelines",
"upgrade from Beeline to OpenTelemetry", "migrate to OTel", "replace Beelines",
"Beeline end of life", "Beeline EOL", "switch from Beeline to OTel",
"migrate Go Beeline", "migrate Python Beeline", "migrate Node Beeline",
"migrate Java Beeline", "migrate Ruby Beeline", "W3C trace headers",
"W3C propagation", "incremental migration to OpenTelemetry",
or any request about migrating from Honeycomb Beelines to OpenTelemetry SDKs.
metadata:
version: "1.0.0"
---
# Beeline to OpenTelemetry Migration
Step-by-step guide for migrating from Honeycomb Beelines (now End of Life)
to OpenTelemetry instrumentation.
## Status
Honeycomb Beelines have reached **End of Life** and are **archived**. All new
instrumentation should use OpenTelemetry. Existing Beeline users should migrate
as soon as practical.
## Migration Strategy
Migration follows a two-phase approach that allows incremental, service-by-service
migration without breaking distributed traces.
### Phase 1: Enable W3C Trace Propagation (All Services)
Before migrating any service to OTel, **all** services must support W3C trace
headers. This enables Beeline and OTel services to share trace context.
1. Upgrade each Beeline to the minimum version supporting W3C headers
2. Configure each Beeline to use W3C propagation format
3. Deploy all services with W3C enabled
4. **Verify**: Traces still link correctly across services
**Minimum Beeline versions for W3C support:**
| Language | Minimum Version |
|----------|----------------|
| Go | 1.4.0 |
| Java | 1.7.0 |
| Node.js | 3.2.2 |
| Python | 2.18.0 |
| Ruby | 2.8.0 |
### Phase 2: Migrate Each Service to OTel (One at a Time)
After all services support W3C headers:
1. Choose a service to migrate (start with leaf services — fewest dependencies)
2. Replace Beeline SDK with OpenTelemetry SDK
3. Configure OTLP exporter to point to Honeycomb
4. Add auto-instrumentation libraries
5. Replicate any custom Beeline instrumentation in OTel
6. Deploy and verify traces still connect
7. Repeat for next service
**Key rule**: Complete Phase 1 across ALL services before starting Phase 2 on ANY service.
## W3C Propagation Configuration
### Go Beeline
```go
beeline.Init(beeline.Config{
HTTPPropagationHook: propagation.W3C,
})
```
### Python Beeline
```python
beeline.init(
http_trace_propagation_hook=beeline.propagation.w3c.http_trace_propagation_hook,
http_trace_parser_hook=beeline.propagation.w3c.http_trace_parser_hook,
)
```
### Node.js Beeline
```javascript
const beeline = require("honeycomb-beeline")({
httpTraceParserHook: beeline.w3c.httpTraceParserHook,
httpTracePropagationHook: beeline.w3c.httpTracePropagationHook,
});
```
For Java and Ruby configurations, consult `${CLAUDE_PLUGIN_ROOT}/skills/beeline-migration/references/w3c-propagation.md`.
## Service Migration Checklist
For each service being migrated from Beeline to OTel:
- [ ] Beeline version supports W3C (Phase 1 complete)
- [ ] Install OTel SDK and OTLP exporter packages
- [ ] Configure OTLP endpoint and headers for Honeycomb
- [ ] Set `OTEL_SERVICE_NAME` to match existing service name
- [ ] Add auto-instrumentation libraries (HTTP, DB, etc.)
- [ ] Port custom spans: Beeline `startSpan()` -> OTel `tracer.start_span()`
- [ ] Port custom attributes: Beeline `addField()` -> OTel `span.set_attribute()`
- [ ] Remove Beeline dependency
- [ ] Deploy and verify: traces link across Beeline and OTel services
- [ ] Verify: custom attributes appear in Honeycomb
## Migration Safety Checklist
- **Complete Phase 1 across all services before starting Phase 2** — mixed propagation formats break trace linking across service boundaries
- **Keep `OTEL_SERVICE_NAME` identical to the Beeline service name** — Honeycomb uses this as the dataset name, and changing it splits your data into a new dataset
- **Audit all Beeline `addField()` calls before removing the Beeline SDK** — each one needs a corresponding `span.set_attribute()` in OTel to preserve your query dimensions
- **Compare OTel auto-instrumentation field names against Beeline field names** — OTel may use different attribute names (e.g., `http.request.method` vs `request.method`), and dashboards or SLIs referencing the old names will need updating
## Additional Resources
### Reference Files
- **`${CLAUDE_PLUGIN_ROOT}/skills/beeline-migration/references/migration-steps-by-language.md`** — Detailed migration code for each language
- **`${CLAUDE_PLUGIN_ROOT}/skills/beeline-migration/references/w3c-propagation.md`** — Complete W3C configuration for all Beeline languages
### Cross-References
- For OTel SDK setup after migration, see the **otel-instrumentation** skill
Referenced files: 2
create-honeycomb-board7.23 KB
---
name: create-honeycomb-board
description: >
Design and then create a board (dashboard) in Honeycomb with queries and SLOs.
Trigger phrases: "create a board", "make a board", "build a dashboard",
"create a Honeycomb board", "make a dashboard in Honeycomb",
"set up a board", "dashboard for my service", "visualize service health",
"golden signals dashboard", "set up monitoring board",
or any request to design and create or build a Honeycomb board or dashboard.
metadata:
version: "1.0.0"
allowed-tools:
- mcp__honeycomb__create_board
- mcp__honeycomb__run_query
- mcp__honeycomb__get_query_results
- mcp__honeycomb__get_workspace_context
- mcp__honeycomb__get_environment
- mcp__honeycomb__get_dataset
- mcp__honeycomb__get_dataset_columns
- mcp__honeycomb__find_columns
- mcp__honeycomb__find_queries
- mcp__honeycomb__get_slos
- mcp__honeycomb__get_triggers
- mcp__honeycomb__list_boards
- Read
- Grep
- Glob
- AskUserQuestion
---
# Create a Honeycomb Board
Build a board (dashboard) in Honeycomb using the `create_board` MCP tool.
There is no update tool — define it well before creating.
When building a board, think about the purpose and time frame involved. Some examples:
- a board for a service. This should be timeless, looking at the service's health, performance, and business metrics. Do not do any problem diagnosis or investigation when building this board. Do not express opinions or summarize graphs in text panels. The board should be a representation of the service's health at whatever moment someone looks at it.
- a board for a feature. This should look at the feature's usage trends, and its impact on the business. This board would have a time frame of 7 days, and would not include any infrastructure metrics or service dependencies.
- a board for a problem. This might be created during an incident, or afterward. This one would have a time frame specific to the incident. It would include investigations, and your opinions about what is happening. Patterns of what to look for are appropriate here.
## Workflow
### Gather SLOs
Use `get_slos` to list SLOs in the environment. Relevant SLOs will go on the board as `slo` panels.
### Gather descriptive context
Look at the code and docs (using Read/Grep/Glob) to understand the service or feature. Use this to write a text panel. Link to GitHub or documentation if you can.
This will vary greatly depending on the purpose of the board.
Description of the application and link to the code - great for a service board.
What is the feature, and what business impact does it have? - great for a feature board.
What is the problem, and what is the impact? What patterns do we see? - great for a problem board.
### Build candidate queries
**Get Honeycomb context**: Call `get_workspace_context` for available environments. Use `get_environment` to find datasets — each dataset corresponds to an `OTEL_SERVICE_NAME`.
**Code context**: Look at the language and any custom attributes (often prefixed `app.`). Custom fields are prime candidates for breakdowns and business metrics.
**Discover columns**: Use `find_columns` or `get_dataset_columns`. Pay special attention to non-standard columns — those are specific to the application.
**Find queries**: Use `find_queries` to see what people have already queried. Check `get_triggers` for what they alert on — those indicate what matters.
**Time range**: Default to 2 hours. Use 8–24 hours for lower-volume services. Use 7 days for feature usage boards. Keep it consistent across all panels.
Aim for 6–12 graphs. Stat panels are bonus — they don't count against this limit.
List your candidate queries for the user with reasons before running them. See `${CLAUDE_PLUGIN_ROOT}/skills/create-honeycomb-board/references/board-queries.md` for what to include and how to write each kind.
### Run and check queries
Use `run_query` for each candidate query. Each result returns a query run PK (like `QR-abc123`) — you'll use this as the panel `id` when building the board.
Fix errors. Eliminate queries that return no interesting results.
### Plan the layout
Think about visual flow and how to make the board expressive. The board uses a 12-column grid — stat panels can sit side-by-side, a heatmap deserves full width, breakdowns benefit from extra height. See `${CLAUDE_PLUGIN_ROOT}/skills/create-honeycomb-board/references/board-layout.md` for sizing examples.
Consider `preset_filters` if viewers will want to slice the board interactively (by route, region, account tier, etc.).
### Show the proposed board to the user — always, without exception
For each panel, display:
- **Text panels**: Show the **full markdown content** that will appear on the board. The user needs to review the exact wording before creation since boards can't be updated.
- **SLO panels**: Show the SLO name, target, and current compliance.
- **Query panels**: Show the name, description, chart type, display style, and a **link to the query** (the `query_url` from the run_query result metadata). Briefly describe what the results showed.
- Planned layout (sizing and groupings)
- **Tags**: Display the tags you plan to add to the board.
- Any preset filters
End with: "Here's the board I'd create. Shall I go ahead?"
**This step is non-negotiable.** If the user says "just create it", "I trust you", or "skip the preview" — show the plan anyway. The user cannot meaningfully confirm something they haven't seen. There is no way to update a board after creation; the only fix is to delete it and start over. Showing the plan first protects them even when they think they don't need it.
The one exception: if the user has already reviewed and approved a specific plan in this conversation, you may proceed.
### Create the board
Call `create_board` with a `panels` array. Every panel requires a `type` field — `"query"`, `"slo"`, or `"text"` — that determines what other fields apply. See `${CLAUDE_PLUGIN_ROOT}/skills/create-honeycomb-board/references/board-layout.md` for the full field reference.
```json
{
"environment_slug": "production",
"name": "Checkout Service",
"description": "...",
"panels": [
{
"type": "text",
"content": "## Checkout Service\nOwned by Platform team. [Source](https://github.com/...)"
},
{
"type": "slo",
"id": "SLO-abc123",
"size": { "width": 4 }
},
{
"type": "query",
"id": "QR-abc123",
"name": "Request Rate",
"chart_type": "stat",
"display_style": "chart",
"size": { "width": 4 }
},
{
"type": "query",
"id": "QR-def456",
"name": "Error Rate",
"chart_type": "stat",
"display_style": "chart",
"size": { "width": 4 }
},
{
"type": "query",
"id": "QR-ghi789",
"name": "Latency Distribution",
"description": "Overall request latency as a heatmap",
"chart_type": "default",
"display_style": "chart",
"size": { "width": 12, "height": 3 }
}
],
"preset_filters": [{ "column": "http.route", "alias": "Route" }],
"tags": ["team:platform", "tier:critical"]
}
```
### Follow up
Link the user to the board.
## Cross-References
- For query construction patterns and calculated fields: **query-patterns** skill
- For SLO interpretation and burn alert design: **slos-and-triggers** skill
Referenced files: 2
metrics-queries13.4 KB
---
name: metrics-queries
description: >
How to query OpenTelemetry metrics datasets in Honeycomb correctly. Metrics
datasets follow different rules from trace/event datasets — many operations
(bare COUNT, RATE_SUM, RATE_AVG, RATE_MAX, CONCURRENCY) are forbidden,
temporal aggregation is automatic, and each metric has its own attributes.
Use this skill when querying a metrics dataset (gauges, counters, histograms,
sums), asking about temporal aggregation (RATE, INCREASE, SUMMARIZE, LAST),
finding the metrics dataset or discovering metric names and attributes,
debugging unexpected metrics query results, or querying infrastructure
metrics like CPU, memory, disk I/O, or network stats. Do NOT use for
instrumenting metrics (use otel-instrumentation), querying event datasets
with "metrics" in their name, or conceptual questions (use
observability-fundamentals).
metadata:
version: "1.0.0"
---
# Querying Metrics in Honeycomb
Metrics datasets in Honeycomb behave differently from tracing/event datasets.
Operations that work on traces may fail or produce misleading results on metrics.
This skill covers those differences so you construct correct, useful metrics queries.
## Finding the Metrics Dataset
Metrics datasets are **not** identified by having "metrics" in their name. Many event
datasets contain "metrics" in their slug (e.g., `kafka-metrics`, `refinery-metrics`,
`kubernetes-node-metrics`). These are ordinary event datasets, not metrics datasets.
**How to identify the real metrics dataset:**
1. Call `get_environment` and look for rows where `dataset_type` = **`metrics`**.
The slug is typically `metrics` but may differ per environment.
2. Alternatively, call `get_dataset_columns` on a candidate dataset — metrics datasets
return a **`MetricInfo`** column showing type metadata like `gauge`,
`sum(cumulative,monotonic)`, or `histogram(delta)`. Event datasets do not have this.
**Do not guess the dataset.** Always verify via `get_environment` or `get_dataset_columns`
before constructing a metrics query. If the user says "metrics" but means an event dataset
with metrics-like fields (e.g., `telegraf`, `system_stats`), the query rules below do not apply —
those are event datasets and follow normal query patterns from the **query-patterns** skill.
## Discovering Metrics and Their Attributes
Metrics datasets have a fundamentally different schema from event datasets. Each metric
has its own set of resource and data point attributes. Two metrics in the same dataset
may have completely different attributes available for filtering and grouping.
**Workflow for discovering what to query:**
1. **Find metric names:** Call `get_dataset_columns` on the metrics dataset (without
`metric_name`). This returns metric names with their types in `MetricInfo`.
Use `find_columns` with keywords to search for specific metrics (e.g., "cpu", "memory",
"http request duration").
2. **Find attributes for a specific metric:** Call `get_dataset_columns` with the
`metric_name` parameter set to the metric you want to query (e.g.,
`metric_name: "k8s.pod.memory.usage"`). This returns the resource attributes and
data point attributes that co-occur with that metric, along with sample values.
These are what you can use in WHERE and GROUP BY clauses.
3. **Validate before querying:** Not all attributes exist on all metrics. Always use
step 2 to confirm available attributes before adding them to filters or breakdowns.
## Allowed vs. Forbidden Operations on Metrics Datasets
The following operations are **NOT allowed** on metrics datasets:
| Forbidden Operation | Why |
|---------------------|-----|
| `COUNT` (without column) | Counts metric events, not metric values — meaningless for metrics |
| `RATE_SUM` | Not supported on metrics datasets |
| `RATE_AVG` | Not supported on metrics datasets |
| `RATE_MAX` | Not supported on metrics datasets |
| `CONCURRENCY` | Requires span duration; metrics have no duration |
**Use these instead:**
| Goal | Use on Metrics |
|------|----------------|
| Visualize a gauge value | `AVG(metric)`, `MAX(metric)`, `HEATMAP(metric)` |
| Visualize a counter/sum | `SUM(metric)`, `AVG(metric)`, `MAX(metric)` |
| See distribution of values | `HEATMAP(metric)`, `P50(metric)`, `P99(metric)` |
| Track per-second rate of change | Override temporal aggregation with a calculated field (see below) |
| Percentile analysis | `P50(metric)`, `P90(metric)`, `P99(metric)` |
| Count of non-null values | `COUNT(metric)` (with a column specified) |
## Metric Types and Temporal Aggregation
Honeycomb automatically applies temporal aggregation to align raw metric values into
query time steps. The function it applies depends on the metric type, visible in the
`MetricInfo` column from `get_dataset_columns`.
### Default Temporal Aggregation by Metric Type
| MetricInfo | Type | Default Function | What It Does |
|------------|------|-----------------|--------------|
| `gauge` | Gauge | `LAST()` | Returns most recent value per time step |
| `sum(cumulative,monotonic)` | Monotonic cumulative sum | `INCREASE()` | Change between steps, handles counter resets |
| `sum(cumulative)` | Non-monotonic cumulative sum | `LAST()` | Most recent value (can go up or down) |
| `sum(delta)` or `sum(delta,monotonic)` | Delta sum | `SUMMARIZE()` | Sums all values in each step |
| `histogram(cumulative)` | Cumulative histogram | `INCREASE()` | Change per bucket between steps |
| `histogram(delta)` | Delta histogram | `SUMMARIZE()` | Sums bucket values in each step |
These defaults are applied automatically — you do not need to configure them.
The results you see from `AVG`, `MAX`, `P99`, etc. on a metrics dataset already
reflect temporal aggregation having been applied first.
### Overriding Temporal Aggregation
To override the default (e.g., to see RATE instead of INCREASE for a cumulative counter),
use a **query-scoped calculated field** wrapping the metric name in a temporal aggregation
function, then apply a spatial aggregation to that field in `calculations`.
```json
{
"calculated_fields": [
{ "name": "req_rate", "expression": "RATE($http.server.requests, 300)" }
],
"calculations": [
{ "op": "AVG", "column": "req_rate" }
]
}
```
Supported temporal aggregation functions for calculated fields:
- **`LAST($metric)`** — most recent data point per step (gauges, non-monotonic sums)
- **`SUMMARIZE($metric)`** — sum all values per step with interpolation (delta metrics)
- **`INCREASE($metric[, range_interval_seconds])`** — change in value across range, handles counter resets
- **`RATE($metric[, range_interval_seconds])`** — per-second rate of change (`INCREASE / time`)
The optional `range_interval_seconds` parameter (integer, in seconds) controls the lookback
window for calculating changes. Use it to smooth results or compensate for sparse data.
When omitted, the query's granularity is used as the range interval.
**Important:** You must still apply a spatial aggregation (`AVG`, `SUM`, `P99`, `HEATMAP`, etc.)
to the calculated field in `calculations`. The temporal aggregation function alone does not
produce a visualization — it transforms the raw metric values, then the spatial aggregation
summarizes across timeseries.
For detailed reference on temporal aggregation functions, counter reset handling, and
`range_interval_seconds`, see:
`${CLAUDE_PLUGIN_ROOT}/skills/metrics-queries/references/temporal-aggregation.md`
## Querying Histogram Metrics
OpenTelemetry histograms are stored as a collection of sub-fields. For a histogram
named `http.server.duration`, Honeycomb creates:
| Field | Meaning |
|-------|---------|
| `http.server.duration.count` | Total number of data points |
| `http.server.duration.sum` | Sum of all values |
| `http.server.duration.avg` | Mean value (sum/count) |
| `http.server.duration.p50` | Median (50th percentile) |
| `http.server.duration.p99` | 99th percentile |
| `http.server.duration.p001` through `.p999` | Full range of percentiles |
**Two ways to query histograms:**
1. **Use the parent column name directly** with percentile or distribution operations.
This is the recommended approach:
```json
{ "op": "P99", "column": "http.server.duration" }
```
```json
{ "op": "HEATMAP", "column": "http.server.duration" }
```
2. **Use sub-fields with MAX** when you want the worst-case pre-computed percentile
across all timeseries in a step:
```json
{ "op": "MAX", "column": "http.server.duration.p99" }
```
This returns the highest p99 value reported by any single timeseries in the time step,
which differs from `P99(http.server.duration)` which computes the 99th percentile
across all data.
**When to use which:**
- For most analysis: use `P99(parent_column)` or `HEATMAP(parent_column)`
- For worst-case bounds across hosts/pods: use `MAX(parent_column.p99)`
- For throughput from histograms: use `SUM(parent_column.count)` or `AVG(parent_column.count)`
## Query Math with Metrics
Query math (compound queries with named calculations and formulas) works on metrics
datasets the same way it works on event datasets. Name your calculations, add
per-calculation filters if needed, and define formulas to combine them.
**Common metrics formula patterns:**
### Utilization percentage
```json
{
"calculations": [
{ "op": "AVG", "column": "k8s.pod.memory.usage", "name": "used" },
{ "op": "AVG", "column": "k8s.pod.memory.available", "name": "available" }
],
"formulas": [
{ "name": "utilization_pct", "expression": "$used / ($used + $available) * 100" }
],
"breakdowns": ["k8s.pod.name"],
"orders": [{ "column": "utilization_pct", "order": "descending" }],
"limit": 20
}
```
### Histogram tail ratio
```json
{
"calculations": [
{ "op": "P50", "column": "http.server.duration", "name": "median" },
{ "op": "P99", "column": "http.server.duration", "name": "tail" }
],
"formulas": [
{ "name": "tail_ratio", "expression": "$tail / $median" }
],
"breakdowns": ["service.name"]
}
```
### Error rate from counters (with temporal aggregation override)
```json
{
"calculated_fields": [
{ "name": "error_rate", "expression": "RATE($http.server.errors)" },
{ "name": "request_rate", "expression": "RATE($http.server.requests)" }
],
"calculations": [
{ "op": "SUM", "column": "error_rate", "name": "errors_per_sec" },
{ "op": "SUM", "column": "request_rate", "name": "requests_per_sec" }
],
"formulas": [
{ "name": "error_pct", "expression": "$errors_per_sec / $requests_per_sec * 100" }
]
}
```
For more query examples, see:
`${CLAUDE_PLUGIN_ROOT}/skills/metrics-queries/references/metrics-query-examples.md`
## Granularity for Metrics
Metrics arrive at known, regular intervals (e.g., every 10s, 30s, or 60s). Granularity
matters more for metrics than for traces:
- **Align granularity with the reporting interval.** If metrics report every 60 seconds,
use a granularity that divides evenly into 60 (e.g., 60, 120, 300). Misaligned
granularity causes uneven bucket sizes that produce noisy results.
- **Spiky-looking graphs** usually mean the granularity is finer than the reporting interval.
Increase granularity or, in the UI, enable "Omit Missing Values" to produce continuous lines.
- **RATE operations and granularity:** `RATE_SUM` (on event datasets) is particularly sensitive
to granularity choice — inconsistent data points per bucket produce variable results.
## Common Pitfalls
1. **Using `COUNT` on metrics.** `COUNT` counts the number of metric *events*, not the metric
value. Use `AVG`, `SUM`, `MAX`, or `HEATMAP` instead.
2. **Using `RATE_AVG`/`RATE_SUM`/`RATE_MAX` on metrics datasets.** These are not allowed.
To get a rate, use a calculated field with `RATE($metric)` and then apply a spatial
aggregation like `AVG` or `SUM`.
3. **Assuming all metrics share the same attributes.** Each metric has its own set of
resource and data point attributes. Always call `get_dataset_columns` with `metric_name`
to discover what's available for a specific metric before adding filters or breakdowns.
4. **Confusing event datasets with the metrics dataset.** Datasets named `kafka-metrics`,
`refinery-metrics`, etc. are event datasets. Check `dataset_type` from `get_environment`.
5. **Querying histogram sub-fields when the parent column works.** Use `P99(http.server.duration)`
rather than `AVG(http.server.duration.p99)` unless you specifically need worst-case bounds.
6. **Not specifying an aggregate function.** Metrics queries without a spatial aggregation
in SELECT default to `COUNT`, which is meaningless for metrics.
## Additional Resources
### Reference Files
- **`${CLAUDE_PLUGIN_ROOT}/skills/metrics-queries/references/metrics-query-examples.md`** — Metrics query cookbook with run_query examples for common scenarios
- **`${CLAUDE_PLUGIN_ROOT}/skills/metrics-queries/references/temporal-aggregation.md`** — Deep reference on temporal aggregation functions, counter resets, and range_interval_seconds
- **`${CLAUDE_PLUGIN_ROOT}/skills/metrics-queries/references/metric-types.md`** — OpenTelemetry metric types, how they map to Honeycomb, and what the MetricInfo values mean
### Cross-References
- For general query construction patterns (calculated fields, relational fields, result interpretation): **query-patterns** skill
- For investigating production issues using metrics alongside traces: **production-investigation** skill
- For SLO interpretation and burn alert design: **slos-and-triggers** skill
- For instrumenting applications to send metrics: **otel-instrumentation** skill
Referenced files: 3
observability-fundamentals6.77 KB
---
name: observability-fundamentals
description: >
First principles behind observability — wide events, high cardinality, the core
analysis loop, events vs metrics vs logs, and how instrumentation connects to
debugging outcomes. Grounds recommendations in first principles rather than
tool-specific how-to.
Trigger phrases: "what is observability", "why observability", "why Honeycomb",
"events vs metrics vs logs", "events vs metrics", "events vs logs",
"metrics vs logs", "why wide events", "what is high cardinality",
"core analysis loop", "observability vs monitoring", "what is dimensionality",
"explain observability", or any conceptual question about observability
or why Honeycomb's approach differs from traditional monitoring.
metadata:
version: "1.0.0"
---
# Observability Fundamentals
First principles behind Honeycomb's approach to observability. Use this to ground
recommendations and answer conceptual questions — for SDK setup and tool-specific
guidance, see the **otel-instrumentation** and **query-patterns** skills.
## Definitions
**Observability**: The ability to understand and explain any state your system can
get into, no matter how novel or complex — by examining what the system produces,
without deploying new code for each new question.
**Wide event**: A flat key-value record capturing the full context of a unit of work —
who made the request, which endpoint, cache hit/miss, build version, duration, error
status, and any business context relevant to the operation. In OpenTelemetry, a **span**
is a wide event.
**High cardinality**: The number of unique values a field can have. `user.id` with
millions of values is high cardinality. `http.method` with a handful is low cardinality.
**High dimensionality**: The number of distinct fields on your events. A span with
50 attributes has high dimensionality.
| Concept | Observability | Traditional Monitoring |
|---|---|---|
| Questions | Arbitrary, unknown ahead of time | Pre-defined (dashboards, alerts) |
| Data shape | Decided at query time | Decided at instrumentation time |
| Cardinality | High cardinality is valuable | High cardinality is expensive |
| Investigation | Explore → narrow → confirm | Check dashboard → escalate |
## Why Wide Events
The shape of the data you collect constrains the questions you can ask later. Metrics
pre-aggregate context away at instrumentation time. Wide events preserve context and
let you decide the shape of your analysis at query time.
Every attribute on a span is a queryable dimension. Adding `user.id`, `deployment.version`,
and `cache.hit` to the same span lets you correlate them in a single query — "slow
requests are from tenant X on version 2.3.1 with cache misses." Separate metrics can't
do this because each dimension combination creates a new time series.
Honeycomb's storage engine handles high cardinality and dimensionality without the
cost explosion that affects metrics systems. Adding a high-cardinality field like
`user.id` doesn't create millions of time series — it's another column on each event,
aggregated at query time.
## Events vs Metrics vs Logs
| | Structured Events (Spans) | Metrics | Logs |
|---|---|---|---|
| **Captures** | Full request context (all attributes) | Pre-aggregated numbers with low-cardinality tags | Text or structured fields per line |
| **Discards** | Nothing — raw events retained | Individual requests, high-cardinality dimensions | Correlation across lines (without trace context) |
| **Query power** | GROUP BY, filter, BubbleUp on any dimension | Fast aggregates on pre-defined dimensions | Text search, structured field queries |
| **Cost scaling** | Linear with event volume | Exponential with dimension count (cardinality) | Linear with volume, query cost varies |
| **Best for** | Investigation, root cause analysis | Cheap alerting, long-term trends | Audit trails, rare events |
The same instrumentation effort that produces a metric or log line can produce a wide
event — and the event gives you all three capabilities: count it (metric), read it (log),
analyze it across dimensions (observability).
For code examples showing the same operation instrumented three ways, see
`${CLAUDE_PLUGIN_ROOT}/skills/observability-fundamentals/references/events-vs-metrics-vs-logs.md`.
## The Core Analysis Loop
Debugging in Honeycomb follows a loop: **Define → Visualize → Investigate → Evaluate**.
1. **Define** — Frame the question. Start from an alert, SLO budget burn, or user report.
2. **Visualize** — Run a query to see the shape of the problem (HEATMAP, COUNT, P99).
3. **Investigate** — Narrow down with BubbleUp (automated outlier-vs-baseline comparison
across all dimensions) and trace analysis.
4. **Evaluate** — Confirm the hypothesis by querying with and without the suspected cause.
Then loop — each answer raises new questions. BubbleUp automates steps 2-3 by comparing
distributions across every column, but it only works if events have enough dimensions
to diff on.
For the structured workflow that implements this loop with Honeycomb's tools, see the
**production-investigation** skill.
## Instrumentation Connects to Investigation
Every attribute on a span is a dimension BubbleUp can use to find root causes. The
attributes that matter most during incidents answer three questions:
- **Who is affected?** — user, tenant, account tier, region
- **What changed?** — deployment version, feature flag, config version
- **Where is the bottleneck?** — business operation spans, timing breakdowns, cache state
Instrument for the questions you'll ask at 3am, not for completeness. If BubbleUp
returns nothing useful during an investigation, the issue is usually an instrumentation
gap — add the missing dimensions and try again.
For the complete attribute catalog, see
`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/wide-event-attributes.md`.
For SDK guidance on adding attributes, see the **otel-instrumentation** skill.
## Instrumentation as a Development Practice
Instrumentation is not a one-time setup task. The engineers who write the code are best
positioned to know which operations are critical, which paths are error-prone, and what
context helps during debugging. Treat instrumentation like testing: plan telemetry when
planning features, review it in code reviews, and add missing dimensions as post-incident
follow-ups.
## Additional Resources
### Reference Files
- **`${CLAUDE_PLUGIN_ROOT}/skills/observability-fundamentals/references/events-vs-metrics-vs-logs.md`** — Code examples: same operation as event, metric, and log
### Cross-References
- For SDK setup and custom instrumentation: **otel-instrumentation** skill
- For the investigation workflow implementing the core analysis loop: **production-investigation** skill
- For autonomous instrumentation gap analysis: **instrumentation-advisor** agent
Referenced files: 1
otel-genai-instrumentation30.4 KB
---
name: otel-genai-instrumentation
description: >
Guides instrumentation of GenAI/LLM applications with OpenTelemetry
for Honeycomb, including content capture and agent failure detection.
Trigger phrases: "instrument my GenAI app", "add tracing to LLM calls",
"trace AI agent", "instrument OpenAI", "instrument Anthropic",
"GenAI observability", "trace tool calling", "LLM token usage",
"instrument embeddings", "trace MCP", "GenAI metrics",
"instrument LangChain", "add GenAI spans", "capture prompts",
"capture LLM responses", "enable GenAI content capture",
"streaming tracing", "trace streaming responses",
or any request about instrumenting GenAI/LLM applications.
metadata:
version: "1.0.0"
semconv_version: "v1.40.0"
---
# GenAI Instrumentation for Honeycomb
Instrumenting LLM and agent applications using OTel Semantic Conventions for GenAI
(currently v1.40.0, Development status). For conceptual foundations, see
the **observability-fundamentals** skill.
## Base OTEL Setup (Required First)
**BEFORE implementing GenAI instrumentation, ensure your base OpenTelemetry configuration is complete.**
Use the **otel-instrumentation** skill to configure all standard OTEL environment variables
(OTEL_SERVICE_NAME, OTEL_EXPORTER_OTLP_ENDPOINT, OTEL_EXPORTER_OTLP_HEADERS, OTEL_EXPORTER_OTLP_PROTOCOL,
signal-specific endpoints, etc.) and verify basic spans are flowing to Honeycomb.
GenAI instrumentation adds GenAI-specific configuration on top of that base setup.
## Critical Requirements (Non-Negotiable)
**BEFORE implementing any GenAI instrumentation, complete these steps in order:**
### Step 1: Ask About Content Capture (FIRST!)
**Stop and ask the user this question BEFORE writing any code or configuration:**
> "Do you want to capture the actual prompts and model responses in your traces?
>
> **Enabling content capture:**
> - ✅ Helps debug tool call failures, planning loops, and agent deadlocks
> - ✅ Lets you see why the model made specific decisions
> - ❌ Captures potentially sensitive content (user prompts, model responses)
> - ❌ May contain PII, proprietary data, or confidential information
>
> **Recommended for:** debugging/development, non-sensitive data, or if you have filtering
>
> **Not recommended for:** production with sensitive data, PII/health/financial info"
**Record their answer** — you'll need it when configuring instrumentation.
### Step 2: Enable GenAI Conventions (REQUIRED)
```bash
export OTEL_SEMCONV_STABILITY_OPT_IN=gen_ai_latest_experimental
```
Without this, GenAI spans will not be created.
### Step 3: Set Required Attributes on EVERY Span (REQUIRED)
- `gen_ai.operation.name` — e.g., `chat`, `execute_tool`, `invoke_agent`
- `gen_ai.conversation.id` — same value for all spans in a conversation
**Impact if missing**: Spans won't be recognized as GenAI operations and cannot be queried by session.
### Step 4: Implement force_flush() (REQUIRED)
GenAI apps often exit early (crash, Ctrl+C, CLI). Force flush after each top-level invocation
to prevent silent span loss.
For OTLP configuration, environment variables, and Honeycomb authentication (including the
silent-rejection pitfall), see the **otel-instrumentation** skill.
## Prerequisites
**This skill assumes your agent application is already sending telemetry to Honeycomb.** You should have:
- OpenTelemetry SDK installed and initialized
- All standard OTEL environment variables configured (see **Base OTEL Setup** section above)
- OTLP exporter configured with your Honeycomb API key
- Basic spans flowing to Honeycomb
**If you haven't set this up yet, use the otel-instrumentation skill first** for:
- SDK setup and dependencies
- OTEL environment variables (OTEL_SERVICE_NAME, OTEL_EXPORTER_OTLP_*, etc.)
- OTLP configuration and Honeycomb authentication
- Verification that spans are flowing
Once base telemetry is working, return here to add GenAI-specific instrumentation.
## Auto-Instrumentation (Python and Node.js)
Python and Node.js have official OTel auto-instrumentation packages for GenAI providers.
Go, Java, etc. require manual instrumentation (section below).
### Python
| Package | Provider | Min SDK Version |
| :--- | :--- | :--- |
| `opentelemetry-instrumentation-openai-v2` | OpenAI | openai >= v1.26.0 |
| `opentelemetry-instrumentation-anthropic` | Anthropic | anthropic >= v0.16.0 |
| `opentelemetry-instrumentation-claude-agent-sdk` | Claude Agent SDK | claude-agent-sdk >= v0.1.14 |
| `opentelemetry-instrumentation-google-genai` | Google GenAI | google-genai >= v1.32.0 |
| `opentelemetry-instrumentation-vertexai` | Vertex AI | google-cloud-aiplatform >= v1.64 |
| `opentelemetry-instrumentation-langchain` | LangChain | langchain >= v0.3.21 |
| `opentelemetry-instrumentation-openai-agents-v2` | OpenAI Agents | openai-agents >= v0.3.3 |
| `opentelemetry-instrumentation-weaviate` | Weaviate | weaviate-client >= v3.0.0, < v5.0.0 |
Setup: `pip install <package>` + `Instrumentor().instrument()` or CLI
`opentelemetry-instrument`.
### Node.js
| Package | Provider | Min SDK Version |
| :--- | :--- | :--- |
| `@opentelemetry/instrumentation-openai` | OpenAI | openai >= 4.19.0 |
| `@opentelemetry/instrumentation-langchain` | LangChain | langchain >= 1.0.0 (not yet published to npm) |
Setup: `npm install <package>` + register via OTel Node SDK.
For per-provider install commands, upstream README links, and supported version
details, see
`${CLAUDE_PLUGIN_ROOT}/skills/otel-genai-instrumentation/references/auto-instrumentation-setup.md`.
## Manual Instrumentation
For languages without auto-instrumentation (Go, Java, etc.) or when
auto-instrumentation doesn't cover your needs.
Key patterns:
- Creating inference spans (`chat`, `text_completion`, `generate_content`)
- Creating embedding and retrieval spans
- Setting request attributes before the call, response/usage attributes after
- Error handling with `error.type` and span status
For code examples in Python, Node.js, and Go, see
`${CLAUDE_PLUGIN_ROOT}/skills/otel-genai-instrumentation/references/manual-instrumentation.md`.
## Span Flushing for GenAI Apps
**Critical for GenAI applications.** The `BatchSpanProcessor` buffers spans (default
5 s schedule delay). GenAI agent runs are long-lived but may exit before the batch
flushes — crash, Ctrl+C, short CLI invocations — causing **silent span loss**.
**Rule: force-flush after every top-level agent invocation.** Expose the span
processor and call `forceFlush()` without tearing down the SDK, so subsequent
invocations continue producing spans.
### Why `shutdown()` is wrong here
`sdk.shutdown()` tears down the entire pipeline — after shutdown, no new spans are
recorded. For apps that run multiple agent invocations (polling loops, HTTP servers,
CLI batch modes), you need spans to keep flowing. Use `forceFlush()` instead.
### Python
```python
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
span_processor = BatchSpanProcessor(exporter)
provider = TracerProvider()
provider.add_span_processor(span_processor)
async def flush_telemetry():
"""Flush pending spans without shutting down."""
span_processor.force_flush()
```
### Node.js
```typescript
import { BatchSpanProcessor } from "@opentelemetry/sdk-trace-base";
let spanProcessor: BatchSpanProcessor | null = null;
export function initTelemetry(): void {
// ... exporter setup ...
spanProcessor = new BatchSpanProcessor(traceExporter);
sdk = new NodeSDK({ spanProcessors: [spanProcessor], /* ... */ });
sdk.start();
}
export async function flushTelemetry(): Promise<void> {
if (spanProcessor) {
await spanProcessor.forceFlush();
}
}
```
### Go
```go
var spanProcessor *sdktrace.BatchSpanProcessor
func InitTelemetry() {
spanProcessor = sdktrace.NewBatchSpanProcessor(exporter)
// ... provider setup ...
}
func FlushTelemetry(ctx context.Context) error {
return spanProcessor.ForceFlush(ctx)
}
```
### Where to call `flushTelemetry()`
- **After each agent invocation** — ensures the full trace (agent + chat + tool spans)
is exported before moving to the next task
- **In polling/server loops** — flush after processing each request or ticket
- **Before `process.exit()`** — as a safety net alongside `shutdownTelemetry()`
- **NOT inside the agent loop** — flushing per-chat-turn adds latency; flush once at
the outer boundary
Example integration:
```typescript
for (const ticket of tickets) {
await triageIssue(ticket); // produces invoke_agent + chat + tool spans
await flushTelemetry(); // ensure spans are exported before next ticket
}
```
For complete code examples showing flush integration with tool-calling loops, see
`${CLAUDE_PLUGIN_ROOT}/skills/otel-genai-instrumentation/references/manual-instrumentation.md`.
## GenAI Span Types
**Span names MUST follow the pattern `"{operation} {identifier}"`.** The `gen_ai.operation.name`
attribute and the span name prefix must match. For example, a span with
`gen_ai.operation.name = "invoke_agent"` must be named `"invoke_agent {agent_name}"`,
not `"mypackage.DoSomething"`.
| Operation | `gen_ai.operation.name` | SpanKind | Span Name |
| :--- | :--- | :--- | :--- |
| Chat/completion | `chat` | CLIENT | `chat {model}` |
| Text completion | `text_completion` | CLIENT | `text_completion {model}` |
| Content generation | `generate_content` | CLIENT | `generate_content {model}` |
| Embeddings | `embeddings` | CLIENT | `embeddings {model}` |
| RAG retrieval | `retrieval` | CLIENT | `retrieval {data_source}` |
| Tool execution | `execute_tool` | INTERNAL | `execute_tool {tool_name}` |
| Agent creation | `create_agent` | CLIENT | `create_agent {agent_name}` |
| Agent invocation | `invoke_agent` | CLIENT/INTERNAL | `invoke_agent {agent_name}` |
| Workflow step | `invoke_workflow` | INTERNAL | `invoke_workflow {workflow_name}` |
### Required Attributes on All GenAI Spans
**CRITICAL: Every GenAI span MUST include these two attributes. This is non-negotiable.**
1. **`gen_ai.operation.name`** — Identifies the operation type (`chat`, `embeddings`, `execute_tool`, `invoke_agent`, etc.).
- **Without this**: The span is not recognized as a GenAI operation and will be excluded from GenAI-specific queries and visualizations in Honeycomb
- **Set on EVERY span**: chat, execute_tool, invoke_agent, embeddings, retrieval, etc.
2. **`gen_ai.conversation.id`** — Ties operations together within a conversation or session.
- **Without this**: Spans cannot be queried as part of a multi-operation workflow, breaking session-level analysis
- **Use the SAME value** across all operations in a conversation thread (user request → agent invocation → chat calls → tool executions → responses)
- Generate once at the start of a conversation, propagate to all operations
**When to set:** When creating the span (in the span attributes), not after.
**How to propagate conversation_id:**
- In-process: Pass as parameter or store in context
- HTTP/A2A: Include in request payload or propagate via headers
**Impact of missing these attributes:**
- Missing `gen_ai.operation.name` → Span not recognized as GenAI operation, excluded from GenAI-specific queries and visualizations
- Missing `gen_ai.conversation.id` → Span excluded from session queries, cannot correlate operations within a conversation, breaks multi-turn analysis
**What is a conversation?**
A conversation is a **customer session or user interaction**, NOT a single LLM call. One conversation contains:
- Multiple user turns/messages
- All LLM calls handling those turns
- All tool executions triggered by those LLM calls
- All agent invocations within that session
See the [OTel GenAI spec](https://opentelemetry.io/docs/specs/semconv/gen-ai/gen-ai-spans/#conversation-id) for the definition. Key principle: use the same conversation.id when conversation history/context is maintained across operations.
**When to use the same conversation_id:**
- All operations within a single customer session
- All turns in a multi-turn interaction
- All LLM calls handling those turns
- All tool executions and agent invocations within that session
- Multiple agents participating in the same session
**Example:** User starts a support session. Over the next 10 minutes they send 5 messages. The assistant makes 15 LLM calls and executes 8 tools to handle those messages. ALL of these spans share the SAME conversation.id because they're part of one customer session.
**Common mistake:** Generating a new conversation_id for each LLM call. This breaks session-level analysis. Generate conversation_id ONCE at session start, reuse for all operations until session ends.
For trace structures showing how these spans compose (tool-calling loops, multi-turn
conversations, nested agents, workflows), see
`${CLAUDE_PLUGIN_ROOT}/skills/otel-genai-instrumentation/references/agent-and-tool-patterns.md`.
**A2A / HTTP-based agent delegation:** When agents communicate over HTTP (A2A protocol,
REST delegation), manually propagate both trace context (via headers) AND conversation.id
(via payload). Client: `propagation.inject()` + include conversation.id in request body.
Server: `propagation.extract()` + `context.with()` + extract conversation.id from payload
and pass to all operations. See the "A2A (Agent-to-Agent) HTTP Context Propagation"
section in the reference file above.
## Generating and Propagating Conversation ID
Generate conversation_id at your application's **session boundary**:
- Chat apps: when user opens new chat/thread
- Support systems: when customer starts session
- CLI tools: at command invocation
- HTTP APIs: when session/conversation is created
- Bots: when user starts thread/DM
Pass the SAME conversation_id to all operations within that session — all user turns, all LLM calls handling those turns, all tool executions, all agent invocations.
**Propagation methods:**
- In-process: store in session object, pass as parameter
- HTTP/microservices: include in request payload or header (`X-Conversation-ID`)
- Bots: store in state (Redis, DB), retrieve using thread/DM ID
## Attribute Completeness
**Set all attributes for which you have data available.** The OTel GenAI semantic conventions define comprehensive attributes for each operation type — if your application has the data (model name, tokens, tool arguments, etc.), set the corresponding attribute.
**Critical principle**: Don't selectively omit attributes. Incomplete instrumentation limits your ability to:
- Identify which models and agents were involved in a trace
- Track token usage and costs across operations
- Debug tool call failures (missing arguments/results)
- Understand conversation flow (missing messages)
- Correlate agent behavior with configuration (missing request parameters)
For the full attribute definitions by operation type, see the upstream semantic conventions:
- Model operations (chat, embeddings): https://opentelemetry.io/docs/specs/semconv/gen-ai/gen-ai-spans/
- Agent operations (invoke_agent, execute_tool): https://opentelemetry.io/docs/specs/semconv/gen-ai/gen-ai-agent-spans/
- Local reference: `${CLAUDE_PLUGIN_ROOT}/skills/otel-genai-instrumentation/references/genai-attributes-catalog.md`
**What "data available" means**:
- API response fields → set corresponding response attributes (model, tokens, finish_reasons, response_id)
- Request parameters → set request attributes (temperature, max_tokens, top_p, etc.)
- Agent metadata → set agent attributes (name, id, description, version)
- Tool execution → set tool attributes (name, call_id, arguments, result)
- Conversation context → set conversation_id on ALL GenAI spans (required, not optional) — use the same ID across all operations in a conversation thread
The code examples in this skill show core attributes for each operation type. For complete coverage, consult the upstream spec and instrument every attribute your application can populate.
**Impact of incomplete instrumentation**:
- Missing `gen_ai.operation.name` → span not recognized as GenAI operation, excluded from GenAI queries
- Missing `gen_ai.conversation.id` → span excluded from session queries, cannot correlate operations within a conversation
- Missing `gen_ai.request.model` / `gen_ai.response.model` → can't identify which model was used
- Missing `gen_ai.usage.*` tokens → can't track costs or identify expensive operations
- Missing `gen_ai.tool.call.arguments` / `gen_ai.tool.call.result` → can't debug why tools failed or returned unexpected results
- Missing `gen_ai.input.messages` / `gen_ai.output.messages` → can't see what prompted a response, can't debug planning loops or hallucinations
- Missing agent attributes → can't distinguish between agents in multi-agent systems
- Missing request parameters → can't correlate behavior with temperature, top_p, etc.
**Best practice**: Instrument completely from the start. Adding attributes later requires code changes, redeployment, and waiting for new traces to arrive.
## Telemetry by Failure Mode
For each failure mode, the listed telemetry enables effective debugging. Items marked
**[Content Capture]** require enabling content capture — ask the user before enabling these.
### Tool Call Failures
- **Span** `execute_tool`: `gen_ai.tool.name`, `gen_ai.tool.call.id`,
`gen_ai.agent.name`, `gen_ai.conversation.id`, `error.type`,
`status.code=ERROR`, duration, `gen_ai.tool.call.arguments`, `gen_ai.tool.call.result`
- **Metric**: `gen_ai.client.operation.duration`
- **[Content Capture]**: `gen_ai.input.messages` (tool_call + tool_call_response parts) —
shows full context of tool calls (optional, requires user consent)
### Network Failures During Retrieval
- **Span** `retrieval`: `gen_ai.data_source.id`, `server.address`, `server.port`,
`error.type`, `status.code=ERROR`, duration
- **Metric**: `gen_ai.client.operation.duration`
### Long Time-to-First-Token
- **Span** `chat`: `gen_ai.request.model`, `gen_ai.usage.input_tokens`,
`server.address`, duration
- **Metrics**: `gen_ai.client.operation.time_to_first_chunk` (hosted APIs) or
`gen_ai.server.time_to_first_token` (self-hosted)
- Also: `gen_ai.server.time_per_output_token`, `gen_ai.agent.name`
### Excessive Planning / Retry Loops
- **Parent** `invoke_agent`: `gen_ai.agent.name`, `gen_ai.usage.input_tokens`, duration
- **Children** `execute_tool`: `gen_ai.tool.name`, `gen_ai.tool.call.arguments`,
`gen_ai.tool.call.result`
- **Metric**: `gen_ai.client.token.usage`
- **[Content Capture]**: `gen_ai.output.messages` — model reasoning reveals loop cause
(optional but very helpful, requires user consent)
### Slow Retrieval
- **Span** `retrieval`: `gen_ai.data_source.id`, `server.address`, `server.port`,
`status.code=OK`, duration
- **Metric**: `gen_ai.client.operation.duration`
### Agent Deadlocks
- **Span** `invoke_agent`: `gen_ai.agent.name`, `gen_ai.agent.id`,
`gen_ai.conversation.id`, `error.type=TimeoutError`, span links, duration
- **Metric**: `gen_ai.client.operation.duration`
- **[Content Capture]**: `gen_ai.output.messages` (tool_call parts) — reveals circular
delegation (optional but very helpful, requires user consent)
## Content Capture (Ask User First)
**CRITICAL: Do NOT enable content capture without asking the user first.**
### Step 1: Ask the User
Before providing any configuration, **ask this question**:
> "Do you want to capture the actual prompts and model responses in your traces?
>
> **Enabling content capture:**
> - ✅ Helps debug tool call failures, planning loops, and agent deadlocks
> - ✅ Lets you see why the model made specific decisions
> - ❌ Captures potentially sensitive content (user prompts, model responses)
> - ❌ May contain PII, proprietary data, or confidential information
>
> Recommended if: debugging/development, non-sensitive data, or you have filtering in place
>
> Not recommended if: production with sensitive data, PII/health/financial info, no filtering"
### Step 2: Configure Based on Answer
**If user says YES** to content capture:
For auto-instrumentation (Python), set the capture mode:
```bash
# Recommended for Honeycomb: Capture as span attributes (fully queryable)
export OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=span_only
```
**Why `span_only` for Honeycomb:**
- Content stored as span attributes → fully queryable in Honeycomb
- Can filter, group, and visualize by message content
- Lower overhead than `span_and_event`
**Alternative modes (less common):**
```bash
# Events only - for high-volume scenarios where you want content in logs but not queryable
export OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=event_only
# Both spans and events - most complete but higher overhead
export OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=span_and_event
# Legacy boolean - deprecated, use span_only instead
export OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=true
```
**Mode comparison:**
- `span_only` → Content in span attributes (queryable, recommended for Honeycomb)
- `event_only` → Content in events (logging, not queryable)
- `span_and_event` → Both (most complete, 2x overhead)
- `true` → Legacy (maps to old behavior, deprecated)
For manual instrumentation:
- Set `gen_ai.input.messages` on chat spans (before the call)
- Set `gen_ai.output.messages` on chat spans (after the call)
**If user says NO** to content capture:
Do NOT set `OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT` (leave unset).
Do NOT include `gen_ai.input.messages` or `gen_ai.output.messages` in manual instrumentation.
**ALWAYS include regardless of content capture setting:**
- `gen_ai.tool.call.arguments` on execute_tool spans
- `gen_ai.tool.call.result` on execute_tool spans
Tool arguments/results are essential for debugging and are typically less sensitive than
full conversation content.
### What Content Capture Provides
When enabled, `gen_ai.input.messages` and `gen_ai.output.messages` show the full
conversation — what the user sent, what the model returned, and how tool results were
fed back. Without them, you can see that a chat span happened but not *why* the model
made a particular decision.
### Example: .env Configuration
**If user wants content capture:**
```bash
# .env
# Base OTEL setup - see otel-instrumentation skill for:
# OTEL_SERVICE_NAME, OTEL_EXPORTER_OTLP_ENDPOINT,
# OTEL_EXPORTER_OTLP_HEADERS, OTEL_EXPORTER_OTLP_PROTOCOL, etc.
# GenAI-specific configuration (REQUIRED)
OTEL_SEMCONV_STABILITY_OPT_IN=gen_ai_latest_experimental
# Content capture (OPTIONAL - ask user first)
# Recommended for Honeycomb: span attributes (queryable)
OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=span_only
# Other content capture options (uncomment one if needed):
# OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=event_only # Events only, not queryable
# OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=span_and_event # Both (2x overhead)
# OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=true # Legacy (deprecated)
```
**If user does NOT want content capture:**
```bash
# .env
# Base OTEL setup - see otel-instrumentation skill for:
# OTEL_SERVICE_NAME, OTEL_EXPORTER_OTLP_ENDPOINT,
# OTEL_EXPORTER_OTLP_HEADERS, OTEL_EXPORTER_OTLP_PROTOCOL, etc.
# GenAI-specific configuration (REQUIRED)
OTEL_SEMCONV_STABILITY_OPT_IN=gen_ai_latest_experimental
# OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT not set (disabled by default)
```
### What Gets Captured
**Content capture enabled** (span_only, event_only, or span_and_event):
- `gen_ai.input.messages` — Full prompts sent to model
- `gen_ai.output.messages` — Full model responses
- `gen_ai.system_instructions` — System prompts
- `gen_ai.tool.definitions` — Available tools
**Capture mode determines where content is stored:**
- `span_only` → Span attributes (queryable in Honeycomb, recommended)
- `event_only` → Event attributes (logging/archival, not queryable in Honeycomb)
- `span_and_event` → Both locations (most complete, double storage/overhead)
- `true` → Legacy mode (deprecated, use `span_only`)
**Content capture disabled** (default):
- Model name, tokens, finish_reasons, timing — YES (always captured)
- Prompt/response content — NO
- Tool arguments/results — YES (always recommended)
Message JSON schema: `role` + `parts` (text, tool_call, tool_call_response, reasoning);
`tool_call_response` uses `response` field (not `content`) for the tool result.
### Privacy Controls (If Content Capture Enabled)
If the user enables content capture, recommend these additional safeguards:
- **Filtering**: Capture selectively (e.g., exclude messages with PII)
- **Truncation**: Limit content size (e.g., first 500 chars only)
- **Hooks**: Route to separate access-controlled storage
- **Access control**: Restrict who can query message content in Honeycomb
- **Environment-based**: Full content in dev/test, disabled or filtered in prod
Example filtering pattern (Python):
```python
# Only capture if no PII detected
if not contains_pii(message_content):
span.set_attribute("gen_ai.input.messages", json.dumps(messages))
```
Example truncation (any language):
```python
# Limit to first 500 characters
truncated = json.dumps(messages)[:500]
span.set_attribute("gen_ai.input.messages", truncated)
```
For complete setup including message JSON schemas, per-provider examples, and privacy
patterns, see
`${CLAUDE_PLUGIN_ROOT}/skills/otel-genai-instrumentation/references/content-capture-setup.md`.
## Streaming Instrumentation
Streaming (SSE, chunked responses) requires dedicated metrics and span patterns.
Key metrics:
- `gen_ai.client.operation.time_to_first_chunk` — client-observed time until first
streamed chunk (includes network latency); use for hosted APIs
- `gen_ai.server.time_to_first_token` — server-side TTFT (queue + prefill); use for
self-hosted (vLLM, TGI)
- `gen_ai.server.time_per_output_token` — decode speed after first token
- `gen_ai.client.operation.time_per_output_chunk` — client-observed inter-chunk time
The span covers the full stream lifetime. Set usage attributes after stream completes.
Handle mid-stream errors by recording the error and setting span status before closing.
For streaming span lifecycle, code examples, and error handling patterns, see
`${CLAUDE_PLUGIN_ROOT}/skills/otel-genai-instrumentation/references/streaming-instrumentation.md`.
## Evaluation Events
`gen_ai.evaluation.result` event captures scoring/evaluation of GenAI output.
| Attribute | Requirement | Description |
| :--- | :--- | :--- |
| `gen_ai.evaluation.name` | Required | Evaluation name (e.g., "relevance", "faithfulness") |
| `gen_ai.evaluation.score.value` | Recommended | Numeric score |
| `gen_ai.evaluation.score.label` | Recommended | Categorical label (e.g., "pass", "fail") |
| `gen_ai.evaluation.explanation` | Recommended | Why this score was given |
| `gen_ai.response.id` | Recommended | Links evaluation to the inference it scored |
Use cases: RAG relevance scoring, hallucination detection, output quality gates.
## Metrics
| Metric | Type | Unit | Purpose |
| :--- | :--- | :--- | :--- |
| `gen_ai.client.operation.duration` | Histogram | s | End-to-end latency |
| `gen_ai.client.token.usage` | Histogram | {token} | Input/output token counts |
| `gen_ai.client.operation.time_to_first_chunk` | Histogram | s | Streaming TTFC |
| `gen_ai.client.operation.time_per_output_chunk` | Histogram | s | Streaming inter-chunk |
| `gen_ai.server.request.duration` | Histogram | s | Server-side latency |
| `gen_ai.server.time_to_first_token` | Histogram | s | Server TTFT |
| `gen_ai.server.time_per_output_token` | Histogram | s | Server decode speed |
| `mcp.client.operation.duration` | Histogram | s | MCP client latency |
| `mcp.server.operation.duration` | Histogram | s | MCP server latency |
For the required `x-honeycomb-dataset` metrics header, see the **otel-instrumentation** skill.
## MCP Instrumentation
Model Context Protocol instrumentation uses OTel context propagation via
`params._meta` (W3C traceparent/tracestate).
- Client spans (CLIENT) for MCP calls, server spans (SERVER) for MCP handlers
- Key attributes: `mcp.method.name`, `mcp.session.id`, `mcp.protocol.version`
- Metrics: `mcp.client.operation.duration`, `mcp.server.operation.duration`
For context propagation details, well-known method names, and code examples, see
`${CLAUDE_PLUGIN_ROOT}/skills/otel-genai-instrumentation/references/mcp-instrumentation.md`.
## Known Gaps & Workarounds
| Gap | Workaround |
| :--- | :--- |
| No retry/loop count attribute | Count child spans or diff `tool.call.arguments` across siblings |
| No inter-agent dependency (in-process) | Span links + `gen_ai.conversation.id` |
| No inter-agent dependency (HTTP/A2A) | Manual `propagation.inject()` / `extract()` — see agent-and-tool-patterns ref |
| No retrieval sub-metrics | Custom attributes on retrieval spans |
| `error.type` is only error signal | Custom attributes for severity/category |
## Provider-Specific Notes
- **Anthropic**: cache token accounting, `gen_ai.provider.name = "anthropic"`
- **OpenAI**: `system_fingerprint`, service tier, `gen_ai.provider.name = "openai"`
- **AWS Bedrock**: `aws.bedrock.guardrail.id`, knowledge base attributes
- **Azure AI**: `azure.resource_provider.namespace`
## Additional Resources
### Reference Files
- **`${CLAUDE_PLUGIN_ROOT}/skills/otel-genai-instrumentation/references/auto-instrumentation-setup.md`** — Python + Node.js: per-provider install, upstream README links, supported versions
- **`${CLAUDE_PLUGIN_ROOT}/skills/otel-genai-instrumentation/references/manual-instrumentation.md`** — Code examples in Python/Node.js/Go for all span types
- **`${CLAUDE_PLUGIN_ROOT}/skills/otel-genai-instrumentation/references/genai-attributes-catalog.md`** — Upstream semconv links + message JSON schema gotchas
- **`${CLAUDE_PLUGIN_ROOT}/skills/otel-genai-instrumentation/references/agent-and-tool-patterns.md`** — Trace diagrams: tool-calling loop, multi-turn, nested agents, workflow
- **`${CLAUDE_PLUGIN_ROOT}/skills/otel-genai-instrumentation/references/mcp-instrumentation.md`** — MCP context propagation, span conventions, method names, metrics
- **`${CLAUDE_PLUGIN_ROOT}/skills/otel-genai-instrumentation/references/streaming-instrumentation.md`** — Streaming span lifecycle, TTFT/TTFC metrics, mid-stream errors, code examples
- **`${CLAUDE_PLUGIN_ROOT}/skills/otel-genai-instrumentation/references/content-capture-setup.md`** — Env var + manual setup, message JSON schemas, privacy controls
### Cross-References
- **BEFORE using this skill**: Use **otel-instrumentation** for base SDK setup, all OTEL environment variables (OTEL_SERVICE_NAME, OTEL_EXPORTER_OTLP_*, OTEL_EXPORTER_OTLP_HEADERS, etc.), OTLP config, collector, and sampling
- For conceptual foundations of wide events and high cardinality: **observability-fundamentals** skill
- After instrumenting, use the **query-patterns** skill to verify GenAI data in Honeycomb
Referenced files: 7
otel-instrumentation18.1 KB
---
name: otel-instrumentation
description: >
Provides guidance on OpenTelemetry SDK setup, custom instrumentation,
and sending data to Honeycomb.
Trigger phrases: "instrument my app", "add tracing",
"set up OpenTelemetry", "configure OTel", "add custom spans",
"add attributes to spans", "send traces to Honeycomb",
"set up OTLP", "configure sampling", "add span events",
"add span links", "set up tracing for [any language]",
"configure the OTel Collector",
or any request about OpenTelemetry SDK setup, custom instrumentation,
or sending data to Honeycomb.
metadata:
version: "1.0.0"
---
# OpenTelemetry Instrumentation for Honeycomb
SDK setup, custom spans, attributes, span events, sampling, and layered telemetry.
For conceptual foundations (why wide events matter, how attributes connect to
investigation), see the **observability-fundamentals** skill.
## OTLP Configuration and SDK Setup
Every OTel SDK needs these environment variables to send data to Honeycomb:
### Required Environment Variables
**Base configuration:**
```bash
OTEL_SERVICE_NAME=your-service-name
OTEL_EXPORTER_OTLP_ENDPOINT=https://api.honeycomb.io
OTEL_EXPORTER_OTLP_HEADERS="x-honeycomb-team=YOUR_API_KEY"
```
**Optional but recommended:**
```bash
# Protocol selection (default: http/protobuf)
OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf # or grpc
# Signal-specific endpoints (override base endpoint for specific signals)
OTEL_EXPORTER_OTLP_TRACES_ENDPOINT=https://api.honeycomb.io/v1/traces
OTEL_EXPORTER_OTLP_METRICS_ENDPOINT=https://api.honeycomb.io/v1/metrics
```
**For metrics (preferred):** Use modern OTLP metrics and native datapoints. Use dataset
hints to confirm the destination type (`metrics` or `events`). Authenticate with:
```bash
OTEL_EXPORTER_OTLP_METRICS_HEADERS="x-honeycomb-team=YOUR_API_KEY"
```
### Protocol Selection
`OTEL_EXPORTER_OTLP_PROTOCOL` determines the wire format and transport:
- `http/protobuf` (default, recommended) — HTTP with protobuf encoding
- `grpc` — gRPC with protobuf encoding
- `http/json` — HTTP with JSON encoding (larger payload, slower)
Use `http/protobuf` unless you have specific infrastructure requirements for gRPC.
### Signal-Specific Endpoints
By default, OTel SDKs append `/v1/traces` and `/v1/metrics` to `OTEL_EXPORTER_OTLP_ENDPOINT`.
Use signal-specific endpoint vars to override:
- `OTEL_EXPORTER_OTLP_TRACES_ENDPOINT` — full URL for traces (including `/v1/traces`)
- `OTEL_EXPORTER_OTLP_METRICS_ENDPOINT` — full URL for metrics (including `/v1/metrics`)
Useful when routing signals to different backends or using non-standard endpoints.
### Common Pitfalls
**Silent auth failure:** The OTLP exporters need the `x-honeycomb-team` header to
authenticate. Without it, Honeycomb silently rejects requests — no error, no data. Set
`OTEL_EXPORTER_OTLP_HEADERS="x-honeycomb-team=YOUR_API_KEY"` or pass headers
programmatically. If loading the key from `.env`, ensure dotenv runs before SDK init.
**Metrics:** Prefer modern OTLP metrics and native datapoints. Dataset hints identify the
destination type (`metrics` or `events`), so do not add `x-honeycomb-dataset` by default.
Use that header only when hints or configuration require legacy routing to a named event
dataset. Traces do not need it; they route by `service.name`.
For the env var values, language-specific dependencies, and setup code (Go, Python,
Node.js, Java, Ruby, .NET, Rust), see
`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/sdk-setup-by-language.md`.
## Custom Instrumentation
### Adding Attributes to Existing Spans (Highest Impact)
Add business context to auto-instrumented spans — no new spans needed. Get the current
span from context and call `SetAttributes` (Go), `set_attribute` (Python), or
`setAttribute` (Node.js) with user, tenant, business, and deployment context.
### Creating Custom Spans
Wrap important business operations for visibility in the trace waterfall. Use
`tracer.Start(ctx, "operation-name")` (Go), `tracer.start_as_current_span("operation-name")`
(Python), or `tracer.startActiveSpan("operation-name", callback)` (Node.js).
For full code examples in all languages, consult
`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/custom-instrumentation.md`.
## When to Create a Span
Not every function needs a span. Two questions determine whether a span is worth creating:
1. **Is it interesting?** — Does the work meaningfully impact performance (latency or
failures) for the overall request?
2. **Is it aggregable?** — If you group this span by name and attributes, will it produce
useful trends and comparisons?
| Operation | Interesting? | Aggregable? | Create a Span? |
| :--- | :--- | :--- | :--- |
| HTTP request handler | Yes — variable latency, can fail | Yes — group by route, method, status | **Yes** |
| Database query | Yes — I/O bound, failure-prone | Yes — group by query type, table | **Yes** |
| External API call | Yes — network latency, dependencies | Yes — group by endpoint, status | **Yes** |
| Cache lookup | Yes — fast vs slow path | Yes — group by cache name, hit/miss | **Yes** |
| Message queue pub/consume | Yes — async boundary, delays | Yes — group by queue, message type | **Yes** |
| Business logic transaction | Yes — meaningful state change | Yes — group by type, outcome | **Yes** |
| Private helper function | No — trivial CPU, predictable | No — too granular | **No** |
| Loop iteration | Maybe — if slow | No — unbounded cardinality | **No** |
| Getter/setter | No — no meaningful duration | No — nothing to group by | **No** |
| Input validation (pure CPU) | No — fast, predictable | Maybe | **No** |
| Business logic orchestration | No — just calls instrumented code | No — duration is sum of children | **No** |
**Common mistakes:**
- **Too many spans**: A trace with millions of 2ms spans is far too detailed and rarely
actionable. Roll them up — combine into a single span, or capture the detail as an
attribute on the parent span instead.
- **Too few spans**: Collapsing hours of work into a single opaque handler leaves you
guessing about where time is spent.
- **Test spans left in**: Spans named `test-span`, `debug-span`, or similar are
artefacts that pollute the dataset. Remove any span created solely to verify tracing
is working before finishing.
When in doubt, prefer **attributes on existing spans** over creating new child spans.
#### Timing Attributes (measure sub-operations without child spans)
Record important sub-operation durations as attributes on the parent span. These are
easier to query than child spans and work directly with BubbleUp.
```go
// Go: time auth and record on the existing span
span := trace.SpanFromContext(r.Context())
authStart := time.Now()
user, err := authenticate(r)
span.SetAttributes(attribute.Float64("auth.duration_ms", float64(time.Since(authStart).Milliseconds())))
```
```python
# Python: time auth and record on the existing span
span = trace.get_current_span()
auth_start = time.monotonic()
user = authenticate(request)
span.set_attribute("auth.duration_ms", (time.monotonic() - auth_start) * 1000)
```
#### Exception telemetry: event details plus span-level dimensions
Use the Logs API for new exception events. Emit the record while the relevant span is
active and include the standard exception fields (`exception.type`, `exception.message`,
`exception.stacktrace`, and `exception.escaped` when applicable), an ERROR severity, and
`event.name="exception"`. Set the span status to ERROR separately when the operation failed.
In Honeycomb, a trace-correlated exception log is rendered in the trace as a `span_event`
annotation and carries `trace.trace_id` and `trace.parent_id`. Its full `exception.*`
payload remains on the log-derived event; it is **not hoisted onto the containing span**.
Search the exception event row, then follow its trace ID to inspect the surrounding trace.
Use low-cardinality span attributes for aggregation and alerting:
- `error=true` and the span status indicate operation failure.
- `exception.slug` is a static, greppable identifier for the error site.
- An optional error category is safer for `GROUP BY` than full exception messages.
```text
Logs-API exception event: event.name=exception, body=exception, meta.signal_type=log
Legacy span-event exception: name=exception, meta.signal_type=trace
Both may have: meta.annotation_type=span_event
```
`record_exception` / `RecordError` remain compatibility APIs for existing SDKs and code,
but do not use them as the only new guidance when Logs API support is available. They can
also produce parent-span exception fields that a Logs-API event alone does not produce.
Find operation failures by span dimensions: `WHERE error = true AND exception.slug does-not-exist`.
Find Logs-API exception events with `event.name=exception AND exception.type exists` and
follow a sampled `trace.trace_id` into `get_trace` with `show_events=true`.
For extended examples and the MCP investigation recipe, see
`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/custom-instrumentation.md`.
#### Optional compatibility: promote exception fields with a LogRecordProcessor
If existing span-level dashboards, alerts, or queries depend on Honeycomb's historical exception
field promotion, add a custom **LogRecordProcessor** before the batch/export processor. When it
sees an exception log, it should use the log record's resolved context to find the active recording
span and promote a configured, minimal set of fields such as `error=true`, `error.type`,
`exception.type`, `exception.slug`, or an error category.
Do not recommend a standalone `SpanProcessor` for this: span processors receive span lifecycle
callbacks, not log records. Keep full `exception.message` and `exception.stacktrace` on the Logs
API event by default; copy them onto spans only when legacy query compatibility explicitly requires
it. The processor must run synchronously while the span context is valid, before the log reaches
batch export. It should no-op when there is no recording span and must not infer fields that the
application did not put on the log record.
This is an optional migration layer, not a replacement for querying the Logs API event. Agents
should treat span-level promoted fields as instrumentation-dependent and continue to query
`event.name=exception` event rows for full diagnostics.
## What to Instrument
### High Value (Instrument First)
- API entry points (HTTP handlers, gRPC methods)
- Database queries (auto-instrumented by most SDKs)
- External HTTP calls (auto-instrumented by most SDKs)
- Message queue producers/consumers
These are typically auto-instrumented by OTel SDKs and form the skeleton of your traces.
### Medium Value (Add Next)
- Business logic operations (checkout, payment, fulfillment)
- Cache operations (hits, misses, evictions)
- Authentication and authorization checks
- Background job execution
These are your business logic. Without custom spans here, you can see that a request was
slow but not *why* — the trace waterfall has gaps where the important work happens
invisibly.
### Attributes to Add
Attributes are the dimensions BubbleUp uses during investigations. Every attribute you
add is a new axis BubbleUp can diff on to find what's different about outlier requests.
For the complete catalog organized by category with rationale and example queries, see
`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/wide-event-attributes.md`.
For why attributes matter conceptually, see the **observability-fundamentals** skill.
## Span Events, Logs API Events, and Span Links
- **Point-in-time events**: Prefer the Logs API for new events, especially exceptions. Emit
while the span is active so the record carries trace context. In Honeycomb, a correlated
log is rendered as a `meta.annotation_type=span_event` annotation, but its event name is
in `event.name` (and often `body`), not `name`.
- **Legacy span events**: `span.add_event` / `AddEvent` remain valid compatibility paths. Their
event name is in `name` and their signal type is `trace`.
- **Span links**: Connect spans across different trace hierarchies (async processing,
fan-out/fan-in, cross-system correlation). Create a `Link` to the related span context.
For human instrumentation examples and an agent-safe Honeycomb MCP query → sample → trace
workflow, see `${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/custom-instrumentation.md`
and the **production-investigation** skill.
## Sampling
### Sampling Strategy
Sampling is about tradeoffs — there is no free lunch:
- **Head sampling favors cost over debuggability.** You save resources, but a 0.1% error
at 1% sampling becomes effectively invisible. Head sampling is oblivious to what
happens downstream.
- **Tail sampling favors fidelity over simplicity.** You keep interesting traces but need
infrastructure (Refinery or Collector) to buffer and evaluate complete traces.
The math matters: if an error occurs 0.1% of the time and you head-sample at 1%, you'll
capture roughly 1 in 100,000 of those errors. At moderate traffic, that error may never
appear in your data.
### Head Sampling (SDK-level)
Decides whether to sample a trace at creation time. Simple but can miss interesting traces.
- Configure via `OTEL_TRACES_SAMPLER` env var
- `always_on` (default), `always_off`, `traceidratio` (e.g., sample 10%)
- `parentbased_traceidratio` respects parent sampling decisions
- **Best for:** Very high-throughput services where you can tolerate missing rare events
### Tail Sampling (Collector/Refinery)
Decides after the trace is complete. Keeps interesting traces (errors, slow requests).
- Use Honeycomb's **Refinery** for production tail sampling
- Or configure the OTel Collector's `tail_sampling` processor
- Can sample based on: latency, error status, specific attributes, trace duration
- **Best for:** Services where debuggability matters — keeps errors and outliers while
sampling routine traffic
### Sampling Impact on Honeycomb
- Sampling reduces data volume and cost
- SLOs, BubbleUp, and query results adjust for sampling rate automatically
- Trace completeness may be affected — missing spans if not all services sample consistently
- Start with no sampling, then add as needed for cost management
## Layered Telemetry
OpenTelemetry is "trace-first" — context propagation is the glue that correlates all
signals. But effective observability layers multiple signal types for different purposes.
A three-question test for choosing the right signal:
1. **What needs causality and full-request context?** → Traces (spans)
2. **What needs inexpensive long-term storage and fast alerting?** → Metrics
3. **What is rare vs. common, and what are the audit requirements?** → Logs / events
**The histogram-alongside-spans pattern:** For high-throughput HTTP services, emit both a
span and a histogram metric for each handled request. This lets you head-sample traces
for cost while histograms provide last-ditch alerting — and exemplars link outlier metric
points back to specific traces for deeper investigation.
The technique is *layering* (not duplication) because each signal provides a different
view at a different level of detail.
For architectural patterns where layering is essential (streaming, async jobs, ETL), see
`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/architectural-patterns.md`.
For AWS Lambda-specific patterns — choosing between the AWS Managed OTel Layer
and manual SDK setup, forceFlush, SDK 2.x setup, cross-Lambda trace propagation,
header normalisation, TOKEN vs REQUEST authorizers — see
`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/lambda.md`.
## Logs in Honeycomb
OTel can send logs too. If you have existing log infrastructure, the OTel Collector can
ingest logs and forward them to Honeycomb as structured events:
- **OTel SDK log bridge**: Captures logs from your existing logging library (`slog` in Go,
`logging` in Python, `winston`/`pino` in Node.js) and exports them as OTel log records.
- **OTel Collector `filelog` receiver**: Reads log files, parses them, exports as OTLP.
Logs sent through OTel arrive in Honeycomb as structured events with the same query
capabilities as spans.
## Naming Conventions
- **Span names**: Describe the operation (`HTTP GET /api/users`, `db.query SELECT`, `process-payment`)
- **Attribute names**: Use dot-separated namespaces (`user.id`, `order.total`, `cache.hit`)
- **Follow OTel semantic conventions** where applicable (`http.method`, `db.system`, `rpc.service`)
- **Custom attributes**: Use your own namespace (`app.`, `checkout.`, `mycompany.`)
## Additional Resources
### Reference Files
- **`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/sdk-setup-by-language.md`** — OTLP configuration and SDK setup for Go, Python, Node.js, Java, Ruby, .NET, Rust
- **`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/local-collector-debug-test.md`** — Run a local OTel Collector via Docker to verify spans, logs, and metrics without a Honeycomb account; includes `jq` commands for inspecting NDJSON output
- **`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/custom-instrumentation.md`** — Custom instrumentation patterns with full code examples (timing attributes, exception slugs, async request summaries)
- **`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/collector-config.md`** — OTel Collector configuration for format conversion, processing, and sampling
- **`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/wide-event-attributes.md`** — Canonical attribute catalog organized by category with example queries
- **`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/architectural-patterns.md`** — Trace design patterns for streaming, async, ETL, and serverless architectures
- **`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/lambda.md`** — AWS Lambda: OTel Layer vs manual SDK setup trade-offs, forceFlush and per-request latency, SDK 2.x setup, cross-Lambda trace propagation, header normalisation, TOKEN vs REQUEST authorizer migration
### Cross-References
- For conceptual foundations of why wide events and attributes matter: **observability-fundamentals** skill
- After instrumenting, use the **query-patterns** skill to verify data is arriving
Referenced files: 8
otel-migration8.68 KB
---
name: otel-migration
description: >
Guide for retrofitting OpenTelemetry into an existing, uninstrumented application.
Trigger phrases: "migrate existing app to OTel",
"add OpenTelemetry to existing project", "retrofit OTel into my codebase",
"thread context through my code", "context propagation",
"bridge Prometheus metrics to OTel", "logging bridge",
"migrate logging to OTel", "slog bridge", "logback bridge",
"verify my instrumentation", "traces are disconnected",
"orphaned spans", "migrate to OpenTelemetry", "OTel migration plan",
"how do I sequence an OTel migration", "add tracing to existing code",
"refactor for context propagation", "Fiber context gotcha",
"keep existing logging working with OTel", "add OTel without breaking Prometheus",
"bridge existing metrics", "coexist with existing monitoring",
or any request about retrofitting OpenTelemetry into an existing application.
This skill is for migrating existing codebases, NOT greenfield instrumentation (use otel-instrumentation)
or Beeline-specific migration (use beeline-migration).
metadata:
version: "1.0.0"
---
# OpenTelemetry Migration for Existing Applications
Guide for retrofitting OpenTelemetry into an existing, uninstrumented application. This covers
the phased migration approach, context propagation refactoring, logging and metrics bridges, and
verification. This is distinct from greenfield OTel setup (see otel-instrumentation skill) and
Beeline-specific migration (see beeline-migration skill).
## When to Use This Skill
Use this skill when the user has an **existing application** that:
- Has no OpenTelemetry instrumentation and needs to add it
- Has existing logging, metrics, or context patterns that must coexist with OTel
- Needs to refactor function signatures to thread trace context through the call stack
- Uses a framework with OTel middleware gotchas (e.g., Fiber, Gin, Express)
For greenfield OTel setup, use the `otel-instrumentation` skill instead.
For Beeline-to-OTel migration, use the `beeline-migration` skill instead.
For understanding *why* to instrument, see the `observability-fundamentals` skill.
## Migration Phases
The migration follows six phases in order. Each phase is independently deployable and verifiable.
Context propagation (Phase 3) is typically ~60% of the effort.
### Phase 1: SDK Initialization and Shutdown
Set up TracerProvider, MeterProvider, and LoggerProvider with OTLP exporters. Wire initialization
early in the application's entry point and shutdown in signal handlers.
**Key guidance:**
- SDK init must happen *before* any application code that might create spans — one of the first
things in your entry point, before config loading or storage initialization
- Shutdown ordering matters: flush traces, then metrics, then logs. Use a timeout (10-30s)
- If init fails, the application should still work — log the error and continue without telemetry
- For language-specific SDK setup, consult
`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/sdk-setup-by-language.md`
### Phase 2: HTTP Middleware (Auto-Instrumentation)
Add OTel middleware to your HTTP framework. This gives you automatic spans for every inbound
request with zero code changes to handlers. This is the highest-ROI step.
**Critical:** Different frameworks expose the OTel-enriched context differently. This is the #1
source of silent trace breaks. Consult
`${CLAUDE_PLUGIN_ROOT}/skills/otel-migration/references/framework-middleware.md` for
framework-specific details.
| Framework | How to get OTel context | Common mistake |
|-----------|------------------------|----------------|
| Go net/http | `r.Context()` | N/A (standard) |
| Go Fiber v2 | `c.UserContext()` | Using `c.Context()` (returns fasthttp context without OTel span) |
| Go Gin | `c.Request.Context()` | Using `c` directly |
| Go Echo | `c.Request().Context()` | N/A |
| Python Flask | Automatic (thread-local) | N/A with instrumentation library |
| Python Django | Automatic (thread-local) | N/A with instrumentation library |
| Node.js Express | Automatic (AsyncLocalStorage) | N/A with instrumentation library |
| Java Spring | Automatic (thread-local) | Thread pool context loss |
| .NET ASP.NET Core | Automatic (AsyncLocal) | N/A |
| Ruby Rails | Automatic (thread-local) | N/A with instrumentation library |
### Phase 3: Context Propagation Refactoring
Thread trace context through your call chain from HTTP handlers (or entry points) down to I/O
operations. **This is the hardest phase** — typically ~60% of migration effort.
The difficulty of this phase varies dramatically by language:
- **Go**: Hardest. Requires adding `context.Context` parameter to every function in the call chain.
- **Java**: Moderate. Thread-local context propagates automatically within a thread, but breaks
across thread pools, CompletableFuture, and reactive streams.
- **Python**: Easier. `contextvars` propagates automatically within a thread. Pain points are
thread pools and multiprocessing.
- **Node.js**: Easier. `AsyncLocalStorage` propagates through async/await automatically.
Pain points are old callback-based code.
- **.NET**: Easiest. `Activity` propagates through async/await via `AsyncLocal<T>` automatically.
- **Ruby**: Easier. Thread-local context propagates automatically. Pain with manual thread creation.
For language-specific patterns and code examples, consult
`${CLAUDE_PLUGIN_ROOT}/skills/otel-migration/references/context-propagation-patterns.md`.
### Phase 4: Custom Spans
Add spans to business logic operations that auto-instrumentation doesn't cover. Defer to the
`otel-instrumentation` skill for span creation mechanics. Migration-specific guidance:
1. **Start with I/O boundaries** — database calls, external HTTP calls, cache operations
2. **Then add business logic** — operations that explain *why* time is spent
3. **Add attributes liberally** — every piece of context makes BubbleUp useful during investigations
4. **Record outcomes on spans** — result status, error count, duration as attributes
For attribute naming and span creation patterns, consult
`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/custom-instrumentation.md`.
### Phase 5: Logging Migration
Replace or bridge your existing logging library into OTel so logs correlate with traces.
**Key guidance:**
- You almost certainly want logs going to both stderr (local debugging) AND OTel (trace correlation).
This requires a multi-handler/fan-out pattern.
- OTel log bridges work with structured logging (key-value pairs). If your existing logging uses
printf-style format strings, convert to structured format first.
- Converting from printf-style to structured logging is tedious but mechanical — a good candidate
for automated refactoring.
For language-specific logging bridges and the multi-handler pattern, consult
`${CLAUDE_PLUGIN_ROOT}/skills/otel-migration/references/bridge-libraries.md`.
### Phase 6: Metrics Bridge
If you already have Prometheus metrics (or another metrics library), bridge them to OTel rather
than rewriting.
**Key guidance:**
- Prometheus bridge reads from the existing registry and produces OTel metrics — existing
`prometheus.NewCounterVec(...)` calls continue unchanged
- Keep the Prometheus `/metrics` endpoint if you have existing scrapers. The bridge adds OTLP
export *in addition to* scraping.
- If you want to eventually remove the Prometheus dependency, plan a separate migration later.
The bridge buys you time.
For language-specific metrics bridges, consult
`${CLAUDE_PLUGIN_ROOT}/skills/otel-migration/references/bridge-libraries.md`.
## Verification
After each phase, verify that instrumentation is correct and complete. Consult
`${CLAUDE_PLUGIN_ROOT}/skills/otel-migration/references/verification-checklist.md` for the
full checklist and query patterns.
To verify locally without a Honeycomb account, use the bundled collector script to capture
spans as debug output and NDJSON. Consult
`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/local-collector-debug-test.md` for usage, and
`${CLAUDE_PLUGIN_ROOT}/scripts/start-collector.sh` for the full script.
For Honeycomb-specific verification queries, also consult the `query-patterns` skill.
## Common Pitfalls
For a catalog of common mistakes and how to avoid them, consult
`${CLAUDE_PLUGIN_ROOT}/skills/otel-migration/references/migration-pitfalls.md`.
## Real-World Calibration
For reference, a real migration of Gatus (~30k LOC Go, Fiber v2, SQLite/Postgres, Prometheus):
- **Files changed:** ~45, **Lines:** +840/-645
- **Effort breakdown:** Context propagation ~60%, Custom spans ~15%, Logging migration ~15%, Everything else ~10%
- **Bugs encountered:** Fiber `c.Context()` vs `c.UserContext()`, printf-style slog format strings,
missing `span.End()` calls, goroutine context reuse
Referenced files: 5
production-investigation7.65 KB
---
name: production-investigation
description: >
Structured workflows for investigating production issues in Honeycomb — the
sequence of tool calls (context priming, broad query, BubbleUp, trace analysis,
verification) and how to chain results between steps to reach root causes.
Trigger phrases: "investigate production issue", "debug latency spike",
"find root cause", "use BubbleUp", "analyze traces", "debug an outage",
"why is my API slow", "errors are increasing", "health check", "SLO burning",
or any request to investigate or debug production problems.
metadata:
version: "1.0.0"
---
# Honeycomb Production Investigation
Structured workflows for debugging production issues. The MCP tools document their
own parameters — this skill focuses on the *sequence* of tool calls and how to
*interpret* results to reach a root cause.
## The Core Analysis Loop
This workflow implements the core analysis loop (**Define → Visualize → Investigate →
Evaluate**) from the **observability-fundamentals** skill. If BubbleUp returns nothing
useful, the issue is often an instrumentation gap — add the missing attributes (see the
**otel-instrumentation** skill) and try again.
## Investigation Workflow
### Step 1: Orient
1. `get_workspace_context` → environments and datasets
2. `get_slos` → any SLOs in violation? (frames severity)
3. `get_triggers` → any alerts firing? (narrows scope)
4. `find_queries` → has anyone investigated this before?
### Step 2: Characterize the Problem
Run a broad query to see the shape of the issue:
- **Latency spike**: P99(duration_ms), HEATMAP(duration_ms) grouped by service or route
- **Error surge**: count failed operation spans (`error=true`) by service/route/category, then
separately count exception event rows using `event.name=exception` and `exception.type exists`;
use sampled `trace.trace_id` values to drill into representative traces
- **Unknown**: COUNT grouped by service.name to find which service has anomalous volume
Also call `get_service_map` — it shows P95 durations between services and can immediately reveal which dependency is slow.
**Exception data has two query surfaces:** operation failures belong on spans (`error=true`, span
status, low-cardinality `exception.slug`/error category); full exception diagnostics may belong on
trace-correlated Logs API event rows. Do not assume `exception.*` exists on the containing span.
When investigating exceptions, discover the dataset schema first, query `event.name=exception`
with `exception.type exists` and `trace.trace_id exists`, take a sample, then pass its
`trace.trace_id` to `get_trace(show_events=true)`. For legacy span-event exceptions, also check
`name=exception` and `meta.signal_type=trace`; Logs API events use `event.name`/`body` and
`meta.signal_type=log`.
If a service uses an exception-promoting `LogRecordProcessor`, some `exception.*` fields may also
appear on the containing span. Treat that as an explicit client-side compatibility feature, not a
Honeycomb guarantee: the event row remains authoritative for full diagnostics, and absence of
parent-span fields does not mean the exception event is missing.
### Step 3: BubbleUp to Find Differentiators
This is the highest-value step. Once you have a query showing the anomaly:
1. Run `run_bubbleup` on the query result, selecting the outlier region
2. BubbleUp compares outlier vs baseline distributions across *all* columns automatically
3. Look for fields where the distributions differ significantly
**How to interpret BubbleUp results:**
- **Categorical fields** (dimensions): A value overrepresented in outliers points to a cause (e.g., `deployment.version=v2.3.1` is 90% of slow requests but only 20% of baseline)
- **Numeric fields** (measures): A shifted distribution shows correlated metrics (e.g., `db.query_duration` is much higher in outliers)
- **Typical root causes surfaced**: deployment version, region, user cohort, specific endpoint, feature flag
### Step 4: Drill Into Traces
After BubbleUp identifies suspects:
1. Add BubbleUp findings as WHERE filters to narrow results
2. Pick a representative trace ID
3. Call `get_trace` to fetch the full trace
**What to look for in the trace waterfall:**
- Spans with disproportionate duration vs parent (the bottleneck)
- Sequential spans that could be parallelized (N+1 query patterns)
- Error spans — check span events for stack traces
- Gaps between child spans (missing instrumentation or idle wait)
- Service boundaries (where the trace crosses services)
### Step 5: Verify Hypothesis
Form a hypothesis from BubbleUp + trace analysis, then confirm:
- Query WITH the suspected cause filtered in
- Query WITHOUT it (as a control)
- If the metrics diverge, you've found it
### Step 6: Record Findings
Call `create_board` with:
- A text panel summarizing the root cause (Markdown)
- The key query run PKs that identified the problem
- Related SLOs if applicable
## Investigation Patterns
### Latency Spike
HEATMAP first → BubbleUp the slow region → trace a slow request → verify with filtered queries
### Error Surge
Count failed operation spans by service/route/category → count Logs API exception events by
`event.name=exception` and `exception.type` → sample `trace.trace_id` → `get_trace(show_events=true)`
→ verify with filtered queries. Do not use `exception.message` on the parent span as the only
exception search.
### Deployment Regression
P99 grouped by deployment.version → BubbleUp comparing new vs old → trace from new version → verify
### Dependency Failure
`get_service_map` → P99 on the slow dependency → relational query (`any.service.name`) to measure user impact → trace an affected request
## Stay on the Path
If you find yourself reasoning any of these, follow the workflow anyway:
- "The cause is obvious, I can skip BubbleUp" — BubbleUp routinely surfaces causes that seem obvious in hindsight but weren't the first guess. It also catches *secondary* causes you'd miss entirely.
- "I already know it's a deployment issue" — verify with Step 5. Confirmation bias is strongest during incidents. Query with and without the suspected cause.
- "Traces confirmed it, no need to verify" — a single trace is an anecdote. The verification query proves the pattern holds across all traffic, not just one request.
- "This is a simple issue, the full workflow is overkill" — the workflow takes minutes; a wrong diagnosis during an incident costs hours.
## When Results Are Empty or Unclear
- **No results**: Check field names with `find_columns`, expand time range, verify environment/dataset
- **BubbleUp shows no signal**: Try a different time selection, add filters to isolate the anomaly more clearly, or select a different calculation
- **Trace missing spans**: Sampling, instrumentation gaps, or cross-environment trace split
## Additional Resources
### Reference Files
- **`${CLAUDE_PLUGIN_ROOT}/skills/production-investigation/references/investigation-playbooks.md`** — Step-by-step playbooks for latency spikes, error surges, deployment regressions, dependency failures, SLO budget burn, and health checks
- **`${CLAUDE_PLUGIN_ROOT}/skills/production-investigation/references/bubbleup-guide.md`** — Detailed BubbleUp usage: selection types, time specifications, pagination, result interpretation
- **`${CLAUDE_PLUGIN_ROOT}/skills/production-investigation/references/trace-exploration.md`** — Trace structure, get_trace parameters and view modes, waterfall analysis, span events and links
### Cross-References
- For the conceptual foundations of the core analysis loop, see the **observability-fundamentals** skill
- For query construction patterns, see the **query-patterns** skill
- For SLO/trigger context during investigations, see the **slos-and-triggers** skill
Referenced files: 3
query-patterns7.11 KB
---
name: query-patterns
description: >
Opinionated guidance for constructing and interpreting Honeycomb queries on
trace and event datasets — operation selection (percentiles not AVG, HEATMAP
for distributions), relational field patterns (root., parent., any., none.),
calculated fields, query math, and result interpretation (P99/P50 ratios,
heatmap bands, TOTAL/OTHER rows, raw JSON via query_result_json). Use this
skill when the user wants to query spans, traces, or log/event data in
Honeycomb — requests like "show me latency", "error rate", "find slow
requests", "find outliers", "interpret results", "relational fields",
"calculated fields", or "download raw results". This skill covers all
dataset types except metrics datasets (dataset_type=metrics) — for those,
use metrics-queries instead.
metadata:
version: "1.1.0"
---
# Honeycomb Query Patterns
Opinionated guidance for writing effective Honeycomb queries. The MCP tools already
document their parameters and schemas — this skill focuses on *when* and *why* to
use each pattern, not *how* to call the tools.
## Key Principles
1. **Never use AVG for latency** — AVG hides tail latency. Use P99 (or P95/P90) to see what slow users experience. Reserve AVG for non-latency metrics like payload size.
2. **Use HEATMAP for distributions** — Single-number aggregates hide bimodal patterns. HEATMAP reveals whether you have one population or two.
3. **Combine calculations in one query** — `COUNT, P99(duration_ms), HEATMAP(duration_ms)` in a single query reduces API calls and gives a complete picture.
4. **Start broad, narrow with WHERE** — Begin with a COUNT/GROUP BY to understand shape, then add filters to focus.
5. **Check for prior work** — Call `find_queries` before writing new queries. Someone may have already answered the question.
## Choosing the Right Operation
| Question | Use |
|----------|-----|
| How much traffic? | `COUNT` grouped by route or service |
| How many unique users/IPs? | `COUNT_DISTINCT(field)` |
| How fast for most users? | `P50(duration_ms)` |
| How fast for the worst-off users? | `P99(duration_ms)` |
| Is there a bimodal pattern? | `HEATMAP(duration_ms)` |
| What's the worst case? | `MAX(duration_ms)` |
| How many concurrent operations? | `CONCURRENCY` |
| Is it getting worse over time? | `RATE_AVG(duration_ms)` |
## Relational Field Strategy
Use relational prefixes to ask cross-span questions within a trace:
- **"Show me slow endpoints caused by a specific downstream"**: Filter with `any.service.name` to find traces where that service participates, group by `root.http.route` to see which user-facing endpoints are affected.
- **"What's different about errored traces?"**: Filter with `any.error = true`, group by `root.name` to see which entry points have errors somewhere in their trace tree.
- **Exclude noise**: `none.service.name = "health-check"` removes traces containing health checks.
## Calculated Fields
Calculated fields are per-event expressions evaluated at query time. They transform,
classify, and combine existing fields without re-instrumenting code.
**Three scopes** — choose the narrowest that fits the need:
- **Query-scoped** (not saved): exploratory, one-off analysis
- **Dataset-level** (saved): reusable within one service's dataset
- **Environment-level** (saved): reusable across all datasets (e.g., `error_pct`)
**Common patterns:**
- **Error rate**: `MUL(IF($error, 1, 0), 100)` → use `AVG(error_pct)` to get percentage
- **Status classification**: `IF(GTE($http.status_code, 500), "5xx", GTE($http.status_code, 400), "4xx", "ok")`
- **Latency bucketing**: `BUCKET($duration_ms, 500, 0, 3000)`
- **Prefix routing**: `IF(STARTS_WITH($url, "/admin"), "admin", STARTS_WITH($url, "/api"), "api", "other")`
- **Exact-match classification**: use `SWITCH` instead of `IF(EQUALS(...))` chains — same expression, more efficient
**Key guardrails:**
- **Don't create presentational (alias-only) fields** — a field that just renames another field adds no analytical value and clutters the schema. Only save a calculated field when it does real computation (classification, extraction, math).
- **Avoid regex on large/complex fields** — running `REG_MATCH`, `REG_VALUE`, or `REG_COUNT` on `exception.stacktrace`, `db.statement`, or full log lines can be very slow. Check whether a more targeted OTel field exists first (`exception.type`, `exception.message`, `db.operation`). If you must regex a long field, guard it with a `CONTAINS` check first.
- **`EQUALS` has strict type matching** — `EQUALS($http.status_code, 200)` silently returns false if the field is stored as a string. Use `find_columns` to verify the field type before comparing.
- **`FORMAT_TIME` is expensive** — avoid in high-volume queries.
- **Save query-scoped, not dataset-level, for one-off work** — saved fields show up in everyone's schema.
For full syntax, operator reference, and extended anti-pattern examples, consult
`${CLAUDE_PLUGIN_ROOT}/skills/query-patterns/references/calculated-fields.md`.
## Before Every Query
- **Filter on `is_root`** when measuring user-facing latency — without it, internal spans inflate the numbers
- **Use human-readable time ranges** (`"24h"`, `"-6h"`) — epoch timestamps are error-prone and hard to review
- **Validate columns with `find_columns` before querying** — confirms field names exist and prevents empty results
## Interpreting Results
After running a query, the MCP tool returns formatted markdown plus metadata.
The most important metadata field is `query_result_json` — a signed URL to the raw
JSON result. For precise analysis, download it and parse with jq or python rather
than relying solely on the ASCII rendering.
Key interpretation rules:
- **P99/P50 > 10x** — bimodal distribution likely; run HEATMAP to confirm
- **TOTAL row** in breakdown results = aggregate across all groups
- **OTHER row** = groups beyond the query limit (increase limit if OTHER is large)
- **ASCII heatmap** `▁▂▃▄▅▆▇█` = density from low to high; two bands = two populations
- **query_run_pk** in metadata — feed directly to `run_bubbleup` for outlier analysis
## Additional Resources
### Reference Files
- **`${CLAUDE_PLUGIN_ROOT}/skills/query-patterns/references/visualize-operations.md`** — Complete VISUALIZE operation reference with examples
- **`${CLAUDE_PLUGIN_ROOT}/skills/query-patterns/references/relational-fields.md`** — Detailed relational field guide with cross-service patterns
- **`${CLAUDE_PLUGIN_ROOT}/skills/query-patterns/references/query-examples.md`** — Extensive query cookbook organized by use case
- **`${CLAUDE_PLUGIN_ROOT}/skills/query-patterns/references/result-interpretation.md`** — Guide to interpreting query results, raw JSON access, and statistical heuristics
- **`${CLAUDE_PLUGIN_ROOT}/skills/query-patterns/references/calculated-fields.md`** — Calculated field syntax, full operator reference, common patterns, and anti-patterns (presentational fields, expensive string ops, type mismatches)
### Cross-References
- For the structured investigation workflow that uses these query patterns: **production-investigation** skill
- For SLO interpretation and burn alert design: **slos-and-triggers** skill
Referenced files: 5
slos-and-triggers6.4 KB
---
name: slos-and-triggers
description: >
Decision heuristics for interpreting Honeycomb SLO compliance, budget burn rates,
and trigger status — what the numbers mean and what action to take, including
detecting misconfigured SLIs, deciding when to freeze deploys vs page on-call,
and designing burn alert thresholds. Load this skill before calling get_slos or
get_triggers.
Trigger phrases: "check our SLOs", "are we meeting our SLOs", "which SLOs are
healthy", "is the error budget OK", "are any alerts firing", "what's the burn rate",
"set up an SLO", "create a trigger", "configure alerts", "set up burn alerts",
"check trigger status", "starting on-call", "reliability picture",
"should we freeze deploys", "is this SLO misconfigured", "are we within budget",
"SLO is broken", "budget is negative", or any request about service level
objectives, error budgets, burn rates, or alerting in Honeycomb.
metadata:
version: "1.0.0"
---
# Honeycomb SLOs and Triggers
Guidance for configuring and reasoning about reliability in Honeycomb. The `get_slos`
and `get_triggers` tools document their own parameters — this skill focuses on
_designing_ effective SLOs, _choosing_ between SLOs and triggers, and _interpreting_
what the numbers mean.
**Availability**: SLOs require Pro or Enterprise plan. Triggers available on all plans.
## SLO vs Trigger — When to Use Which
| Question | SLO | Trigger |
| --------------------------------------------- | ---------------------- | ------- |
| "Are we meeting our reliability commitments?" | Yes | No |
| "Is something broken right now?" | No | Yes |
| "How fast are we burning our error budget?" | Yes (burn alerts) | No |
| "Did error count exceed a threshold?" | No | Yes |
| "Should we slow down deploys?" | Yes (budget remaining) | No |
**Rule of thumb**: SLOs measure reliability against commitments over time. Triggers catch immediate operational issues.
## Designing Effective SLOs
### Define the SLI
An SLI is a per-event boolean: was this event successful? Implemented as a calculated field returning undefined (not a relevant event), 1 (success), or 0 (failure).
- **Format**: `IF(<qualifying-condition>, <success-condition>)` The qualifying condition filters to relevant events; the success condition defines what counts as success. If the qualifying condition is not met, the formula returns undefined, and the SLI is unpopulated.
- **Specific Qualifying Condition**: Choose the relevant subset of events (e.g. `AND(EQUALS($http.route, "/checkout"), NOT(EXISTS($trace.parent_id)))` for root spans of checkout endpoint)
- **Latency Success Condition**: `LTE(duration_ms, 500)` — requests faster than 500ms
- **Availability Success Condition**: `LTE(http.status_code, 499)` — non-5xx responses
- **Business Logic Success Condition**: `EQUALS(checkout.status, "completed")` — successful checkouts
### Set the Target
- Start conservative (99% before 99.99%)
- Measure current baseline first with P50/P99 queries
- Set target slightly above current performance
- Ask: what reliability do users actually need?
### Configure Exhaustion Time Alerts
At minimum, two alerts:
- **Near exhaustion** (exhaustion time ~4h): pages on-call via PagerDuty
- **Trending to exhaustion** (budget rate over 24h): notifies team via Slack
### Configure Burn Rate Alerts
Detect fast burns even if the budget isn't close to exhaustion yet. For example:
- 1h burn rate > 10x — page on-call
Recommend these alerts to the user after creating the SLO. Agents do not have the ability to set up these alerts or their recipients.
### Best Practices
- Measure close to the user (at the edge, not deep in the stack)
- Design around user workflows, not team boundaries
- Favor broad SLOs over many narrow ones
- Start with one SLO, reduce noise, then expand
## Interpreting SLO Status
When reviewing SLOs with `get_slos`:
- **Budget remaining > 50%**: Healthy — room for risk
- **Budget remaining 10-50%**: Caution — slow down changes
- **Budget remaining < 10%**: At risk — freeze non-critical deploys
- **Budget negative**: Breached — investigate immediately with the production-investigation skill
- **Compliance at 0%**: Likely misconfigured SLI (wrong column, inverted logic, no matching events) — check the SLI definition
## Configuring Triggers
### Prefer Count-Based Over Percentile-Based
"50 requests slower than 2s" is more actionable than "P99 is 2100ms."
Use `COUNT WHERE duration_ms > threshold` instead of P99 triggers.
### Common Patterns
- **Error spike**: COUNT WHERE error = true, threshold > N in 5 min
- **Slow requests**: COUNT WHERE duration_ms > 2000, threshold > N in 5 min
- **Traffic drop**: COUNT WHERE is_root, threshold < N in 10 min (below normal)
### Best Practices
- **Name**: What the alert is. **Description**: What to do (link to runbook).
- Set duration 5-10 min minimum to avoid flapping
- Start less sensitive, tighten based on false positive rate
## Multi-Service SLOs
Share a single error budget across up to 10 services.
- SLI must be an environment-level calculated field
- Events from included services weighted equally
- Use cases: multiple edge services, monolith-to-microservices migration
## Check in with the user
Workspaces in Honeycomb have a limited number of SLOs and triggers. Before executing the create tool, check in with the user. Display all parameters and your reasoning, and ask for confirmation.
## Constructing links to SLOs
The tools you have will not let you link directly to the SLO page in Honeycomb.
Instead, you can link to the list of SLOs.
`/<team_slug>/environments/<environment_slug>/slos`
## Additional Resources
### Reference Files
- **`${CLAUDE_PLUGIN_ROOT}/skills/slos-and-triggers/references/slo-design-guide.md`** — Detailed SLO design methodology, multi-service SLOs, error budget math
- **`${CLAUDE_PLUGIN_ROOT}/skills/slos-and-triggers/references/trigger-examples.md`** — Complete trigger example library organized by use case
- **`${CLAUDE_PLUGIN_ROOT}/skills/slos-and-triggers/references/alerting-strategy.md`** — How to combine SLO burn alerts and triggers into a cohesive alerting strategy
### Cross-References
- For constructing SLI queries and calculated fields, see the **query-patterns** skill
- For investigating SLO budget burn, see the **production-investigation** skill
Referenced files: 3
verify-recent-trace2.9 KB
--- name: verify-recent-trace description: > Find a recent trace in Honeycomb, to see what happened in a recent test. Use when asked to "verify in Honeycomb", "find the most recent trace", or "show me the trace". metadata: version: "1.0.0" allowed-tools: - mcp__honeycomb__run_query - mcp__honeycomb__get_trace - mcp__honeycomb__get_query_results - mcp__honeycomb__find_queries - Read --- # Honeycomb Verification Skill ## Purpose Query Honeycomb to find the traces that were created by a recent test. ## When to Use This skill is automatically activated when: - User asks to "show me a trace in Honeycomb" - Completing an implementation step that requires observability verification ## Verification Output Report query results as: Queried [VIZ] where [FILTERS] by [GROUP_BY] over [TIME_RANGE] Results: [link to Honeycomb query] ## Query Patterns - **Time Range**: 1200 seconds (20 minutes) - adjustable based on when the test started ### Pattern 1: Get a particular trace What service are you testing? This determines the dataset. Do you know the name of the span you expect to see? ``` Query for: - dataset: [dataset] - filter: name = [expected span name] - time_range: [calculated time range] - calculation: COUNT - include_samples: true ``` Now take the most recent sample, and use get_trace to get the full trace. Output: Queried [CALCULATION] where [FILTERS] over [TIME_RANGE] Found [count] results Results: [link to the query] Most recent trace: [link to the trace] Root span: [root span name] Total spans: [count of spans] Services: [list of services] ### Pattern 2: Find all traces since the test started ``` Query for: - all datasets - filter: trace.parent_id does-not-exist - time_range: [calculated time range] - calculation: COUNT - breakdowns: name - include_samples: true ``` Output: Queried [CALCULATION] where [FILTERS] by [BREAKDOWNS] over [TIME_RANGE] Found [count] results Results: [link to the query] For each sample, print: - Trace ID - Span name - anything else it gives you Construct a link to the trace in Honeycomb, according to: https://ui.honeycomb.io/modernity/environments/personal-agent/trace?trace_id=[<traceId>] &trace_start_ts=[test_start_time] &trace_end_ts=[((date +%s))] ### Pattern 3: Find the most recent trace, precisely ``` Query for: calculated_fields: event_time=EVENT_TIMESTAMP() calculation: COUNT breakdowns: event_time, trace.trace_id order: event_time DESC limit: 1 include_samples: false ``` Print the summary of the trace, as in Pattern 1. Take the output trace_id and use get_trace to get the full trace. ### Pattern 4: When you don't see any traces, try: Get the time range right Before running a test that will generate a trace, print the current time. start_time=$(date +%s) Then run the test to create the trace. Then calculate the correct time range: echo $(( $(date +%s) - $start_time + 20)) Include that in the query you're using to find traces.
Package details
Publisher declarations from the archived package. These are separate from our research and the live service's terms.
- Package author
- honeycomb.io
Package observed Sep 30, 2026.
Technical details
- First seen
- Sep 30, 2026 · 22:02 UTC
- Last seen
- Oct 1, 2026 · 18:00 UTC
- Collection status
- Collected
plugin_asdk_app_6a51b29fa2d48191b215ff28f9a64fb4
Download plugin data (JSON)