← HoneycombCONTENT HISTORY

Update to Honeycomb

Snapshot Sep 30, 2026 · 22:51 UTC · version 1.0.0

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "metrics-queries",
  "description": "How to query OpenTelemetry metrics datasets in Honeycomb correctly. Metrics datasets follow different rules from trace/event datasets — many operations (bare COUNT, RATE_SUM, RATE_AVG, RATE_MAX, CONCURRENCY) are forbidden, temporal aggregation is automatic, and each metric has its own attributes. Use this skill when querying a metrics dataset (gauges, counters, histograms, sums), asking about temporal aggregation (RATE, INCREASE, SUMMARIZE, LAST), finding the metrics dataset or discovering metric names and attributes, debugging unexpected metrics query results, or querying infrastructure metrics like CPU, memory, disk I/O, or network stats. Do NOT use for instrumenting metrics (use otel-instrumentation), querying event datasets with \"metrics\" in their name, or conceptual questions (use observability-fundamentals).\n",
  "included_files": [
    {
      "relative_path": "references/metric-types.md",
      "size_in_bytes": 7756
    },
    {
      "relative_path": "references/metrics-query-examples.md",
      "size_in_bytes": 7837
    },
    {
      "relative_path": "references/temporal-aggregation.md",
      "size_in_bytes": 6766
    }
  ],
  "skill_md_contents": "---\nname: metrics-queries\ndescription: >\n  How to query OpenTelemetry metrics datasets in Honeycomb correctly. Metrics\n  datasets follow different rules from trace/event datasets — many operations\n  (bare COUNT, RATE_SUM, RATE_AVG, RATE_MAX, CONCURRENCY) are forbidden,\n  temporal aggregation is automatic, and each metric has its own attributes.\n  Use this skill when querying a metrics dataset (gauges, counters, histograms,\n  sums), asking about temporal aggregation (RATE, INCREASE, SUMMARIZE, LAST),\n  finding the metrics dataset or discovering metric names and attributes,\n  debugging unexpected metrics query results, or querying infrastructure\n  metrics like CPU, memory, disk I/O, or network stats. Do NOT use for\n  instrumenting metrics (use otel-instrumentation), querying event datasets\n  with \"metrics\" in their name, or conceptual questions (use\n  observability-fundamentals).\nmetadata:\n  version: \"1.0.0\"\n---\n\n# Querying Metrics in Honeycomb\n\nMetrics datasets in Honeycomb behave differently from tracing/event datasets.\nOperations that work on traces may fail or produce misleading results on metrics.\nThis skill covers those differences so you construct correct, useful metrics queries.\n\n## Finding the Metrics Dataset\n\nMetrics datasets are **not** identified by having \"metrics\" in their name. Many event\ndatasets contain \"metrics\" in their slug (e.g., `kafka-metrics`, `refinery-metrics`,\n`kubernetes-node-metrics`). These are ordinary event datasets, not metrics datasets.\n\n**How to identify the real metrics dataset:**\n\n1. Call `get_environment` and look for rows where `dataset_type` = **`metrics`**.\n   The slug is typically `metrics` but may differ per environment.\n2. Alternatively, call `get_dataset_columns` on a candidate dataset — metrics datasets\n   return a **`MetricInfo`** column showing type metadata like `gauge`,\n   `sum(cumulative,monotonic)`, or `histogram(delta)`. Event datasets do not have this.\n\n**Do not guess the dataset.** Always verify via `get_environment` or `get_dataset_columns`\nbefore constructing a metrics query. If the user says \"metrics\" but means an event dataset\nwith metrics-like fields (e.g., `telegraf`, `system_stats`), the query rules below do not apply —\nthose are event datasets and follow normal query patterns from the **query-patterns** skill.\n\n## Discovering Metrics and Their Attributes\n\nMetrics datasets have a fundamentally different schema from event datasets. Each metric\nhas its own set of resource and data point attributes. Two metrics in the same dataset\nmay have completely different attributes available for filtering and grouping.\n\n**Workflow for discovering what to query:**\n\n1. **Find metric names:** Call `get_dataset_columns` on the metrics dataset (without\n   `metric_name`). This returns metric names with their types in `MetricInfo`.\n   Use `find_columns` with keywords to search for specific metrics (e.g., \"cpu\", \"memory\",\n   \"http request duration\").\n\n2. **Find attributes for a specific metric:** Call `get_dataset_columns` with the\n   `metric_name` parameter set to the metric you want to query (e.g.,\n   `metric_name: \"k8s.pod.memory.usage\"`). This returns the resource attributes and\n   data point attributes that co-occur with that metric, along with sample values.\n   These are what you can use in WHERE and GROUP BY clauses.\n\n3. **Validate before querying:** Not all attributes exist on all metrics. Always use\n   step 2 to confirm available attributes before adding them to filters or breakdowns.\n\n## Allowed vs. Forbidden Operations on Metrics Datasets\n\nThe following operations are **NOT allowed** on metrics datasets:\n\n| Forbidden Operation | Why |\n|---------------------|-----|\n| `COUNT` (without column) | Counts metric events, not metric values — meaningless for metrics |\n| `RATE_SUM` | Not supported on metrics datasets |\n| `RATE_AVG` | Not supported on metrics datasets |\n| `RATE_MAX` | Not supported on metrics datasets |\n| `CONCURRENCY` | Requires span duration; metrics have no duration |\n\n**Use these instead:**\n\n| Goal | Use on Metrics |\n|------|----------------|\n| Visualize a gauge value | `AVG(metric)`, `MAX(metric)`, `HEATMAP(metric)` |\n| Visualize a counter/sum | `SUM(metric)`, `AVG(metric)`, `MAX(metric)` |\n| See distribution of values | `HEATMAP(metric)`, `P50(metric)`, `P99(metric)` |\n| Track per-second rate of change | Override temporal aggregation with a calculated field (see below) |\n| Percentile analysis | `P50(metric)`, `P90(metric)`, `P99(metric)` |\n| Count of non-null values | `COUNT(metric)` (with a column specified) |\n\n## Metric Types and Temporal Aggregation\n\nHoneycomb automatically applies temporal aggregation to align raw metric values into\nquery time steps. The function it applies depends on the metric type, visible in the\n`MetricInfo` column from `get_dataset_columns`.\n\n### Default Temporal Aggregation by Metric Type\n\n| MetricInfo | Type | Default Function | What It Does |\n|------------|------|-----------------|--------------|\n| `gauge` | Gauge | `LAST()` | Returns most recent value per time step |\n| `sum(cumulative,monotonic)` | Monotonic cumulative sum | `INCREASE()` | Change between steps, handles counter resets |\n| `sum(cumulative)` | Non-monotonic cumulative sum | `LAST()` | Most recent value (can go up or down) |\n| `sum(delta)` or `sum(delta,monotonic)` | Delta sum | `SUMMARIZE()` | Sums all values in each step |\n| `histogram(cumulative)` | Cumulative histogram | `INCREASE()` | Change per bucket between steps |\n| `histogram(delta)` | Delta histogram | `SUMMARIZE()` | Sums bucket values in each step |\n\nThese defaults are applied automatically — you do not need to configure them.\nThe results you see from `AVG`, `MAX`, `P99`, etc. on a metrics dataset already\nreflect temporal aggregation having been applied first.\n\n### Overriding Temporal Aggregation\n\nTo override the default (e.g., to see RATE instead of INCREASE for a cumulative counter),\nuse a **query-scoped calculated field** wrapping the metric name in a temporal aggregation\nfunction, then apply a spatial aggregation to that field in `calculations`.\n\n```json\n{\n  \"calculated_fields\": [\n    { \"name\": \"req_rate\", \"expression\": \"RATE($http.server.requests, 300)\" }\n  ],\n  \"calculations\": [\n    { \"op\": \"AVG\", \"column\": \"req_rate\" }\n  ]\n}\n```\n\nSupported temporal aggregation functions for calculated fields:\n- **`LAST($metric)`** — most recent data point per step (gauges, non-monotonic sums)\n- **`SUMMARIZE($metric)`** — sum all values per step with interpolation (delta metrics)\n- **`INCREASE($metric[, range_interval_seconds])`** — change in value across range, handles counter resets\n- **`RATE($metric[, range_interval_seconds])`** — per-second rate of change (`INCREASE / time`)\n\nThe optional `range_interval_seconds` parameter (integer, in seconds) controls the lookback\nwindow for calculating changes. Use it to smooth results or compensate for sparse data.\nWhen omitted, the query's granularity is used as the range interval.\n\n**Important:** You must still apply a spatial aggregation (`AVG`, `SUM`, `P99`, `HEATMAP`, etc.)\nto the calculated field in `calculations`. The temporal aggregation function alone does not\nproduce a visualization — it transforms the raw metric values, then the spatial aggregation\nsummarizes across timeseries.\n\nFor detailed reference on temporal aggregation functions, counter reset handling, and\n`range_interval_seconds`, see:\n`${CLAUDE_PLUGIN_ROOT}/skills/metrics-queries/references/temporal-aggregation.md`\n\n## Querying Histogram Metrics\n\nOpenTelemetry histograms are stored as a collection of sub-fields. For a histogram\nnamed `http.server.duration`, Honeycomb creates:\n\n| Field | Meaning |\n|-------|---------|\n| `http.server.duration.count` | Total number of data points |\n| `http.server.duration.sum` | Sum of all values |\n| `http.server.duration.avg` | Mean value (sum/count) |\n| `http.server.duration.p50` | Median (50th percentile) |\n| `http.server.duration.p99` | 99th percentile |\n| `http.server.duration.p001` through `.p999` | Full range of percentiles |\n\n**Two ways to query histograms:**\n\n1. **Use the parent column name directly** with percentile or distribution operations.\n   This is the recommended approach:\n   ```json\n   { \"op\": \"P99\", \"column\": \"http.server.duration\" }\n   ```\n   ```json\n   { \"op\": \"HEATMAP\", \"column\": \"http.server.duration\" }\n   ```\n\n2. **Use sub-fields with MAX** when you want the worst-case pre-computed percentile\n   across all timeseries in a step:\n   ```json\n   { \"op\": \"MAX\", \"column\": \"http.server.duration.p99\" }\n   ```\n   This returns the highest p99 value reported by any single timeseries in the time step,\n   which differs from `P99(http.server.duration)` which computes the 99th percentile\n   across all data.\n\n**When to use which:**\n- For most analysis: use `P99(parent_column)` or `HEATMAP(parent_column)`\n- For worst-case bounds across hosts/pods: use `MAX(parent_column.p99)`\n- For throughput from histograms: use `SUM(parent_column.count)` or `AVG(parent_column.count)`\n\n## Query Math with Metrics\n\nQuery math (compound queries with named calculations and formulas) works on metrics\ndatasets the same way it works on event datasets. Name your calculations, add\nper-calculation filters if needed, and define formulas to combine them.\n\n**Common metrics formula patterns:**\n\n### Utilization percentage\n```json\n{\n  \"calculations\": [\n    { \"op\": \"AVG\", \"column\": \"k8s.pod.memory.usage\", \"name\": \"used\" },\n    { \"op\": \"AVG\", \"column\": \"k8s.pod.memory.available\", \"name\": \"available\" }\n  ],\n  \"formulas\": [\n    { \"name\": \"utilization_pct\", \"expression\": \"$used / ($used + $available) * 100\" }\n  ],\n  \"breakdowns\": [\"k8s.pod.name\"],\n  \"orders\": [{ \"column\": \"utilization_pct\", \"order\": \"descending\" }],\n  \"limit\": 20\n}\n```\n\n### Histogram tail ratio\n```json\n{\n  \"calculations\": [\n    { \"op\": \"P50\", \"column\": \"http.server.duration\", \"name\": \"median\" },\n    { \"op\": \"P99\", \"column\": \"http.server.duration\", \"name\": \"tail\" }\n  ],\n  \"formulas\": [\n    { \"name\": \"tail_ratio\", \"expression\": \"$tail / $median\" }\n  ],\n  \"breakdowns\": [\"service.name\"]\n}\n```\n\n### Error rate from counters (with temporal aggregation override)\n```json\n{\n  \"calculated_fields\": [\n    { \"name\": \"error_rate\", \"expression\": \"RATE($http.server.errors)\" },\n    { \"name\": \"request_rate\", \"expression\": \"RATE($http.server.requests)\" }\n  ],\n  \"calculations\": [\n    { \"op\": \"SUM\", \"column\": \"error_rate\", \"name\": \"errors_per_sec\" },\n    { \"op\": \"SUM\", \"column\": \"request_rate\", \"name\": \"requests_per_sec\" }\n  ],\n  \"formulas\": [\n    { \"name\": \"error_pct\", \"expression\": \"$errors_per_sec / $requests_per_sec * 100\" }\n  ]\n}\n```\n\nFor more query examples, see:\n`${CLAUDE_PLUGIN_ROOT}/skills/metrics-queries/references/metrics-query-examples.md`\n\n## Granularity for Metrics\n\nMetrics arrive at known, regular intervals (e.g., every 10s, 30s, or 60s). Granularity\nmatters more for metrics than for traces:\n\n- **Align granularity with the reporting interval.** If metrics report every 60 seconds,\n  use a granularity that divides evenly into 60 (e.g., 60, 120, 300). Misaligned\n  granularity causes uneven bucket sizes that produce noisy results.\n- **Spiky-looking graphs** usually mean the granularity is finer than the reporting interval.\n  Increase granularity or, in the UI, enable \"Omit Missing Values\" to produce continuous lines.\n- **RATE operations and granularity:** `RATE_SUM` (on event datasets) is particularly sensitive\n  to granularity choice — inconsistent data points per bucket produce variable results.\n\n## Common Pitfalls\n\n1. **Using `COUNT` on metrics.** `COUNT` counts the number of metric *events*, not the metric\n   value. Use `AVG`, `SUM`, `MAX`, or `HEATMAP` instead.\n2. **Using `RATE_AVG`/`RATE_SUM`/`RATE_MAX` on metrics datasets.** These are not allowed.\n   To get a rate, use a calculated field with `RATE($metric)` and then apply a spatial\n   aggregation like `AVG` or `SUM`.\n3. **Assuming all metrics share the same attributes.** Each metric has its own set of\n   resource and data point attributes. Always call `get_dataset_columns` with `metric_name`\n   to discover what's available for a specific metric before adding filters or breakdowns.\n4. **Confusing event datasets with the metrics dataset.** Datasets named `kafka-metrics`,\n   `refinery-metrics`, etc. are event datasets. Check `dataset_type` from `get_environment`.\n5. **Querying histogram sub-fields when the parent column works.** Use `P99(http.server.duration)`\n   rather than `AVG(http.server.duration.p99)` unless you specifically need worst-case bounds.\n6. **Not specifying an aggregate function.** Metrics queries without a spatial aggregation\n   in SELECT default to `COUNT`, which is meaningless for metrics.\n\n## Additional Resources\n\n### Reference Files\n- **`${CLAUDE_PLUGIN_ROOT}/skills/metrics-queries/references/metrics-query-examples.md`** — Metrics query cookbook with run_query examples for common scenarios\n- **`${CLAUDE_PLUGIN_ROOT}/skills/metrics-queries/references/temporal-aggregation.md`** — Deep reference on temporal aggregation functions, counter resets, and range_interval_seconds\n- **`${CLAUDE_PLUGIN_ROOT}/skills/metrics-queries/references/metric-types.md`** — OpenTelemetry metric types, how they map to Honeycomb, and what the MetricInfo values mean\n\n### Cross-References\n- For general query construction patterns (calculated fields, relational fields, result interpretation): **query-patterns** skill\n- For investigating production issues using metrics alongside traces: **production-investigation** skill\n- For SLO interpretation and burn alert design: **slos-and-triggers** skill\n- For instrumenting applications to send metrics: **otel-instrumentation** skill\n"
}

SHA-256: 0287f04345c5c96f0654aeec1194422b0e9fa24d350d5cc5cc67531edfbe0a7d