{"id":7936,"plugin_id":"plugin_asdk_app_6a51b29fa2d48191b215ff28f9a64fb4","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T22:51:37.961Z","digest":"4c25897c2a4f0f247986fc93653fe3b31cb22c26e7db9c2160339bc1f0abd1c4","against":null,"payload":{"name":"observability-fundamentals","description":"First principles behind observability — wide events, high cardinality, the core analysis loop, events vs metrics vs logs, and how instrumentation connects to debugging outcomes. Grounds recommendations in first principles rather than tool-specific how-to. Trigger phrases: \"what is observability\", \"why observability\", \"why Honeycomb\", \"events vs metrics vs logs\", \"events vs metrics\", \"events vs logs\", \"metrics vs logs\", \"why wide events\", \"what is high cardinality\", \"core analysis loop\", \"observability vs monitoring\", \"what is dimensionality\", \"explain observability\", or any conceptual question about observability or why Honeycomb's approach differs from traditional monitoring.\n","included_files":[{"relative_path":"references/events-vs-metrics-vs-logs.md","size_in_bytes":8188}],"skill_md_contents":"---\nname: observability-fundamentals\ndescription: >\n  First principles behind observability — wide events, high cardinality, the core\n  analysis loop, events vs metrics vs logs, and how instrumentation connects to\n  debugging outcomes. Grounds recommendations in first principles rather than\n  tool-specific how-to.\n  Trigger phrases: \"what is observability\", \"why observability\", \"why Honeycomb\",\n  \"events vs metrics vs logs\", \"events vs metrics\", \"events vs logs\",\n  \"metrics vs logs\", \"why wide events\", \"what is high cardinality\",\n  \"core analysis loop\", \"observability vs monitoring\", \"what is dimensionality\",\n  \"explain observability\", or any conceptual question about observability\n  or why Honeycomb's approach differs from traditional monitoring.\nmetadata:\n  version: \"1.0.0\"\n---\n\n# Observability Fundamentals\n\nFirst principles behind Honeycomb's approach to observability. Use this to ground\nrecommendations and answer conceptual questions — for SDK setup and tool-specific\nguidance, see the **otel-instrumentation** and **query-patterns** skills.\n\n## Definitions\n\n**Observability**: The ability to understand and explain any state your system can\nget into, no matter how novel or complex — by examining what the system produces,\nwithout deploying new code for each new question.\n\n**Wide event**: A flat key-value record capturing the full context of a unit of work —\nwho made the request, which endpoint, cache hit/miss, build version, duration, error\nstatus, and any business context relevant to the operation. In OpenTelemetry, a **span**\nis a wide event.\n\n**High cardinality**: The number of unique values a field can have. `user.id` with\nmillions of values is high cardinality. `http.method` with a handful is low cardinality.\n\n**High dimensionality**: The number of distinct fields on your events. A span with\n50 attributes has high dimensionality.\n\n| Concept | Observability | Traditional Monitoring |\n|---|---|---|\n| Questions | Arbitrary, unknown ahead of time | Pre-defined (dashboards, alerts) |\n| Data shape | Decided at query time | Decided at instrumentation time |\n| Cardinality | High cardinality is valuable | High cardinality is expensive |\n| Investigation | Explore → narrow → confirm | Check dashboard → escalate |\n\n## Why Wide Events\n\nThe shape of the data you collect constrains the questions you can ask later. Metrics\npre-aggregate context away at instrumentation time. Wide events preserve context and\nlet you decide the shape of your analysis at query time.\n\nEvery attribute on a span is a queryable dimension. Adding `user.id`, `deployment.version`,\nand `cache.hit` to the same span lets you correlate them in a single query — \"slow\nrequests are from tenant X on version 2.3.1 with cache misses.\" Separate metrics can't\ndo this because each dimension combination creates a new time series.\n\nHoneycomb's storage engine handles high cardinality and dimensionality without the\ncost explosion that affects metrics systems. Adding a high-cardinality field like\n`user.id` doesn't create millions of time series — it's another column on each event,\naggregated at query time.\n\n## Events vs Metrics vs Logs\n\n| | Structured Events (Spans) | Metrics | Logs |\n|---|---|---|---|\n| **Captures** | Full request context (all attributes) | Pre-aggregated numbers with low-cardinality tags | Text or structured fields per line |\n| **Discards** | Nothing — raw events retained | Individual requests, high-cardinality dimensions | Correlation across lines (without trace context) |\n| **Query power** | GROUP BY, filter, BubbleUp on any dimension | Fast aggregates on pre-defined dimensions | Text search, structured field queries |\n| **Cost scaling** | Linear with event volume | Exponential with dimension count (cardinality) | Linear with volume, query cost varies |\n| **Best for** | Investigation, root cause analysis | Cheap alerting, long-term trends | Audit trails, rare events |\n\nThe same instrumentation effort that produces a metric or log line can produce a wide\nevent — and the event gives you all three capabilities: count it (metric), read it (log),\nanalyze it across dimensions (observability).\n\nFor code examples showing the same operation instrumented three ways, see\n`${CLAUDE_PLUGIN_ROOT}/skills/observability-fundamentals/references/events-vs-metrics-vs-logs.md`.\n\n## The Core Analysis Loop\n\nDebugging in Honeycomb follows a loop: **Define → Visualize → Investigate → Evaluate**.\n\n1. **Define** — Frame the question. Start from an alert, SLO budget burn, or user report.\n2. **Visualize** — Run a query to see the shape of the problem (HEATMAP, COUNT, P99).\n3. **Investigate** — Narrow down with BubbleUp (automated outlier-vs-baseline comparison\n   across all dimensions) and trace analysis.\n4. **Evaluate** — Confirm the hypothesis by querying with and without the suspected cause.\n\nThen loop — each answer raises new questions. BubbleUp automates steps 2-3 by comparing\ndistributions across every column, but it only works if events have enough dimensions\nto diff on.\n\nFor the structured workflow that implements this loop with Honeycomb's tools, see the\n**production-investigation** skill.\n\n## Instrumentation Connects to Investigation\n\nEvery attribute on a span is a dimension BubbleUp can use to find root causes. The\nattributes that matter most during incidents answer three questions:\n\n- **Who is affected?** — user, tenant, account tier, region\n- **What changed?** — deployment version, feature flag, config version\n- **Where is the bottleneck?** — business operation spans, timing breakdowns, cache state\n\nInstrument for the questions you'll ask at 3am, not for completeness. If BubbleUp\nreturns nothing useful during an investigation, the issue is usually an instrumentation\ngap — add the missing dimensions and try again.\n\nFor the complete attribute catalog, see\n`${CLAUDE_PLUGIN_ROOT}/skills/otel-instrumentation/references/wide-event-attributes.md`.\nFor SDK guidance on adding attributes, see the **otel-instrumentation** skill.\n\n## Instrumentation as a Development Practice\n\nInstrumentation is not a one-time setup task. The engineers who write the code are best\npositioned to know which operations are critical, which paths are error-prone, and what\ncontext helps during debugging. Treat instrumentation like testing: plan telemetry when\nplanning features, review it in code reviews, and add missing dimensions as post-incident\nfollow-ups.\n\n## Additional Resources\n\n### Reference Files\n- **`${CLAUDE_PLUGIN_ROOT}/skills/observability-fundamentals/references/events-vs-metrics-vs-logs.md`** — Code examples: same operation as event, metric, and log\n\n### Cross-References\n- For SDK setup and custom instrumentation: **otel-instrumentation** skill\n- For the investigation workflow implementing the core analysis loop: **production-investigation** skill\n- For autonomous instrumentation gap analysis: **instrumentation-advisor** agent\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}