← HoneycombCONTENT HISTORY

Update to Honeycomb

Snapshot Sep 30, 2026 · 22:51 UTC · version 1.0.0

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "slos-and-triggers",
  "description": "Decision heuristics for interpreting Honeycomb SLO compliance, budget burn rates, and trigger status — what the numbers mean and what action to take, including detecting misconfigured SLIs, deciding when to freeze deploys vs page on-call, and designing burn alert thresholds. Load this skill before calling get_slos or get_triggers. Trigger phrases: \"check our SLOs\", \"are we meeting our SLOs\", \"which SLOs are healthy\", \"is the error budget OK\", \"are any alerts firing\", \"what's the burn rate\", \"set up an SLO\", \"create a trigger\", \"configure alerts\", \"set up burn alerts\", \"check trigger status\", \"starting on-call\", \"reliability picture\", \"should we freeze deploys\", \"is this SLO misconfigured\", \"are we within budget\", \"SLO is broken\", \"budget is negative\", or any request about service level objectives, error budgets, burn rates, or alerting in Honeycomb.\n",
  "included_files": [
    {
      "relative_path": "references/alerting-strategy.md",
      "size_in_bytes": 2859
    },
    {
      "relative_path": "references/slo-design-guide.md",
      "size_in_bytes": 5339
    },
    {
      "relative_path": "references/trigger-examples.md",
      "size_in_bytes": 4305
    }
  ],
  "skill_md_contents": "---\nname: slos-and-triggers\ndescription: >\n  Decision heuristics for interpreting Honeycomb SLO compliance, budget burn rates,\n  and trigger status — what the numbers mean and what action to take, including\n  detecting misconfigured SLIs, deciding when to freeze deploys vs page on-call,\n  and designing burn alert thresholds. Load this skill before calling get_slos or\n  get_triggers.\n  Trigger phrases: \"check our SLOs\", \"are we meeting our SLOs\", \"which SLOs are\n  healthy\", \"is the error budget OK\", \"are any alerts firing\", \"what's the burn rate\",\n  \"set up an SLO\", \"create a trigger\", \"configure alerts\", \"set up burn alerts\",\n  \"check trigger status\", \"starting on-call\", \"reliability picture\",\n  \"should we freeze deploys\", \"is this SLO misconfigured\", \"are we within budget\",\n  \"SLO is broken\", \"budget is negative\", or any request about service level\n  objectives, error budgets, burn rates, or alerting in Honeycomb.\nmetadata:\n  version: \"1.0.0\"\n---\n\n# Honeycomb SLOs and Triggers\n\nGuidance for configuring and reasoning about reliability in Honeycomb. The `get_slos`\nand `get_triggers` tools document their own parameters — this skill focuses on\n_designing_ effective SLOs, _choosing_ between SLOs and triggers, and _interpreting_\nwhat the numbers mean.\n\n**Availability**: SLOs require Pro or Enterprise plan. Triggers available on all plans.\n\n## SLO vs Trigger — When to Use Which\n\n| Question                                      | SLO                    | Trigger |\n| --------------------------------------------- | ---------------------- | ------- |\n| \"Are we meeting our reliability commitments?\" | Yes                    | No      |\n| \"Is something broken right now?\"              | No                     | Yes     |\n| \"How fast are we burning our error budget?\"   | Yes (burn alerts)      | No      |\n| \"Did error count exceed a threshold?\"         | No                     | Yes     |\n| \"Should we slow down deploys?\"                | Yes (budget remaining) | No      |\n\n**Rule of thumb**: SLOs measure reliability against commitments over time. Triggers catch immediate operational issues.\n\n## Designing Effective SLOs\n\n### Define the SLI\n\nAn SLI is a per-event boolean: was this event successful? Implemented as a calculated field returning undefined (not a relevant event), 1 (success), or 0 (failure).\n\n- **Format**: `IF(<qualifying-condition>, <success-condition>)` The qualifying condition filters to relevant events; the success condition defines what counts as success. If the qualifying condition is not met, the formula returns undefined, and the SLI is unpopulated.\n- **Specific Qualifying Condition**: Choose the relevant subset of events (e.g. `AND(EQUALS($http.route, \"/checkout\"), NOT(EXISTS($trace.parent_id)))` for root spans of checkout endpoint)\n- **Latency Success Condition**: `LTE(duration_ms, 500)` — requests faster than 500ms\n- **Availability Success Condition**: `LTE(http.status_code, 499)` — non-5xx responses\n- **Business Logic Success Condition**: `EQUALS(checkout.status, \"completed\")` — successful checkouts\n\n### Set the Target\n\n- Start conservative (99% before 99.99%)\n- Measure current baseline first with P50/P99 queries\n- Set target slightly above current performance\n- Ask: what reliability do users actually need?\n\n### Configure Exhaustion Time Alerts\n\nAt minimum, two alerts:\n\n- **Near exhaustion** (exhaustion time ~4h): pages on-call via PagerDuty\n- **Trending to exhaustion** (budget rate over 24h): notifies team via Slack\n\n### Configure Burn Rate Alerts\n\nDetect fast burns even if the budget isn't close to exhaustion yet. For example:\n\n- 1h burn rate > 10x — page on-call\n\nRecommend these alerts to the user after creating the SLO. Agents do not have the ability to set up these alerts or their recipients.\n\n### Best Practices\n\n- Measure close to the user (at the edge, not deep in the stack)\n- Design around user workflows, not team boundaries\n- Favor broad SLOs over many narrow ones\n- Start with one SLO, reduce noise, then expand\n\n## Interpreting SLO Status\n\nWhen reviewing SLOs with `get_slos`:\n\n- **Budget remaining > 50%**: Healthy — room for risk\n- **Budget remaining 10-50%**: Caution — slow down changes\n- **Budget remaining < 10%**: At risk — freeze non-critical deploys\n- **Budget negative**: Breached — investigate immediately with the production-investigation skill\n- **Compliance at 0%**: Likely misconfigured SLI (wrong column, inverted logic, no matching events) — check the SLI definition\n\n## Configuring Triggers\n\n### Prefer Count-Based Over Percentile-Based\n\n\"50 requests slower than 2s\" is more actionable than \"P99 is 2100ms.\"\nUse `COUNT WHERE duration_ms > threshold` instead of P99 triggers.\n\n### Common Patterns\n\n- **Error spike**: COUNT WHERE error = true, threshold > N in 5 min\n- **Slow requests**: COUNT WHERE duration_ms > 2000, threshold > N in 5 min\n- **Traffic drop**: COUNT WHERE is_root, threshold < N in 10 min (below normal)\n\n### Best Practices\n\n- **Name**: What the alert is. **Description**: What to do (link to runbook).\n- Set duration 5-10 min minimum to avoid flapping\n- Start less sensitive, tighten based on false positive rate\n\n## Multi-Service SLOs\n\nShare a single error budget across up to 10 services.\n\n- SLI must be an environment-level calculated field\n- Events from included services weighted equally\n- Use cases: multiple edge services, monolith-to-microservices migration\n\n## Check in with the user\n\nWorkspaces in Honeycomb have a limited number of SLOs and triggers. Before executing the create tool, check in with the user. Display all parameters and your reasoning, and ask for confirmation.\n\n## Constructing links to SLOs\n\nThe tools you have will not let you link directly to the SLO page in Honeycomb.\nInstead, you can link to the list of SLOs.\n\n`/<team_slug>/environments/<environment_slug>/slos`\n\n## Additional Resources\n\n### Reference Files\n\n- **`${CLAUDE_PLUGIN_ROOT}/skills/slos-and-triggers/references/slo-design-guide.md`** — Detailed SLO design methodology, multi-service SLOs, error budget math\n- **`${CLAUDE_PLUGIN_ROOT}/skills/slos-and-triggers/references/trigger-examples.md`** — Complete trigger example library organized by use case\n- **`${CLAUDE_PLUGIN_ROOT}/skills/slos-and-triggers/references/alerting-strategy.md`** — How to combine SLO burn alerts and triggers into a cohesive alerting strategy\n\n### Cross-References\n\n- For constructing SLI queries and calculated fields, see the **query-patterns** skill\n- For investigating SLO budget burn, see the **production-investigation** skill\n"
}

SHA-256: f7a5b1268df033377430d62f6fbc6b06ec9efa3817aa57615adc921a1b961b57