← Files Platform Engineering CopilotARCHIVED FILE
skills/platform-engineering/references/observability_telemetry_patterns.md
1.62 KB · Oct 2, 2026 · 00:37 UTC
# Observability and Telemetry Patterns ## Goal Telemetry should answer operational questions. ## Signals ### Traces Useful for request paths, dependency latency, and distributed causality. ### Metrics Useful for aggregations, SLOs, rates, saturation, and alerts. ### Logs Useful for detailed events, errors, and diagnostic context. Correlate signals using stable request/trace/service/deployment identifiers. ## OpenTelemetry OpenTelemetry provides a vendor-neutral model/tooling for traces, metrics, and logs. Use when it fits the stack to reduce instrumentation fragmentation and preserve backend choice. ## Service telemetry Common dimensions: - service name; - environment; - version/deployment; - region/zone; - status/error category; - dependency. Be careful with: - user IDs; - request URLs with unbounded values; - high-cardinality labels; - secrets/tokens; - payload bodies. ## Golden signals Useful starting concepts: - latency; - traffic; - errors; - saturation. Use business-specific signals where platform metrics alone miss user impact. ## Dashboards A dashboard should support a decision: - Is service healthy? - Is rollout safe? - Where is latency? - Are errors localized? - Is capacity saturated? Avoid dashboards that only display everything available. ## Alerts Alert on conditions requiring action. Link: - runbook; - owner; - service; - environment; - useful context. ## Cost Telemetry can become expensive. Control: - sampling; - log level; - retention; - metric cardinality; - duplicate collection; - export destinations. Do not reduce incident-critical signals merely to cut cost without understanding the risk.
SHA-256: 955bb4df42fc6e091488f201830fcc04830bb0fa4afe64ef4aa45f5b93bc63f3