← Files Go: Production EngineeringARCHIVED FILE
skills/go-production-operations/references/incident-contract.md
1.17 KB · Oct 5, 2026 · 18:32 UTC
# Incident contract An operator should be able to distinguish: - no admission versus slow processing; - caller cancellation versus internal deadline; - dependency rejection versus local overload; - retry attempts versus original operations; - known failure versus ambiguous outcome; - unready drain versus crashed process; - dropped telemetry versus no events. During recovery, ramp traffic, retries, reconnects, and cache fill so the recovery path does not reproduce overload. ## Telemetry under failure - Bound metric label domains; a tenant, user, raw URL, request ID, or error string can create an unbounded series set. - Bound exporter queues and batches. Decide whether request-path emission drops, samples, or blocks, and cap any blocking with the operation budget. - Count dropped, truncated, sampled, and rejected telemetry outside the failing export path where possible. - Treat logs, metric attributes, span events, and exception bodies as data egress paths subject to classification and size limits. - Flush within the shutdown budget. Do not close the exporter while owned producers can still enqueue, and do not let an unreachable collector block process exit indefinitely.
SHA-256: 44a417ad20df76c1ad67cc2b03e697cdd67c8a873f1c2d43b75739a2752cbb65