← Go: Distributed SystemsCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to Go: Distributed Systems
Snapshot Sep 30, 2026 · 23:15 UTC · version 0.4.0
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"description": "Use for cross-service Go deadlines, retries, admission, shedding, and recovery. Do not use for commits, brokers, or local speed.",
"included_files": [
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 264
},
{
"relative_path": "evals.json",
"size_in_bytes": 8360
},
{
"relative_path": "references/overload-signals-and-scope.md",
"size_in_bytes": 3206
},
{
"relative_path": "references/overload.md",
"size_in_bytes": 1552
},
{
"relative_path": "references/retries.md",
"size_in_bytes": 3166
},
{
"relative_path": "skill.json",
"size_in_bytes": 3272
}
],
"name": "go-service-resilience",
"skill_md_contents": "---\nname: go-service-resilience\ndescription: \"Use for cross-service Go deadlines, retries, admission, shedding, and recovery. Do not use for commits, brokers, or local speed.\"\nlicense: Apache-2.0\ncompatibility: \"Go 1.24 or newer. Policies are transport-neutral; verify client and dependency behavior against repository versions.\"\n---\n\n# Go service resilience\n\nDesign remote calls so a small dependency failure does not become unbounded work. Timeouts, retries, queues, and concurrency limits form one feedback system; configuring them independently can amplify the outage they were meant to survive.\n\n## Map the call path\n\nInspect the caller, transport, dependency contract, deployment topology, and existing resilience layer. Record:\n\n1. the end-to-end latency or deadline objective;\n2. every queue and attempt along the call graph;\n3. which layer owns retries;\n4. whether the operation is safe to replay and how identity is preserved;\n5. maximum in-flight work and downstream capacity;\n6. which failures are transient, permanent, overload, or ambiguous;\n7. the fallback’s correctness and capacity cost.\n\nDo not add retries before finding existing retries in clients, proxies, service meshes, jobs, and callers.\n\n## Allocate one deadline budget\n\nPropagate the caller’s context. Derive a shorter deadline only when this layer owns a sub-budget. Never extend the caller’s deadline by replacing the context.\n\nBefore each attempt, account for:\n\n```text\nremaining budget > queue wait + attempt bound + possible backoff + response margin\n```\n\nIf not, fail without starting work the caller can no longer use. A transport timeout covers one phase; it does not replace an end-to-end deadline.\n\nUse separate policies for interactive requests, batch work, and long-lived streams. “30 seconds everywhere” is not a resilience design.\n\n## Decide whether retry is legal\n\nRetry only if all conditions hold:\n\n- the failure class is plausibly transient;\n- the operation can be replayed with the same semantic identity;\n- this layer is the designated retry owner;\n- the dependency has capacity for retry traffic;\n- attempts and elapsed time are bounded;\n- the remaining caller budget can complete another useful attempt.\n\nApplication idempotency matters more than HTTP method names. A `GET` can trigger a broken side effect; a `POST` can be replay-safe under a durable idempotency contract.\n\nAmbiguous outcomes require reconciliation or same-identity replay, not a new operation.\n\nRead [references/retries.md](references/retries.md) for classification, backoff, jitter, and hedging. Treat every hedge as another live attempt inside one operation budget, and verify actual Go client support instead of assuming a cross-language service-config feature exists.\n\n## Control amplification\n\nIf a five-deep call chain makes three total attempts at every layer, one user request can create up to 243 leaf attempts. If “retry three times” means three retries after the first attempt, the bound is 4^5 = 1,024. Prefer one retry owner near the layer that understands replay semantics and user budget.\n\nUse exponential backoff with jitter to desynchronize callers, but remember:\n\n- backoff does not reduce the first retry wave;\n- a high retry cap still creates excess load;\n- retrying overload responses can keep a dependency overloaded;\n- unbounded queued retries consume memory and expire before execution.\n\nApply a retry budget or token limit so retries remain a controlled fraction of normal traffic. Observe attempts per original operation, not only raw request count.\n\n## Bound concurrency and queues\n\nAdmission control should happen before allocating expensive state or spawning work. Choose a bound from dependency capacity and tolerated queue latency, then choose overload behavior:\n\n- reject early and cheaply;\n- shed lower-priority work;\n- degrade optional work;\n- queue only within an explicit memory and latency limit.\n\nSeparate concurrency pools only when isolation matches real failure domains. Per-dependency bulkheads can stop one slow dependency from consuming every worker, but too many pools strand capacity and complicate fairness.\n\nBound both logical operations and active dependency attempts. An operation-lifecycle permit bounds callers sleeping in backoff; an attempt permit bounds active dependency load and is released before backoff, then reacquired cancellation-aware. Holding scarce dependency capacity during sleep starves fresh work and recovery probes; releasing it without an operation bound creates unbounded sleepers.\n\nRead [references/overload.md](references/overload.md) for load shedding, circuit behavior, and recovery. Read [references/overload-signals-and-scope.md](references/overload-signals-and-scope.md) when admission spans replicas, tenants, priorities, proxies, or protocol retry signals.\n\n## Treat circuit breakers as state machines\n\nA circuit breaker can reduce futile load, but it introduces shared state, thresholds, probes, and synchronized recovery. Add one only when simpler bounded concurrency, deadlines, and retry control are insufficient.\n\nDefine:\n\n- which failures count;\n- sample size and open threshold;\n- open duration and probe concurrency;\n- behavior for callers while open;\n- how recovery avoids a thundering herd;\n- per-instance versus shared scope.\n\nDo not use a circuit breaker to hide incorrect timeouts or as a substitute for dependency health signals.\n\n## Validate fallbacks\n\nA fallback is another production path. Verify:\n\n- its data freshness and consistency semantics;\n- authorization and privacy equivalence;\n- capacity during the same outage;\n- whether it turns a hard failure into silently wrong data;\n- how callers and telemetry distinguish degraded results.\n\nCaches can be latency optimizations or capacity dependencies. A cache that all instances fall through simultaneously can make recovery worse. Define stampede control and stale-data policy explicitly.\n\n## Preserve error semantics\n\nReturn errors that let the owning boundary decide:\n\n- retryable versus permanent when the API intentionally exposes that contract;\n- overload versus dependency unavailability;\n- deadline exceeded versus caller cancellation;\n- ambiguous outcome versus confirmed rejection.\n\nDo not string-match error messages. Do not log at every layer. Include attempt metadata in telemetry without wrapping the same cause into an unreadable chain.\n\n## Recovery and rollout\n\nDesign for recovery, not only steady failure:\n\n- ramp traffic or probes so cold caches and reconnect storms do not re-trigger overload;\n- randomize periodic work and reconnects;\n- drain retry queues under a controlled rate;\n- keep health checks cheap and separate readiness from liveness;\n- test configuration changes because retry/timeout policy is production code.\n\n## Finish with evidence\n\nWithin the request’s authority, test the feedback system:\n\n- latency beyond the attempt and caller deadline;\n- transient failure followed by recovery;\n- permanent and ambiguous failures;\n- partial dependency capacity and overload rejection;\n- retry storms from many callers;\n- fallback saturation and stale data;\n- recovery with cold connections or caches.\n\nRecord attempts per operation, in-flight work, queue time, rejected work, dependency latency, and terminal outcomes. A passing unit test for a backoff function does not validate overload behavior.\n\n## Output contract\n\nFor implementation, keep retry ownership, replay identity, deadline allocation, and capacity bounds close to the call. For review, give the amplification path or failure schedule and quantify the maximum attempts/in-flight work where possible. Do not prescribe circuit breakers or retries by default.\n"
}SHA-256 of public snapshot: 44451046d4b9f7dcecb787f59e0a65c20948923533765d6b5e7cf86852503189