← Plugin catalog
Developer Tools
Go: Distributed Systems
Ashwin Gopalsamy v0.4.0
Publisher description
From the marketplace listing
Invariant-driven Go concurrency, consistency, messaging, resilience, coordination, and distributed-change review skills.
Language: English · Automatically detected from descriptions.
Files & skills
File archives
Plugin package49 files · 712 KBBrowse files →
Skill instructions
go-concurrency-lifecycle7.13 KB
---
name: go-concurrency-lifecycle
description: "Use for in-process Go concurrency ownership, cancellation, races, leaks, synchronization, and bounds. Do not use for broker semantics."
license: Apache-2.0
compatibility: "Go 1.24 or newer. Guidance targets the stable Go 1.25 and 1.26 families and degrades to older repository versions when required."
---
# Go concurrency lifecycle
Make concurrent code explain who owns each goroutine, what can block it, and how it terminates. Treat channels, mutexes, atomics, and contexts as tools with different contracts—not as a hierarchy of idiomatic preference.
## Establish the contract
Inspect the call path and repository conventions before proposing a pattern. Write down only the invariants relevant to the change:
1. **Ownership:** which component starts the goroutine and waits for or stops it?
2. **Lifetime:** is it request-scoped, operation-scoped, component-scoped, or process-scoped?
3. **Termination:** enumerate every blocking point and the event that releases it.
4. **Failure:** where does an error go, and which sibling work should it cancel?
5. **Capacity:** what bounds goroutines, queued work, memory, and downstream concurrency?
6. **State:** which data is shared, and what establishes happens-before ordering?
Do not add concurrency until the expected latency or ownership benefit justifies the extra state space.
## Choose the synchronization mechanism
Choose from the invariant, not a slogan:
| Need | Default starting point |
| --- | --- |
| Protect a small in-memory invariant | `sync.Mutex` guarding the data |
| Publish independent read-mostly snapshots | immutable value plus `atomic.Pointer` |
| Transfer ownership or coordinate a stream | channel with documented producer and close owner |
| Wait for a fixed set of goroutines | `sync.WaitGroup` or an existing repository error group |
| Cancel work derived from an operation | propagated `context.Context` |
| Bound parallel calls | fixed worker count or semaphore acquired before spawning |
Channels do not make shared state disappear. Mutexes do not make lifecycle disappear. A buffer is capacity, not correctness.
Read [references/synchronization.md](references/synchronization.md) when selecting between mutexes, atomics, and channels or when proving publication safety.
## Design the lifecycle
### Operation-scoped work
- Accept the caller’s context; do not replace it with `context.Background()`.
- Derive cancellation only when this layer owns the shorter lifetime, and call the cancel function.
- Start sibling goroutines only after defining how the first failure affects the others.
- Wait before returning if the goroutines access operation-owned memory or resources.
- Preserve the primary error; do not turn expected cancellation of siblings into the reported cause.
### Component-scoped work
- Make `Start`/`Run` and `Stop`/context ownership visible at the component boundary.
- Decide whether repeated start or stop is invalid, idempotent, or supported; encode that state.
- Reject new work before draining accepted work during shutdown.
- Bound shutdown with the caller’s deadline, but do not silently abandon resource owners.
### Detached work
Detached work is valid only when its lifetime, failure reporting, and resource ownership are intentionally process-scoped. A context is not automatically required for a short, non-blocking goroutine; conversely, passing a context does not prevent a leak if blocking operations ignore it.
Read [references/lifecycles.md](references/lifecycles.md) for worker, pipeline, and shutdown patterns. Read [references/supervision-and-failure.md](references/supervision-and-failure.md) when goroutines can fail, panic, cancel siblings, or outlive the caller.
## Bound work before spawning
Prefer admission control before goroutine creation. If the caller owns the operation or must observe failure, join the work and propagate its error; a detached goroutine is valid only under the process-owned contract above. This fragment demonstrates admission only, not a complete request-scoped lifecycle:
```go
select {
case slots <- struct{}{}:
case <-ctx.Done():
return ctx.Err()
}
go func() {
defer func() { <-slots }()
process(ctx, item)
}()
```
If a goroutine is created first and then waits for a slot, overload still creates unbounded goroutines. If the caller must observe admission failure, return it synchronously rather than hiding it in the goroutine.
For a queue, state the overload policy: block, reject, shed oldest/newest, or persist elsewhere. Never infer safety from an arbitrary buffer size.
## Review shared state
For every mutable value reachable by multiple goroutines:
- identify all readers and writers;
- name the lock, channel transfer, atomic operation, or immutable publication that orders them;
- keep the protected invariant adjacent to its synchronization field;
- do not copy values containing synchronization primitives;
- do not call unknown or blocking code while holding a lock unless the invariant requires it and the latency is bounded;
- consider compound invariants—a collection of individually atomic fields can still be inconsistent.
Treat the race detector as runtime evidence over executed paths, not as a proof that unexecuted paths are safe.
## Review channels as protocols
For each channel, record:
- producer set and consumer set;
- who closes it, if anyone;
- whether close means end-of-stream, cancellation, or broadcast;
- whether a send or receive can remain blocked after a peer exits;
- the capacity rationale and overload behavior.
Only a sender with exclusive knowledge that no future send can occur should close a channel. Many channels never need closing because their lifetime is bounded by the owning object.
## Diagnose before changing
- **Race:** start from the reported accesses and find the missing ordering edge.
- **Deadlock:** capture goroutine states; map locks and channel waits as a wait-for graph.
- **Leak:** compare goroutine profiles across steady load and after shutdown; locate the first blocking frame owned by the application.
- **Throughput collapse:** measure queueing, contention, scheduler delay, and downstream saturation before changing worker counts.
Go 1.27’s announced goroutine leak profile is prerelease as of this skill version. Do not prescribe it for stable toolchains until the repository adopts a released version.
## Finish with evidence
Within the authority of the request:
1. inspect the diff for every new `go`, channel, mutex, atomic, and `WaitGroup` operation;
2. exercise cancellation, early consumer exit, error, overload, and shutdown paths;
3. run focused tests under `-race` when the environment supports it;
4. use deterministic synchronization or `testing/synctest` on Go 1.25+ rather than timing sleeps;
5. report what was not exercised and why.
Do not claim race freedom, leak freedom, or deadlock freedom solely because a test passed.
## Output contract
For implementation, make ownership and termination legible in the code without architecture ceremony. For review, report concrete findings with the violated invariant, failure schedule, and smallest sufficient correction. Do not produce a generic concurrency checklist when the code has no concurrent path.
Referenced files: 6
go-data-consistency4.14 KB
--- name: go-data-consistency description: "Use for Go transaction, cache, migration, and DB/broker consistency. Do not use for broker ACK or replay." license: Apache-2.0 compatibility: "Go 1.24 or newer; database, driver, and cache semantics must be verified against deployed versions." --- # Go data consistency Start from the invariant and legal concurrent outcomes. A transaction is a correctness boundary, not a repository convenience. ## Define the durable unit Identify reads that justify writes, constraints, isolation behavior, transaction owner, external effects, idempotency identity, and what a caller sees when commit outcome is unknown. Keep all statements protecting one invariant on the same transaction handle. ## Enforce close to state Prefer unique, foreign-key, check, exclusion, or conditional-write constraints where the database can enforce the invariant atomically. Use row locks, advisory locks, optimistic versions, or serializable isolation only after naming the anomaly they prevent. Isolation labels differ across engines. After `BeginTx`, use the transaction handle exclusively. Defer rollback as cleanup, return commit errors, close rows, check iteration errors, propagate context, and remember `sql.DB` is a concurrent pool whose limits can queue and time out callers. ## Retry and ambiguity Replay the whole transaction from fresh reads only for stable retryable error classes, within a bounded budget, with external side effects excluded or independently idempotent. A connection failure during commit can leave the outcome unknown. Resolve by durable operation identity or reconciliation; never report a definite rollback without evidence. ## Cache consistency State whether the cache is authoritative, derived, or optional. Define invalidation/order behavior, stale-read tolerance, stampede control, and recovery after cache loss. A version embedded only in the cached value cannot reveal that the authoritative version advanced. Write-through is safe only if it participates in the authoritative commit. For correctness-dependent reads, use authoritative reads, a reader-visible generation check, or durable ordered propagation such as transactional outbox/CDC; do not rely on best-effort invalidation. ## Migrations Use expand, mixed-version deployment, bounded restartable backfill, switchover, and delayed contract. Verify old writer/new reader, new writer/old reader, rollback, and partial progress. Treat backfill progress as resumability metadata, not completeness proof. Partition by a stable key, update conditionally so a stale batch cannot overwrite a newer write, validate the authoritative rows before cutover, and remove the old representation only after old binaries, queued work, consumers, and rollback paths can no longer use it. Verify engine-version lock, rewrite, and failed-DDL behavior before choosing an online mechanism. ## Revisioned projections For a derived cache or controller built from snapshot plus watch, bind a consistent snapshot revision to a watch starting at the following revision. Apply the source's atomic revision unit before advancing the local checkpoint. Treat cancellation, compaction, and restored-cluster lineage as explicit resynchronization states; build and catch up a private replacement generation before publishing it. Read [references/revisioned-watch-resync.md](references/revisioned-watch-resync.md) for the etcd-specific evidence and portable procedure. Read [references/isolation-and-ambiguity.md](references/isolation-and-ambiguity.md) for transaction counterexamples, [references/online-changes.md](references/online-changes.md) for expand/backfill/cutover failure schedules, and [references/replica-read-authority.md](references/replica-read-authority.md) when correctness depends on asynchronous replica visibility or failover. For cursor APIs over changing or replicated data, read [references/pagination-under-mutation.md](references/pagination-under-mutation.md) and choose explicitly between live traversal and stable-snapshot semantics. ## Output contract Name the invariant, transaction boundary, permitted anomalies, and ambiguous-outcome policy. For review, show the violating interleaving and smallest repair.
Referenced files: 8
go-distributed-coordination2.97 KB
--- name: go-distributed-coordination description: "Use for Go cross-process ownership, leases, fencing, failover, sagas, and recovery. Do not use for local mutexes." license: Apache-2.0 compatibility: "Go 1.24 or newer; coordination guarantees are specific to the deployed service and client version." --- # Go distributed coordination A lease is time-bounded evidence, not permanent ownership. A paused or partitioned former owner can resume after a successor takes over. ## Define the safety property Identify the resource, operations requiring exclusion or order, authoritative store, owner identity, lease duration, clock assumptions, renewal path, takeover rule, and irreversible effects. Decide whether the requirement is safety, liveness, or both. ## Fence stale owners For irreversible writes, obtain a monotonically increasing fencing token and make the authoritative store reject effects from older tokens. A lease check performed before the write is insufficient when the holder can pause between check and effect. Use storage-level compare-and-swap, version predicates, or transaction constraints where possible. Process-local mutexes and leader flags cannot coordinate replicas. If the target cannot compare fencing tokens—such as some payment, email, or filesystem effects—a lease alone cannot guarantee exclusion. Use target-enforced stable idempotency, route the effect through a fenceable authoritative outbox, or state the weaker guarantee and duplicate-recovery requirement explicitly. ## Design lease lifecycle Bound acquisition and renewal; propagate cancellation; stop admission before expiry; treat renewal uncertainty as loss of authority; make release best-effort rather than the sole takeover mechanism. Observe lease age, renewal latency, token, owner, failed fences, and takeover count without high-cardinality leakage. When ownership depends on timestamps or survives serialization/restart, read [references/time-and-authority.md](references/time-and-authority.md). Go's monotonic reading protects elapsed comparisons only inside one process; it is not serialized and cannot establish cross-replica authority. ## Sagas and compensation Persist each workflow transition and command identity. Make steps and compensations idempotent. Compensation is a new business action, not rollback: it can fail, race with late success, or be legally impossible. Define forward recovery, manual exception handling, and terminal evidence. Read [references/fencing-and-sagas.md](references/fencing-and-sagas.md) for failure schedules. For regional writer promotion, coordination-cluster recovery, or failback, read [references/multi-region-failover.md](references/multi-region-failover.md). Routing a request to one region is not proof that every old writer and external effect has lost authority. ## Output contract State the stale-owner schedule, authoritative fence, takeover behavior, and recovery path. Do not prescribe distributed locks when a local constraint or partitioned owner is sufficient.
Referenced files: 6
go-message-processing9.3 KB
--- name: go-message-processing description: "Use for Go broker delivery, ACK, replay, ordering, poison recovery, outbox, or inbox. Do not use for in-process channels." license: Apache-2.0 compatibility: "Go 1.24 or newer. Broker guarantees and client APIs are version-specific and must be verified against the deployed system." --- # Go message processing Assume delivery can be duplicated, delayed, reordered, and interrupted at every boundary unless the deployed broker contract proves otherwise. “Exactly once” is scoped; it does not automatically make a database mutation, payment call, or email happen once. ## Model the state machine Choose the relevant path before loading details: producer contract, consumer effect, or outbox relay. Inspect consumer groups, retry/dead-letter policy, claiming, and acknowledgment only for consuming paths; inspect publication state and relay ownership for a producer/outbox path. In either case inspect the broker and client version, schema, database transaction path, and stable identity. A consuming path uses a state machine such as: ```text received -> admitted -> claimed/deduplicated -> effect committed -> acknowledged ``` For each transition, ask what happens if the process crashes immediately before and after it. State: 1. the stable message or operation identity; 2. the semantic payload fingerprint associated with that identity; 3. the ordering key and where ordering is guaranteed; 4. the durable effect and its transaction boundary; 5. when acknowledgment becomes safe; 6. retry limits, delay, and poison-message destination; 7. concurrency, memory, and downstream capacity bounds. Do not start from a library callback signature; start from these failure semantics. ## Choose the delivery contract ### At-most-once Acknowledge before the effect or accept loss on crash. Use only when loss is cheaper than duplication and that trade-off is explicit. ### At-least-once Perform the effect before acknowledgment. Crash between effect and acknowledgment causes redelivery, so the effect must be idempotent or deduplicated durably. ### Broker “exactly once” Name its boundary. It may cover producer records, broker offsets, a region, or one transaction API while excluding external databases and services. Preserve application-level idempotency whenever the effect escapes that boundary. Read [references/delivery-and-ordering.md](references/delivery-and-ordering.md) for broker guarantees, ordering, and acknowledgment trade-offs. ## Make processing durably idempotent For a local SQL effect, prefer one transaction that: 1. inserts or claims the message identity under a unique constraint; 2. verifies an existing identity has the same semantic fingerprint; 3. applies the domain state transition conditionally; 4. records the terminal result needed for replay; 5. commits before acknowledgment. On duplicate identity with the same fingerprint, return the recorded outcome or no-op according to the protocol. On the same identity with a different fingerprint, reject and alert; silently treating it as the original operation can apply the wrong command. An in-memory cache, local mutex, or process-local singleflight can reduce duplicate concurrent work but cannot provide durable deduplication across crashes, replicas, or retention windows. Use `go-data-consistency` for isolation and commit ambiguity inside this unit. ## Coordinate database state and publication If one operation changes database state and publishes an event, two independent commits create a gap: - database commits, publish fails: state exists without event; - publish succeeds, database rolls back: event describes nonexistent state. Use a transactional outbox when the database is the source of truth: 1. write domain state and an outbox row in one transaction; 2. relay committed rows with stable event identities; 3. make relay publication retryable; 4. mark progress without assuming publish acknowledgment is infallible; 5. keep consumers idempotent because the relay can republish. Use an inbox/deduplication record for inbound effects when the same local transaction can guard processing. Read [references/outbox-and-inbox.md](references/outbox-and-inbox.md) for relay and retention decisions. For initial loads, replica rebuilds, or change-data-capture bootstrap, read [references/cdc-snapshot-handoff.md](references/cdc-snapshot-handoff.md). A table scan and a later stream position do not form a safe cutover unless one source-consistent boundary prevents gaps and resolves snapshot/stream collisions. ## Acknowledge from the owner The component that knows the durable outcome owns acknowledgment. Do not acknowledge merely because a callback returned or a message entered an in-memory worker queue. Verify client-library concurrency rules: - whether acknowledgments must occur on the poll/receive goroutine; - whether processing can outlive a lease and how it is extended; - whether cancellation stops fetching, processing, or both; - whether partition revocation waits for or fences in-flight work; - whether one failed item can block a batch acknowledgment. If acknowledgment result itself can fail, retain enough durable processing state to handle redelivery. For cumulative offsets, track completion per partition and commit only through the highest contiguous completed offset; later completion must never skip unfinished lower offsets. On rebalance, stop admission and either drain within the revocation budget or fence unfinished ownership before committing progress. For consumer-group assignment, revocation, or cooperative rebalance changes, read [references/rebalance-ownership.md](references/rebalance-ownership.md). Bind every worker and completion frontier to one assignment generation; a stale worker must not advance a successor's progress or perform an unfenced effect merely because it was admitted earlier. ## Preserve ordering only where needed Global ordering is expensive and often unavailable. Define the business key whose events must be serialized, such as account ID or aggregate ID. Then verify: - the producer assigns the same key consistently; - the broker orders within the documented scope; - the consumer does not reintroduce reordering through parallel workers; - retries of one key do not unnecessarily block unrelated keys; - sequence/version checks detect stale or missing transitions when required. Ordering does not replace idempotency: the same event can appear twice in order. ## Bound concurrency and apply backpressure Bound before accepting more work than can be safely retained: - maximum in-flight messages and bytes; - per-key concurrency when ordering matters; - downstream database and RPC capacity; - lease/visibility deadline relative to processing latency; - shutdown drain time. Pausing broker fetch is often safer than accumulating an unbounded Go channel. A large prefetch can cause synchronized lease expiry and redelivery during slowdown. Use `go-concurrency-lifecycle` for in-process ownership and `go-service-resilience` for downstream attempt policy. ## Handle poison and terminal failures Classify failures: - **transient:** bounded retry with backoff and jitter; - **permanent input/schema:** quarantine or dead-letter with diagnostic context; - **business rejection:** record a terminal outcome, usually do not retry unchanged; - **dependency ambiguity:** reconcile using operation identity before replay; - **systemic overload:** reduce admission; do not accelerate retries. Dead-lettering is not resolution. Record original identity, schema/version, failure class, attempt count, timestamps, and safe diagnostic context. Provide a replay process that preserves or intentionally replaces identity and cannot bypass current validation. For quarantine evidence, ordering impact, retention, access control, and bounded redrive, read [references/poison-and-redrive.md](references/poison-and-redrive.md). Broker-assigned identity and enqueue time may change during redrive; preserve the application identity needed for deduplication and audit. ## Schema and compatibility Treat messages as public persisted contracts: - include an explicit event type and schema/version strategy; - prefer additive evolution while old producers and consumers coexist; - distinguish absent from zero when semantics require it; - retain unknown fields only when the encoding and compatibility contract support it; - do not couple business behavior to Go struct names or package paths; - make replay of historical events part of compatibility review. ## Finish with failure injection Within the request’s authority, exercise crashes or injected errors: - before and after durable effect commit; - before, during, and after acknowledgment; - duplicate delivery concurrently across replicas; - same identity with conflicting payload; - out-of-order and missing sequence; - poison message and dead-letter failure; - shutdown with in-flight work; - downstream overload and lease expiry. Do not claim exactly-once business effects from a happy-path integration test. State the broker guarantee, application guarantee, effect boundary, and untested crash points separately. ## Output contract For implementation, make the state machine and durable identity visible. For review, describe the crash point or interleaving that violates the business invariant and the minimum durable correction. Avoid broker-specific code until repository dependencies identify the broker and version.
Referenced files: 8
go-service-resilience7.5 KB
--- name: go-service-resilience description: "Use for cross-service Go deadlines, retries, admission, shedding, and recovery. Do not use for commits, brokers, or local speed." license: Apache-2.0 compatibility: "Go 1.24 or newer. Policies are transport-neutral; verify client and dependency behavior against repository versions." --- # Go service resilience Design remote calls so a small dependency failure does not become unbounded work. Timeouts, retries, queues, and concurrency limits form one feedback system; configuring them independently can amplify the outage they were meant to survive. ## Map the call path Inspect the caller, transport, dependency contract, deployment topology, and existing resilience layer. Record: 1. the end-to-end latency or deadline objective; 2. every queue and attempt along the call graph; 3. which layer owns retries; 4. whether the operation is safe to replay and how identity is preserved; 5. maximum in-flight work and downstream capacity; 6. which failures are transient, permanent, overload, or ambiguous; 7. the fallback’s correctness and capacity cost. Do not add retries before finding existing retries in clients, proxies, service meshes, jobs, and callers. ## Allocate one deadline budget Propagate the caller’s context. Derive a shorter deadline only when this layer owns a sub-budget. Never extend the caller’s deadline by replacing the context. Before each attempt, account for: ```text remaining budget > queue wait + attempt bound + possible backoff + response margin ``` If not, fail without starting work the caller can no longer use. A transport timeout covers one phase; it does not replace an end-to-end deadline. Use separate policies for interactive requests, batch work, and long-lived streams. “30 seconds everywhere” is not a resilience design. ## Decide whether retry is legal Retry only if all conditions hold: - the failure class is plausibly transient; - the operation can be replayed with the same semantic identity; - this layer is the designated retry owner; - the dependency has capacity for retry traffic; - attempts and elapsed time are bounded; - the remaining caller budget can complete another useful attempt. Application idempotency matters more than HTTP method names. A `GET` can trigger a broken side effect; a `POST` can be replay-safe under a durable idempotency contract. Ambiguous outcomes require reconciliation or same-identity replay, not a new operation. Read [references/retries.md](references/retries.md) for classification, backoff, jitter, and hedging. Treat every hedge as another live attempt inside one operation budget, and verify actual Go client support instead of assuming a cross-language service-config feature exists. ## Control amplification If a five-deep call chain makes three total attempts at every layer, one user request can create up to 243 leaf attempts. If “retry three times” means three retries after the first attempt, the bound is 4^5 = 1,024. Prefer one retry owner near the layer that understands replay semantics and user budget. Use exponential backoff with jitter to desynchronize callers, but remember: - backoff does not reduce the first retry wave; - a high retry cap still creates excess load; - retrying overload responses can keep a dependency overloaded; - unbounded queued retries consume memory and expire before execution. Apply a retry budget or token limit so retries remain a controlled fraction of normal traffic. Observe attempts per original operation, not only raw request count. ## Bound concurrency and queues Admission control should happen before allocating expensive state or spawning work. Choose a bound from dependency capacity and tolerated queue latency, then choose overload behavior: - reject early and cheaply; - shed lower-priority work; - degrade optional work; - queue only within an explicit memory and latency limit. Separate concurrency pools only when isolation matches real failure domains. Per-dependency bulkheads can stop one slow dependency from consuming every worker, but too many pools strand capacity and complicate fairness. Bound both logical operations and active dependency attempts. An operation-lifecycle permit bounds callers sleeping in backoff; an attempt permit bounds active dependency load and is released before backoff, then reacquired cancellation-aware. Holding scarce dependency capacity during sleep starves fresh work and recovery probes; releasing it without an operation bound creates unbounded sleepers. Read [references/overload.md](references/overload.md) for load shedding, circuit behavior, and recovery. Read [references/overload-signals-and-scope.md](references/overload-signals-and-scope.md) when admission spans replicas, tenants, priorities, proxies, or protocol retry signals. ## Treat circuit breakers as state machines A circuit breaker can reduce futile load, but it introduces shared state, thresholds, probes, and synchronized recovery. Add one only when simpler bounded concurrency, deadlines, and retry control are insufficient. Define: - which failures count; - sample size and open threshold; - open duration and probe concurrency; - behavior for callers while open; - how recovery avoids a thundering herd; - per-instance versus shared scope. Do not use a circuit breaker to hide incorrect timeouts or as a substitute for dependency health signals. ## Validate fallbacks A fallback is another production path. Verify: - its data freshness and consistency semantics; - authorization and privacy equivalence; - capacity during the same outage; - whether it turns a hard failure into silently wrong data; - how callers and telemetry distinguish degraded results. Caches can be latency optimizations or capacity dependencies. A cache that all instances fall through simultaneously can make recovery worse. Define stampede control and stale-data policy explicitly. ## Preserve error semantics Return errors that let the owning boundary decide: - retryable versus permanent when the API intentionally exposes that contract; - overload versus dependency unavailability; - deadline exceeded versus caller cancellation; - ambiguous outcome versus confirmed rejection. Do not string-match error messages. Do not log at every layer. Include attempt metadata in telemetry without wrapping the same cause into an unreadable chain. ## Recovery and rollout Design for recovery, not only steady failure: - ramp traffic or probes so cold caches and reconnect storms do not re-trigger overload; - randomize periodic work and reconnects; - drain retry queues under a controlled rate; - keep health checks cheap and separate readiness from liveness; - test configuration changes because retry/timeout policy is production code. ## Finish with evidence Within the request’s authority, test the feedback system: - latency beyond the attempt and caller deadline; - transient failure followed by recovery; - permanent and ambiguous failures; - partial dependency capacity and overload rejection; - retry storms from many callers; - fallback saturation and stale data; - recovery with cold connections or caches. Record attempts per operation, in-flight work, queue time, rejected work, dependency latency, and terminal outcomes. A passing unit test for a backoff function does not validate overload behavior. ## Output contract For implementation, keep retry ownership, replay identity, deadline allocation, and capacity bounds close to the call. For review, give the amplification path or failure schedule and quantify the maximum attempts/in-flight work where possible. Do not prescribe circuit breakers or retries by default.
Referenced files: 6
review-go-distributed-change5.32 KB
--- name: review-go-distributed-change description: "Use to review Go diffs/PRs for distributed failures in transactions, brokers, retries, leases, or effects. Do not use for fintech." license: Apache-2.0 compatibility: "Go 1.24 or newer; repository and deployed-system guarantees control the review." --- # Review a Go distributed change Find the failure schedule that violates a system invariant. ## Resolve only missing contracts Treat semantics stated by the prompt, repository, and deployed-system documentation as the review contract. Do not load auxiliary skills or references when those sources already define the relevant transaction, broker, remote-effect, retry, or lease behavior. When a missing contract blocks a finding, load only the matching focused skill: `go-data-consistency` for storage/commit semantics, `go-message-processing` for broker acknowledgement or ordering, `go-service-resilience` for client retry/timeout policy, or `go-distributed-coordination` for leases and fencing. Stop once the missing contract is resolved. Read [references/schedule-catalog.md](references/schedule-catalog.md) only when the changed path contains an effect boundary not covered below or a concrete counterexample is still needed. ## Build the state/effect graph Trace admission, local reads, decisions, durable writes, remote calls, publication, acknowledgement, response, and recovery. Mark transaction boundaries, goroutine owners, retry owners, ordering keys, leases, queues, and version transitions. ## Test causal schedules - concurrent same/conflicting identities; - cancellation before admission, during work, and after a durable effect; - crash immediately before and after commit, publish, and acknowledgement; - timeout with unknown remote result; - redelivery, duplication, delay, and reordering; - retry at several layers under dependency overload; - lease expiry while an owner is paused; - mixed versions during deploy and rollback; - recovery with cold caches, reconnects, and queued retries. Use only schedules reachable in the changed path. Quantify maximum goroutines, queue entries, attempts, held connections, and duplicate effects where possible. ## Close every effect boundary Before finalizing, give every applicable boundary an explicit disposition: | Boundary | Required disposition | |---|---| | Database commit | State how an unknown commit is resolved by stable operation identity or an authoritative read. | | Remote effect | State whether replay can duplicate the effect and name target-enforced idempotency, fencing, or reconciliation. | | Publication | Preserve one logical event identity across ambiguous publication and retries. | | Consumer acknowledgement or cumulative offset | State when it is safe, how failure or an unknown result replays, and which durable outcome makes replay harmless; never infer this from an outbox alone. | | Lease renewal or release | Treat an unknown renewal as uncertain authority, stop unsafe work before expiry, enforce a monotonic fence at every authoritative effect, and prevent a stale release from revoking a successor. | For every retry, enforce one end-to-end deadline and distinguish transient faults from permanent failures and ambiguous outcomes before replay. When retries or redeliveries exist at several layers, compute the configured maximum leaf attempts as their product; if a bound is unknown, leave it symbolic rather than inventing one. Assign retry ownership to one bounded layer. For each goroutine, name its owner, derive cancellation from that lifecycle, bound each blocking operation, and join before the owner returns. Work that can outlive the request or lease must have a deliberate lifecycle and must not perform an unfenced effect after authority is lost. ## Completeness gate Do not emit the final review while either applicable disposition is only implied: - If consumer acknowledgement appears, state when it becomes safe, how a failed or unknown acknowledgement causes replay, and which durable outcome makes that replay harmless. Scope any exactly-once statement to the boundary that actually provides it. - If any retry appears, state one end-to-end deadline, transient and permanent error classes, and reconciliation before replaying an ambiguous outcome. - If two or more layers can retry or redeliver, state the product bound and select one retry owner. A numeric bound without the ownership repair is incomplete. - If a lease appears, state the paused-owner takeover schedule, the authoritative fence, behavior after an unknown renewal, and how the renewer derives cancellation from its owner, bounds each renewal call, and terminates before release or return. ## Findings Tie each finding to a changed line and include invariant, trigger schedule, state consequence, and smallest correction. Do not recommend retries without replay safety, locks without an authoritative boundary, or “exactly once” without naming its scope. ## Output contract Lead with correctness and availability findings. Route monetary, ledger, settlement, or compliance consequences to `review-go-fintech-change` for domain adjudication. Audit the final prose—not only the analysis—against every applicable completeness-gate clause. Add any missing clause to the nearest causal finding; do not rely on an inbox/outbox recommendation or attempt count to imply acknowledgement replay, retry ownership, deadline, or error classification.
Referenced files: 4
Package details
Publisher declarations from the archived package. These are separate from our research and the live service's terms.
- Package license
- Apache-2.0
- Package author
- Ashwin Gopalsamy
- Keywords
- go, golang, distributed-systems, concurrency, messaging, consistency, resilience, coordination
Declared capabilities
- Design failure-safe Go systems
- Diagnose distributed failures
- Review distributed changes
- Gophers
- Golang
Package observed Oct 2, 2026.
Technical details
- First seen
- Sep 30, 2026 · 22:02 UTC
- Last seen
- Oct 2, 2026 · 18:00 UTC
- Collection status
- Collected
plugins_6a931917afc48191a2ce571d737eb104
Download plugin data (JSON)