{"id":18863,"plugin_id":"plugins_6a8ceaf162b88191851ca2442d67e12d","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:14:58.731Z","digest":"587f65b0a25349beb890e1df2eedd7108d63521e6700e78e999d20a3630668aa","against":null,"payload":{"description":"Deploys, schema migrations, rollback and infrastructure. Use when \"can I ship this on Friday\", \"this migration is scary\", \"what is the blast radius\", \"prevent accidental deletion of the database\", or a change drops a column. Covers expand/contract, canary rollout, kill switches, prevent_destroy, tested backups. For an incident that already happened use retro.","included_files":[],"name":"ops","skill_md_contents":"---\nname: ops\ndescription: >-\n  Deploys, schema migrations, rollback and infrastructure. Use when \"can I ship this on Friday\", \"this migration is scary\", \"what is the blast radius\", \"prevent accidental deletion of the database\", or a change drops a column. Covers expand/contract, canary rollout, kill switches, prevent_destroy, tested backups. For an incident that already happened use retro.\n---\n\n# Poka-Yoke for Deploys and Infrastructure\n\nOperations is where irreversible mistakes concentrate. Code mistakes are usually recoverable, git remembers, a revert ships in twenty minutes. A dropped table, a deleted bucket, a rotated\ncredential, or a terminated stateful node is not recoverable by any amount of engineering\nafter the fact.\n\nSo the governing question in this mode is different from the rest of the plugin. Not \"can this\nbe done wrong?\" but: **when this is done wrong, how much is affected, and can it be undone?**\nThose two axes, blast radius and reversibility, determine every device below.\n\n## Answer these four first\n\nBefore any framework or table, establish these. They are what an operator actually needs, and\nthey are the things most often left out: an answer that skips them is not useful no matter\nhow well organized the rest is. Say each one plainly, in a sentence, before going deeper.\n\n1. **What here is irreversible, and what restores it?** Name the specific unrecoverable step: a dropped column, a deleted bucket, a rotated key. Then say what would restore it: a\n   backup, a snapshot, a rebuild. **If the answer is \"nothing\", say so explicitly.** An\n   irreversible step with no stated restore path is the single most important thing you can\n   tell someone, and it is the first thing to get lost in a longer answer.\n2. **What breaks during the rollout window?** Deploys are not atomic. For a period, old code\n   runs against the new state. Say what happens in that window, usually this is the actual\n   outage, not the change itself.\n3. **Can the irreversible part ship separately?** Most changes are a reversible part and an\n   irreversible part stapled together. Splitting them is nearly always available and nearly\n   always right; say so concretely rather than in general.\n4. **If it goes wrong, who is available and how fast is rollback?** Timing questions are about\n   staffing and recovery speed, not superstition. A change that reverts in two minutes is fine\n   on a Friday afternoon; one that needs a four-hour restore with two people asleep is not.\n\nCover all four even when the answer is brief. If you only have room for a little, spend it\nhere rather than on the taxonomy: the rungs below are how to *think* about the fix, but these\nfour are what the person has to know before they ship.\n\n## The blast radius ladder\n\nMost ops poka-yoke is not about preventing the bad change. It is about ensuring the bad change\nreaches 1% of traffic instead of 100%. You cannot prevent every bad deploy; you can make bad\ndeploys cheap.\n\n| Rung | Device | What it buys |\n|---|---|---|\n| **1 Control** | The dangerous operation is impossible in this environment, `prevent_destroy` on stateful resources, deletion protection on the database, no human write access to prod, immutable infrastructure | The mistake cannot be made at all |\n| **1 Control** | Progressive rollout with automatic rollback on error-rate: the bad version is withdrawn before most users see it | The mistake is capped and self-healing |\n| **2 Warning** | Required plan review, a deploy that prints what it will destroy and demands typed confirmation, alerts wired to the rollout | The mistake is visible at the moment of decision |\n| **3 Detection** | Post-deploy smoke tests, monitoring, an on-call human | The mistake is found after users find it |\n| **0** | A runbook step that says \"double-check the environment first\" | Nothing |\n\n## Reversibility is the highest-leverage property\n\nBefore adding any gate, ask whether the operation can be made reversible instead: a\nreversible operation needs far weaker devices, because the cost of the mistake collapses.\n\n- **Soft delete and retention windows** on anything user-facing. S3 versioning plus MFA delete,\n  database point-in-time recovery, trash with a 30-day window.\n- **Deletion protection flags** on databases, buckets, clusters, and load balancers. These\n  cost nothing and stop the single most expensive class of cloud mistake.\n- **Backups that have actually been restored.** An untested backup is a belief, not a device, and this is the most common false sense of protection in the industry. Restore drills on a\n  schedule, timed, into a real environment. If nobody has restored it, treat the data as\n  unbacked when you assess blast radius.\n- **Immutable artifacts** so rolling back means redeploying a known-good image, not rebuilding\n  and hoping the build is reproducible.\n\n## Schema migrations: expand and contract\n\nCo-deploying a destructive schema change with the code that depends on it is an outage, not a\nrisk, during the rollout window old code necessarily runs against the new schema.\n\nThe pattern, one deploy per step:\n\n1. **Expand**: add the new column/table, nullable, with no code depending on it.\n2. **Backfill**: in batches, resumable, throttled, with progress recorded so a failure resumes\n   rather than restarts.\n3. **Dual-write**: new code writes both old and new; both remain readable.\n4. **Switch reads**: behind a flag, so switching back is instant.\n5. **Contract**: drop the old column, in a later deploy, once nothing references it.\n\nSteps 1–4 are reversible: the old column stays readable throughout, so rolling the deploy\nback is enough. Only step 5 is not, which is exactly why it gets its own deploy and its own\ngate. The device that makes this stick is a CI check that refuses any `DROP`, `TRUNCATE`, or\n`ALTER ... DROP` in a changed migration unless the PR carries an explicit approval label, see\n`guardrails` for the gate itself.\n\nMigration-specific hazards worth checking every time: a lock taken on a large table during\npeak traffic; a backfill with no batch limit; an index created without `CONCURRENTLY`; a\n`NOT NULL` added without a default on a populated table; a rename, which is a drop and an add\nwearing a disguise.\n\n## Feature flags and kill switches\n\nA kill switch is a poka-yoke for a change you cannot fully test in advance. It converts \"roll\nback a deploy\" (minutes, and impossible if the migration already ran) into \"flip a boolean\"\n(seconds). Ship risky changes dark, behind a flag, then enable progressively.\n\nTwo things make flags devices rather than debt:\n\n- **The off path must be tested**, not just the on path. A kill switch whose disabled branch\n  was never exercised is a second untested code path shipped at your worst moment.\n- **Flags need an expiry.** A permanent flag is a permanent untested branch and a permanent\n  source of \"works for some users only\" bugs. Track age and remove them; stale flags are the\n  standard way this device turns into a hazard.\n\n## Infrastructure as code\n\n- **`prevent_destroy` on every stateful resource**: databases, buckets, volumes, DNS zones.\n  One line, and it turns the worst cloud accident into a failed plan. It is Control against the\n  accident, not against intent: removing the block is another one-line change, so review has to\n  read a diff that deletes a `prevent_destroy` as itself a destructive change.\n- **Plan review as a required check**, with the plan output posted to the PR. A human approving\n  a diff they cannot see is rung zero.\n- **Fail the plan on unexpected destruction**: a check that counts destroy actions and blocks\n  the apply unless the change is explicitly labeled as intentionally destructive. Terraform\n  will happily replace a database to change one immutable attribute, and the plan says so in\n  a line people skim past.\n- **Separate state and credentials per environment**, so a misconfigured shell cannot point a\n  staging apply at production. Environment confusion is a mistake of *context*, and the device\n  is making the contexts physically incapable of touching each other.\n- **No console access for routine work.** Manual changes drift from code and are invisible to\n  review; drift detection turns that into a Warning at minimum.\n\n## Production access\n\nThe strongest device is not needing access: good observability, safe read-only debugging\ntools, and self-service runbooks remove most reasons a human ever holds a prod shell.\n\nWhere access is genuinely needed: time-boxed and audited, read-only by default, write access\nrequiring a second approver, and a shell prompt that makes the environment impossible to\nmistake. Environment confusion, running the staging command against prod, is a top-tier\nops mistake and it is fixed by making prod look and feel different, not by remembering.\n\nWrap dangerous scripts so the safe form is the easy one: dry-run by default with `--apply` to\ncommit, print the affected count before acting, refuse to run against prod without an explicit\nflag, and refuse an empty or wildcard target.\n\n## Auditing an ops setup\n\nWork through these, and report using the finding structure from `audit`:\n\n1. **What is irreversible today?** Every resource whose loss is unrecoverable. Which have\n   deletion protection? When was the backup last *restored*, not last taken?\n2. **What is the blast radius of a bad deploy?** All users at once, or 1%? Is rollback\n   automatic on an error-rate signal, or does it need a human who is asleep?\n3. **Can a migration and its dependent code land together?** Is anything stopping it?\n4. **Can a staging command reach production?** Shared credentials, shared state, an ambiguous\n   prompt, a `--env` flag defaulting to prod.\n5. **What has no kill switch?** Anything risky that can only be withdrawn by a full deploy.\n6. **Which flags are older than 90 days?**\n\nPropose devices before applying them, and never apply infrastructure changes without explicit\napproval: an `apply` is exactly the class of outward-facing, hard-to-reverse action that\nbelongs to the user, not to you. For anything you cannot run yourself (console settings,\nbranch protection, IAM), hand over the exact steps or CLI command.\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}