← Poka-YokeCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to Poka-Yoke
Snapshot Sep 30, 2026 · 23:14 UTC · version 0.2.0
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"description": "Stop an AI agent damaging your repo: PreToolUse hooks, permission deny rules, protected paths, verification gates. Use when \"claude keeps force pushing\", \"CLAUDE.md says X but it still does Y\", \"stop the agent touching prod or .env\", or making a repo safe for unattended agent work. For AI features you ship to users use llm.",
"included_files": [],
"name": "agent-guardrails",
"skill_md_contents": "---\nname: agent-guardrails\ndescription: >-\n Stop an AI agent damaging your repo: PreToolUse hooks, permission deny rules, protected paths, verification gates. Use when \"claude keeps force pushing\", \"CLAUDE.md says X but it still does Y\", \"stop the agent touching prod or .env\", or making a repo safe for unattended agent work. For AI features you ship to users use llm.\n---\n\n# Poka-Yoke for AI-Written Code\n\nAn agent is a fast, tireless operator with no memory of yesterday and a strong prior toward\nappearing successful. That is the exact profile Shingo designed poka-yoke for, except an\nagent makes mistakes faster than any human, and never learns from the ones you correct in\nconversation.\n\nThe governing insight: **instructions to an agent are rung zero.** A line in CLAUDE.md saying\n\"never commit to main\" is training, and training degrades, under long contexts, compaction,\nand subagents that never read the file. A PreToolUse hook that denies the push is a device. If\nyou have been repeating the same correction to an agent, that is the signal to stop writing\ninstructions and install a device.\n\n## A complete answer covers all five\n\n**The diagnosis is not the answer.** \"Instructions are not enforcement\" is the right insight,\nand it is satisfying to write, but someone asking *\"what am I doing wrong?\"* has a repo they\nneed to fix: not a question about their prose. Explaining why the rules fail and stopping\nthere leaves them exactly where they started. State the insight in a sentence, then spend the\nrest of the answer on the replacement.\n\nReplacing an instruction with a device is not one step, it is five, and stopping after the\nfirst leaves the person with a rule that looks enforced and is not. Naming the deny rule is\nthe easy part and the least of it. Cover every one of these, briefly, before adding depth:\n\n1. **The deny rule, with real syntax.** Show the actual `permissions.deny` entry for their\n case, `\"Bash(git push --force:*)\"`: not a description of one. A pattern they have to\n invent themselves is a step where this fails.\n2. **A hook where a pattern is not enough.** Deny rules match strings. Anything conditional: a `DELETE` without a `WHERE`, an edit allowed in one directory but not another, a\n production hostname, needs a `PreToolUse` hook that inspects the call and returns a deny.\n Say which of their two rules needs which.\n3. **What the deny message says.** The agent reads it and acts on it, so a bare refusal\n produces a workaround, often a worse one. The message must name what was blocked, why, and\n what to do instead. This is the one place prose belongs in a device.\n4. **Where the config lives, so it applies to everyone.** `.claude/settings.json`, committed.\n A rule in `settings.local.json` protects one machine, which is the same failure as\n documenting it: the protection exists only where someone remembered to set it up.\n5. **Proof that it fires.** Run the blocked action and confirm the denial *and* its message,\n then run the legitimate neighbouring action and confirm it still works. Untested hooks fail\n open more often than people expect: a regex that does not match the real command string is\n a hook that does nothing while looking like protection. **An unverified device is worse\n than no device, because it creates confidence without protection.**\n\nSteps 3 and 5 are the ones most often dropped, and they are what separate a device that works\nfrom one that merely exists.\n\n## The three failure modes, and the device for each\n\n**1. The agent does something destructive.** Force-push, `rm -rf`, dropping a table, editing\n`.env`, running against production, `git checkout .` over uncommitted work, `--no-verify`.\nThese are irreversible and fast. Device: **deny at the tool boundary**: a hook or permission\nrule that refuses the call before it executes. This is Control and it is the only rung that\nmatters for irreversible actions.\n\n**2. The agent writes code that looks right and isn't.** Plausible-but-wrong is an agent's\ncharacteristic defect: correct-looking imports of things that don't exist, tests that assert\nnothing, error handling that swallows, a stub that returns a hardcoded value. Device: **the\ntype checker and the test suite as required gates**, plus lint rules against silent failure.\nEverything in `guardrails` applies here with extra force, because the volume of\ngenerated code is higher and human review attention per line is lower.\n\n**3. The agent reports success it didn't achieve.** \"All tests pass\" when the suite wasn't\nrun; \"done\" with the build broken. Device: **verification the agent cannot fake**: a Stop\nhook that actually runs the tests, or a CI gate. Never accept a claim of completion that only\nexists as text.\n\n## Devices, strongest first\n\n### Deny rules in settings.json\n\nThe cheapest device and the first thing to install. Permission denies are evaluated before the\ntool runs and need no scripting:\n\n```jsonc\n{\n \"permissions\": {\n \"deny\": [\n \"Bash(git push --force:*)\",\n \"Bash(git push -f:*)\",\n \"Bash(git commit --no-verify:*)\",\n \"Read(./.env)\",\n \"Read(./.env.*)\",\n \"Edit(./.env)\",\n \"Edit(./migrations/**)\",\n \"Bash(terraform apply:*)\"\n ]\n }\n}\n```\n\nReading `.env` matters as much as writing it: an agent that reads a secret can echo it into a\nlog, a commit, or a message to a third-party service. Deny the read.\n\nA deny entry matches the **start** of the command, so it only holds where the dangerous form\nis the prefix. That is why `rm -rf` is not on this list: `\"Bash(rm -rf /:*)\"` would leave\n`rm -fr /`, `rm -Rf /` and `cd / && rm -rf *` untouched while looking like coverage.\nRecursive delete needs the hook below, see the `rm` pattern in\n`../../assets/devices/claude-hooks/guard_dangerous_commands.py`.\n\nPut team-wide rules in `.claude/settings.json` (committed) and personal ones in\n`.claude/settings.local.json` (gitignored), otherwise the rules exist only on the machine of\nwhoever set them up, which is the same failure as documenting them.\n\n### PreToolUse hooks for anything conditional\n\nWhen the rule needs logic, \"block `DELETE` without a `WHERE`\", \"block edits to\n`schema.prisma` unless a migration exists\", \"block production hostnames in a connection\nstring\": a hook script inspects the call and returns a deny with a reason.\n\nTemplates in `../../assets/devices/claude-hooks/`. The critical detail:\n**the deny message is read by the agent and is your only chance to redirect it.** A bare\n\"denied\" produces a workaround attempt, often a creative and worse one. A message that says\nwhat was blocked, why, and what to do instead produces the right action. Write it as you would\nwrite an error message for a colleague:\n\n> Blocked: `DELETE` without a `WHERE` clause on `users`. Unbounded deletes are irreversible\n> here. Add a `WHERE` clause, or if a full truncate is genuinely intended, ask the user to\n> confirm and run it themselves.\n\n### Stop hooks that verify completion\n\nRun the type check and the test suite when the agent tries to finish. This converts \"tests\npass\" from a claim into a fact, and it is the single highest-value hook in most repos.\n\n### Machine-checkable CLAUDE.md\n\nAnything in CLAUDE.md that *can* be a check should be one; what remains should be facts the\nagent needs rather than rules you hope it follows.\n\n- \"Always run `make fmt` before committing\" → a pre-commit hook.\n- \"Never use `any`\" → a lint rule with a required check.\n- \"Don't edit generated files\" → a deny rule, plus a header in the generated files.\n- \"Use `pnpm`, not `npm`\" → a deny on `Bash(npm install:*)` with a message naming `pnpm`.\n\nWhat legitimately stays as prose: architecture, domain vocabulary, where things live, why\npast decisions were made. Facts, not commands.\n\n### Make the safe path the easy path\n\nAgents follow the shortest route to a working answer. If `make test` runs the right thing with\nthe right env, it gets used; if the correct invocation is a fifteen-flag command documented in\na wiki, it does not. Every ergonomic improvement here is a poka-yoke: a `make check` that\nbundles fmt + lint + types + tests, a `.env.example` with every key present, a devcontainer or\na single setup script. Ambiguity is where agents improvise, and improvisation is where damage\ncomes from.\n\n## A caution about over-restriction\n\nDeny rules that block ordinary work produce an agent that spends its turns fighting the\nharness, and a user who turns the rules off. Aim the strong devices at **irreversible and\noutward-facing** actions, force-push, prod, secrets, destructive SQL, deletion, publishing, and leave ordinary editing and reading alone. Reversibility is the right axis: git makes most\ncode changes cheap to undo, so they do not need a gate. A rotated credential and a dropped\ntable do not.\n\n## Verify each device\n\nSame discipline as any other guardrail, and easy to check here: try the blocked action and\nconfirm the denial and its message, then confirm the legitimate neighbouring action still\nworks. Untested hooks fail open surprisingly often: a regex that doesn't match the real\ncommand string is a hook that does nothing while looking like protection.\n\nLeave a `poka-yoke:` marker comment on each rule naming what it prevents, and show the user\neach config before writing it. Hooks execute code on their machine on every tool call; that is not a change to\nmake on someone's behalf unseen.\n"
}SHA-256 of public snapshot: 042f3b159b47f45c74aaef219a44e6b56b73d80608a4a914d9be9099ad39dde0