← Poka-YokeCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to Poka-Yoke
Snapshot Sep 30, 2026 · 23:14 UTC · version 0.2.0
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"description": "AI features you ship to users: structured output, tool schemas, prompt injection, evals. Use when \"the model returns bad JSON\", \"it hallucinates\", \"stop it calling the wrong tool\", \"add evals\", or an LLM feature can trigger refunds, emails or writes. Covers schema-constrained output, idempotent tool calls, confirmation gates. For agents editing your repo use agent-guardrails.",
"included_files": [],
"name": "llm",
"skill_md_contents": "---\nname: llm\ndescription: >-\n AI features you ship to users: structured output, tool schemas, prompt injection, evals. Use when \"the model returns bad JSON\", \"it hallucinates\", \"stop it calling the wrong tool\", \"add evals\", or an LLM feature can trigger refunds, emails or writes. Covers schema-constrained output, idempotent tool calls, confirmation gates. For agents editing your repo use agent-guardrails.\n---\n\n# Poka-Yoke for LLM Features\n\nThis is about AI features **you ship to users**: not about agents editing your repo, which is\n`agent-guardrails`.\n\nThe defining property of an LLM is that it is a component with a non-zero error rate on every\ncall, and no amount of prompt engineering drives that to zero. This is not a defect to fix; it\nis the material you are building with. Shingo's framing fits perfectly: you do not make the\noperator more careful, you build the jig.\n\nWhich means the central discipline here: **prompt instructions are rung zero.** \"Always respond\nwith valid JSON,\" \"never make up a citation,\" \"do not reveal the system prompt\". These are\nrequests to an unreliable component, and they are the LLM equivalent of a comment saying \"be\ncareful.\" They help, they are worth writing, and they are not devices. A device is something\noutside the model that constrains what it can produce or what its output can reach.\n\n## Building, not reviewing\n\nMost of the time this mode is reached *while someone is building the thing*, not afterwards.\nThat changes the deliverable. They asked for the feature, so produce the feature, working, complete,\nin their stack. Do not hand back a severity table when the person is mid-feature; a list of\nfindings about code they have not written yet is not useful to them.\n\nThen add a short closing note, three or four lines, covering:\n\n- which misuses the shape you chose makes impossible, and at which rung,\n- what you left possible on purpose, and why that tradeoff is the right one here.\n\nThat closing note is what stops the device being undone in six months by someone who cannot\nsee why it is there. It is also the difference between mistake-proofing and a code generator:\nthe reasoning travels with the code.\n\nWhen the code already exists and they are asking what is wrong with it, switch to the audit\nvoice, ranked findings with the mistake, the consequence, and the device. Match the mode to\nwhere they are in the work, not to this file's default.\n\n## The boundary: nothing the model says is trusted until something checks it\n\nDraw the same line you would draw around any external, untrusted input, because that is\nexactly what model output is, and doubly so when the model has read user-supplied text.\n\n### Structured output over prose parsing (Control, contact lens)\n\nNever regex a model's prose. Use the provider's constrained/structured output mode with a\nschema, then validate the parsed result against that schema yourself:\n\n```python\nclass Extraction(BaseModel):\n model_config = ConfigDict(extra=\"forbid\")\n sentiment: Literal[\"positive\", \"neutral\", \"negative\"]\n confidence: float = Field(ge=0.0, le=1.0)\n```\n\nConstrained decoding makes malformed output largely unrepresentable, and the schema check\ncatches the rest. That removes the whole class of parse failures, malformed JSON, missing\nfields, invented enum values.\n\nTwo things the schema still cannot tell you: whether the values are *correct*, and what to do\nwhen validation fails. Decide the failure path explicitly, retry once with the error fed\nback, then fall back to a deterministic path or return a clear failure. A silent default here\nis `except: pass` with a language model attached.\n\n### Enumerate rather than generate wherever possible\n\nThe strongest device in this whole mode: if the output is a choice from a known set, have the\nmodel choose an ID from a list you supply and reject anything not in it. A model asked to\nproduce a category name will invent one eventually; a model choosing among five IDs cannot.\nApplies to routing, classification, tool selection, and picking a record, and it converts an\nopen-ended generation problem into a closed-set one that a `Literal` type enforces.\n\n### Ground factual claims, and make ungrounded output impossible to render\n\nFor anything retrieval-backed, require the response to cite retrieved chunk IDs, then verify\neach cited ID actually exists in what you retrieved and drop or flag claims that don't\nresolve. That check establishes that a citation resolves, not that the chunk it points at\nsupports the claim, where the claim is consequential, add an entailment check or human review\non top. Prompting for citations is rung zero; *verifying* them is a real device. Show the\nsource in the UI so the user can check. This is the interface half of the same device.\n\nWhen retrieval returns nothing relevant, the correct behavior is to say so. A model handed no\ncontext will answer anyway, and that answer is invention. Check for the empty-context case in\ncode, before the call, and short-circuit.\n\n## Side effects: the model proposes, the system disposes\n\nThe most expensive LLM bugs are not wrong text. They are actions. Refunds issued, emails sent,\nrecords deleted, all because a model decided to.\n\n- **Split tool calls by reversibility.** Read-only tools execute freely. Anything irreversible\n or outward-facing, payment, email, deletion, publishing, external writes, requires a human\n confirmation that names the specific action and its parameters. This is the same ladder as\n everywhere else; irreversible actions need Control.\n- **Make the tool schema tight.** Enums instead of free strings, required parameters instead of\n optional ones, ranges on numbers, and no \"extra context\" free-text field the model can use\n to smuggle in intent. A wide tool schema is a wide attack surface and a wide mistake surface.\n- **Validate arguments server-side, always.** The model is a client, and a client's input is\n never trusted. `refund(amount)` must re-check the amount against the actual order: the\n model saying `9999` is not authorization.\n- **Idempotency keys on every effectful tool call**, backed by a unique constraint. Agent loops\n retry; retries double-charge. This is hazard M2 with a higher retry rate than any human path.\n- **Scope credentials to the user, not to the service.** If the tool runs with service-level\n access, a prompt injection reaches everything. Pass the requesting user's authorization\n through, so the model cannot exceed what that user could do, see `authz`.\n\n## Prompt injection is a boundary problem, not a prompt problem\n\nAny text the model reads, user input, retrieved documents, web pages, emails, tool results, can carry instructions. No system prompt reliably prevents this, and treating it as a prompt\nengineering problem is why it keeps happening.\n\nThe devices are structural: keep untrusted content clearly delimited and labeled as data;\nnever let model output flow into a privileged action without validation or confirmation;\nscope permissions so a successful injection has a small blast radius; and treat any model\noutput that will be rendered as HTML, executed as SQL, or passed to a shell exactly as you\nwould treat user input from an attacker, because functionally it is.\n\nThe load-bearing question is not \"can the model be tricked?\" (yes) but \"**what can the model\nreach if it is tricked?**\"\n\n## Bounds: cost and loops\n\nAn agent loop with no cap is an unbounded resource operation, hazard F7 with a billing\naccount attached. Set a maximum step count, a token budget per request, and a wall-clock\ntimeout, all enforced in your code rather than requested in the prompt. Alert on cost per\nuser, and cap it per tenant so one runaway conversation cannot become a five-figure invoice.\n\n## Evals are the detection rung, and they are load-bearing\n\nYou cannot unit-test a probabilistic component, but you can measure it, and without\nmeasurement you have no idea whether a prompt change helped.\n\n- **A held-out eval set with assertions**, run in CI on every prompt, model, or retrieval\n change. Prompts are code with no type checker. This is the only gate they have.\n- **Assert on the structured fields**, which are checkable, rather than on prose similarity.\n This is another reason structured output pays for itself.\n- **Every production failure becomes an eval case.** This is the `retro` loop applied\n to a component that cannot be fixed, only constrained: you cannot patch the model, so the\n regression test *is* the fix, and it must cover the class rather than the one input.\n- **Pin the model version.** A provider updating a model underneath you is an unannounced\n deploy of your most unpredictable component. Pin it, and re-run evals before moving.\n\n## Auditing an LLM feature\n\n1. **Where does model output go?** Trace each path. Which reach a database, an API, a shell,\n the DOM, or a user as fact? Each needs a check at that boundary.\n2. **What is parsed from prose that could be structured?**\n3. **Which tools have irreversible effects, and what gates them?**\n4. **What untrusted text enters the context, and what could an instruction in it reach?**\n5. **What happens when the model fails**: malformed output, refusal, timeout, rate limit,\n empty retrieval? Is there a deterministic fallback, or does it fail silently?\n6. **What bounds exist on steps, tokens, and cost?**\n7. **Is there an eval suite, does CI run it, and does a regression block the merge?**\n\nReport with the structure from `audit`, and be honest about rungs, with a\nprobabilistic component, most in-model devices are Warning at best, and only the checks\n*outside* the model reach Control.\n"
}SHA-256 of public snapshot: 9aa24d4796858bebde402cf8bcd3a921350b224414c8772f81e12ec689fefb55