← get-fableCONTENT HISTORY

Update to get-fable

Snapshot Sep 30, 2026 · 23:14 UTC · version 1.5.1

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "description": "Diagnose repeated command failures, stale build caches, branch drift, or contradictory evidence before attempting further code edits. Use when commands fail repeatedly, tests stay red after attempted fixes, build output contradicts source code, or the execution path is confused — even if the user does not explicitly say \"fable-recover\" (e.g. \"still failing\", \"why did this fail again\", \"stuck in a failure loop\", \"diagnose this error\"). Do NOT use for routine first-time test failures in fresh TDD (use fable-tdd).",
  "included_files": [
    {
      "relative_path": "agents/openai.yaml",
      "size_in_bytes": 373
    },
    {
      "relative_path": "evals/scenarios.json",
      "size_in_bytes": 4337
    },
    {
      "relative_path": "examples/recovering-stale-test-cache.md",
      "size_in_bytes": 540
    },
    {
      "relative_path": "references/attribution-ladder.md",
      "size_in_bytes": 1731
    },
    {
      "relative_path": "references/diagnostic-falsification-playbook.md",
      "size_in_bytes": 2915
    },
    {
      "relative_path": "skill.package.json",
      "size_in_bytes": 463
    },
    {
      "relative_path": "templates/recovery-diagnosis.template.md",
      "size_in_bytes": 926
    }
  ],
  "name": "fable-recover",
  "skill_md_contents": "---\nname: fable-recover\ndescription: \"Diagnose repeated command failures, stale build caches, branch drift, or contradictory evidence before attempting further code edits. Use when commands fail repeatedly, tests stay red after attempted fixes, build output contradicts source code, or the execution path is confused — even if the user does not explicitly say \\\"fable-recover\\\" (e.g. \\\"still failing\\\", \\\"why did this fail again\\\", \\\"stuck in a failure loop\\\", \\\"diagnose this error\\\"). Do NOT use for routine first-time test failures in fresh TDD (use fable-tdd).\"\nversion: 1.3.0\npack: core\ninputs:\n  - failure_evidence\nrequires:\n  - failure_streak\nproduces:\n  - revised_hypothesis\n  - bounded_repair\ngates:\n  - diagnosis_changed\nfallback: fable-discover\nmutatesWorkspace: false\nparallelSafe: false\nneural_links:\n  precursors:\n    - fable-tdd\n    - fable-execute\n    - fable-verify\n  continuations:\n    - fable-discover\n    - fable-plan\n    - fable-execute\n  lateral_peers:\n    - fable-discover\n  recovery: fable-discover\n---\n\n# Fable Recover\n\nStop spending mutations on a failing hypothesis. Rebuild causal confidence before changing code again.\n\n## Mission\nRecovery exists for the moment an agent is most likely to become expensive and irrational: the same task has failed more than once, output contradicts expectations, or edits appear to have no effect.\n\nThe goal is not to \"try something different.\" The goal is to explain why previous attempts failed, falsify competing causes, and issue one repair that is justified by new evidence.\n\n## Activate When\n- the same command/test/behavior fails after two materially similar implementation attempts;\n- output appears stale or unaffected by known source changes;\n- tests/runtime disagree;\n- evidence contradicts a load-bearing assumption;\n- failures alternate or depend on timing/environment;\n- the agent is about to repeat a command/patch without new diagnostic information.\n\n## Do Not Activate When\n- first failure is a trivial syntax/type error introduced by the current edit;\n- architecture is simply unknown and no repeated failure occurred (`fable-discover`);\n- expected RED in TDD is being observed correctly;\n- a narrow review finding already has an obvious bounded repair.\n\n## Failure Classification\nStart by classifying the evidence, not the code.\n\n| Class | Typical signals | First probes |\n| --- | --- | --- |\n| Harness | assertion never reached, fixture/mock/setup error | prove test path and fixture |\n| Environment | local/CI/OS/env-specific | versions, env, cwd, permissions |\n| Artifact/cache | source changes not reflected | entrypoint, build timestamps, cache, dist |\n| Execution path | edited code never executes | tracing/instrumentation/registration |\n| Dependency/version | unexpected API/runtime semantics | lockfile, resolved version, official source |\n| Data/state | only some fixtures/accounts/orders fail | minimal failing state, persistence boundaries |\n| Concurrency/timing | intermittent/order-sensitive | deterministic coordination, shared state |\n| Product logic | harness/path proven, assertion consistently wrong | isolate algorithm/branch |\n| Invariant/design | local fixes move failure elsewhere | identify violated cross-component rule |\n\nDo not jump to product logic until cheaper external explanations are falsified.\n\n## Recovery Protocol\n\n### Stage 1 — Freeze mutation\nNo new production edits until the diagnosis changes. Preserve the failing state and collect exact evidence.\n\nRecord:\n- command/action;\n- exact error/output;\n- workspace/commit/mutation generation;\n- attempts already made and what differed;\n- expected observation.\n\n### Stage 2 — Reproduce minimally\nFind the smallest reliable reproduction. If the failure is flaky, capture seeds/order/time/environment and work on determinism before another fix.\n\n### Stage 3 — Build a hypothesis queue\nCreate 2-5 plausible causes ranked by:\n- ability to explain all observed evidence;\n- probability given recent changes;\n- cost/safety of falsification.\n\nEach hypothesis must predict an observation that would distinguish it.\n\nBad: \"maybe cache.\"\n\nGood: \"CLI executes stale `dist/cli.js`; if true, source timestamp will be newer than dist and direct source invocation will show new behavior.\"\n\n### Stage 4 — Walk the attribution ladder\nUse the cheapest separating probes first:\n\n1. **Harness** — does the test/probe reach the intended assertion/path with realistic inputs?\n2. **Environment/artifact** — correct branch, cwd, env, version, build, cache, process?\n3. **Execution path** — is edited code actually reached? which implementation is registered?\n4. **Data/dependency** — does input/version/state differ from assumptions?\n5. **Product logic** — with above proven, isolate the wrong branch/algorithm.\n6. **Invariant/design** — if local logic is individually reasonable but system remains wrong, identify the violated system contract.\n\nThe ladder is guidance, not ritual. Skip a rung only when existing evidence already proves it.\n\n### Stage 5 — Instrument or bisect when observation is weak\nUse narrow temporary diagnostics, binary search/bisect, toggling one variable, or comparing known-good/bad states.\n\nChange one diagnostic dimension at a time so the result is interpretable.\n\n### Stage 6 — Falsify, do not accumulate guesses\nAfter each probe:\n- reject hypothesis;\n- strengthen hypothesis;\n- or revise the queue.\n\nDo not keep contradicted explanations alive as \"maybe still related.\"\n\n### Stage 7 — Form the revised diagnosis\nA valid diagnosis explains:\n- why the observed failure occurred;\n- why previous attempts did not fix it;\n- what evidence distinguishes it from alternatives;\n- what smallest repair should change the outcome.\n\n### Stage 8 — Issue one bounded repair\nReturn to `fable-execute` or `fable-tdd` with one repair and one expected proof. If diagnosis reveals architecture uncertainty, route to plan/discover instead.\n\n## Decision Rules\n- Never repeat an unchanged failed command unless a named environmental/state variable changed or the rerun is explicitly measuring nondeterminism.\n- Do not delete caches/build artifacts reflexively before recording evidence; destructive cleanup can erase the clue that proves staleness.\n- If source changes have no runtime effect, prove the executed artifact/path before editing logic again.\n- If CI-only failure exists, compare environment/version/parallelism first; do not assume CI is \"random.\"\n- If failure is data-specific, minimize the failing data/state before broad refactor.\n- If a dependency/version hypothesis emerges, route external semantic verification to `fable-research`.\n- If local patches shift failure between components, suspect a shared invariant/design and return to `fable-plan`.\n- Similar failed fixes count as a failure streak; superficial syntax changes do not reset diagnostic responsibility.\n\n## Invariants\n- Recovery is diagnostic/read-only until a revised diagnosis exists.\n- Every new probe is chosen to distinguish hypotheses.\n- Contradicted hypotheses are removed.\n- Diagnostic instrumentation is temporary and cleaned after repair.\n- Final repair is bounded and tied to a predicted observable outcome.\n\n## Failure Taxonomy of Recovery Itself\n### Blind cleanup\nCache/build reset makes problem disappear but root cause is unknown. Record as unresolved unless causal evidence is obtained.\n\n### Hypothesis sprawl\nLong list of possibilities with no discriminating probes. Rank and test the cheapest separator.\n\n### Mutation during diagnosis\nAgent edits product while still uncertain, invalidating the failing state. Revert/restore diagnostic baseline where safe and restart evidence collection.\n\n### Confirmation bias\nOnly probes supporting first theory are run. Add at least one falsifier for the leading hypothesis.\n\n### Non-minimal reproduction\nHuge suite/system creates too many confounders. Isolate smaller path before interpreting results.\n\n## Anti-Patterns\n- third/fourth patch with same causal theory;\n- \"clear cache and see\" without recording before/after evidence;\n- rerunning flaky command until green;\n- blaming environment without comparing environments;\n- adding broad logging everywhere;\n- changing multiple diagnostic variables at once;\n- preserving disproved assumptions;\n- solving symptom while unable to explain previous failures.\n\n## Recovery Packet\n\n```text\nFailure/reproduction:\nAttempts already made:\nHypothesis queue:\nProbe → observation → hypothesis effect:\nRevised diagnosis:\nWhy prior attempts failed:\nBounded repair:\nExpected proof after repair:\nResidual uncertainty:\nNext Skill:\n```\n\n## Completion Criteria\nRecovery completes only when:\n- a reliable enough reproduction exists or nondeterminism is explicitly characterized;\n- leading alternative causes were falsified with evidence;\n- diagnosis materially differs from the failed assumption/attempt;\n- proposed repair is bounded and predicts a concrete changed observation;\n- execution can resume without another blind retry.\n\n## Progressive Resources\n- Deep guide: `references/diagnostic-falsification-playbook.md`\n- Existing ladder: `references/attribution-ladder.md`\n- Example: `examples/recovering-stale-test-cache.md`\n"
}

SHA-256 of public snapshot: b7130692389edc01e6ef647ba9d486c8f0f77ca7eb4e42a0b0c87a12a5f0dd99